Modulate launches Velma AI to detect platform risks via acoustic signals
Modulate has developed Velma, an AI voice intelligence platform that analyzes acoustic signals beyond simple transcription to detect risks like fraud, churn, deepfakes, and compliance issues in real-time conversations. The platform and its associated APIs (including Deepfake Detection and PII Redaction) are designed for enterprise use, offering features for fraud prevention, AI agent guardrails, and trust & safety in various voice-enabled environments.
Key Takeaways
- Velma platform detects sarcasm, emotion, and speaker dynamics that traditional text-based transcription pipelines typically discard.
- Integrated Deepfake Detection API claims a 98.9% accuracy rate for identifying synthetic audio streams.
- Activision is currently using the technology to identify and mitigate disruptive player behavior within its multiplayer titles.
- Platform compliance includes ISO 27001 certification and adherence to HIPAA, GDPR, and the EU AI Act.
Why It Matters
Real-time voice moderation is shifting from simple keyword filtering to native acoustic analysis, helping platforms manage the liability risks of unscripted live audio. As streaming services integrate more social and interactive voice features, the ability to isolate deepfakes and PII in sub-second intervals becomes a core infrastructure requirement rather than a secondary safety feature. This development forces a move away from the traditional 'transcription-then-LLM' workflow, which often loses critical contextual signals like stress or intent. Watch for whether major cloud providers respond by acquiring these specialized audio-native models to bolster their own generic speech-to-text offerings.
Additional Context
The rollout of Velma occurs as the regulatory landscape for synthetic media tightens significantly. Per reports from StackCyber and 5Calls in May 2026, the federal 'TAKE IT DOWN Act' now mandates that interactive computer services implement 48-hour notice-and-takedown windows for non-consensual deepfake content. Failure to comply can result in federal criminal penalties, placing immense pressure on streaming and social platforms to adopt automated detection tools capable of operating at scale. This follows the reintroduction of the 'NO FAKES Act' in April 2025, which aims to establish a federal property right over an individual's voice and likeness. Investment in the sector has accelerated alongside these legal shifts. According to data from Tracxn and VoxCloneAI, venture funding for voice AI surged to $2.1 billion in 2025, an eightfold increase from 2022 levels. By mid-2026, the sector had already attracted over $550 million in new capital across specialized players like ElevenLabs, which reached an $11 billion valuation, and Vapi, which secured a $50 million Series B in May 2026. This capital influx is driving a move toward performance-based service level agreements (SLAs) in enterprise contracts, particularly around latency and detection accuracy floors. In the streaming industry specifically, AI has transitioned from experimental use to a foundational stack layer. Reporting from Streaming Media and MwareTV in early 2026 indicates that nearly 67% of Fortune 500 firms now utilize production-grade voice AI, with platforms like Netflix and Twitch increasingly relying on real-time moderation to protect brand safety during live events. As low-level infrastructure such as CDN and encoding becomes commoditized, market differentiation is increasingly defined by the 'AI features layer,' which handles semantic search, live translation, and the proactive detection of toxicity.
Read full article at modulate.ai
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source