Modulate and Scam.ai partner to combat deepfakes in video conferencing
Modulate and Scam.ai have announced a strategic partnership to integrate Modulate's synthetic voice detection technology into the Scam.ai platform. The collaboration aims to provide a unified solution for identifying AI-generated deepfakes across audio, video, and image communication channels.
Key Takeaways
- Modulate's Velma platform uses an Ensemble Listening Model to analyze tone, timing, and behavioral cues to distinguish human speech from AI.
- Scam.ai expands its existing image and document verification capabilities to include real-time video stream analysis for face-swap detection.
- The integrated solution provides confidence scores to help users identify multi-vector attacks involving synthetic documents, voices, and video avatars.
- Modulate currently holds the top ranking on the Speech Deepfake Detection Arena leaderboard on Hugging Face.
Why It Matters
The integration of specialized audio and visual detection tools addresses the growing threat of multi-modal deepfakes that bypass single-layer security protocols. As streaming infrastructure increasingly supports high-fidelity, real-time communication, the ability to verify biological presence becomes a critical requirement for enterprise video platforms. This partnership signals a shift toward holistic verification stacks where synthetic voice detection must work in tandem with computer vision to maintain trust in remote environments. Industry observers should monitor the adoption rates of these detection APIs by major teleconferencing providers as they face rising pressure to mitigate AI-enabled social engineering and corporate espionage.
Additional Context
Modulate has positioned its voice analysis technology as a critical layer in the emerging deepfake detection stack, and the partnership with Scam.ai reflects a broader industry trend toward multi-modal verification. In early 2025, Modulate raised a $3.2 million seed round led by General Catalyst to scale its synthetic voice detection platform, signaling investor confidence that audio-based deepfake identification would become a standalone market category. The company's Velma product, which analyzes vocal micro-patterns to distinguish human speech from AI-generated audio, has been deployed across financial services and government communication channels where voice impersonation poses direct monetary risk.
The regulatory environment is tightening around synthetic media detection, creating tailwinds for companies like Modulate and Scam.ai. In March 2025, the U.S. Senate Commerce Committee advanced the DEFIANCE Act, which would create a federal cause of action for victims of non-consensual deepfake imagery, establishing legal precedent that could extend to audio and video deepfakes used in corporate fraud. Meanwhile, the Federal Communications Commission proposed rules in January 2025 requiring telecom carriers to implement caller ID authentication frameworks that could incorporate synthetic voice detection as part of broader anti-robocall enforcement. These regulatory moves create compliance pressure on enterprise communication platforms to integrate detection tools proactively rather than reactively.
On the technical side, independent benchmarking of deepfake detection systems remains limited but is growing. A January 2025 study from the University of Buffalo found that state-of-the-art voice cloning models can fool commercial detection systems up to 40% of the time when adversarial perturbations are applied, underscoring why layered approaches combining audio and visual analysis outperform single-modality solutions. Scam.ai's platform, which already handles video and image forensics, adds Modulate's audio layer to create what the companies describe as a unified detection API. Competing efforts include Microsoft's Video Authenticator tool, which the company expanded in late 2024 to cover real-time video streams in Teams calls, and Intel's FakeCatcher system, which analyzes photoplethysmography signals in video to detect synthetic faces with 96% accuracy in controlled tests. The convergence of these approaches suggests that enterprise video platforms will increasingly bundle multi-modal detection as a default feature rather than an add-on.
Read full article at eejournal.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source