Profuz Digital adds ElevenLabs and face detection to media asset platform
Profuz Digital is set to showcase an AI-enhanced update to its LAPIS media management platform at IBC2026. The new release features integrated speech recognition, face detection, and content fingerprinting, alongside support for the ElevenLabs API to improve broadcast asset management and localization workflows.
Key Takeaways
- Integrated ElevenLabs API support enables specialized AI-driven speech generation and audio localization.
- New multi-source metadata analysis combines speech recognition, face detection, and content fingerprinting in layered processing steps.
- Subtitle generation is enhanced through SubtitleNEXT integration, supporting live subtitle stream generation and automated transcription.
- Storage and lifecycle upgrades include native S3 Glacier integration and automated file duplication checking.
- Current enterprise users of the LAPIS platform include Canal+ France, the Council of Europe, and Bulgarian National Radio.
Why It Matters
The integration of high-fidelity audio engines like ElevenLabs directly into media asset management (MAM) platforms signals a shift toward fully automated, multi-modal localization pipelines. By layering face detection and content fingerprinting on top of standard transcription, Profuz is addressing the metadata bottleneck that slows down large-scale archive searches. For the broader ecosystem, this move forces a convergence between production hardware and generative AI services, where the MAM serves as an active orchestration layer rather than a passive repository. Watch for whether major broadcasters adopt these automated workflows for Tier 1 live content or keep them restricted to archival and promotional materials.
Additional Context
The integration of specialized AI audio into media management arrives as the global AI transcription and localization market faces rapid expansion. Per market analysis from Sonix in January 2026, the AI meeting and media transcription segment is projected to grow at a 25.6% compound annual growth rate through 2034, driven by enterprise demands for 99% accuracy and real-time processing speeds. By late 2025, leading tools like OpenAI’s Whisper had already set new baselines for multilingual accuracy, making automated speech-to-text a standard requirement for broadcast asset management systems like Profuz LAPIS. ElevenLabs has simultaneously solidified its position in the professional media sector through aggressive high-profile partnerships. In November 2025, the company launched its "Iconic Marketplace" to ethically license voices from actors such as Michael Caine for creative projects. According to reports from Pulse2 in May 2026, ElevenLabs reached an annual recurring revenue of $500 million in early 2026, fueled by enterprise deployments in dubbing and marketing. These developments have transformed AI voice from an experimental tool into a foundational piece of the media technology stack. The upcoming IBC 2026 exhibition at the Amsterdam RAI (September 11-14) is expected to emphasize "Agentic AI"—systems capable of autonomous task distribution across multiple processing engines. As noted by Profuz Digital leadership following IBC 2025, the industry focus has shifted from simple AI implementation to intelligent orchestration, where MAM systems automatically select the best AI engine for specific tasks based on language or content genre. This trend aligns with recent Trusted Partner Network (TPN) assessments for localization platforms, as broadcasters demand both high automation and strict content security in cloud-based workflows.
Read full article at tvnewscheck.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source