ElevenLabs launches Dubbing V2 with emotion-aware voice cloning and API
ElevenLabs has launched Dubbing V2, an AI tool designed for video localization that includes voice cloning, emotion preservation, and a similarity slider. The company also announced plans to release an API to support automated, high-volume localization workflows.
Key Takeaways
- Dubbing V2 supports videos up to 180 minutes in length and requires a minimum of 11 seconds for accurate voice cloning.
- A new 'similarity slider' allows users to adjust the closeness of the voice clone to optimize for either identity preservation or natural target-language phonetics.
- Internal processing now separates the soundtrack from dialogue, allowing the AI to recreate performance characteristics like pacing, pauses, and delivery style.
- ElevenLabs confirmed an upcoming API release and specialized partner programs for creators and enterprise marketing teams to scale automated dubbing.
Why It Matters
The release shifts AI dubbing from mere literal translation to performance-matched localization, directly addressing the 'flat' robotic tone that has historically limited viewer retention in non-native markets. By automating the preservation of emotional nuance—previously a manual task for professional voice directors—ElevenLabs is lowering the barrier for mid-market streaming services to globalize their libraries at a fraction of traditional costs. This move intensifies competition with platform-native tools like YouTube’s auto-dubbing, forcing third-party vendors to compete on technical fidelity rather than just accessibility. Watch for enterprise adoption rates of the Dubbing V2 API as a signal for when localized-by-default becomes the standard for B2B video and niche OTT services.
Additional Context
The launch of Dubbing V2 follows a period of rapid financial and technical escalation for ElevenLabs. Per Bloomberg and Sifted (July 2026), the startup is currently in early talks for a secondary share sale that could value the company at approximately $22 billion, doubling the $11 billion valuation reached during its $500 million Series D round in February 2026. This trajectory reflects a broader industry pivot toward AI-driven localization as 'table stakes' for digital distribution. In early 2026, YouTube expanded its AI auto-dubbing tool to 27 languages and integrated expressive speech and lip-sync features to enhance visual realism for all eligible creators (per YouTube Blog, February 2026). Competitive pressure in the B2B space is also mounting as streaming infrastructure providers embed AI intelligence directly into live pipelines. Per Parks Associates (May 2026), companies like NVIDIA and Eluvio showcased at NAB 2026 that automated localization is moving toward 'agentic AI' systems, which coordinate end-to-end production tasks with minimal human intervention. Furthermore, enterprise leaders such as Synthesia and CAMB.AI are actively deploying voice-preserving technology for high-stakes implementations, including localized live sports commentary and large-scale corporate training. ElevenLabs' focus on performance and emotional fidelity targets the premium segment of this market, aiming to convert its reported $500 million in annual recurring revenue into a dominant 70% enterprise revenue split by 2027.
Read full article at youtube.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source