Soniox TTS v2 launches with 60 languages and $0.70 hourly pricing
Soniox has released TTS v2, an AI model that supports over 60 languages, real-time streaming, and voice cloning from seconds of audio. The model enables developers to utilize audio tags for emotional control and supports precise synchronization via character-level timestamps.
Key Takeaways
- New model tts-rt-v2 enables high-fidelity voice cloning from only a few seconds of source audio.
- Integrated audio tags allow programmatic control of emotions and delivery styles like whispering or hesitation within a single passage.
- API pricing is set at $0.70 per generated hour, targeting high-volume developers in the US, Europe, and Japan.
- Support for 60+ languages includes 'language mixing,' allowing natural switching between different tongues in a single utterance.
Why It Matters
The release of Soniox TTS v2 intensifies the race for real-time, low-latency synthetic speech in the B2B sector. By offering granular emotional tags and character-level timestamps, Soniox addresses the technical hurdles of synchronizing AI voices with interactive characters and automated agents. This launch places significant price pressure on incumbents, arriving as the voice cloning market is projected to reach $3 billion in 2026 according to Mordor Intelligence. The inclusion of precise pronunciation for technical and financial codes suggests a strategic focus on mission-critical enterprise applications over simple creative narration. Industry observers should monitor if competitors match this sub-dollar hourly pricing to retain developer interest in the specialized voice API market.
Additional Context
The launch of Soniox TTS v2 arrives just as regulatory scrutiny of synthetic media reaches a critical turning point. Per the European Commission, the EU AI Act’s transparency obligations (Article 50) went live on August 2, 2026, mandating that all AI-generated audio content be clearly disclosed to listeners. This regulatory shift has forced providers to embed machine-readable watermarks and disclosure tools directly into their APIs to avoid administrative fines that can reach €15 million or 3% of global turnover.
Simultaneously, the competitive landscape for high-fidelity voice synthesis has consolidated around speed and emotional range. In May 2026, OpenAI confirmed the acquisition of voice cloning startup Weights.gg, integrating its technology to bolster the OpenAI Realtime API. Market data from ElevenLabs and HeyGen indicate that sub-200ms latency has become the new industry standard for low-latency voice agents, while the broader voice cloning sector is growing at a 25.8% CAGR toward an estimated $29.8 billion by 2036, per Future Market Insights.
In the U.S., legislative momentum continues to build around the NO FAKES Act, which aims to establish a federal right against unauthorized digital replicas of a person's voice. This follows the 2024 ELVIS Act in Tennessee, the first state law to explicitly protect human voices from unauthorized AI cloning. As a result, commercial providers like Soniox now emphasize explicit consent and identity preservation in their documentation to mitigate legal risks for enterprise clients using short-sample cloning technology.
Read full article at techgenyz.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source