Google has launched two new text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, designed for enterprise applications including audiobooks and automated video production. The models feature support for up to 130 languages, custom voice cloning, and integrated content provenance tools like SynthID and C2PA.
The launch of these models provides streaming and video production teams with high-fidelity tools to automate localization and narration at scale. By integrating SynthID and C2PA standards, Google is addressing the growing enterprise requirement for verifiable content provenance in synthetic media. This move intensifies competition in the AI audio space, as Google leverages its cloud infrastructure to offer lower-latency inference compared to specialized startups. The inclusion of oratory cues for pacing and non-lexical vocalizations suggests a shift toward more naturalistic AI narration for long-form content. Watch for how these models are integrated into Google Vids to streamline B2B video workflows.
Google's Gemini 3.8 Flash TTS and Flash-Lite TTS models are part of a broader Gemini Audio family that has expanded rapidly in September 2026. The TTS launch follows Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, which Google released on September 15 for real-time voice reasoning, giving developers a full pipeline from live speech understanding to expressive speech generation within the same model family. Google has positioned the TTS models as complementary to these live audio models, targeting different use cases: Live for conversational agents and TTS for content production at scale.
On benchmarks, Google's Flash TTS secured the number-one position on Artificial Analysis's Pronunciation Robustness benchmark at 89.5%, pushing xAI's TTS offering down to third place. Google now holds three of the top four slots on that leaderboard, a concentration that matters for enterprise buyers evaluating voice-AI contracts based on objective quality scores. The company also reported first and second place finishes on Hume AI's Overall Quality Index and Voice Design Benchmark, with Flash TTS scoring 71.4 on the latter and leading in accent modeling at 60.8.
For video production workflows specifically, Google has named integration partners including HeyGen, Wondercraft, and Ollang, which are integrating the new TTS models to accelerate global dubbing and localize media with regional accents. Developer platforms Agora, LiveKit, Pipecat, and Vercel also support the Gemini API for building speech generation experiences. Flash-Lite TTS is rolling out directly inside Google Vids, Google's enterprise video creation tool, while Flash TTS targets higher-fidelity applications such as audiobooks and interactive media through Google AI Studio and the Gemini API.
Google has launched its Gemini Flash TTS and Flash-Lite TTS models, designed to provide high-fidelity, multilingual audio for enterprise applications like Google Vids. By integrating SynthID watermarking and C2PA standards, these models offer verifiable content provenance, helping production teams automate localization and narration while meeting growing industry requirements for synthetic media authentication.
Google has introduced Gemini 3.8 Flash TTS and Flash-Lite TTS, which are designed to provide high-fidelity synthetic audio for enterprise video production workflows.
The Flash TTS model supports up to 130 languages, while the cost-optimized Flash-Lite TTS model supports 101 languages.
Google integrates SynthID watermarking and C2PA records directly into the TTS pipeline to provide inaudible tracking and metadata for AI-generated audio files.
The new Gemini TTS models secured the top two spots on Hume AI Inc. benchmarks for audio quality and human-evaluated performance.
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source