Google launches Gemini 3.5 Live Translate with 70-language speech-to-speech support
Google has launched Gemini 3.5 Live Translate, a new speech-to-speech AI model supporting over 70 languages which is rolling out across consumer, enterprise, and developer products. The model will expand real-time translation features in Google Meet for Workspace customers and the Google Translate mobile app. It is also available via the Gemini Live API, allowing developers to build multilingual voice applications for various use cases including live broadcasting, with partners like Agora and LiveKit already integrating it.
Key Takeaways
- Supports over 70 languages and 2,000 language-pair combinations, a significant increase from the previous five-language limit in Google Meet.
- Preserves speaker prosody—including pitch, pacing, and intonation—rather than outputting flat synthetic audio.
- Integrated directly into Google Meet for enterprise, Google Translate for consumers, and available for developers via the Gemini Live API.
- Includes a new Android-specific 'listening mode' that streams private translated audio through the device earpiece without requiring headphones.
- Ecosystem partners including Agora and LiveKit are already integrating the model for live broadcasting and multilingual voice applications.
Why It Matters
The move shifts real-time translation from a disjointed speech-to-text-to-speech relay to a continuous, native-audio model. For the streaming industry, this reduces technical barriers for global live broadcasts and virtual events by minimizing the delay between original and translated speech. By decoupling Live Translate from general-purpose agents, Google is targeting high-utility, low-latency niches like customer support and international webinars. This heightens pressure on competitors like OpenAI and Zoom, who recently launched similar dedicated translation interpreters. Watch for the emergence of 'listening mode' style implementations in professional streaming apps to enable private, real-time localized audio feeds during high-stakes live events.
Additional Context
The launch of Gemini 3.5 Live Translate follows a month of intense competition in the voice AI sector. Per Slator and official reports, OpenAI released its own trio of models—including GPT-Realtime-Translate—on May 7, 2026. This OpenAI offering supports 70 input languages and is explicitly positioned as an interpreter specialist rather than a general assistant, mirroring Google’s strategic branding. These developments signal a move across the industry to unbundle specific voice tasks like transcription and translation from broader, more expensive LLM reasoning tasks to reduce operational latency. Simultaneously, the video conferencing market has standardized real-time translation as a core feature. As of June 2026, Zoom’s AI Companion 3.0 supports real-time translation in over 50 languages, integrated directly into its 'conversational work surface' for both live captions and post-meeting summaries. According to Verbit, by mid-2026, multilingual streaming has transitioned from an experimental feature to a consumer expectation, particularly for global sports and educational content. This shift is reflected in the infrastructure layer, as AWS Media Services recently updated its portfolio to automate AI-driven live dubbing and CEA-608 caption embedding at scale, per AWS product documentation from April 2026. Industry benchmarks are also evolving to keep pace with these native audio models. Unlike historical text-based metrics, new evaluations like AutoMQM are being adopted to assess error rates in continuous audio streams. However, as noted in recent Slator reporting, both Google and OpenAI have faced criticism for withholding specific latency and accuracy data, making it difficult for enterprise developers to compare Gemini 3.5 against competitive stacks. Despite the lack of raw benchmarks, the rapid adoption by platforms like Agora and Grab suggests that the industry is prioritizing the fluidity and naturalness of speaker-replicated audio over pure linguistic accuracy.
Read full article at slator.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source