Google Gemini 3.5 Transcribe hits 2.6% error rate for audio processing
Google has launched Gemini 3.5 Transcribe, a speech-to-text model offering two distinct APIs for real-time and recorded audio processing. The model supports over 85 languages and features a 70% improvement in transcription speed over the previous Chirp 3 model, with reported word error rates of 2.6% for non-streaming and 4.0% for streaming.
Key Takeaways
- Dual-endpoint architecture separates real-time streaming from recorded audio processing to optimize for latency versus metadata.
- Reported word error rates reach 2.6% for non-streaming and 4.0% for streaming tasks, verified by Artificial Analysis.
- Smart mode automatically removes fillers and self-corrections but cannot be used simultaneously with diarization or timestamps.
- Live sessions are currently capped at 10 minutes, while recorded file processing supports up to one hour of audio.
Why It Matters
The launch of Google Gemini 3.5 Transcribe provides streaming platforms with a high-accuracy alternative for live captioning and real-time voice agents. By achieving a 70% speed increase over Chirp 3, Google is narrowing the gap between automated transcription and human-level performance for global media localization. This release pressures competitors like Agora and LiveKit to deepen their managed-service integrations as developers weigh the trade-offs between verbatim accuracy and structured 'smart' formatting. Industry observers should monitor the upcoming Chrome integration and the potential removal of the 10-minute limit for live streaming sessions.
Additional Context
Google's Gemini 3.5 Transcribe enters a crowded field of speech-to-text providers that have been racing to capture real-time captioning and voice-agent workloads. In early 2025, LiveKit launched its Agents framework with built-in speech-to-text pipelines supporting multiple model backends, positioning itself as an open-source orchestration layer that lets developers swap transcription engines without re-architecting their infrastructure. Pipecat, the open-source voice AI framework, has similarly expanded its supported STT providers, giving streaming developers a modular path that reduces lock-in to any single vendor's transcription model.
On the business side, Google has been aggressive about pricing and availability to drive adoption of its speech models. Google Cloud announced in mid-2025 that Chirp 2 would be available through the Cloud Speech-to-Text V2 API with pay-as-you-go pricing starting at $0.016 per minute, a rate that undercut several competing managed transcription services. The company has also been integrating its speech models directly into the Gemini API ecosystem, bundling transcription with its Interactions API for conversational AI use cases. Meanwhile, Agora raised $100 million in a Series C round in late 2024 to expand its real-time engagement platform, signaling investor confidence in the voice and video infrastructure layer that sits alongside transcription services.
Independent benchmarking has become a key differentiator as transcription accuracy claims proliferate. Artificial Analysis published a speech-to-text leaderboard in 2025 that evaluated models on word error rate, latency, and language coverage across standardized test sets, providing developers with third-party comparisons rather than relying solely on vendor-reported figures. Google's claimed 2.6% average WER across 85 languages would place Gemini 3.5 Transcribe near the top of such rankings, though independent verification on streaming workloads specifically remains limited. The model's 70% speed improvement over Chirp 3 is particularly relevant for live captioning pipelines, where end-to-end latency from audio capture to rendered text must stay under one second to meet broadcast and accessibility standards. Vercel's AI SDK has already added support for multiple STT providers, and the platform's edge runtime architecture enables transcription inference to run close to end users, reducing round-trip delays that compound in global streaming deployments.
Read full article at marktechpost.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source