Google Gemini 3.5 Transcribe launches with sub-second latency for live streaming
Google has launched Gemini 3.5 Transcribe in public preview, featuring separate APIs for real-time streaming and recorded audio processing. The service offers sub-second latency for live applications and includes speaker attribution and word-level timestamps for recorded content.
Key Takeaways
- Live API provides continuous bidirectional streaming for voice agents and real-time captioning
- Interactions API delivers word-level timestamps and speaker labels for recorded call logs and meetings
- Model supports more than 85 languages and handles alphanumeric entities like order IDs and postal codes
- Performance testing shows word error rates of 4.0% for streaming and 2.6% for non-streaming tasks
Why It Matters
The introduction of specialized APIs for live and recorded audio signals a shift toward intent-aware transcription that cleans filler words and resolves self-corrections in real time. For the streaming ecosystem, this reduces the technical friction of integrating low-latency captioning and voice-driven navigation into media applications. By improving time-to-final transcription by 70% over the previous Chirp 3 model, Google is positioning its AI stack as a viable alternative for high-volume enterprise workflows. The critical signal to watch is the release of product-specific pricing, which will determine if the service is cost-competitive against established speech-to-text providers for production-scale deployments.
Additional Context
Google's speech-to-text capabilities are entering an increasingly crowded market where cloud providers and specialized vendors are racing to deliver production-ready transcription for media and enterprise workflows. In May 2026, Google published new documentation on optimizing websites for generative AI features in Search, signaling a broader strategy to make its AI stack the default layer for content discovery and processing. The Gemini 3.5 Transcribe launch extends that strategy into real-time audio, positioning Google against dedicated speech AI providers that have built deep integrations with cloud infrastructure. Deepgram, one of the most direct competitors, has been aggressively expanding its enterprise footprint through cloud-native deployments. Deepgram's integration with AWS IAM temporary delegation provides scoped, time-bound access for support engineers directly to SageMaker endpoints, enabling customers to run real-time speech-to-text and voice agent workloads inside their own VPCs without routing data to external regions. The company's Flux model achieves sub-300 millisecond end-to-end latency for streaming transcription, directly competing with the sub-second latency claims Google is making for Gemini 3.5 Transcribe in live applications. The competitive dynamics around speech AI are also being shaped by how major platforms approach AI-driven content optimization and agentic search. Akamai introduced AI Brand Presence to help organizations optimize their content for AI search and traffic, combining AI-optimized context delivery with visibility dashboards that track which AI models are consuming site content. Akamai reported an 85% increase in citations and a 364% surge in brand presence for general searches after deploying the technology on its own website, with a 133% presence jump in ChatGPT compared to competitors. This illustrates the broader ecosystem shift where AI models are not just transcribing content but actively consuming and reshaping how it is discovered, creating new demand for high-quality, low-latency transcription pipelines that can feed both human and machine audiences.
Read full article at superpowerdaily.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source