Google Gemma 4 outperforms OpenAI GPT-5 in voice AI latency benchmarks
LiveKit has published a benchmark report comparing the performance and cost-efficiency of various LLMs for voice AI agent scenarios, such as hotel receptionists. The data indicates that Google's Gemma 4 31B model outperformed several OpenAI GPT models in terms of time-to-first-token latency and operational cost.
Key Takeaways
- Gemma 4 31B recorded a median Time-to-First-Token (TTFT) of 136ms, nearly five times faster than GPT-5.4's 626ms.
- Operational costs for Gemma 4 31B sit at $0.40 per 1M input tokens, compared to $5.00 for GPT-5.4 and $10.00 for GPT-5.5.
- OpenAI's GPT-4.1 Nano is the only model to underprice Gemma 4 on input ($0.10) but suffers from a 379ms median TTFT.
- Google's Gemini 3.5 Flash showed the highest median latency in the group at 775ms despite an 86% pass rate.
- Gemma 4 maintained quality parity with a pass rate of 82%, matching the performance of much larger models like GPT-5.4.
Why It Matters
The sub-150ms TTFT achieved by Gemma 4 moves open-weight models from experimental to production-ready for real-time voice applications where human-like response times are critical. For the streaming and communications ecosystem, this shifts the competitive focus from raw reasoning capabilities to 'time-to-ear' metrics, where Google now holds a distinct infrastructure lead. Developers can now deploy frontier-grade voice agents with 90% lower token costs than OpenAI’s flagship tiers. Watch for OpenAI to respond with further 'Nano' or 'Instant' tier optimizations to reclaim the low-latency crown as voice-to-voice interaction becomes the primary interface for agentic streaming services.
Additional Context
The release of Gemma 4 in April 2026 marked a strategic shift for Google DeepMind, adopting the Apache 2.0 license to encourage commercial adoption and local deployment. Per Layer3 Labs (July 2026), the Gemma 4 family includes variants ranging from 2B parameters for mobile to 31B dense models for workstations, with the 31B variant ranking #3 on the Arena AI open-source leaderboard. This open-weight approach directly challenges OpenAI's closed-source dominance, particularly as developers seek to avoid vendor lock-in and high per-token API fees in high-volume production environments.
OpenAI has countered by diversifying its GPT-5 lineup, which debuted in August 2025. According to AI Pricing Guru (August 2026), the current GPT-5.6 family consists of the Luna, Terra, and Sol tiers, with Luna targeting high-volume workloads at $0.20 per million input tokens. While OpenAI retains a lead in complex reasoning and long-context tasks—supporting windows up to 1.05M tokens—LiveKit's benchmarks suggest that for specialized voice tasks, the larger parameter counts and complex routing of the GPT-5 series introduce latency penalties that open-weight models like Gemma 4 avoid.
Industry standards for voice AI now emphasize sub-second round-trip latency to maintain conversational flow. Per Trillet (August 2026), human turn-taking gaps typically average 200ms to 400ms; when AI response time exceeds 800ms, user retention drops significantly. Infrastructure providers like Telnyx and Retell are increasingly co-locating LLM inference with telephony stacks to shave milliseconds off the total path. LiveKit's data confirms that model selection is now the primary bottleneck, with Gemma 4's 136ms TTFT providing the necessary 'headroom' for speech-to-text and text-to-speech processing within a 500ms total budget.
Read full article at livekit.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source