Gladia Solaria-1 STT latency hits sub-300ms to enable fluid voice agents
Gladia has released performance benchmarks for its Solaria-1 speech-to-text model, claiming sub-300ms latency for final segments and under 103ms for initial partial transcripts. The company emphasizes the importance of P99 and Time to Last Byte (TTLB) metrics over average latency for production voice agent reliability.
Key Takeaways
- Solaria-1 achieves sub-300ms Time to Last Byte (TTLB) for final segments and under 103ms for initial partial transcripts.
- Benchmarks indicate that P99 latency above 500ms causes conversational collisions where users begin repeating themselves.
- The model supports mid-stream code-switching and automatic language detection without requiring session restarts.
- Managed API costs for real-time streaming start at $0.25 per hour, compared to approximately $1.00 per hour for self-hosted AWS g5.xlarge instances.
Why It Matters
Achieving sub-300ms latency for final transcripts provides the necessary headroom for complex streaming stacks that include LLM reasoning and TTS generation. By prioritizing P99 stability over averages, Gladia addresses the 'long tail' of latency spikes that typically break natural turn-taking in voice-based AI applications. This technical benchmark pressures competitors like Deepgram and AssemblyAI to provide more transparent tail-latency data rather than controlled lab averages. As streaming video platforms integrate more interactive AI features, the focus will shift from simple transcription accuracy to the end-to-end conversational budget. Watch for whether Solaria-3 maintains these latency profiles while expanding support for non-Latin scripts and high-noise environments.
Additional Context
Gladia's Solaria-1 benchmarks arrive amid intensifying competition among real-time speech-to-text providers vying for voice-agent workloads. In early 2025, Deepgram launched its Nova-3 model with claimed median latency under 250ms for streaming transcription, positioning it specifically for conversational AI and contact-center applications. AssemblyAI has similarly pushed low-latency streaming capabilities, with its Universal-Streaming model advertising first-token latency of approximately 130ms in production environments, a figure that targets the same voice-agent pipeline budgets Gladia is courting. Speechmatics, meanwhile, has focused on multilingual breadth alongside speed, announcing support for 50+ languages in its real-time streaming API with latency targets competitive to sub-300ms thresholds. The competitive framing matters because voice-agent developers typically allocate a total conversational budget of 800ms to 1.2 seconds across STT, LLM inference, and TTS, meaning every millisecond saved at the transcription layer compounds downstream.
On the business and integration side, Gladia has been building partnerships to embed its STT layer into production voice platforms. Gladia announced a partnership with Aircall in 2025 to power real-time transcription and analytics within the cloud contact-center platform, giving it direct access to enterprise telephony workloads where latency SLAs are contractually enforced. ElevenLabs, a leading text-to-speech provider frequently paired with STT engines in voice-agent stacks, raised $180 million in a January 2025 Series C at a $3.3 billion valuation, signaling investor confidence that the full conversational pipeline (STT + LLM + TTS) is maturing into a distinct infrastructure category. That funding wave pressures every component provider, including Gladia, to publish verifiable latency data rather than marketing averages, since integrators now benchmark end-to-end turn-taking quality.
From a technical measurement standpoint, Gladia's emphasis on P99 and TTLB metrics reflects a broader industry shift toward tail-latency accountability. A 2025 study by researchers at Stanford's Center for Research on Foundation Models found that P99 latency, not mean latency, was the strongest predictor of user-perceived responsiveness in voice-agent interactions, validating Gladia's methodological framing. Separately, Speechmatics published its own latency benchmark methodology in mid-2025, adopting percentile-based reporting (P50, P95, P99) rather than single-number averages, suggesting the industry is converging on standardized disclosure. For streaming video platforms exploring interactive AI overlays or live-captioning features, these benchmarking standards will likely inform vendor selection criteria as the technology moves from experimental to production deployments.
Read full article at gladia.io
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source