StreamingMemeStreamingMemeBuyers Guide
AboutLeaderboardsEventsSubmit News
Subscribe

Daily Brief

The streaming industry in your inbox every morning.

Daily Brief

The streaming industry in your inbox every morning.

StreamingMemeStreamingMeme

StreamingMeme is the streaming technology industry news aggregator.

Explore

Buyers GuideLeaderboardsEventsSubmit News

Stay updated

Weekly digest of new companies and streaming news.

Categories

Encoding & SoftwareVideo Delivery & CDNStreaming PlatformsAI for VideoProduction HardwareBusiness NewsMonetization & Ad TechRegulatory & Policy

© 2026 StreamingMeme. All rights reserved.

AboutPrivacy PolicyTermsContact
EncodingCDNPlatformsAI & VideoHardwareBusinessAd TechPolicyIBC Guide
← AI for Video
AI & VideoTechnical DevelopmentAugust 28, 2026

Gladia Solaria-1 STT latency hits sub-300ms to enable fluid voice agents

Gladia Solaria-1 STT latency hits sub-300ms to enable fluid voice agents
Gladia

Gladia has released performance benchmarks for its Solaria-1 speech-to-text model, claiming sub-300ms latency for final segments and under 103ms for initial partial transcripts. The company emphasizes the importance of P99 and Time to Last Byte (TTLB) metrics over average latency for production voice agent reliability.

Key Takeaways

  • Solaria-1 achieves sub-300ms Time to Last Byte (TTLB) for final segments and under 103ms for initial partial transcripts.
  • Benchmarks indicate that P99 latency above 500ms causes conversational collisions where users begin repeating themselves.
  • The model supports mid-stream code-switching and automatic language detection without requiring session restarts.
  • Managed API costs for real-time streaming start at $0.25 per hour, compared to approximately $1.00 per hour for self-hosted AWS g5.xlarge instances.

Why It Matters

Achieving sub-300ms latency for final transcripts provides the necessary headroom for complex streaming stacks that include LLM reasoning and TTS generation. By prioritizing P99 stability over averages, Gladia addresses the 'long tail' of latency spikes that typically break natural turn-taking in voice-based AI applications. This technical benchmark pressures competitors like Deepgram and AssemblyAI to provide more transparent tail-latency data rather than controlled lab averages. As streaming video platforms integrate more interactive AI features, the focus will shift from simple transcription accuracy to the end-to-end conversational budget. Watch for whether Solaria-3 maintains these latency profiles while expanding support for non-Latin scripts and high-noise environments.

Additional Context

Gladia's Solaria-1 benchmarks arrive amid intensifying competition among real-time speech-to-text providers vying for voice-agent workloads. In early 2025, Deepgram launched its Nova-3 model with claimed median latency under 250ms for streaming transcription, positioning it specifically for conversational AI and contact-center applications. AssemblyAI has similarly pushed low-latency streaming capabilities, with its Universal-Streaming model advertising first-token latency of approximately 130ms in production environments, a figure that targets the same voice-agent pipeline budgets Gladia is courting. Speechmatics, meanwhile, has focused on multilingual breadth alongside speed, announcing support for 50+ languages in its real-time streaming API with latency targets competitive to sub-300ms thresholds. The competitive framing matters because voice-agent developers typically allocate a total conversational budget of 800ms to 1.2 seconds across STT, LLM inference, and TTS, meaning every millisecond saved at the transcription layer compounds downstream.

On the business and integration side, Gladia has been building partnerships to embed its STT layer into production voice platforms. Gladia announced a partnership with Aircall in 2025 to power real-time transcription and analytics within the cloud contact-center platform, giving it direct access to enterprise telephony workloads where latency SLAs are contractually enforced. ElevenLabs, a leading text-to-speech provider frequently paired with STT engines in voice-agent stacks, raised $180 million in a January 2025 Series C at a $3.3 billion valuation, signaling investor confidence that the full conversational pipeline (STT + LLM + TTS) is maturing into a distinct infrastructure category. That funding wave pressures every component provider, including Gladia, to publish verifiable latency data rather than marketing averages, since integrators now benchmark end-to-end turn-taking quality.

From a technical measurement standpoint, Gladia's emphasis on P99 and TTLB metrics reflects a broader industry shift toward tail-latency accountability. A 2025 study by researchers at Stanford's Center for Research on Foundation Models found that P99 latency, not mean latency, was the strongest predictor of user-perceived responsiveness in voice-agent interactions, validating Gladia's methodological framing. Separately, Speechmatics published its own latency benchmark methodology in mid-2025, adopting percentile-based reporting (P50, P95, P99) rather than single-number averages, suggesting the industry is converging on standardized disclosure. For streaming video platforms exploring interactive AI overlays or live-captioning features, these benchmarking standards will likely inform vendor selection criteria as the technology moves from experimental to production deployments.


Read full article at gladia.io

Enjoy our coverage?

Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.

Add as preferred source

Related Articles

Fora Soft: Fora Soft benchmarks cascaded AI pipelines for 800ms live video translation
VentureBeat: Runway uses model distillation to slash real-time video latency by 90%
Wowza: Decoupling inference from delivery infrastructure optimizes custom AI video workflows
Sports Video Group: Fox Sports Scales AR Operations for 104-Match FIFA World Cup
University of Ottawa (uO Research): AI-assisted super-resolution cuts cloud gaming bandwidth by 56%
Get this in your inbox → Subscribe

Newest

about 15 hours ago
Pulse 2.0: Verizon scales Google Cloud AI partnership to automate network and marketing
about 15 hours ago
stackcompass.dev: EU AI Act content labeling mandates three-tier taxonomy for synthetic media
about 15 hours ago
GetDeploying: Salad undercuts Vast.ai on RTX 5090 distributed GPU cloud pricing
about 15 hours ago
Pipeline Publishing: NVIDIA data shows 89% of operators increasing telecom AI-native architectures spend
about 15 hours ago
MediaPost: FTC weighs lawsuit against YouTube content moderation and demonetization policies
about 15 hours ago
Pulse 2.0: Superstep Capital backs Zencore ZenAI Factory launch for Google Cloud
1 day ago
Kyiv Post: Ukraine petitions ITU to block Russian Rassvet satellites over its territory
1 day ago
CryptoSlate: IREN AI cloud revenue hits $128M amid $639M hardware impairment
1 day ago
Shattered Media: AWS Lambda SnapStart latency drops to 90ms for Java workloads
1 day ago
Content+Technology: AMWA and EBU advance Dynamic Media Facility roadmap at IBC2026
1 day ago
ScanX: Twelve states sue to block $110 billion Warner Bros. Paramount Skydance merger
1 day ago
Cyber Security News: Malvertising infrastructure threats now drive 45.9% of PropellerAds campaign rejections
1 day ago
Ad-hoc-news.de: Innovid Q2 2026 earnings show narrowed losses on $114.5M revenue
1 day ago
Ad-hoc-news.de: Navitas Semiconductor Claros acquisition targets AI data center power delivery
1 day ago
Glitchwire: RIAA and SAG-AFTRA AI music labeling framework creates major label loophole
1 day ago
Marktechpost: Google Gemini Omni 1.1 Flash adds 40-second video scene extension
1 day ago
Medium: Pipecat voice AI framework launches to solve real-time streaming interruption challenges
1 day ago
IoT Portal: RISC-V RVA23 profile mandates vector extensions for efficient edge AI silicon
1 day ago
Reuters: ESPN US Open RedZone brings whip-around coverage to 16 tennis courts
1 day ago
groundcover: Groundcover analysis reveals eBPF monitoring performance overhead reaches 41% in high-concurrency workloads

Upcoming Events

Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
Sep
29–1
SCTE TechExpoAtlanta
Sep
29–30
SportsPro AI+TechLondon
View all events →

Top Sources

  1. 1.PPC Land73
  2. 2.TVNewsCheck61
  3. 3.SiliconANGLE54
  4. 4.Sports Video Group50
  5. 5.AdExchanger41
  6. 6.Advanced Television40
  7. 7.Beet.TV38
  8. 8.MediaPost35
Full leaderboards →

Newest

about 15 hours ago
Pulse 2.0: Verizon scales Google Cloud AI partnership to automate network and marketing
about 15 hours ago
stackcompass.dev: EU AI Act content labeling mandates three-tier taxonomy for synthetic media
about 15 hours ago
GetDeploying: Salad undercuts Vast.ai on RTX 5090 distributed GPU cloud pricing
about 15 hours ago
Pipeline Publishing: NVIDIA data shows 89% of operators increasing telecom AI-native architectures spend
about 15 hours ago
MediaPost: FTC weighs lawsuit against YouTube content moderation and demonetization policies
about 15 hours ago
Pulse 2.0: Superstep Capital backs Zencore ZenAI Factory launch for Google Cloud
1 day ago
Kyiv Post: Ukraine petitions ITU to block Russian Rassvet satellites over its territory
1 day ago
CryptoSlate: IREN AI cloud revenue hits $128M amid $639M hardware impairment
1 day ago
Shattered Media: AWS Lambda SnapStart latency drops to 90ms for Java workloads
1 day ago
Content+Technology: AMWA and EBU advance Dynamic Media Facility roadmap at IBC2026
1 day ago
ScanX: Twelve states sue to block $110 billion Warner Bros. Paramount Skydance merger
1 day ago
Cyber Security News: Malvertising infrastructure threats now drive 45.9% of PropellerAds campaign rejections
1 day ago
Ad-hoc-news.de: Innovid Q2 2026 earnings show narrowed losses on $114.5M revenue
1 day ago
Ad-hoc-news.de: Navitas Semiconductor Claros acquisition targets AI data center power delivery
1 day ago
Glitchwire: RIAA and SAG-AFTRA AI music labeling framework creates major label loophole
1 day ago
Marktechpost: Google Gemini Omni 1.1 Flash adds 40-second video scene extension
1 day ago
Medium: Pipecat voice AI framework launches to solve real-time streaming interruption challenges
1 day ago
IoT Portal: RISC-V RVA23 profile mandates vector extensions for efficient edge AI silicon
1 day ago
Reuters: ESPN US Open RedZone brings whip-around coverage to 16 tennis courts
1 day ago
groundcover: Groundcover analysis reveals eBPF monitoring performance overhead reaches 41% in high-concurrency workloads

Upcoming Events

Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
Sep
29–1
SCTE TechExpoAtlanta
Sep
29–30
SportsPro AI+TechLondon
View all events →

Top Sources

  1. 1.PPC Land73
  2. 2.TVNewsCheck61
  3. 3.SiliconANGLE54
  4. 4.Sports Video Group50
  5. 5.AdExchanger41
  6. 6.Advanced Television40
  7. 7.Beet.TV38
  8. 8.MediaPost35
Full leaderboards →