NVIDIA's New Nemotron ASR Transcribes 40 Languages with Configurable Latency
NVIDIA has released Nemotron 3.5 ASR, a 600M-parameter, open-weights streaming Automatic Speech Recognition (ASR) model capable of transcribing 40 language-locales in real-time. This model offers configurable latency from 80ms to 1.12s at inference time and is self-hostable, providing flexibility for integrators.
Key Takeaways
- Nemotron 3.5 ASR is a 600M-parameter open-weights streaming ASR model, released under OpenMDW-1.1.
- It transcribes 40 language-locales using a single checkpoint, without per-language models.
- Latency is configurable at inference time from 80ms to 1.12s via `att_context_size`, with no retraining required.
- The Cache-Aware FastConformer-RNNT architecture processes each audio frame once, reporting 17x concurrent streams on an H100 versus buffered approaches.
- Fine-tuning the open weights reduced WER by 31-32% on Greek and Bulgarian at the 80ms setting.
Why It Matters
NVIDIA's Nemotron 3.5 ASR offers a single, configurable model for real-time, multilingual speech transcription, which reduces operational complexity for developers. Its open-weights approach, coupled with demonstrable fine-tuning gains, allows custom domain adaptation not always available with closed API solutions. This could shift competitive dynamics for companies offering ASR services by lowering the barrier to entry for high-performance, real-time transcription across diverse languages. Watch for adoption rates among enterprise streaming platforms and specialized media companies requiring self-hosted, fine-tuned ASR capabilities.
Additional Context
The launch of Nemotron 3.5 ASR arrives as the speech-to-text market enters a phase of extreme specialization and efficiency. Per recent benchmarks from Smallest.ai in May 2026, streaming ASR performance now hinges on balancing sub-300ms latency with accent resilience, where models like OpenAI’s GPT-4o Mini and Deepgram’s Nova-3 compete for dominance in European and Asian locales. NVIDIA’s move toward 'cache-aware' inference addresses a critical bottleneck: traditional buffered streaming often recomputes overlapping audio chunks, consuming excessive VRAM and compute cycles. According to NVIDIA technical benchmarks from early 2026, this architectural shift allows a single H100 to sustain up to 560 concurrent streams at a 320ms chunk size, roughly 3x the capacity of previous-generation systems. Simultaneously, competitors are pushing deeper into specialized enterprise features. ElevenLabs released Scribe v2 in January 2026, targeting the market with 90+ languages and 'negative latency' predictive punctuation, while Deepgram’s Nova-3, launched in early 2025, remains a leader in English performance with reported 54% WER reductions over legacy models per Deepgram’s own data. However, NVIDIA’s open-weights strategy provides a unique leverage point for streaming developers. Unlike closed APIs, the OpenMDW-1.1 license allows for specialized fine-tuning on proprietary corpora, a necessity for regulated industries like healthcare or finance where data privacy is paramount. This release effectively commoditizes high-end multilingual streaming, placing the burden of differentiation back on vendor-specific features like diarization, entity detection, and sentiment analysis.
Read full article at marktechpost.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source