Speechmatics launches Melia speech-to-text model to lead in multilingual code-switching
Speechmatics has launched Melia, a new multilingual speech-to-text model capable of handling code-switching across 56+ languages. The model is available in production preview via batch processing and is positioned as a cost-effective solution for contact center analytics and broadcast captioning.
Key Takeaways
- Melia beats Deepgram and Microsoft on 91% of FLEURS language benchmarks and AssemblyAI on 77%.
- Pricing starts at $0.129 per hour, making it Speechmatics' lowest-priced transcription option.
- Internal tests on noisy monolingual audio show a 5% word error rate reduction compared to the Standard model.
- The model includes language metadata for every transcript, aiding automated routing and compliance reporting.
Why It Matters
The launch of Melia addresses a critical gap in automated transcription: the inability of most models to handle 'Spanglish' or other mixed-language audio without manual configuration. By providing native code-switching across 56+ languages in a single pass, Speechmatics reduces the orchestration overhead for global streamers and news organizations managing multi-regional content. This puts pressure on rivals like Deepgram and AssemblyAI to broaden their multilingual accuracy beyond core high-resource languages. As streaming services expand deeper into non-English markets, the cost-to-accuracy ratio of these specialized models will determine which STT providers win the B2B infrastructure layer. Watch for the release of Melia's real-time variant later this year to challenge low-latency incumbents in live broadcast.
Additional Context
The release of Melia comes as competition in the speech-to-text (STT) sector shifts from raw English accuracy to specialized multilingual capabilities. Per Coval and Northflank reports from early 2026, word error rates on clean English audio have largely plateaued across top providers, forcing vendors to differentiate on latency, cost, and handling real-world 'messy' audio. Microsoft recently entered the proprietary fray with MAI-Transcribe-1 in April 2026, claiming significant performance gains over open-source baselines like Whisper Large v3. Meanwhile, Deepgram’s release of its Flux model in early 2025 and subsequent multilingual updates in April 2026 signaled a trend toward folding complex conversational logic, such as turn-taking and code-switching, directly into the recognition layer. Market analysis from Mordor Intelligence and Polaris indicates the speech-to-speech translation and transcription market is on track to surpass $760 million in 2026, driven by a 17% CAGR in AI-driven tools. Speechmatics, a Series B company with roughly $90.6 million in total funding per Tracxn data, has maintained its position by focusing on deployment flexibility, including on-premises and on-device options. In July 2026, AssemblyAI countered the trend by launching Universal-3.5 Pro, which also emphasizes native code-switching across 18 high-demand languages. This concentration of product launches suggests that B2B customers in the media and contact center sectors are now prioritizing models that can accurately capture global dialects without requiring multiple language-specific integrations.
Read full article at speechmatics.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source