Speechmatics Melia 1 code-switching model cuts Arabic error rates by half
Speechmatics has updated its Melia 1 multilingual speech-to-text model to improve code-switching accuracy for Arabic, Mandarin, and Tamil. The company claims the model achieves a 15.1% mixed error rate on Arabic-English code-switching and provides faster batch processing speeds.
Key Takeaways
- Melia 1 achieved a 15.1% mixed error rate on Arabic-English audio, compared to 33.2% for the next best model tested.
- Switch point F1 accuracy reached 0.928 for Mandarin and Tamil, indicating high reliability in detecting mid-sentence language changes.
- Batch processing speeds now allow a 60-minute recording to be transcribed in under 20 seconds.
- Alphanumeric string handling for reference numbers and postcodes has been improved across all supported languages.
Why It Matters
The ability to accurately transcribe code-switching is critical for streaming and communication platforms operating in multilingual hubs like Singapore and Abu Dhabi. By reducing error rates in Arabic and Tamil by significant margins, Speechmatics addresses a technical gap where traditional models often drop technical terms or fail to recognize language transitions. This performance puts pressure on incumbents like Amazon and AssemblyAI to improve their handling of non-Western language pairs. As streaming services expand localized content and global contact centers demand better analytics, tracking these specialized accuracy metrics will become the standard for enterprise-grade speech AI. Watch for the model's upcoming expansion into US Hispanic and Indic language markets.
Additional Context
Speechmatics operates in an increasingly crowded multilingual speech recognition market where several vendors are pushing code-switching and low-resource language capabilities. In March 2025, AssemblyAI launched Universal-Streaming, a model that processes audio in real time with sub-300ms latency across 99 languages, positioning it as a direct competitor for enterprise transcription workloads that require fast turnaround. Soniox, another rival benchmarked in Speechmatics' own testing, released its v3 ASR engine in early 2025 with claimed state-of-the-art word error rates on multilingual benchmarks including FLEURS, targeting the same enterprise and developer segments that Speechmatics serves with Melia 1.
The business case for accurate code-switching transcription is expanding beyond media into regulated industries. Amazon Transcribe, which Speechmatics benchmarks against, added support for 100+ languages and dialects in its 2024 platform update, including Arabic and Mandarin variants, signaling that hyperscalers view multilingual ASR as a core cloud service differentiator. Meanwhile, the global speech recognition market is projected to reach $26.8 billion by 2030, growing at a 17.2% CAGR according to Grand View Research, driven partly by demand for multilingual customer service automation in the Middle East and Southeast Asia, regions where code-switching between Arabic, English, and Tamil is routine.
On the technical side, independent benchmarks provide context for how Melia 1's claimed 15.1% mixed error rate compares to the broader field. The FLEURS benchmark, maintained by Google Research, evaluates ASR systems across 102 languages and has become a standard reference for multilingual model comparison, though it does not specifically measure code-switching accuracy. For that task, a 2024 study from the Qatar Computing Research Institute evaluated Arabic-English code-switching systems and found that transformer-based models still struggle with switch-point detection, with F1 scores below 0.7 on conversational data, which makes Speechmatics' reported 0.714 switch-point F1 score notable as it exceeds what recent academic work has achieved on similar data. Yahia Abaza, who leads Speechmatics' research efforts, has emphasized that real-world deployment in contact centers and media workflows requires models trained on authentic code-switched audio rather than synthetic mixing, a distinction that separates Melia 1's approach from simpler language-identification pipelines. As , the demand for such high-fidelity transcription will only increase.
Read full article at speechmatics.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source