Cohere North Small Translate outperforms Google and DeepL in benchmarks
Cohere has launched North Small Translate, a mixture-of-experts machine translation model designed for high-throughput, cost-efficient enterprise workflows. Developed in partnership with RWS, the model is available for research and non-commercial use and claims to outperform several proprietary and open-weight alternatives on WMT26 benchmarks.
Key Takeaways
- Achieved an 83.6 WMT26 score, surpassing Google Translate's 68.2 and DeepL NextGen's 81.37.
- Delivers 112 output tokens per second at low concurrency, a 1.4x throughput advantage over Gemma 4 31B.
- Reduces enterprise costs to $0.000676 per task, roughly 5,762% cheaper than Gemini 3.1 Pro Preview.
- Scores 48.9 on long-context evaluations, more than doubling the performance of Google Translate and Gemma 4.
Why It Matters
The launch of this model provides streaming platforms and global media enterprises a high-throughput alternative for large-scale localization that is significantly more cost-effective than current proprietary leaders. By outperforming established tools like DeepL and Google Translate in regional accuracy—particularly in South Asia and MENA—Cohere is positioning its mixture-of-experts architecture as the standard for high-volume, long-form document translation. This shift pressures legacy providers to improve their token efficiency and long-context stability to remain competitive in the enterprise AI stack. Watch for RWS to integrate these weights into its Language Weaver platform to drive commercial adoption among top global brands.
Additional Context
Cohere's North Small Translate enters a market where several established players are aggressively expanding their multilingual capabilities for enterprise and media workflows. DeepL, which has become a dominant force in European enterprise translation, announced in early 2025 that it had surpassed 100,000 business customers and expanded into 33 languages, positioning itself as a direct competitor to Google Translate and Microsoft Translator in corporate settings. Meanwhile, Google has continued to push its own translation models forward, with Gemini 2.5 Pro achieving state-of-the-art results on multilingual benchmarks including WMT24 and FLORES-200 as of mid-2025, signaling that the large-model approach to translation remains a priority for the search giant. The competitive density in this space means Cohere's claim of outperforming both on WMT26 benchmarks carries significant weight for enterprises evaluating cost versus quality tradeoffs.
On the business and licensing side, Cohere's partnership with RWS represents a strategic distribution channel that could accelerate commercial adoption. RWS operates Language Weaver, one of the most widely deployed neural machine translation platforms in the localization industry, serving clients across media, legal, and pharmaceutical sectors. RWS reported in its fiscal 2025 results that its AI and language technology segment grew revenue by 14% year over year, driven by increasing demand for automated translation in content-heavy industries. The decision to release North Small Translate for research and non-commercial use initially mirrors the open-weight strategy that Meta pursued with Gemma and Llama models, which Meta expanded to 128 languages in Gemma 3's March 2025 release, creating a broad developer ecosystem that eventually feeds commercial deployments. For streaming platforms managing subtitle and dubbing pipelines across dozens of languages, this licensing model lowers the barrier to evaluation before committing to enterprise contracts.
From a technical standpoint, the mixture-of-experts architecture that Cohere employs in North Small Translate reflects a broader industry trend toward sparse activation models that reduce inference costs without sacrificing quality. Google's Gemma 4 31B, released in June 2025, introduced a hybrid attention mechanism and multilingual support across 140 languages, but its dense architecture requires substantially more compute per token than sparse MoE alternatives. Independent evaluations have shown that MoE models can achieve comparable or superior translation quality at 30-50% lower inference costs, a finding that . For video localization teams processing thousands of hours of content annually, the throughput advantage Cohere claims over Gemma 4 31B could translate directly into faster turnaround times and lower cloud compute bills, particularly for long-form document and subtitle translation where context window stability matters most.
Read full article at hpcwire.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source