Cohere has launched North Small Translate, a 218-billion parameter mixture-of-experts model optimized for enterprise machine translation across 50 languages. The model is designed for secure, localized processing of sensitive documents and is available for commercial use via Cohere's Model Vault environment.
The release of this specialized model signals a shift away from general-purpose LLMs for high-stakes enterprise tasks like document translation. By utilizing a non-reasoning architecture, Cohere addresses the specific technical failure of long-context drift that often plagues broader models in professional settings. For the streaming and media ecosystem, this technology offers a path toward more reliable automated localization of sensitive internal documents and safety manuals without exposing data to third-party APIs. The emphasis on sovereign AI suggests that control over data residency is becoming as critical as raw performance for global organizations. Watch for whether Cohere expands this mixture-of-experts approach to real-time subtitling or dubbing workflows.
Cohere has been building its enterprise AI portfolio aggressively over the past year, with translation as a key vertical. In April 2025, Cohere released Command A, its flagship enterprise LLM optimized for agentic workflows and multilingual tasks, which supports 23 languages and was positioned as a competitor to OpenAI and Anthropic in corporate deployments. The company's broader strategy around sovereign AI has been reinforced by partnerships with governments and regulated industries. In March 2025, Cohere announced a collaboration with the UK government's AI Safety Institute to evaluate frontier model safety, signaling its push into compliance-heavy sectors where data residency requirements are strict. Nick Frosst, Cohere's CEO and co-founder, has repeatedly framed the company's differentiation around on-premises and private-cloud deployment options that hyperscalers like Google cannot easily match for sensitive workloads. The competitive landscape for enterprise machine translation has intensified significantly. DeepL, which raised $300 million in a January 2024 funding round that valued the company at $2 billion, launched DeepL Voice in late 2024 to target real-time speech translation for enterprise meetings, expanding beyond its document-translation roots into live communication scenarios. Google, meanwhile, announced in May 2025 that its Gemini-powered translation capabilities had been integrated across Google Workspace, supporting over 130 languages with context-aware document translation built directly into Docs and Sheets. For Cohere, the differentiation lies not in language count but in deployment sovereignty: North Small Translate runs inside a customer's own infrastructure through Model Vault, avoiding the data-residency concerns that prevent many regulated enterprises from sending documents to Google Cloud Translation or DeepL's API endpoints. On the technical side, mixture-of-experts architectures have become the dominant approach for balancing model quality against inference cost in production translation systems. Meta's NLLB-200 project demonstrated that sparse MoE models could cover 200 languages with a single architecture, though it was released as a research model rather than a commercial product. Cohere's earlier Tiny Aya model, released in November 2024 as a 3.9-billion parameter multilingual model trained on 50 languages, served as a proof of concept for the company's multilingual data pipeline and evaluation methodology. The jump from Tiny Aya's 3.9 billion parameters to North Small Translate's 218 billion parameters (with sparse activation) reflects a broader industry trend where MoE models achieve dense-model quality at a fraction of the compute cost, making on-premises deployment of high-quality translation economically viable for the first time in enterprise settings. As competition in this space grows, AI video dubbing tools continue to set the pace for automated localization features.
Cohere has launched North Small Translate, a 218-billion parameter mixture-of-experts model designed for secure enterprise translation. By utilizing a non-reasoning architecture, it achieves superior performance in long-context tests compared to Google Translate. This release matters because it offers organizations a sovereign AI solution for sensitive document translation without third-party API exposure.
It is a 218-billion parameter mixture-of-experts model designed for secure, localized enterprise translation across 50 languages.
In benchmarks, it achieved a WMT26 score of 83.60, surpassing Google Translate, and scored 48.9 on long-context tests compared to Google's 21.3.
Commercial deployment is restricted to the Cohere Model Vault environment, which allows the model to run inside a customer's own infrastructure to ensure data sovereignty.
The architecture uses 25 billion active parameters to reduce compute and memory footprints while maintaining high-quality translation performance.
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source