Microsoft MAI-Transcribe-2 launch cuts speech recognition costs by 72 percent
Microsoft has released MAI-Transcribe-2, a speech-recognition model priced at $0.10 per hour that supports 60 languages and includes features like speaker diarization and keyword biasing. The model is being positioned as a faster, more cost-effective alternative to existing solutions from OpenAI, Google, and ElevenLabs, reflecting Microsoft's strategy to reduce reliance on third-party AI providers.
Key Takeaways
- Pricing dropped to $0.10 per audio hour, a 72% reduction from the $0.36 rate of the previous version released five months ago
- Performance benchmarks show the model is 10 times faster than OpenAI's GPT-Transcribe and 7 times faster than ElevenLabs' Scribe v2
- New features include Hinglish and Spanglish code-switching support and a 'verbatim' mode for legal and compliance teams
- The model achieved a 5.2% average word error rate across 60 languages on the FLEURS multilingual benchmark
Why It Matters
This release marks a critical pivot in Microsoft's AI strategy, moving from hosting partner models to deploying high-efficiency in-house alternatives that improve margins. By bundling premium features like diarization into a low base rate, Microsoft is commoditizing transcription services that were previously high-margin moats for specialized vendors. Within the streaming and enterprise ecosystem, this creates immediate pressure on providers like ElevenLabs and Google to justify higher price points for similar speech-to-text workloads. The rapid three-release cadence in five months suggests Microsoft has stabilized its architecture for specific modalities. Watch for whether Microsoft begins replacing OpenAI models within the Teams and Nuance product lines to further reduce internal GPU expenditures.
Additional Context
Microsoft has been rapidly expanding its portfolio of proprietary AI models under the MAI brand, signaling a broader strategic shift away from dependence on OpenAI. In May 2025, Microsoft launched its first MAI models on Azure AI Foundry, including MAI-Voice-1 for text-to-speech generation, marking the company's initial foray into building foundation models independently. The transcription model MAI-Transcribe-2 represents the third release in this series within five months, demonstrating an accelerated cadence that positions Microsoft as both a platform provider and a direct competitor to the model vendors it hosts. Mustafa Suleyman, who leads Microsoft AI after the company's acquisition of Inflection AI's team in March 2024, has overseen this push toward proprietary model development.
The pricing pressure from MAI-Transcribe-2 arrives amid intensifying competition in the speech recognition market. ElevenLabs raised $180 million in a Series C round in January 2025 at a $3.3 billion valuation, reflecting investor confidence in specialized audio AI despite the threat of commoditization from larger platform players. Meanwhile, Google DeepMind released Gemini 2.5 Flash in March 2025 with native audio transcription capabilities built into its multimodal architecture, bundling speech-to-text into a broader model offering rather than pricing it as a standalone service. These moves illustrate how transcription is shifting from a standalone product category into a feature layer within larger AI platforms, compressing margins for specialized vendors.
On the technical front, Microsoft has emphasized latency and throughput as differentiators for MAI-Transcribe-2. The model processes audio at speeds that Microsoft claims exceed OpenAI's Whisper large-v3 model by a factor of three on standard benchmarks, though independent third-party benchmarking has not yet been published. For streaming platforms processing large volumes of content for accessibility compliance, the combination of 60-language support and speaker diarization at $0.10 per hour could significantly reduce the cost of generating subtitles and transcripts at scale. The FCC's 2024 order requiring streaming services to provide accurate closed captions for all on-demand content has created regulatory demand for high-volume transcription, making cost-per-hour a critical metric for platform operators evaluating vendors.
Read full article at venturebeat.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source