Zendesk AI voice translation eliminates third-party apps for contact centers
Zendesk has announced a native real-time voice translation capability for its contact center platform, supporting 13 languages for bidirectional communication. The feature is currently in a closed early access program with general availability expected in Q1 2027.
Key Takeaways
- General availability for the native voice translation feature is scheduled for Q1 2027 following an October early access program.
- Initial language support includes 13 options such as English, Mandarin, Spanish, and Hindi, with plans for further expansion.
- Administrators can configure translation by specific queues and choose to record original audio, translated audio, or both.
- The early access version excludes support for video calls, multiparty calls, and supervisor barge-in functionality.
Why It Matters
This launch marks a transition from fragmented third-party overlays to native infrastructure for real-time audio processing. By embedding translation directly into the agent desktop, Zendesk reduces the technical debt and licensing complexity typically associated with multilingual support. For the broader streaming and communications ecosystem, this signals a move toward unified AI stacks where context—such as customer history and intent—is preserved across translated sessions. The industry should monitor the Q1 2027 commercial rollout to see if usage-based pricing models can compete with established point solutions like Krisp and Sanas.
Additional Context
Zendesk's entry into native voice translation places it in direct competition with specialized AI audio companies that have built dedicated real-time translation and noise-suppression products for contact centers. Krisp, which raised a $25 million Series B in 2024, has expanded beyond noise cancellation into real-time speech translation and accent conversion for enterprise contact centers, partnering with multiple CCaaS platforms to offer its AI voice processing as a middleware layer. Sanas, meanwhile, has focused specifically on accent conversion and real-time voice adaptation, securing partnerships with major BPO providers to reduce language friction in outsourced support operations. Zendesk's decision to build translation natively rather than integrate with these vendors signals a strategic bet that contact center buyers prefer unified platforms over best-of-breed point solutions.
The broader CCaaS market is consolidating around AI-native capabilities as a differentiator. Zendesk acquired Klaus in 2024 for its AI-powered quality assurance technology, and the company has since integrated Klaus's conversation intelligence features directly into its agent workspace. Competitors are pursuing similar strategies: Five9 announced in early 2025 that its AI Agent Studio would support multilingual voice interactions natively, while Genesys Cloud has embedded real-time translation through partnerships with third-party providers. The licensing economics matter here. Zendesk's approach of bundling translation into its existing contact center subscription contrasts with per-minute or per-seat pricing models that standalone vendors like Sanas and Krisp typically charge, potentially lowering total cost of ownership for mid-market buyers who lack procurement teams to negotiate separate AI vendor contracts.
On the technical side, real-time voice translation in contact centers faces latency and accuracy challenges that differ from text-based translation. Research from Stanford's Center for Research on Foundation Models found that speech-to-speech translation systems still exhibit 15-20% higher error rates than cascaded speech-to-text-to-text-to-speech pipelines, though end-to-end models are closing the gap. Zendesk's 13-language support at launch is narrower than what some competitors offer. Sanas supports over 30 languages for accent conversion, and Krisp's translation layer covers more than 20 languages. The closed early access program Zendesk has initiated will be a critical test of whether native integration can maintain sub-500-millisecond latency while preserving conversational context, a threshold that .
Read full article at nojitter.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source