PolyAI launches Dialog-RSN-1 audio-native model achieving sub-300ms latency
PolyAI has launched Dialog-RSN-1, an audio-native dialog model designed for enterprise voice agents that processes raw audio to improve latency and emotional nuance. Operating on a sub-300 millisecond response time, the system uses a combined architecture for signal processing while maintaining a separate text-to-speech output for enterprise control.
Key Takeaways
- Dialog-RSN-1 fuses speech recognition, turn-taking, and response generation into a single audio-native LLM.
- The model targets sub-300ms latency on A100 GPUs using 8B to 30B parameter architectures.
- Audio-native input perception allows the system to detect caller hesitation and frustration that text-based models miss.
- A hybrid approach uses a separate text-to-speech system to prevent the cost and branding issues of integrated speech-to-speech models.
Why It Matters
PolyAI is addressing a critical bottleneck in enterprise voice agents: the latency and loss of signal inherent in traditional 'cascaded' architectures. By reasoning over raw audio rather than intermediate transcripts, Dialog-RSN-1 offers human-like turn-taking speed without sacrificing the brand control companies require for customer-facing roles. This launch intensifies competition among agentic platform players like Five9 and Salesforce, who are also collapsing the stack to improve performance. As enterprises move toward 'agentic' service models, the market is pivoting from simple deflection to nuanced, high-stakes conversation. Watch for the public API opening and the release of their open-source 'Dialog-Eval' framework as benchmarks for real-time voice quality.
Additional Context
The launch of Dialog-RSN-1 comes as the enterprise voice AI sector shifts from experimental pilots to operational infrastructure. According to a July 2025 Research and Markets report, the global conversational AI market is projected to grow 192% to $49.8 billion by 2031. This growth is driven by a move away from scripted bots toward autonomous 'agentic' systems capable of reasoning. Major incumbents are responding by unifying their technology stacks; Salesforce introduced the Agentforce Contact Center in March 2026, which natively combines voice and CRM data to provide AI agents with full customer context.
Competitive pressure is also mounting from high-profile speech-to-speech models. OpenAI’s GPT-Realtime-2, released in May 2026, and Google’s Gemini 3.1 Flash Live, launched in March 2026, both offer low-latency audio-to-audio processing. However, industry benchmarking from Speko in early 2026 suggests these models often force a trade-off between speed and granular control over voice output. PolyAI’s decision to keep audio awareness on the input side only is a direct tactical response to these enterprise branding and cost concerns.
PolyAI’s rapid product velocity is backed by significant capital, including an $86 million Series D closed in December 2025 co-led by Georgian, Hedosophia, and Khosla Ventures. The company, which now serves over 200 enterprises including Marriott and FedEx, was ranked the fastest-growing AI firm in the 2025 Deloitte UK Technology Fast 50. Per Deloitte, PolyAI achieved a three-year revenue growth rate of 2,274%, reflecting high demand for voice-native automation that can handle complex industrial and financial service queries.
Read full article at cmswire.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source