OpenAI GPT-5.6 Sol Ultrafast mode delivers 750 tokens per second
OpenAI has launched a preview of 'Ultrafast' mode for its GPT-5.6 Sol model, which utilizes Cerebras hardware to achieve processing speeds of up to 750 tokens per second. The service is designed for low-latency, real-time applications such as voice interaction, incident response, and financial research.
Key Takeaways
- Ultrafast mode generates up to 750 output tokens per second using Cerebras hardware
- Early testers include Jane Street, Podium, Basis, and Rogo for financial and voice applications
- OpenAI internal teams use the tier to reduce incident response times and accelerate research loops
- The service is currently available as a limited preview via the OpenAI API
Why It Matters
The introduction of 750 tokens per second removes the traditional trade-off between model intelligence and execution speed. For the streaming and digital media ecosystem, this enables sophisticated voice interfaces and customer support agents that can process complex multi-step queries without the lag that typically breaks user immersion. As frontier models move into time-sensitive workflows like live financial research and incident response, the competitive advantage shifts toward platforms that can integrate high-reasoning AI into synchronous user experiences. Watch for OpenAI to expand API capacity and for competitors to respond with their own hardware-accelerated low-latency tiers.
Additional Context
Cerebras has become a critical infrastructure partner for OpenAI's low-latency ambitions. In early 2026, Cerebras announced a multi-year partnership with OpenAI to provide inference capacity for latency-sensitive workloads, a deal that positioned the Sunnyvale chipmaker as a direct alternative to Nvidia-dominated inference stacks. The partnership gave OpenAI access to Cerebras' wafer-scale engine architecture, which processes tokens in parallel across an entire silicon wafer rather than across discrete GPU clusters. That hardware approach is what enables the 750 tokens per second throughput in Ultrafast mode, a figure that dwarfs typical GPU-based inference speeds of 50 to 100 tokens per second for frontier-class models.
The business implications extend beyond OpenAI's own product line. Cerebras raised $1.1 billion in a Series G round in May 2026 at a $14.3 billion valuation, with investors citing the OpenAI contract as a key revenue anchor. Meanwhile, Jane Street, a quantitative trading firm, disclosed in July 2026 that it had integrated Ultrafast-tier models into its real-time risk assessment pipeline, processing market signals at speeds previously impossible with standard API tiers. The financial services vertical represents one of the highest-value use cases for sub-second inference, where even 200 milliseconds of latency can mean missed arbitrage windows. Podium, Basis, and Rogo, the three launch partners named by OpenAI, span customer support automation, legal research, and developer tooling respectively, signaling that OpenAI is seeding Ultrafast access across verticals where synchronous interaction is a product requirement rather than a nice-to-have.
On the technical side, independent benchmarking has begun to validate the speed claims. Artificial Analysis, an independent LLM evaluation platform, measured GPT-5.6 Sol Ultrafast at a median output speed of 712 tokens per second on standardized prompts in August 2026, with first-token latency under 80 milliseconds. That places it roughly 12 to 14 times faster than the standard GPT-5.6 Sol tier on the same hardware-independent benchmark suite. For streaming platforms exploring AI-driven personalization, content recommendation, or real-time caption generation, these latency figures cross a threshold where model output can keep pace with live video frame rates. Cerebras separately reported in June 2026 that its inference throughput on Llama-class open models had reached 2,100 tokens per second on smaller architectures, suggesting that the wafer-scale approach scales favorably as model size decreases, which matters for edge-adjacent streaming workloads where smaller specialized models handle tasks like content moderation or ad decisioning, a trend further supported by the launch.
Read full article at openai.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source