Cerebras AI chip production scales to power OpenAI Ultrafast mode
Cerebras Systems is powering a new 'Ultrafast' service tier for OpenAI's GPT-5.6 Sol, utilizing its Wafer-Scale Engine architecture to achieve 750 tokens per second. The company is scaling production capacity through partnerships with TSMC and contract manufacturers to meet a $25.4 billion backlog.
Key Takeaways
- OpenAI's Ultrafast mode uses Cerebras hardware to reach 750 output tokens per second for GPT-5.6 Sol.
- Cerebras secured a $25.4 billion backlog and expects 2026 core revenue to reach $885 million.
- Manufacturing capacity is expanding through partnerships with TSMC, Flex, Sanmina, and Rocket EMS.
- The Wafer-Scale Engine uses 44GB of on-chip SRAM to eliminate the latency of off-chip data transfers.
Why It Matters
The deployment of Cerebras hardware for OpenAI's high-end tier signals a shift toward specialized silicon for real-time AI applications where latency is the primary constraint. By keeping model weights entirely on-chip, Cerebras avoids the supply-chain bottlenecks currently affecting Nvidia, specifically high-bandwidth memory and advanced packaging shortages. For the streaming and digital media ecosystem, these speeds enable more responsive conversational interfaces and real-time content personalization that were previously limited by GPU processing delays. Watch for the unveiling of the CS-4 system at the upcoming Supernova event to see if the company can maintain its trajectory of doubling processing speeds annually.
Additional Context
Cerebras Systems has been expanding its manufacturing ecosystem to support the scale required by its Wafer-Scale Engine chips. The company's CS-3 and CS-4 systems rely on TSMC's advanced process nodes, and Cerebras confirmed in early 2026 that it had secured multi-year wafer supply agreements with TSMC to underpin its $25.4 billion order backlog. Contract manufacturing partners Flex and Sanmina handle board-level assembly and system integration, while Rocket EMS provides additional capacity for the growing number of inference-focused deployments. This multi-vendor approach mirrors how Nvidia diversified its supply chain after 2023 shortages, though Cerebras avoids the high-bandwidth memory bottleneck entirely by embedding all SRAM directly on the wafer.
Read full article at investors.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source