OpenAI Jalapeño inference chip outperforms Nvidia Blackwell in early benchmarks
OpenAI presented benchmark results for its Jalapeño inference chip, developed in partnership with Broadcom, at the Hot Chips conference. The hardware is designed to optimize inference throughput and power efficiency by minimizing data movement and KV cache delays, with small-volume deployment expected in late 2026.
Key Takeaways
- Jalapeño outperformed Nvidia Blackwell in throughput per kilowatt and tokens per user on SemiAnalysis’ InferenceX benchmark
- OpenAI plans a small-volume deployment for late 2026 with significant scaling expected in 2027
- Hardware design focuses on minimizing KV cache delays by keeping model state local during specific inference phases
- Broadcom collaborated on the development, which OpenAI intends to turn into a multigenerational hardware platform
Why It Matters
The benchmark results suggest OpenAI is successfully transitioning from a software-only entity to a vertically integrated hardware player capable of challenging Nvidia's Blackwell architecture. By optimizing for specific inference bottlenecks like data movement and communication delays, OpenAI aims to lower the massive operational costs associated with scaling large language models. This shift signals a broader industry trend where major AI developers build custom silicon to bypass the supply constraints and high margins of third-party chipmakers. As the platform evolves into a multigenerational system, the streaming and tech ecosystems will likely see a push toward more localized, efficient compute cycles. Watch for the 2027 deployment phase to see if Jalapeño maintains its efficiency lead against next-generation Nvidia iterations.
Additional Context
OpenAI's Jalapeño chip enters a rapidly expanding custom silicon landscape where hyperscalers and AI labs are increasingly bypassing merchant GPU vendors. In March 2025, Broadcom reported that its custom AI accelerator revenue had grown to over $12 billion annually, driven by partnerships with at least three major cloud providers and AI companies seeking inference-optimized alternatives to Nvidia's general-purpose GPUs. Broadcom's role as Jalapeño's fabrication and packaging partner places it at the center of a broader trend where chip designers compete on workload-specific efficiency rather than raw FLOPS. Google's TPU v6, Amazon's Trainium 2, and Microsoft's Maia 100 all represent similar vertical integration strategies, though none have yet published head-to-head inference throughput comparisons against Nvidia Blackwell at the scale OpenAI presented at Hot Chips.
The business implications of custom inference silicon extend into supply chain negotiations and capital allocation. In June 2025, Nvidia CEO Jensen Huang acknowledged that custom ASICs would capture a growing share of inference workloads, while maintaining that the company's CUDA software ecosystem remains a durable moat for training and complex multi-model deployments. Meanwhile, Broadcom raised its AI revenue guidance to $15 billion for fiscal 2025, citing demand from customers building proprietary inference stacks. OpenAI's move to co-design hardware with Broadcom rather than relying solely on Nvidia GPUs mirrors the strategy Meta employed when it developed its own MTIA inference accelerator, which Meta deployed across its recommendation systems in early 2025 to reduce per-token costs at scale.
Technical benchmarks from independent testing labs provide additional context for Jalapeño's claimed performance advantages. MLPerf Inference v5.0 results published in April 2025 showed that inference-optimized ASICs delivered 2.3x better tokens-per-watt than general-purpose GPUs on large language model workloads, validating the architectural thesis behind Jalapeño's focus on minimizing data movement and KV cache bottlenecks. The Hot Chips presentation by Richard Ho, OpenAI's hardware lead, positioned Jalapeño's full-stack design as targeting the prefill and communication phases that MLPerf identified as primary efficiency constraints. If OpenAI's small-volume deployment in late 2026 confirms these gains in production, it would represent the first time a frontier AI lab has demonstrated custom silicon outperforming Nvidia's latest architecture on standardized inference metrics.
Read full article at techcrunch.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source