OpenAI Jalapeño chip benchmarks show 3.6x latency cut for AI agents
OpenAI has released performance benchmarks for its custom Jalapeño inference chip, developed in partnership with Broadcom. The hardware is designed to improve throughput and latency for agentic AI workloads by optimizing memory bandwidth and reducing data movement between cores.
Key Takeaways
- Jalapeño reduced end-to-end latency by 1.7 to 3.6 times across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 models.
- Hardware design keeps model state and KV cache local to minimize data movement between cores and chips.
- The 700-watt rated chip demonstrated 8.6 to 104.3 times more work per watt at specific time-between-tokens operating points.
- Performance on highly interactive workloads was measured at 2.1 to 4.1 times faster than comparison systems.
Why It Matters
The release of these benchmarks signals a shift toward hardware specifically optimized for agentic workflows rather than general-purpose compute. By reducing the latency compounding that occurs when models are called sequentially, OpenAI addresses a primary bottleneck in deploying responsive AI agents at scale. For the streaming and broader tech ecosystem, this vertical integration reduces reliance on third-party hardware providers while potentially lowering the operational costs of high-token-volume applications. As OpenAI continues to cut API pricing, this custom silicon provides the necessary infrastructure margin to maintain profitability. Watch for whether Broadcom scales this architecture for other enterprise partners or keeps the design exclusive to OpenAI's internal clusters.
Additional Context
Broadcom has positioned itself as the leading custom AI accelerator partner for hyperscalers and AI labs seeking alternatives to Nvidia's dominant GPU ecosystem. In early 2025, Broadcom reported that its AI-related revenue reached $12.2 billion in fiscal 2024, driven by custom XPU programs with three major customers, a figure analysts expect to grow as additional clients move from tape-out to volume production. The Jalapeño chip represents the most publicly visible example of Broadcom's custom silicon division working with an AI-native company rather than a traditional cloud provider, signaling that the foundry model is expanding beyond Google's TPU and Meta's MTIA programs into the inference-first AI lab segment.
The competitive landscape for custom inference silicon has intensified considerably over the past year. Google announced in April 2025 that its seventh-generation Ironwood TPU delivers 4,614 teraflops of FP8 compute per chip, positioning it as the company's most powerful inference-focused processor to date. Meanwhile, Amazon Web Services confirmed in March 2025 that its Trainium2 chips are being used by Anthropic to train and serve Claude models at scale, demonstrating that custom silicon is becoming a standard path for frontier AI labs seeking cost and performance advantages over merchant GPUs. OpenAI's Jalapeño benchmarks enter this crowded field with a specific focus on agentic workloads, differentiating from the throughput-first designs of competitors.
Independent analysis of inference chip economics suggests that memory bandwidth, not raw compute, is the binding constraint for multi-step agentic tasks. SemiAnalysis published a detailed breakdown in June 2025 showing that inference workloads with high KV-cache reuse patterns benefit disproportionately from HBM bandwidth improvements over FLOPS gains, a finding that aligns with OpenAI's stated design philosophy for Jalapeño. The chip's emphasis on reducing data movement between cores directly addresses the memory-wall bottleneck that SemiAnalysis identified as the primary limiter of tokens-per-second in long-context, multi-turn inference scenarios. For streaming applications that increasingly rely on agentic AI for content recommendation, metadata generation, and real-time personalization, these bandwidth-first architectures could meaningfully reduce per-token costs at production scale.
Read full article at thenewstack.io
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source