Nvidia launches Vera Rubin platform to resolve agentic AI bottlenecks
Nvidia has launched the Vera Rubin platform, a comprehensive AI infrastructure stack that includes the new 88-core Arm-based Olympus CPU. Designed specifically to handle agentic AI workloads, the hardware emphasizes memory bandwidth, low-latency data movement, and increased performance-per-watt for cloud and infrastructure providers.
Key Takeaways
- Vera CPU features 88 custom Arm-based Olympus cores designed for high single-thread responsiveness and branch-intensive AI orchestration.
- The 1.2 TB/s LPDDR5X memory subsystem and 1.8 TB/s NVLink-C2C interconnect bridge the CPU-GPU bandwidth gap for tool-calling pipelines.
- Internal testing claims 1.8x higher performance on specific agent workloads compared to unnamed x86 systems.
- Launch partners including CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure have begun deploying the Vera Rubin NVL72 rack-scale systems.
- Platform shifts from chiplet-constrained designs to a monolithic Scalable Coherency Fabric providing 3.4 TB/s on-die bandwidth.
Why It Matters
Nvidia is re-evaluating the CPU's role in the AI data center as generative AI shifts toward autonomous agents. Unlike typical inference, agentic AI requires constant loops between the GPU (reasoning) and CPU (orchestration, tool execution, and code handling), making standard high-core-count server CPUs a performance bottleneck. By integrating custom silicon across the entire stack—from the Olympus core to Spectrum-X networking—Nvidia is moving from a component supplier to a system architect. This strategy forces competitors like AMD and Intel to defend their server market share by proving x86 architectures can match these purpose-built, agent-optimized interconnect speeds. For streaming and B2B platforms, this suggests a future where real-time, agent-driven personalization will rely on this specialized high-bandwidth hardware.
Additional Context
The Vera Rubin rollout marks the centerpiece of Nvidia's move to an annual data center architecture cadence, following the earlier 2024 Blackwell launch. Per Forbes in July 2026, analysts estimate the Vera CPU could generate $20 billion in revenue for Nvidia this year alone, potentially capturing a significant portion of the $200 billion server CPU market traditionally dominated by Intel and AMD. TradingKey reported in July 2026 that early shipments were delivered in June to foundational model leaders including OpenAI and Anthropic, with OpenAI reportedly planning to deploy the systems at scale starting this quarter to power more complex reasoning agents. Simultaneous reports from CoreWeave identify a massive efficiency gain, claiming the Vera Rubin NVL72 racks deliver 10x the token output per megawatt compared to the previous-generation Grace Blackwell systems. This performance metric is critical as cooling and power constraints become the primary limiters for hyperscale expansion. While Nvidia maintains a lead in integrated systems, competition is intensifying as AMD prepares its Helios rack-scale systems for H2 2026. Furthermore, major hyperscalers including Microsoft and Meta continue to develop in-house silicon, such as the Maia and Artemis chips, to reduce long-term dependency on Nvidia's premium-priced platform bundles.
Read full article at siliconangle.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source