Nvidia launches Vera CPU with custom Olympus core for agentic AI
Nvidia has announced its new Vera CPU, which utilizes the custom Olympus microarchitecture to minimize latency in agentic AI loops. The monolithic chip is designed to improve data throughput and reduce bottlenecks for attached AI accelerators in high-performance computing environments.
Key Takeaways
- Custom Olympus microarchitecture replaces Arm Neoverse cores, featuring a 10-wide decode front end for higher single-threaded performance.
- LPDDR5X memory subsystem delivers 1.2TB/sec bandwidth, providing roughly three times more memory bandwidth per core than competing server platforms.
- Early benchmarks claim 2X faster agentic sandbox startup and up to 6X better performance in streaming data processing via partnership with Redpanda and HPE.
- Second-generation Scalable Coherent Fabric links the monolithic die, reportedly offering three times greater core-to-core bandwidth than chiplet designs.
Why It Matters
Nvidia is pivoting from using off-the-shelf Arm designs to fully custom silicon to address the specific latency requirements of AI agents. By optimizing for the "agent loop" rather than raw core density, Nvidia aims to prevent the host CPU from becoming a bottleneck for its high-margin GPUs. For the streaming industry, such optimizations in data processing and orchestration could significantly reduce the overhead of real-time AI-driven metadata generation and content moderation. This move tightens Nvidia's vertically integrated stack, forcing rivals like Intel and AMD to prove their general-purpose architectures can maintain per-thread responsiveness under heavy agentic workloads. Tracking independent SPEC CPU 2026 benchmarks against AMD’s upcoming 256-core "Venice" will be the next critical performance signal.
Additional Context
The Vera CPU arrives as Nvidia accelerates its hardware release schedule to an annual cadence. Per SemiAnalysis (February 2026), the chip is a cornerstone of the broader "Vera Rubin" platform, which succeeds the Blackwell generation. While previous Grace CPUs primarily functioned as memory managers for GPUs, Vera is positioned as a standalone competitor in the server market. Nvidia founder Jensen Huang recently emphasized at GTC Taipei that "AI agents will be the largest users of computing," signaling a shift toward autonomous software that requires high-per-core performance for code execution and tool invocation rather than just massive parallel throughput. Competitors are already reframing the benchmark debate. Following Nvidia's internal data release, AMD issued a response in June 2026 claiming its 192-core EPYC "Turin" processors deliver up to 2.37 times higher performance per watt at the rack level. AMD argues that Nvidia’s focus on single-thread latency ignores the core density requirements of hyperscale data centers. Furthermore, per HotHardware (June 2026), AMD’s upcoming "Venice" architecture is projected to offer a 27% per-core performance advantage over Vera, potentially challenging Nvidia's claims of single-threaded supremacy in the agentic era. Despite the competitive pressure, Nvidia has secured early ecosystem validation. Per Company Reporting (May 2026), OpenAI and Anthropic are among the labs planning to integrate Vera into their "AI factories." Additionally, cloud providers including Oracle Cloud, CoreWeave, and Lambda have committed to deploying Vera-based systems starting in late 2026. This integration is bolstered by a 1.8 TB/s NVLink-C2C interface, which establishes a coherent memory space between the Vera CPU and Rubin GPUs—a proprietary interconnectivity that remains a unique moat against x86-based competitors.
Read full article at forbes.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source