NVIDIA benchmarks show Vera CPU outpacing x86 in AI-native storage
NVIDIA released benchmark data for its Vera CPU and BlueField-4 STX storage processor, demonstrating performance improvements in encryption, compression, and data integrity checking compared to x86 CPUs. The company claims these gains are specifically engineered to reduce latency and power consumption in storage paths supporting agentic AI workflows.
Key Takeaways
- Vera CPU delivered 1.43x higher AES-128 encryption throughput and 1.29x faster decryption than the x86 baseline.
- Data integrity checking via CRC32C showed a 3.67x performance lead, while Reed-Solomon recovery throughput reached 3.26x.
- A two-stage write pipeline combining compression and encryption outperformed x86 by 3.21x.
- Hardware features include 88 custom Olympus cores, 164MB unified L3 cache, and 1.2 TB/s SOCAMM2 LPDDR5X memory bandwidth.
Why It Matters
NVIDIA is moving to own the storage-processing layer to prevent general-purpose CPUs from stalling high-speed GPU clusters. In agentic AI, where every step requires retrieving and securing persistent memory, traditional x86 storage paths create compounding latency. By integrating Vera into the BlueField-4 STX architecture, NVIDIA provides a specialized offload engine for the encryption and compression tasks that scale with context windows. This shift forces storage vendors to choose between standard hardware or deep integration with the NVIDIA Rubin ecosystem. Watch for Tier-1 storage providers to announce native Vera BlueField-4 support in H2 2026 to stay competitive in AI-native data centers.
Additional Context
The benchmark data follows NVIDIA’s full production launch of the Vera Rubin platform at GTC Taipei in June 2026. Per Wikipedia (January 2026) and NVIDIA Newsroom (March 2026), the platform consists of seven co-designed chips—including the Rubin GPU and Vera CPU—manufactured on TSMC’s 3nm process. While the Rubin GPU focuses on a vendor-stated 50 petaflops of FP4 inference, the Vera CPU is explicitly branded as the 'CPU for agents.' According to Digital Applied (June 2026), Vera’s Olympus cores utilize a novel 'Spatial Multithreading' (SMT-X) approach that partitions core resources physically rather than time-slicing them, allowing for more predictable performance during the branch-heavy orchestration loops common in AI agents.
Industry adoption for this architecture is scaling quickly among hyperscalers. Per Civo (May 2026) and NVIDIA (May 2026), early adopters including AWS, Google Cloud, Microsoft Azure, and Oracle Cloud have already planned deployments. Furthermore, manufacturing partners such as Dell Technologies, HPE, and Supermicro are building standalone Vera systems. Tom’s Hardware (July 2026) noted that Vera represents NVIDIA’s first in-house CPU core design, moving away from the stock Arm Neoverse designs used in previous Grace modules. This level of vertical integration aims to address the 'KV cache' bottleneck, where expanding context windows for large language models outgrow GPU memory and must be offloaded to storage without transiting a traditional, slower host CPU.
Competitive pressure is mounting as AMD and Intel prepare rival AI infrastructure updates. Per ServeTheHome (July 2026), the release of the Vera whitepaper coincided with AMD’s 'Advancing AI' event, highlighting the race to provide host processors that can keep pace with 3.6 TB/s NVLink interconnects. While Vera features lower L3 cache totals than some high-end x86 Xeon or EPYC chips, its 1.2 TB/s memory bandwidth is designed to sustain thousands of parallel sandbox environments. With partner availability for BlueField-4 STX expected in the second half of 2026, the industry is shifting toward a model where the entire rack, rather than the individual chip, serves as the primary unit of compute.
Read full article at developer.nvidia.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source