NVIDIA begins volume production of Vera Rubin NVL72 rack platform
NVIDIA has initiated full-scale production of its Vera Rubin NVL72 rack-scale AI platform, partnering with cloud providers like Google, Microsoft, and Oracle to deploy high-density compute for agentic AI and video intelligence. The new architecture is designed to address power-constrained data centers with a reported tenfold increase in tokens-per-megawatt efficiency compared to its Grace Blackwell predecessor.
Key Takeaways
- Vera Rubin NVL72 integrates 72 GPUs and 36 Vera CPUs using a custom 88-core Olympus design optimized for agentic workloads.
- CoreWeave confirmed a 10x throughput-per-megawatt gain on DeepSeek-R1 benchmarks, matching NVIDIA’s pre-production targets.
- The 45°C liquid-cooling inlet supports chiller-free operation, significantly reducing water consumption in new AI facilities.
- Sixth-generation NVLink provides 260 TB/s aggregate rack bandwidth, double the capacity of the previous Blackwell interconnect.
- Microsoft and Mistral signed a multibillion-dollar agreement to deploy the architecture for sovereign AI services across Europe.
Why It Matters
NVIDIA is shifting the industry’s primary benchmark from raw FLOPS to tokens-per-megawatt, a critical pivot as power availability becomes the chief bottleneck for data center expansion. By integrating proprietary CPUs, DPUs, and switching into a unified rack-scale product, NVIDIA is locking hyperscalers into a full-stack architecture that is difficult to replicate with off-the-shelf components. For the streaming and video intelligence sector, these gains in orchestration speed and lower token costs make real-time agentic video processing economically viable at scale. Watch for upcoming MLPerf results in late 2026 to see if independently verified data matches CoreWeave’s early 10x efficiency claims.
Additional Context
The transition to Vera Rubin marks NVIDIA’s most aggressive move into a one-year product cadence, following the Grace Blackwell launch by only 15 months. Per SemiAnalysis (July 2026), the platform’s performance leap is largely attributed to the use of HBM4 memory, which provides up to 22 TB/s of bandwidth compared to the 8 TB/s found in Blackwell Ultra. This increase in memory throughput is vital for sustaining the large context windows and mixture-of-experts (MoE) architectures currently dominating frontier AI development. While Blackwell delivered a 3x to 4x jump over the Hopper architecture, Rubin represents a more fundamental redesign of the compute-to-memory ratio.
Simultaneously, the platform’s focus on European sovereign AI reflects growing regulatory pressure. Per ioplus.nl (July 2026), the Microsoft-Mistral partnership allows European regulated sectors—including finance and healthcare—to process data in fully disconnected Azure Local environments. This deployment model is specifically designed to satisfy regional data residency laws while utilizing thousands of Vera Rubin GPUs physically located within European borders. Mistral’s goal is to reach 1 gigawatt of operational compute capacity by 2030, a roadmap that relies heavily on the efficiency metrics NVIDIA is touting with the new hardware.
Competitors are also responding to the efficiency-first landscape. Google Cloud is simultaneously scaling its TPU v5 infrastructure, while AMD has committed to delivering its MI455X architecture to compete on UALink interconnect standards. However, NVIDIA’s dominance is currently anchored by its supply chain reaching over 350 factory sites. Per Datacentre Magazine (July 2026), the physical requirements for Rubin are also reshaping the data center market; the high power density of the NVL72 makes liquid cooling a prerequisite for deployment, effectively making legacy air-cooled facilities obsolete for the highest-tier AI workloads.
Read full article at 4rfv.co.uk
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source