CoreWeave and Nvidia integrate Vera Rubin to optimize agentic AI inference
CoreWeave and Nvidia are promoting integrated rack-scale systems powered by the Nvidia Vera Rubin platform to optimize AI infrastructure for continuous agentic workflows. By co-managing compute, networking, and storage at the rack level, CoreWeave aims to reduce token costs and improve inference throughput for production-grade AI environments.
Key Takeaways
- Vera Rubin platform delivers 10x better inference throughput per watt at one-tenth the cost per million tokens versus Blackwell.
- Integrated rack-scale infrastructure treats compute, networking, and storage as a unified system rather than discrete components.
- New hardware control layers including 'Racky' rack manager and 'Valvey' valve assembly coordinate liquid cooling and real-time observability.
- CoreWeave Kubernetes Service and 'Mission Control' provide a software layer to manage workloads across hundreds of thousands of GPUs.
- The Vera Rubin NVL72 rack pairs 72 Rubin GPUs with 36 Vera CPUs to support continuous reasoning and agentic planning loops.
Why It Matters
The shift toward agentic AI moves infrastructure requirements from bursty, single-shot inference to continuous, stateful reasoning loops. CoreWeave’s rack-scale approach effectively makes the 'rack the computer,' addressing the direct business impact of token economics as agents become the primary enterprise product. By validating power, liquid cooling, and networking within a unified control plane, CoreWeave provides the low-latency reliability required for mission-critical production systems. This move intensifies the competition among cloud providers to go beyond raw GPU counts, focusing instead on integrated systems engineering that lowers total cost of ownership for trillion-parameter, high-context AI sessions. Watch for CoreWeave’s deployment of Vera Rubin in late 2026 as a benchmark for enterprise-wide agentic scalability.
Additional Context
The Vera Rubin platform entered full production in June 2026, marking a pivotal transition in AI chip design where interconnect bandwidth and memory capacity now supersede raw compute FLOPS as the primary performance metrics. Per SiliconReport (July 2026), the Rubin NVL72 rack features 20.7 terabytes of HBM4 memory pooled across 72 GPUs, specifically designed to eliminate the bottlenecks reasoning models face when moving data between units. Nvidia’s Q1 FY2027 earnings report highlights this trend, with CEO Jensen Huang noting that the rapid ramp of the predecessor Blackwell platform was driven by the urgent need for lower token generation costs at inference, which currently accounts for approximately two-thirds of all AI compute according to Deloitte. CoreWeave has positioned itself as a lead partner for this transition through a January 2026 strategic expansion that included a $2 billion investment from Nvidia. Per Forbes (January 2026), this partnership aims to build 5 gigawatts of 'AI factories' by 2030, with CoreWeave serving as a primary testing ground for Nvidia’s first custom Arm-based CPUs, the Vera series. This deeper technical alignment includes the integration of CoreWeave's proprietary Mission Control software into Nvidia’s reference architectures. Meanwhile, the specialized cloud market is facing a 'hardware super-cycle' characterized by surging memory costs; Indiatimes (July 2026) reported that server DRAM and HBM prices have risen nearly 95% this year, making rack-level efficiency the only viable path to maintaining predictable margins for inference services.
Read full article at siliconangle.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source