CoreWeave deploys multi-plane networking to scale agentic AI workloads
CoreWeave outlines its architectural approach to scaling infrastructure for agentic AI workloads, emphasizing the use of NVIDIA Spectrum-X Ethernet and BlueField-3 DPUs. The company details five key strategies, including non-blocking network design, multi-plane topology, and hardware-accelerated load balancing, to manage unpredictable traffic patterns at fleet scale.
Key Takeaways
- Multi-plane topology enables scaling to 128,000 GPUs without adding a third tier or increasing hop latency
- NVIDIA ConnectX-9 SuperNICs move load balancing to hardware to react to congestion at line rate
- BlueField-3 DPUs offload networking and security tasks to preserve host CPU cycles for agent orchestration
- Non-blocking design maintains a 1:1 uplink ratio to prevent jitter during synchronized training bursts
- GPU Straggler Detection in CoreWeave Mission Control automates the isolation of underperforming nodes
Why It Matters
The shift from rhythmic training workloads to unpredictable, multi-hop agentic chains requires a fundamental redesign of the streaming and AI backend. By moving networking logic to dedicated silicon like the BlueField-3 DPU, CoreWeave ensures that complex agent tool calls do not compete with infrastructure overhead for limited compute resources. This architectural approach signals a move away from general-purpose cloud environments toward specialized fabrics that prioritize low tail-latency over simple aggregate bandwidth. As streaming platforms increasingly integrate AI agents for personalized discovery and real-time metadata generation, the ability to maintain lossless performance under bursty conditions will become a competitive necessity. Watch for the deployment of NVIDIA Vera Rubin NVL72 systems to serve as the benchmark for next-generation agentic performance.
Additional Context
CoreWeave has rapidly expanded its GPU cloud footprint to meet surging demand for AI inference and training workloads. In March 2025, CoreWeave raised $1.5 billion in its initial public offering on the Nasdaq, pricing shares at $40 each and valuing the company at approximately $23 billion. The capital raise was earmarked for data center buildout and GPU procurement, with NVIDIA serving as both a key supplier and a strategic investor. NVIDIA holds a stake in CoreWeave and has committed to purchasing cloud services from the company through 2032, a relationship that gives CoreWeave early access to new silicon generations including the Vera Rubin platform.
The competitive landscape for GPU cloud infrastructure has intensified as hyperscalers and independent providers race to secure capacity for agentic AI workloads. In May 2025, NVIDIA announced that its Spectrum-X Ethernet platform had been adopted by multiple cloud providers for AI networking, positioning it as an alternative to InfiniBand for large-scale training and inference clusters. The platform pairs ConnectX SuperNICs with BlueField DPUs to offload networking, storage, and security functions from host CPUs. CoreWeave announced in April 2025 that it would deploy NVIDIA's next-generation Blackwell Ultra GB300 systems across its data centers, targeting enterprises running multi-step reasoning workloads that require sustained low-latency interconnects between thousands of GPUs.
Technical benchmarks for multi-plane Ethernet fabrics in AI environments have begun to emerge from independent testing. In a study published in early 2025, researchers at Meta demonstrated that RoCEv2 over Ethernet could match InfiniBand performance for collective communication patterns used in large model training, provided that adaptive routing and priority-based flow control were properly configured. This finding supports CoreWeave's architectural bet that Ethernet-based fabrics can serve both training and inference without the operational complexity of maintaining separate InfiniBand and Ethernet networks. NVIDIA's own published benchmarks for Spectrum-X show up to 1.6x improvement in effective bandwidth for AI workloads compared to standard Ethernet switches, a figure CoreWeave's non-blocking leaf-to-spine design is intended to preserve at fleet scale.
Read full article at coreweave.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source