CoreWeave deploys multi-plane networking for agentic AI across 50 data centers
CoreWeave outlines its infrastructure strategy for supporting agentic AI workloads, emphasizing the use of NVIDIA Spectrum-X Ethernet and BlueField-3 DPUs. The company details a multi-plane, non-blocking network design intended to manage unpredictable, latency-sensitive traffic patterns at scale across 50 data centers.
Key Takeaways
- Multi-plane topology enables the fabric to scale to 128,000 GPUs without adding latency-inducing network tiers
- NVIDIA ConnectX-9 SuperNICs move load balancing to hardware to react to congestion at line rate
- Non-blocking 1:1 uplink ratios ensure full line rate for every node across 50 data centers
- BlueField-3 DPUs offload networking and security tasks to preserve host CPU cycles for agent orchestration
Why It Matters
The shift from rhythmic training workloads to bursty, multi-hop agentic chains requires a fundamental redesign of the AI backend. By moving load balancing and telemetry into dedicated silicon like the ConnectX-9, CoreWeave reduces the jitter that typically accumulates during complex model-to-model calls. This infrastructure strategy signals a move away from general-purpose cloud environments toward specialized fabrics that treat networking as a first-class citizen alongside compute. As streaming platforms increasingly integrate AI agents for real-time content discovery and metadata generation, the industry must track whether these multi-plane designs become the standard for maintaining sub-millisecond response times at fleet scale.
Additional Context
CoreWeave's multi-plane networking strategy sits within a broader competitive landscape of AI cloud providers racing to build specialized fabrics for inference-heavy workloads. In March 2025, CoreWeave raised $1.5 billion in its initial public offering on Nasdaq, valuing the GPU cloud provider at roughly $23 billion and giving it capital to expand its data center footprint beyond the 50 facilities it currently operates. The company has since announced plans to invest approximately $12 billion in additional data center capacity through 2026, a buildout that will require the same multi-plane, non-blocking fabric architecture described in its engineering blog to maintain performance consistency at scale.
NVIDIA's Spectrum-X platform, which CoreWeave uses as its Ethernet foundation, has become a key differentiator in the AI networking market. At GTC 2025 in March, NVIDIA announced that Spectrum-X Ethernet had been adopted by more than 40 cloud and enterprise customers, including hyperscalers and neoclouds seeking alternatives to InfiniBand for inference and agentic workloads. The platform pairs ConnectX-8 SuperNICs with BlueField-3 DPUs to offload congestion control and telemetry from host CPUs, a design philosophy that CoreWeave has extended with its own multi-plane topology. Meanwhile, NVIDIA confirmed at Computex 2025 that its Vera Rubin NVL72 rack-scale system would ship with Spectrum-X Ethernet as the default interconnect, signaling that the Ethernet-based approach is becoming the company's primary path for next-generation AI clusters.
On the technical side, independent benchmarking of multi-plane Ethernet fabrics for AI inference remains limited, but early data suggests meaningful latency advantages over traditional Clos networks. A Broadcom-commissioned study published in February 2025 found that AI inference workloads experienced 35% lower tail latency on Ethernet fabrics with dedicated congestion management compared to standard data center networks, though the study did not test CoreWeave's specific multi-plane configuration. The distinction matters because agentic AI chains, where one model's output feeds another in rapid succession, amplify tail latency in ways that single-model inference does not. CoreWeave's decision to maintain 1:1 oversubscription ratios across all planes directly addresses this amplification effect, and its use of BlueField-3 DPUs for in-network telemetry gives operators visibility into per-flow behavior that traditional switch-based monitoring cannot provide at the same granularity.
Read full article at coreweave.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source