DDN targets GPU efficiency for massive deployments like xAI and Salesforce
DataDirect Networks (DDN) CEO Alex Bouzari discussed the role of specialized AI data infrastructure in optimizing GPU utilization for large-scale deployments during the RAISE Summit 2026. The infrastructure is highlighted for its role in enabling sovereign AI projects and managing high-scale workloads for clients like xAI and Salesforce.
Key Takeaways
- DDN software at Salesforce resulted in a 1.5x speed increase for model training and a 42% reduction in training costs.
- Data sovereignty is driving DDN's involvement in a dozen national AI projects requiring localized data residency.
- NVIDIA has utilized DDN for its internal supercomputers for eight years, a partnership CEO Jensen Huang calls essential.
- The infrastructure manages massive scales, including hundreds of thousands of GPUs for xAI's Colossus cluster.
Why It Matters
The streaming and AI industries are shifting focus from raw compute acquisition to economic efficiency. As large-scale clusters like xAI’s struggle with low utilization rates, the data layer has emerged as the critical bottleneck. For streaming providers using large language models for recommendation or content generation, infrastructure that maximizes GPU uptime is becoming a competitive necessity to control ballooning cloud costs. Moving forward, the industry metric for success will likely pivot from total GPU count to cost-per-token and effective utilization. Watch for whether DDN's data residency features accelerate the adoption of sovereign AI clouds in highly regulated markets.
Additional Context
The emphasis on GPU efficiency arrives as large-scale deployments face significant performance hurdles. Per The Information and Business Insider in May 2026, xAI’s massive 550,000-GPU fleet in Memphis was reportedly operating at just 11% Model FLOPs Utilization (MFU). This underscores a broader industry crisis where coordination overhead and data bottlenecks leave billions of dollars in silicon sitting idle. In contrast, hyperscalers like Meta and Google typically maintain utilization rates between 43% and 46%, setting a high benchmark for third-party infrastructure providers like DDN to match in private or sovereign data centers. To address these gaps, DDN launched the Infinia 2.4 platform at the RAISE Summit in July 2026. The new release introduces distributed Key-Value (KV) Cache acceleration, which per internal DDN benchmarks, can provide up to 1.5x higher GPU utilization by moving the cache layer closer to the data source. This technology is designed to integrate specifically with NVIDIA’s Dynamo and DSX blueprint architectures, streamlining the path between storage and compute to reduce 'time-to-first-token' in generative AI workloads. Competing platforms are also racing to solve the data-readiness problem. At the same RAISE Summit in July 2026, Hammerspace demonstrated its own high-performance data platform aimed at automating orchestration to eliminate the wait times between data ingestion and AI execution. As the market for AI factories matures, the competition between software-defined storage vendors is increasingly defined by their ability to reduce cost-per-token and ensure sovereign control, particularly as European nations seek independent AI infrastructure that bypasses traditional US-based hyperscale clouds.
Read full article at siliconangle.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source