NVIDIA open-sources cuFile APIs with Google, Intel, and Meta
NVIDIA has open-sourced its cuFile APIs in collaboration with Google, Intel, and Meta to standardize GPU-to-storage data access. The initiative, dubbed Storage-Next, involves over 40 industry partners and aims to reduce infrastructure bottlenecks that cause GPU idle time during large-scale AI inference.
Key Takeaways
- NVIDIA moved its cuFile APIs to a neutral GitHub organization with Google, Intel, and Meta serving as inaugural maintainers.
- The new Storage-Next initiative aligns over 40 vendors, including Micron, KIOXIA, and DDN, to standardize GPU-driven storage behavior.
- Internal benchmarks claim the Vera CPU in BlueField-4 STX provides 3.21x higher throughput than x86 CPUs in compression and encryption pipelines.
- The CMX Context Memory Storage platform creates a dedicated tier for KV cache to optimize long-context agentic inference.
Why It Matters
NVIDIA is shifting its strategy to address the growing data bottleneck between storage and accelerators, which currently leaves GPUs idle for up to 50% of inference runtime. By open-sourcing the interface, NVIDIA attempts to replicate its CUDA success—making its architecture the industry baseline while retaining the performance advantage on its proprietary silicon. This move forces competitors like AMD to compete on a software standard defined by NVIDIA. Strategists should monitor if the 12 storage providers codesigning STX systems, such as Dell and HPE, can maintain differentiation as their roles shift toward becoming qualified enclosure suppliers for NVIDIA's DOCA-driven architecture.
Additional Context
The announcement at FMS 2026 follows a period of intense competition in AI-native networking and storage. Per AMD, July 2026, the company recently introduced its Helios rack-scale platform featuring the Pensando Salina DPU and Vulcano AI NIC to provide an alternative data-path stack for OEMs. AMD claims its Helios architecture generates up to 30% more tokens per dollar than NVIDIA's Vera Rubin NVL72 solutions, specifically targeting the efficiency of agentic AI workloads that require persistent memory access across multiple services. Industry data highlights the urgency of these infrastructure pivots. According to the Futurum 1H 2026 Data Intelligence Decision Maker Survey, over 56% of enterprises are already employing quantization or distillation techniques to manage escalating inference costs. This focus on cost containment has turned the KV cache—which stores intermediate model calculations—into a critical product category. While NVIDIA's CMX aims to define this tier, rivals like WEKA reported in May 2026 that their software-based augmented memory solutions can deliver 4x to 5x higher token throughput from existing hardware, illustrating that the battle for the inference memory hierarchy is being fought across both hardware and software layers.
Read full article at futurumgroup.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source