Solidigm deploys 122TB SSDs to solve agentic AI inference bottlenecks
Solidigm is deploying 122TB SSDs and liquid-cooled storage solutions to address data-access bottlenecks in agentic AI inference and large-scale enterprise GPU clusters. The company positions these high-capacity, cooling-integrated drives as a critical 'intelligence layer' for maintaining constant GPU token generation density.
Key Takeaways
- 122.88TB D5-P5336 SSDs provide high-density storage aimed at keeping expensive GPU clusters at 100% utilization
- Newly launched cold-plate-cooled SSDs enable fully fanless Nvidia GPU server designs as racks transition to liquid cooling
- Agentic AI sessions require managing 5GB to 10GB of context data to generate up to 40,000 tokens per prompt
- Solidigm and parent SK Hynix are positioning storage as a strategic memory extension rather than a commodity tier
Why It Matters
The shift from training to agentic inference transforms storage from simple plumbing into a critical extension of the memory hierarchy. As AI agents move from single prompts to long-duration reasoning chains, GPUs face idle time waiting for massive KV cache and context data to load from legacy infrastructure. High-density, liquid-cooled SSDs mitigate these bottlenecks, allowing for the extreme token density required by sovereign and enterprise AI factories. This technical pivot suggests that infrastructure ROI in the streaming and media AI sector will increasingly depend on storage throughput rather than raw compute alone. Watch for the adoption of Nvidia's CMX context memory platform as a signal for broader storage-layer standardization.
Additional Context
The transition to 'agentic AI' is driving a structural shift in data center economics where the primary constraint is no longer model size, but the ability to serve context data at scale. Per Forbes (June 2026), while token generation costs have dropped significantly, total enterprise AI expenditures are rising due to the massive volumes of persistent session data required for autonomous agents. This trend has created a specific demand for KV cache offloading, a process where large language model context is moved from limited GPU memory to high-speed NVMe storage to sustain performance during complex, multi-turn interactions. Market volatility is already reflecting this infrastructure pressure. Per TechRadar (February 2026), the price of Solidigm’s 122.88TB D5-P5336 SSD surged nearly 200% to $37,128 in less than a year, driven by tight NAND supply and the aggressive build-out of enterprise AI clusters. To address the resulting thermal challenges, NVIDIA introduced its CMX Context Memory Storage Platform in early 2026, which uses BlueField-4 DPUs to manage context data between SSDs and compute units. According to TrendForce (June 2026), these architectural shifts are expected to sustain a structural memory shortage through the end of the year as hyperscalers prioritize liquid-cooled, high-density storage pods.
Read full article at siliconangle.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source