VAST Data and AMD link AI software to external GPU memory
VAST Data and AMD have expanded their architectural collaboration to integrate VAST's AI Operating System with AMD's Instinct GPUs and EPYC processors. The partnership leverages persistent key-value caching to manage large-scale inference context outside of GPU memory, aiming to reduce recomputation costs for generative AI workloads.
Key Takeaways
- VAST reported a 9x improvement in time-to-first-token (TTFT) and 9.7x higher token throughput on MI355X systems by offloading context memory.
- The partnership utilizes sixth-generation AMD EPYC 'Venice' processors to power VAST's next-generation CBox and EBox storage platforms.
- Integrated PCIe Gen 6 support provides 2x the I/O bandwidth compared to previous hardware for data warehouse and event streaming services.
- Native data lifecycle policies enable automated expiration and deletion of sensitive user information stored within the persistent KV cache.
Why It Matters
This collaboration addresses the 'memory wall' hindering long-form generative AI by moving inference state from expensive HBM3E to scalable persistent storage. By integrating with AMD’s ROCm software and Instinct accelerators, VAST is positioning its platform as a necessary context-management layer rather than just high-performance storage. For the broader ecosystem, this move enables AMD to field a competitive enterprise inference stack against NVIDIA’s native Dynamo and NIXL caching tiers. Success will likely be measured by how well this architecture performs in high-concurrency environments like the AMD Helios rack-scale systems slated for production in late 2026.
Additional Context
The VAST-AMD partnership is a direct response to the massive data footprints generated by modern Large Language Models (LLMs). According to reported data, a single 128,000-token context window can generate between 20 GB and 50 GB of KV cache depending on precision. As enterprise use cases shift toward 'agentic AI'—systems that maintain long-term session history across multiple turns—this data demand quickly exceeds the 288 GB HBM3E capacity of high-end accelerators like the AMD Instinct MI355X. Per industry reporting in mid-2026, managing this state has become a primary operational bottleneck for AI cloud providers. While VAST and AMD focus on persistent storage, others are exploring tiered memory hierarchies. Vultr, which joined the Vultr Cloud Alliance in early 2026 per official company releases, has been validating architectures that span both AMD and NVIDIA ecosystems. This 'two-lane' GPU strategy allows providers to compare VAST’s software-defined context management against hardware-native solutions. Recent benchmarks from competitive clusters using NVIDIA’s BlueField-4 DPUs and Spectrum-X networking have shown that while local RAM offload provides lower latency for single-node tasks, shared persistent namespaces are required for the multi-node scaling typical of enterprise 'AI factories.' Beyond performance, the collaboration highlights a growing focus on data governance for AI. Unlike transient GPU memory, persistent KV cache stores potentially sensitive user prompts and history for long periods. Per VAST’s July 2026 announcement, the integration of native data lifecycle policies allows these caches to inherit existing enterprise security and retention rules. This mechanism addresses regulatory requirements that point-product caching solutions often overlook, providing a compliance bridge for highly regulated sectors such as finance and healthcare as they move models from small-scale testing into production inference environments.
Read full article at hyperframeresearch.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source