VAST Data and AMD integrate KV cache offloading to boost inference
VAST Data is expanding its partnership with AMD to integrate sixth-generation EPYC processors and Instinct GPUs into its AI Operating System. The collaboration, which also involves DriveNets, introduces a reference architecture focused on optimizing AI inference workloads and key-value cache management for agentic AI applications.
Key Takeaways
- Early testing with AMD Instinct MI355X processors delivered a ninefold improvement in time-to-first-token during concurrent agentic workloads.
- Integrated 6th-Gen AMD EPYC processors into VAST’s CBox and EBox platforms to improve I/O and storage performance.
- Launched an AI infrastructure reference architecture combining AMD Helios rack-scale systems with DriveNets networking and VAST software.
- Automated lifecycle management for KV caches enables the expiration and deletion of sensitive data within persistent AI conversations.
Why It Matters
Inference efficiency is becoming a primary bottleneck as AI deployments shift from one-off queries to persistent, context-heavy agents. By offloading KV cache data from limited GPU memory to disaggregated storage, VAST and AMD allow enterprises to scale concurrent users without proportional hardware costs. For the streaming ecosystem, this signals a transition toward "AI factories" where generative AI performance is defined by data movement architecture rather than raw compute alone. Watch for independent validation of these throughput claims as major cloud providers begin deploying AMD Helios rack-scale infrastructure in late 2026.
Additional Context
The expansion follows AMD’s July 2026 launch of the 6th-Gen EPYC 'Venice' family, which introduced up to 256 cores and 512 threads per socket. Per Tom’s Hardware (July 2026), these processors utilize the Zen 6 microarchitecture and support PCIe Gen-6, doubling the I/O bandwidth available for high-speed storage clusters. AMD has positioned these chips specifically for 'agentic sandboxes' where CPUs handle the logic and tool-calling functions of AI agents, complementing the raw throughput of the Instinct GPU line.
Simultaneously, AMD announced its Helios rack-scale systems, which compete directly with NVIDIA’s GB200 NVL72. According to Fierce Network (July 2026), the Helios architecture leverages open-standard Ethernet networking via the AMD Pensando Pollara 400 NIC. This contrast to NVIDIA’s proprietary InfiniBand interconnects is a central pillar of the VAST and AMD joint strategy, targeting cloud providers like Oracle Cloud Infrastructure and hyperscalers seeking more flexible, vendor-neutral hardware stacks.
The focus on KV cache offloading is particularly relevant as model context windows expand. Per VAST Data's own technical documentation (July 2026), externalizing the KV cache to its Disaggregated Shared Everything (DASE) architecture allows the system to reuse context across long conversations without recomputing data for every new request. This technical shift is intended to reduce the 'cost per token' in production environments, a metric that multi-agent AI token costs executives from Meta and OpenAI highlighted as critical during the AMD Advancing AI 2026 keynote. As agentic AI traffic scaling continues to challenge existing infrastructure, these hardware-level optimizations will become standard for high-concurrency video and data platforms.
Read full article at pulse2.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source