StreamingMemeStreamingMemeBuyers Guide
AboutLeaderboardsEventsSubmit News
Subscribe

Daily Brief

The streaming industry in your inbox every morning.

Daily Brief

The streaming industry in your inbox every morning.

StreamingMemeStreamingMeme

The independent buyers guide and news aggregator for the streaming technology industry.

Explore

Buyers GuideLeaderboardsEventsSubmit News

Stay updated

Weekly digest of new companies and streaming news.

Categories

Encoding & SoftwareVideo Delivery & CDNStreaming PlatformsAI for VideoProduction HardwareBusiness NewsMonetization & Ad TechRegulatory & Policy

© 2026 StreamingMeme. All rights reserved.

AboutPrivacy PolicyTermsContact
EncodingCDNPlatformsAI & VideoHardwareBusinessAd TechPolicy
← AI for Video
AI & VideoStrategic PartnershipJuly 27, 2026

VAST Data and AMD link AI software to external GPU memory

VAST Data and AMD link AI software to external GPU memory
HyperFRAME Research

VAST Data and AMD have expanded their architectural collaboration to integrate VAST's AI Operating System with AMD's Instinct GPUs and EPYC processors. The partnership leverages persistent key-value caching to manage large-scale inference context outside of GPU memory, aiming to reduce recomputation costs for generative AI workloads.

Key Takeaways

  • VAST reported a 9x improvement in time-to-first-token (TTFT) and 9.7x higher token throughput on MI355X systems by offloading context memory.
  • The partnership utilizes sixth-generation AMD EPYC 'Venice' processors to power VAST's next-generation CBox and EBox storage platforms.
  • Integrated PCIe Gen 6 support provides 2x the I/O bandwidth compared to previous hardware for data warehouse and event streaming services.
  • Native data lifecycle policies enable automated expiration and deletion of sensitive user information stored within the persistent KV cache.

Why It Matters

This collaboration addresses the 'memory wall' hindering long-form generative AI by moving inference state from expensive HBM3E to scalable persistent storage. By integrating with AMD’s ROCm software and Instinct accelerators, VAST is positioning its platform as a necessary context-management layer rather than just high-performance storage. For the broader ecosystem, this move enables AMD to field a competitive enterprise inference stack against NVIDIA’s native Dynamo and NIXL caching tiers. Success will likely be measured by how well this architecture performs in high-concurrency environments like the AMD Helios rack-scale systems slated for production in late 2026.

Additional Context

The VAST-AMD partnership is a direct response to the massive data footprints generated by modern Large Language Models (LLMs). According to reported data, a single 128,000-token context window can generate between 20 GB and 50 GB of KV cache depending on precision. As enterprise use cases shift toward 'agentic AI'—systems that maintain long-term session history across multiple turns—this data demand quickly exceeds the 288 GB HBM3E capacity of high-end accelerators like the AMD Instinct MI355X. Per industry reporting in mid-2026, managing this state has become a primary operational bottleneck for AI cloud providers. While VAST and AMD focus on persistent storage, others are exploring tiered memory hierarchies. Vultr, which joined the Vultr Cloud Alliance in early 2026 per official company releases, has been validating architectures that span both AMD and NVIDIA ecosystems. This 'two-lane' GPU strategy allows providers to compare VAST’s software-defined context management against hardware-native solutions. Recent benchmarks from competitive clusters using NVIDIA’s BlueField-4 DPUs and Spectrum-X networking have shown that while local RAM offload provides lower latency for single-node tasks, shared persistent namespaces are required for the multi-node scaling typical of enterprise 'AI factories.' Beyond performance, the collaboration highlights a growing focus on data governance for AI. Unlike transient GPU memory, persistent KV cache stores potentially sensitive user prompts and history for long periods. Per VAST’s July 2026 announcement, the integration of native data lifecycle policies allows these caches to inherit existing enterprise security and retention rules. This mechanism addresses regulatory requirements that point-product caching solutions often overlook, providing a compliance bridge for highly regulated sectors such as finance and healthcare as they move models from small-scale testing into production inference environments.


Read full article at hyperframeresearch.com

Get this in your inbox → Subscribe

Enjoy our coverage?

Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.

Add as preferred source

Related Articles

The Cloudflare Blog: Cloudflare and Anthropic add Claude agents to Cloudflare Sandboxes
YourStory: Frammer AI and Cineverse tap 71,000 titles for short-form video
NVIDIA Newsroom: NAVER Scales AI Infrastructure with NVIDIA DSX for Gigawatt AI Factories
TM Broadcast: Fox names AWS preferred AI provider to automate live sports highlights

Newest

about 2 hours ago
AOL: UK Government considers complete Freeview switch-off between 2034 and 2044
about 2 hours ago
Broadcast: Games of the Future 2026 secures global streaming and broadcast distribution
about 2 hours ago
Investing.com: Alphabet upgraded as Google Cloud revenue surges 82% on AI demand
about 3 hours ago
The Desk: Phynd launches ad-supported cloud gaming beta on LG webOS
about 3 hours ago
Kalkine Media: Adveritas hits A$16.3M recurring revenue, shifts toward cash flow breakeven
about 8 hours ago
MediaPost: Microsoft launches Project Perception to defend programmatic supply chains from AI-driven fraud
about 8 hours ago
Digiday: IAB Redefining Media Types Standard targets automated video ad transparency
about 8 hours ago
AdExchanger: Streaming ad tech consolidation turns independent platforms into proprietary gardens
about 8 hours ago
VentureBeat: Moonshot AI releases Kimi K3 weights with $20M revenue licensing threshold
about 8 hours ago
The Fast Mode: AMD and South Korea Partner to Build Heterogeneous Sovereign AI Infrastructure
about 8 hours ago
Hyper.ai: Google DeepMind and UC Riverside launch framework to trace synthetic video
about 8 hours ago
Exame: Brazil launches TV 3.0 with 4K VVC and interactive IP layers
about 8 hours ago
Advanced Television: Roku and Fire TV solidify gatekeeper status as OS influence grows
about 8 hours ago
AdExchanger: Google mandates biometric passkeys for Ads API as AI costs reshape agency deals
1 day ago
Hackernoon: Production voice pipeline solves African language latency and hallucination problems
1 day ago
Design & Reuse: Stricter ETSI secure boot standards mandate hardware-level chain of trust
1 day ago
SiliconANGLE: Dell and AMD target cloud token costs with modular AI inference
1 day ago
GlobeNewswire: Kaltura serves 7 million concurrent World Cup viewers using microservices architecture
1 day ago
Master of Code Global: Multimodal AI latency framework tackles processing bottlenecks in enterprise pipelines
1 day ago
Tech Xplore: Google and UC Riverside unveil SAGA tool to trace AI video origins

Upcoming Events

Jul
29–30
Buffer-Free VideoSeattle
Aug
17–20
SET EXPOSao Paulo
Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
View all events →

Top Sources

  1. 1.Sports Video Group105
  2. 2.SiliconANGLE93
  3. 3.AdExchanger66
  4. 4.Tech Times65
  5. 5.YouTube62
  6. 6.TechCrunch56
  7. 7.PPC Land51
  8. 8.arXiv50
Full leaderboards →

Newest

about 2 hours ago
AOL: UK Government considers complete Freeview switch-off between 2034 and 2044
about 2 hours ago
Broadcast: Games of the Future 2026 secures global streaming and broadcast distribution
about 2 hours ago
Investing.com: Alphabet upgraded as Google Cloud revenue surges 82% on AI demand
about 3 hours ago
The Desk: Phynd launches ad-supported cloud gaming beta on LG webOS
about 3 hours ago
Kalkine Media: Adveritas hits A$16.3M recurring revenue, shifts toward cash flow breakeven
about 8 hours ago
MediaPost: Microsoft launches Project Perception to defend programmatic supply chains from AI-driven fraud
about 8 hours ago
Digiday: IAB Redefining Media Types Standard targets automated video ad transparency
about 8 hours ago
AdExchanger: Streaming ad tech consolidation turns independent platforms into proprietary gardens
about 8 hours ago
VentureBeat: Moonshot AI releases Kimi K3 weights with $20M revenue licensing threshold
about 8 hours ago
The Fast Mode: AMD and South Korea Partner to Build Heterogeneous Sovereign AI Infrastructure
about 8 hours ago
Hyper.ai: Google DeepMind and UC Riverside launch framework to trace synthetic video
about 8 hours ago
Exame: Brazil launches TV 3.0 with 4K VVC and interactive IP layers
about 8 hours ago
Advanced Television: Roku and Fire TV solidify gatekeeper status as OS influence grows
about 8 hours ago
AdExchanger: Google mandates biometric passkeys for Ads API as AI costs reshape agency deals
1 day ago
Hackernoon: Production voice pipeline solves African language latency and hallucination problems
1 day ago
Design & Reuse: Stricter ETSI secure boot standards mandate hardware-level chain of trust
1 day ago
SiliconANGLE: Dell and AMD target cloud token costs with modular AI inference
1 day ago
GlobeNewswire: Kaltura serves 7 million concurrent World Cup viewers using microservices architecture
1 day ago
Master of Code Global: Multimodal AI latency framework tackles processing bottlenecks in enterprise pipelines
1 day ago
Tech Xplore: Google and UC Riverside unveil SAGA tool to trace AI video origins

Upcoming Events

Jul
29–30
Buffer-Free VideoSeattle
Aug
17–20
SET EXPOSao Paulo
Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
View all events →

Top Sources

  1. 1.Sports Video Group105
  2. 2.SiliconANGLE93
  3. 3.AdExchanger66
  4. 4.Tech Times65
  5. 5.YouTube62
  6. 6.TechCrunch56
  7. 7.PPC Land51
  8. 8.arXiv50
Full leaderboards →