AMD and Moonshot AI optimize Instinct MI355X for agentic video workloads
AMD and Moonshot AI have co-engineered an agentic AI serving stack optimized for AMD Instinct MI355X GPUs using a new Unified Memory & Bandwidth Pool. The solution aims to solve memory-bottleneck issues in persistent, multi-turn agentic workloads by enhancing KV-cache offloading and scheduler awareness.
Key Takeaways
- AMD Instinct MI355X serves as the hardware base for the Kimi K2.6 agentic model on the ROCm platform.
- Unified Memory & Bandwidth Pool (UMBP) enables a multi-tier KV cache across HBM, host DRAM, and SSDs.
- Integrated scheduler-aware caching delivered a 3.2X reduction in p99 Time to First Token (TTFT).
- Optimizations include zero-CU SDMA restores, Delta Tokenization (DeltaTok), and intra-turn tool-engine overlap.
Why It Matters
As AI shifts from simple chat to persistent, multi-turn 'agentic' workflows, memory management becomes the primary bottleneck for video and coding applications. This shift moves the performance yardstick from raw compute (FLOPS) to KV-cache reuse and scheduling efficiency. By integrating the scheduler directly with a multi-tier memory pool, AMD is providing a blueprint for serving frontier models like Kimi K2.6 that outgrow local HBM. Competitively, this positions the MI355X as a viable alternative to NVIDIA’s Blackwell architecture for long-horizon autonomous tasks. Watch for whether this integrated stack becomes the standard for high-bandwidth memory (HBM) offloading in enterprise streaming and automated video production environments.
Additional Context
The collaboration arrives as the industry grapples with the transition to 'test-time scaling' or 'thinking' models, where output token counts are surging by more than 5X annually, according to recent reporting from TrendForce in June 2026. This trend is putting unprecedented pressure on memory architectures, as reasoning-heavy models require massive KV-caches to maintain context over thousands of coordinated steps. NVIDIA addressed this in early 2026 with its CMX Context Memory Storage Platform managed by BlueField-4 DPUs, indicating that the 'memory bottleneck' is now the central front in the AI infrastructure war. Technically, the Kimi K2.6 model used in this validation represents a significant scaling milestone for open-weight models. Per DeepInfra and Miraflow reporting from April 2026, K2.6 utilizes a 1-trillion parameter Mixture-of-Experts (MoE) architecture capable of orchestrating up to 300 sub-agents in a single run. This 'Agent Swarm' capability is specifically designed for long-horizon tasks, such as autonomous software engineering and multimodal content generation, which frequently exceed the 288GB HBM3E capacity of individual MI355X accelerators. On the software side, AMD has accelerated its ROCm release cycle to achieve feature parity with NVIDIA's CUDA ecosystem. ROCm 7.14, released in July 2026, brought production-level support for SGLang and the MoRI communication fabric, which are critical components of the joint AMD-Moonshot stack. According to Forbes, recent MLPerf 6.0 benchmarks show the MI355X achieving within 80-90% of local performance compared to NVIDIA’s B300 GPU, suggesting that software-led optimizations like UMBP are narrowing the gap in real-world agentic throughput.
Read full article at amd.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source