vLLM adds day-0 support for 2.8T Kimi K3 MoE model
vLLM has introduced day-0 production support for Moonshot AI’s Kimi K3, a 2.8-trillion-parameter Mixture-of-Experts model. The integration provides optimized kernel backends and caching strategies to improve latency for agentic workloads on NVIDIA Blackwell and AMD MI355X hardware.
Key Takeaways
- Kimi K3 utilizes a 2.8T parameter architecture with 896 total experts, activating 16 per token.
- Integrated DSpark speculative decoding delivers 370 tok/s on 16 NVIDIA GB300 GPUs, a 3.14x speedup over standard inference.
- Native supports for NVIDIA Blackwell (B300) and AMD MI355X hardware at launch.
- New hybrid prefix caching redesigned to support recurrent state in the Kimi Delta Attention stack.
- Model features a 1-million-token context window with native vision and multimodal reasoning capabilities.
Why It Matters
The day-0 support for Kimi K3 marks a critical shift in the weight-open ecosystem, moving toward 3T-class models that rival proprietary streaming agents. For streaming engineers, the optimization of hybrid linear attention (KDA) and recurrent state caching solves the memory bottlenecks typically associated with million-token context windows. This enables highly personalized, long-horizon agentic workflows—such as real-time video editing or deep research agents—to run on standardized B300 and MI355X clusters. As vLLM standardizes these kernels, the barrier to deploying frontier-level intelligence for enterprise-grade video applications drops significantly. Competitive focus will likely shift to the efficiency of disaggregated prefill/decode architectures for high-throughput multimodal serving.
Additional Context
The release of Kimi K3 weights follows a growing trend of high-parameter Mixture-of-Experts (MoE) models dominating the open-weight landscape in 2026. Per Microsoft, October 2025, the deployment of NVIDIA GB300 NVL72 superclusters has set a new standard for serving multitrillion-parameter systems, providing the 1.5x FP4 performance jump necessary for Kimi K3's MXFP4 native weights. This hardware-software co-design is now standard, with vLLM's implementation directly leveraging NVIDIA’s Blackwell Ultra platform for reasoning tasks.
Simultaneously, the competitive pressure among open-weight developers is intensifying. Per 36kr, July 2026, the AMD team completed day-one adaptation for the MI355X chip to ensure cross-vendor compatibility, while cloud providers like Modal have already launched K3 hosting services. This rapid ecosystem uptake mirrors the release cycles of other frontier models throughout 2026, such as DeepSeek V4 and Llama 4, which have pushed MoE architectures toward 1,000-expert configurations to balance scaling efficiency with inference cost.
Industry benchmarks suggest Kimi K3 is specifically targeting high-entropy coding and agentic tasks. According to Inferact, July 2026, the open-sourced DSpark speculator was trained on vLLM hidden states to achieve full numerical parity, specifically to improve the 1M-token context performance required for large-scale codebase navigation. As of July 27, 2026, Moonshot AI has released these weights under a modified MIT license, allowing enterprises to self-host the model for GDPR-compliant workloads without routing data through proprietary external APIs.
Read full article at vllm.ai
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source