vLLM prepares day-0 support for 2.8T parameter Kimi K3 release
The vLLM project and Moonshot AI are collaborating to provide day-0 open-source serving support for the new 2.8-trillion-parameter Kimi K3 model. The integration includes specialized kernels and architectural optimizations for NVIDIA and AMD hardware to handle Kimi K3's high-scale inference requirements and 1-million-token context window.
Key Takeaways
- Kimi K3 employs a massive 2.8T parameter sparse MoE architecture activating 16 of 896 experts per token.
- Engineers implemented KDA-aware prefix caching in vLLM to separate physical state-block size from prefix-match granularity.
- Day-0 support includes tailored FlashKDA integration and fused decode kernels for NVIDIA and AMD GPU stacks.
- Optimization work includes a new MLA module specialized for prefill/decode-disaggregated (PD) production environments.
Why It Matters
The collaboration marks a critical advance in high-performance open-weight model serving. By securing day-0 vLLM support, Moonshot AI ensures that its 2.8T parameter model — currently the largest open-weight system — is immediately deployable at scale. This integration solves the memory and latency bottlenecks inherent in Kimi Delta Attention (KDA) and multi-trillion MoE architectures, moving these complex frontier designs into the standard open-source stack. For the streaming ecosystem, this provides a verified blueprint for serving ultra-long context multimodal models without proprietary lock-in. Watch for initial production throughput benchmarks on NVIDIA GB200 and AMD MI300 clusters following the July 27 weight release.
Additional Context
Moonshot AI announced Kimi K3 on July 16, 2026, positioning the 2.8-trillion-parameter model as the largest open-weight AI system to date. Per Tom's Hardware, the model currently ranks third on major intelligence indices, trailing only the closed-source Claude Fable 5 and GPT-5.6 Sol, while outperforming most competitors in frontend coding benchmarks with an Elo of 1,679. The model's architecture is a significant departure from the previous Kimi K2 line, utilizing Kimi Delta Attention (KDA) and Attention Residuals to achieve what Moonshot claims is a 2.5x improvement in scaling efficiency. These architectural shifts allow for a 1-million-token context window, matching the capacity of the current market leaders.
Recent reporting from VentureBeat in July 2026 indicates that while K3 is currently accessible via API, the move to release open weights by July 27 is intended to capture the market for self-hosted enterprise agents. The high operational costs of 3T-class models remains a primary hurdle; industry analysts at Artificial Analysis estimate that production deployments require supernode configurations of at least 64 accelerators. Moonshot has addressed these economic constraints through extreme Mixture-of-Experts (MoE) sparsity, activating only 1.8% of the total parameters (approximately 50 billion) during a standard forward pass. This enables API pricing of roughly $3 per million input tokens, effectively undercutting American proprietary rivals for long-horizon reasoning tasks.
Read full article at vllm.ai
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source