Alibaba Qwen3.8-Flash-Next achieves 16K tokens per second on NVIDIA GB300
Alibaba has released Qwen3.8-Flash-Next, a 125B-parameter multimodal mixture-of-experts model optimized for long-context tasks like agentic coding. NVIDIA has validated the model on its GB300 NVL72 platform, providing deployment recipes through NeMo and various inference engines.
Key Takeaways
- Hybrid architecture uses Gated DeltaNet and Qwen Sparse Attention to maintain fixed-size recurrent states across 1M-token contexts
- NVIDIA GB300 NVL72 platform enables 130 TB/s all-to-all communication for low-latency inference
- Benchmarks show 7.6x prefill and 4.9x decoding speedups compared to standard full-attention models
- Deployment support includes NVIDIA NeMo AutoModel for fine-tuning and open-source recipes for SGLang and vLLM
Why It Matters
The validation of Qwen3.8-Flash-Next on the GB300 NVL72 platform marks a significant shift toward high-throughput, long-context AI for complex video and coding workflows. By combining Gated DeltaNet with sparse attention, Alibaba addresses the KV cache memory bottleneck that typically degrades performance in massive context windows. For the streaming ecosystem, this infrastructure enables more sophisticated automated metadata tagging and tool-driven content processing at scale. The integration with NVIDIA NeMo ensures that B2B developers can move from local RTX workstations to rack-scale production without changing their underlying software stack. Watch for Qwen4 architecture benchmarks to see if these hybrid attention gains hold as parameter counts scale further.
Additional Context
Alibaba's Qwen model family has become one of the most widely adopted open-weight LLM ecosystems in the AI infrastructure space. The company's Qwen3 series, released in April 2025, introduced a hybrid thinking mode that allows models to switch between extended reasoning and rapid response, and Qwen3 models quickly gained traction across major cloud providers including AWS, Azure, and Google Cloud within weeks of launch. NVIDIA has consistently validated Qwen releases on its latest hardware, positioning the partnership as a reference architecture for enterprises deploying open-weight models at scale. The GB300 NVL72 platform, part of NVIDIA's Blackwell Ultra generation, represents the company's current flagship rack-scale system for inference-heavy workloads, and Alibaba's decision to target it first for Qwen3.8-Flash-Next signals the model's orientation toward production agentic pipelines rather than research benchmarks. The competitive landscape for long-context inference is intensifying as multiple model providers race to demonstrate million-token capabilities on NVIDIA hardware. Nvidia is working on AI deals worth more than $750 billion, including a partnership with SK Group to do more than $500 billion in business, underscoring the scale of capital flowing into AI compute infrastructure that platforms like GB300 NVL72 are designed to serve. Meanwhile, Cerebras filed for an IPO with a reported $10 billion contract with OpenAI, signaling that alternative architectures are beginning to challenge NVIDIA's dominance in AI acceleration. For Alibaba, the strategic value of NVIDIA validation lies in ensuring Qwen3.8-Flash-Next remains deployable on the most widely available enterprise GPU infrastructure, even as competitors explore wafer-scale and custom silicon approaches. On the inference software side, the ecosystem around Qwen3.8-Flash-Next includes multiple serving frameworks that NVIDIA has optimized for Blackwell-class hardware. Deepgram's integration with AWS SageMaker demonstrates how real-time AI endpoints now run inside customer VPCs with sub-second latency using bidirectional streaming, a deployment pattern that mirrors the low-latency, high-throughput requirements NVIDIA targets with its NeMo and SGLang recipes for Qwen models. The 16K tokens-per-second throughput figure reported for Qwen3.8-Flash-Next on GB300 NVL72 places it among the fastest open-weight models for agentic coding tasks, where tool-call latency and context retention directly affect developer productivity. For streaming and media companies evaluating , these throughput numbers suggest that million-token context windows for metadata extraction, script analysis, and for multi-step editorial workflows are approaching production viability on commercially available rack-scale systems.
Read full article at developer.nvidia.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source