NVIDIA GB300 NVL72 achieves 4K tokens per second with Alibaba Qwen3.8
NVIDIA has announced support for the new Alibaba Qwen3.8-2.4T-A95B open-weight model on its GB300 NVL72 hardware platform. The release focuses on high-throughput, agentic-workload performance and long-context inference, offering developers specific optimization recipes for deployment.
Key Takeaways
- Alibaba released open weights for Qwen3.8-2.4T-A95B, a 2.4-trillion parameter model with 95 billion activated per token.
- NVIDIA GB300 NVL72 platform delivers over 350 tokens per second per user in FP8 precision.
- The model uses a hybrid attention architecture to maintain bounded memory for contexts up to one million tokens.
- Configurable reasoning controls (low, high, xhigh) allow developers to trade compute for inference depth per request.
Why It Matters
The collaboration marks a critical milestone for high-throughput AI factories, providing an open-weight alternative to closed frontier models for large-scale video and document analysis. By optimizing the Qwen3.8-2.4T-A95B for Blackwell Ultra architecture, NVIDIA is solving the memory and compute bottlenecks that previously limited long-context agentic workflows. For the streaming industry, this suggests a move toward more autonomous content metadata generation and complex tool-use workflows that exceed the capabilities of standard chat models. Industry leaders should track how this combination performs on multimodal benchmarks like VideoMME to gauge its impact on automated media production.
Additional Context
The deployment of Alibaba Qwen3.8 on NVIDIA GB300 NVL72 occurs amid a massive infrastructure buildout centered on the Blackwell Ultra architecture. Per HPE in November 2025, the GB300 NVL72 racks were first made available to order with a price tag exceeding $3 million per unit. These systems are engineered to provide 20 petaflops of FP4 compute performance, specifically to handle the trillion-parameter models that are becoming the new baseline for enterprise AI. Microsoft Azure also signaled the scale of this shift in October 2025, announcing production clusters featuring over 4,600 GB300 NVL72 units to support OpenAI and other frontier model developers.
Alibaba’s Qwen series has rapidly become a dominant force in the open-weight ecosystem, with over 700 million cumulative downloads reported as of January 2026. The Qwen3.8-Max model, which this hardware update supports, is positioned as a direct competitor to proprietary systems like Anthropic’s Fable 5. According to Alibaba Cloud reports from August 2026, the model has demonstrated significant proficiency in long-horizon tasks, such as an autonomous coding demo where it operated for 16 days to build a self-evolving engineering system. For streaming professionals, the model’s performance on the VideoMME benchmark—scoring 90.4 with subtitles—highlights its potential for sophisticated, real-time video understanding at scale at scale.
Read full article at developer.nvidia.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source