Alibaba debuts 2.4T parameter Qwen3.8-Max capable of 100-hour video analysis
Alibaba has introduced Qwen3.8-Max, a 2.4 trillion parameter LLM featuring a Gated DeltaNet attention mechanism that enables linear scaling for processing long-form data, including up to 100 hours of video. The model will be available via Alibaba Cloud and will be released as an open-source project, including a more compact 27B parameter version, next week.
Key Takeaways
- Qwen3.8-Max features 2.4 trillion parameters, a seven-fold increase over the Qwen3.5 model released in February.
- The architecture utilizes a Gated DeltaNet module that enables linear hardware scaling for long-context prompts.
- Internal testing showed the model successfully completing a 16-day autonomous coding project and a 500-step chip design task.
- Alibaba will open-source the model weights next week, alongside a smaller 27-billion parameter version.
- The model scored 1,668 on the Frontend Code Arena benchmark, trailing Anthropic's Claude Opus 5 by only 37 points.
Why It Matters
The launch of Qwen3.8-Max represents a significant technical leap for video-native AI workloads by moving beyond the quadratic scaling limitations of standard attention mechanisms. For the streaming industry, this suggests a future where high-fidelity, long-form content analysis and metadata generation can be performed at a fraction of current compute costs. Strategically, Alibaba’s decision to open-source a model of this scale pressures Western labs like Anthropic and Meta to maintain their lead in proprietary performance while competing with a high-utility, open-weight alternative. Watch for the public weights release on Hugging Face next week to see how independent developers leverage the 27B version for edge-based video processing.
Additional Context
The release of Qwen3.8-Max follows a period of rapid architectural convergence in the Chinese AI sector. Per AI Modeling and Marktechpost (May 2026), several top labs including Alibaba and Moonshot AI have adopted a 3:1 hybrid attention ratio, pairing Gated DeltaNet or similar linear mechanisms with standard full-attention layers. This specific configuration is reported to reduce KV cache memory requirements by up to 75% for million-token contexts, solving the primary bottleneck for processing massive video files. This shift is critical as the industry moves from simple chatbots toward agentic AI video production capable of multi-day autonomous tasks.
Alibaba has concurrently reorganized its internal AI structures to support this push. According to CRN Asia (July 2026), the company established the Alibaba Token Hub Business Group under CEO Eddie Wu to consolidate its model-as-a-service and Qwen business units. This structural change aligns with Alibaba's stated goal of investing RMB 380 billion ($53 billion) into cloud and AI infrastructure by 2028. Recent expansions include new availability zones in Japan and Malaysia, specifically designed to host the high-compute demands of frontier models like the Qwen3 series.
Competitive pressure within the domestic Chinese market remains intense. Forbes (August 2026) noted that Alibaba’s stock rose 4.5% in premarket trading following the announcement, as investors reacted to Qwen3.8-Max appearing to close the gap with U.S. models like Anthropic's Claude Fable 5. This release comes just days after rival Moonshot AI launched Kimi K3, a 2.8 trillion parameter model, signaling an 'arms race' of scale among Chinese providers who are increasingly prioritizing open-weight availability to gain global market share over closed-source American competitors.
Read full article at siliconangle.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source