Tencent releases Apache-licensed Hy3 model to challenge Zhipu AI
Tencent has released Hy3, a 295B-parameter Mixture-of-Experts model, under an Apache 2.0 license to facilitate wider enterprise adoption. The model is optimized for production reliability and deployment on export-compliant hardware, positioning it as a competitor to Zhipu AI's GLM-5.2 for agentic search and orchestration tasks.
Key Takeaways
- Hy3 features 295 billion total parameters with 21 billion active per token across 192 experts.
- Internal testing reports a 5.4% hallucination rate, down from 12.5% in the April preview.
- The model includes a 3.8B-parameter multi-token prediction layer to accelerate speculative decoding.
- Recommended serving configuration targets the Nvidia H20-3e to ensure compliance with U.S. export restrictions.
- Tencent is offering the model for free on the OpenRouter platform for a two-week period.
Why It Matters
The move shifts the competitive landscape by removing regional licensing barriers that previously blocked many Western enterprises from using Chinese frontier models. By prioritizing production reliability and lower memory requirements (under 300GB in FP8) over raw coding benchmarks, Tencent is positioning Hy3 as a more attainable, self-hostable alternative to larger models like the 744B-parameter GLM-5.2. For the streaming and video ecosystem, this provides a cost-effective path for high-accuracy metadata tagging, agentic content search, and complex tool orchestration without requiring flagship-scale H200 infrastructure. Watch for independent evaluations from platforms like Artificial Analysis to verify Tencent's internal reliability claims against Western benchmarks.
Additional Context
The release of Hy3 follows a period of rapid iteration for Tencent’s artificial intelligence division. Per TechNode (July 2026), the official launch comes just ten weeks after a preview release, during which the model was refined by 50 internal product teams. Tencent has reportedly doubled its AI spending to over $5 billion for 2026 to compete with domestic rivals like Alibaba and ByteDance. The model is already integrated into Tencent’s enterprise suite, including the WorkBuddy agent and the Yuanbao chatbot. Competitive pressure in the Chinese open-weight market has intensified throughout 2026. Per Interconnects.ai (June 2026), Zhipu AI released its 744B-parameter GLM-5.2 under an MIT license in mid-June, specifically targeting the 1-million-token context window frontier. While GLM-5.2 maintains a lead in coding-specific tasks, the hardware requirements for such large models often necessitate 8-way H200 nodes, which are difficult to source under current trade limitations. Hardware constraints remain a primary driver of model architecture in the region. Per TrendForce (July 2025), Nvidia significantly revised its China-exclusive H20 lineup to include HBM3e-boosted variants like the H20-3e used for Hy3. These chips are designed to stay under U.S. performance density caps while offering enough memory bandwidth for inference. By sizing Hy3 to fit within the memory limits of eight H20-3e GPUs, Tencent ensures its flagship model can be deployed at scale despite ongoing geopolitical restrictions.
Read full article at venturebeat.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source