Chinese open-weight AI models capture nearly 30% of production token volume
Data from Vercel's AI Gateway indicates that open-weight AI models, led by Chinese developers, accounted for 29% of production token volume in June 2026. While expenditures remain dominated by US frontier model providers, streaming and media enterprises are increasingly adopting lower-cost models for specific video generation and back-office tasks.
Key Takeaways
- Open-weight models processed 29% of Vercel AI Gateway tokens in June but represented less than 4% of total spending.
- DeepSeek became the third-largest provider by volume at 22.6%, trailing only Anthropic and Google.
- ByteDance’s Seedance accounted for nearly 50% of video-related spending despite generating only one-third of total videos.
- Leading US labs retained 95% of revenue, with Anthropic alone capturing 61% of total spend on 32% of tokens.
- Frontier model pricing rose roughly 12% in June, offsetting the shift toward lower-cost open-weight alternatives.
Why It Matters
The streaming industry is shifting toward a tiered AI strategy where commodity tasks migrate to low-cost open-weight models while high-stakes applications remain on premium frontier systems. For video engineering teams, this means the 'dual-stack' approach is becoming standard: using proprietary models like Anthropic's Claude for complex back-office logic and orchestration, while leveraging Chinese models like DeepSeek for high-volume inference at one-tenth the cost. The dominance of ByteDance in video spend suggests enterprises are willing to pay a premium for specific architectural advantages in media generation. As the performance gap narrows, expect US labs to face increasing pressure to justify their high token premiums through specialized features rather than general intelligence.
Additional Context
The surge in adoption reflects a broader market trend where Chinese AI labs are achieving near-frontier performance despite ongoing hardware constraints. Per the South China Morning Post (July 2026), DeepSeek’s V4 Flash has emerged as a high-volume leader by focusing on efficient inference, utilizing domestic hardware like Huawei’s Ascend chips to sidestep U.S. export controls. Analysts at Goldman Sachs noted in July 2026 that Chinese models have reached a 'critical point' of intelligence, allowing global enterprises—including major financial institutions—to host models like Alibaba’s Qwen series locally to manage spiraling API costs. In the video sector, competition has intensified as vendors move beyond simple visual output toward ready-to-use commercial assets. Per CNET (July 2026), ByteDance recently unveiled Seedance 2.5, which supports 30-second native single-shot video and 4K resolution, moving the benchmark past the 15-second ceiling maintained by rivals for much of the previous year. This release coincided with local-editing features that allow creators to modify specific frames without full regeneration, a critical requirement for ad-supported streaming platforms seeking to automate localized promotional content. Simultaneously, the cost of being 'wrong' in AI remains the primary revenue defensive for US labs. Vercel data highlights that back-office agents remain the most expensive workload per token, as enterprises prioritize the reliability of Anthropic and OpenAI for high-stakes automation. However, per the Financial Times (July 2026), even this moat is thinning; some startups have reportedly migrated entire production stacks from Claude to DeepSeek, citing operational cost reductions of up to 90% while maintaining acceptable accuracy for non-core business logic.
Read full article at computing.co.uk
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source