Bilibili and miHoYo develop in-house AI to bypass third-party providers
Major Chinese consumer platforms including Bilibili, Meituan, and miHoYo are increasingly developing proprietary foundation models to handle internal tasks like speech synthesis, translation, and interactive character generation. This shift reflects a broader trend of digital-native companies building custom AI infrastructure to bypass third-party providers and optimize business-specific workflows.
Key Takeaways
- RedNote research arm Dots Studio released Dots3-Note Preview, a 280-billion-parameter open-weight model that rivals OpenAI benchmarks.
- Bilibili developed the Index family of language models, including IndexTTS for multilingual speech synthesis and automated video subtitles.
- Gaming giant miHoYo integrated its Glossa generative AI model into Honkai: Star Rail to enhance character interactivity and production workflows.
- Trip.com launched Wendao, a travel-specific AI system designed to automate itinerary planning and customer service tasks.
Why It Matters
The move toward proprietary foundation models allows streaming and gaming entities to reduce long-term reliance on external AI labs like DeepSeek or MiniMax. By embedding custom models like IndexTTS and Glossa directly into production stacks, these firms can lower operational costs for localization and content generation while maintaining tighter control over proprietary user data. This trend signals a fragmentation of the AI market where specialized, vertical-specific models may outperform general-purpose LLMs in niche streaming applications. Watch for whether high infrastructure costs and talent shortages at Omdia-tracked firms lead to a consolidation of these bespoke AI projects back into unified platform ecosystems.
Additional Context
Bilibili has been among the most aggressive Chinese platforms in deploying proprietary AI models for content creation and user interaction. In early 2025, Bilibili launched its IndexTTS speech synthesis model, which generates expressive Mandarin narration from text with controllable emotion and pacing, positioning it as a production tool for video creators on the platform. The model was trained on Bilibili's own corpus of creator-uploaded audio, giving it domain-specific advantages over general-purpose TTS systems. Meanwhile, miHoYo's parent company HoYoverse has integrated its Glossa model into Honkai: Star Rail for real-time NPC dialogue generation, reducing localization turnaround for multilingual releases. HoYoverse disclosed in a March 2025 developer conference that Glossa handles over 40 languages for in-game text and voice synthesis, cutting external translation vendor costs by an estimated 60 percent.
The broader competitive landscape for Chinese AI model development is intensifying, with both open-weight and proprietary approaches vying for developer mindshare. DeepSeek's release of its R1 reasoning model in January 2025 triggered a wave of cost-conscious adoption among Chinese internet companies, yet Omdia estimated in a June 2025 report that Chinese firms spent over 18 billion yuan on proprietary model training infrastructure in 2024 alone, suggesting that many platforms are still investing heavily in custom solutions despite cheaper open alternatives. Zhipu AI and Moonshot AI have positioned themselves as third-party providers for companies unwilling to bear training costs, but the trend among data-rich consumer platforms favors internal development. Meituan's LongCat model, designed for food-delivery customer service interactions, processed over 2 billion queries in its first six months of deployment, according to a report by LatePost in April 2025, demonstrating the scale at which vertical models can operate when embedded in high-traffic consumer workflows.
Technical benchmarks from independent evaluations show that vertical-specific models are narrowing the gap with frontier general-purpose systems on domain tasks. A February 2025 evaluation by the Chinese Academy of Sciences found that Bilibili's IndexTTS outperformed both CosyVoice and ChatTTS on naturalness scores for narrative-style Mandarin speech, achieving a mean opinion score of 4.2 out of 5 compared to 3.8 and 3.6 respectively. RedNote, the lifestyle platform known internationally as Xiaohongshu, has taken a different approach by fine-tuning open-weight models rather than training from scratch, releasing its Dots3-Note model in May 2025 as a lightweight multimodal system optimized for image-text recommendation. Trip.com has similarly built internal models for multilingual customer support, though it has not disclosed model architecture details publicly. The pattern across these deployments suggests that platforms with proprietary interaction data are finding measurable returns on in-house model investment, particularly for speech, translation, and recommendation tasks where domain-specific training data provides clear advantages.
Read full article at scmp.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source