Zhipu AI releases SCAIL-2 to industrialize controlled character video generation
Zhipu AI, in collaboration with Tsinghua University, has released SCAIL-2, a generative AI video model that utilizes an end-to-end architecture, abandoning explicit intermediate representations like "stick-figures." This model aims to digitize motion capabilities and reform the digital content production pipeline for film and television, allowing AI to directly interpret visual context for more nuanced motion understanding.
Key Takeaways
- Replaces rigid pose estimators with direct latent space splicing for pixel-level motion transfer.
- Trained on MotionPair-60K, a dataset of 60,000 synthesized motion pairs for high-fidelity control.
- Integrated with the ComfyUI workflow ecosystem to target professional digital content creators.
- Supports zero-shot execution for complex tasks including animal-driven animation and first-person perspectives.
- Utilizes a unified Transformer architecture to reduce inference latency and optimize computing power allocation.
Why It Matters
SCAIL-2 addresses the primary bottleneck in AI video production: the lack of granular, predictable control required for professional film and television. By assetizing performance through reusable visual vectors rather than abstracted skeletons, Zhipu is attempting to shift the production pipeline from manual character rigging to intention-driven creation. For the broader industry, this moves Zhipu from a tool provider to an infrastructure layer offering a standardized production protocol for digital assets. Competitively, it positions Zhipu as a peer to North American giants like Runway by providing open-source access to high-end control mechanisms. Watch for Zhipu's integration into major game engines and traditional VFX software as a signal of industrial adoption.
Additional Context
The release of SCAIL-2 follows a period of massive capital expansion for Zhipu AI. Per Forbes and Bloomberg (June 2026), the company is currently seeking to raise 15 billion yuan ($2.1 billion) through a secondary listing on Shanghai’s STAR Market, just months after its January 2026 debut on the Hong Kong Stock Exchange. This capital influx is specifically earmarked for foundation model R&D, making it one of the largest single research commitments by a Chinese AI firm. Zhipu’s market capitalization surged past $80 billion in early 2026, reflecting investor confidence in its transition from a pure-play large language model developer to a multimodal ecosystem leader. Zhipu’s rapid iteration cycle—which included the launch of the 744-billion parameter GLM-5.1 in April 2024 per Presenc AI—has allowed it to challenge closed-source leaders like OpenAI and Anthropic on performance benchmarks. According to TechNode (May 2026), Zhipu is increasingly optimizing its models for domestic hardware, such as Huawei’s Ascend and Cambricon chips, to insulate its operations from U.S. export controls on advanced GPUs. This hardware-agnostic strategy is coupled with an aggressive open-source policy, including the release of model weights under the MIT License to capture the global developer market. Market analysts note that Zhipu is now considered one of the 'six AI tigers' in China, competing directly with ByteDance’s Seedance and Alibaba’s Qwen series for dominance in agentic and multimodal workflows (per SCMP, June 2026). The focus on controlled video generation through SCAIL-2 aligns with a broader industry shift toward 'AI for production,' where the priority has moved from generating short, creative clips to building reliable toolchains for professional digital asset generation and autonomous software engineering.
Read full article at eu.36kr.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source