MirroS Code-as-World outperforms Gemini-3.1 Flash in physical video reasoning
MirroS has released Code-as-World, an open-source paradigm that converts video into executable MuJoCo physics simulations using an agentic discovery loop. The model, which outperforms Gemini-3.1 Flash on physical reasoning benchmarks, provides a framework for generating verified training data with exact physical labels for video intelligence applications.
Key Takeaways
- The framework uses an agentic discovery loop to recover executable world representations from real footage in up to five rounds.
- MirroS released the 4B and 9B model checkpoints under the Apache 2.0 license, fine-tuned from Qwen3.5.
- Training utilized 73,335 image-space QA pairs and nearly 2,600 executable worlds to provide exact physical labels.
- The system represents scenes through a triple structure of composition, evolution, and appearance to allow for re-simulation.
Why It Matters
The release of MirroS Code-as-World marks a shift from purely generative video models toward systems that understand the underlying physics of a scene. By translating pixels into executable MuJoCo programs, developers can now generate verified training data with precise physical labels like mass and friction, which are absent in raw video. This technical development challenges closed-model dominance, as the 9B variant already exceeds the physical reasoning accuracy of Gemini-3.1 Flash and ChatGPT-5.1. For the streaming and AI ecosystem, this provides a path toward more realistic synthetic content and better automated metadata extraction. Watch for whether MirroS expands this framework beyond rigid-body physics into fluid or soft-body simulations.
Additional Context
MuJoCo, the physics engine at the center of MirroS Code-as-World, has become a cornerstone of robotics and embodied AI research since Google DeepMind open-sourced it in 2022. Google DeepMind released MuJoCo as a free, open-source physics engine in May 2022, positioning it as a foundational tool for simulating rigid-body dynamics in reinforcement learning and robotics research. The engine's adoption has expanded well beyond robotics into video understanding, where researchers use it to generate ground-truth physical parameters for training data. NVIDIA has also integrated MuJoCo into its Isaac platform for synthetic data generation, and the broader ecosystem of physics-aware video models continues to grow as labs seek to bridge the gap between pixel-level understanding and physical reasoning.
The competitive landscape for physical video reasoning has intensified in 2026, with multiple labs pursuing code-based world representations. Google's Gemini-3.1 Flash, which Code-as-World outperforms on the MRA benchmark, was released in early 2026 as a lightweight multimodal model optimized for real-time inference tasks, and its physical reasoning capabilities represent the baseline that newer frameworks now target. Meanwhile, Qwen3.5 from Alibaba has emerged as another contender in open-weight multimodal reasoning, and the vLLM inference framework has become a standard deployment layer for models like these, enabling efficient serving of large language and vision models at scale. The fact that a 9B-parameter open-source model can exceed proprietary systems on physics benchmarks signals a broader trend toward specialized small models outperforming general-purpose large models on narrow technical tasks.
On the technical side, the agentic loop approach used by Code-as-World reflects a wider movement in AI research toward iterative self-correction for code generation. NVIDIA Research published work in 2025 describing how agent frameworks combine code agency with LLM agency to produce verified outputs through multi-step reasoning, a pattern directly relevant to how Code-as-World iteratively refines MuJoCo programs until they match observed video dynamics. The streaming and video intelligence industry stands to benefit from this approach because executable physics representations enable such as object mass, velocity, and collision forces, which traditional computer vision pipelines cannot provide. For applications ranging from sports analytics to autonomous vehicle training data, the ability to convert raw video into physically accurate simulations without manual annotation represents a significant reduction in data preparation costs.
Read full article at marktechpost.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source