Induction Labs Photon-1 trains on 18 years of raw video
Induction Labs has released Photon-1, a 106B-parameter mixture-of-experts model that utilizes next-latent-token prediction to learn task policies from raw computer-use video without action labels. The research demonstrates significant compression gains using finite scalar quantization and improved efficiency compared to standard multimodal baselines for task simulation.
Key Takeaways
- Photon-1 utilizes finite scalar quantization (FSQ) to compress video frames into 960 tokens (~2.2 KB), achieving a reported 100x efficiency gain over OCR for state detection.
- The model required 30,000 H200 GPU-hours for pretraining, roughly 27x less compute than the estimated requirements for Google's Gemini 3.1 Flash-Lite.
- Induction Labs sustains a 40% end-to-end Model Flops Utilization (MFU) using custom PyTorch fused kernels and a differential latent encoder.
- Despite training only on desktop video, Photon-1 outperformed LLM baselines in checkers and billiard physics simulation after specific downstream finetuning.
Why It Matters
The removal of the 'action label' requirement solves the primary data bottleneck for training autonomous video agents, enabling models to learn complex logic from passive internet-scale content. By predicting future latent states rather than raw pixels, Photon-1 achieves a 3x reduction in serving costs, making high-parameter digital agents more economically viable for B2B workflow automation. This shifts the technical frontier from multimodal supervised learning to pure self-supervised video world-modeling. Watch for whether Induction Labs transitions from research results to an open-weight release or a commercial API to challenge existing agentic frameworks.
Additional Context
Induction Labs emerged from stealth in 2025 as part of a growing cohort of San Francisco-based startups focused on autonomous 'computer-use' agents. Per PitchBook, August 2025, the firm received early backing from Y Combinator, which reportedly acquired its first dedicated GPU cluster specifically to support the intensive compute requirements of the 'imagination model' architecture. The team, led by 19-year-old researcher Jonathan Li, is betting that passive observation of human input can mirror the pretraining success seen in large language models while avoiding the manual labeling costs that have hampered previous robotics-focused video models. Technically, Photon-1 builds on a surge of research into next-latent-token prediction (NLTP) as a way to inject a 'recurrent inductive bias' into standard transformers. Recent academic work, such as the NextLat framework presented at NeurIPS in late 2025, demonstrated that predicting future latent states helps models form coherent 'belief states' about their environment. This approach allows transformers to plan multi-step actions without the myopic bias often seen in simple next-token prediction. Industry interest in this method has intensified as companies like Google and DeepMind explore Finite Scalar Quantization (FSQ) as a drop-in replacement for older vector quantization (VQ) methods due to its resistance to codebook collapse. Competitive activity in the world-model space is also heating up. In July 2026, the Reactor research team released 'Open Dreamer,' an open-source world-model pipeline for video tokenization and action-conditioned dynamics. Similarly, Black Forest Labs recently announced FLUX 3, which integrates multimodal flow models for robot action prediction. Induction Labs’ claim of beating Gemini 3.1 Flash-Lite on internal benchmarks indicates that specialized, efficient MoE architectures are reaching parity with generalized frontier models in narrow tasks like desktop simulation and physics modeling.
Read full article at marktechpost.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source