StreamingMemeStreamingMemeBuyers Guide
AboutLeaderboardsEventsSubmit News
Subscribe

Daily Brief

The streaming industry in your inbox every morning.

Daily Brief

The streaming industry in your inbox every morning.

StreamingMemeStreamingMeme

The independent buyers guide and news aggregator for the streaming technology industry.

Explore

Buyers GuideLeaderboardsEventsSubmit News

Stay updated

Weekly digest of new companies and streaming news.

Categories

Encoding & SoftwareVideo Delivery & CDNStreaming PlatformsAI for VideoProduction HardwareBusiness NewsMonetization & Ad TechRegulatory & Policy

© 2026 StreamingMeme. All rights reserved.

AboutPrivacy PolicyTermsContact
EncodingCDNPlatformsAI & VideoHardwareBusinessAd TechPolicy
← AI for Video
AI & VideoTechnical DevelopmentJuly 26, 2026

Induction Labs Photon-1 trains on 18 years of raw video

Induction Labs Photon-1 trains on 18 years of raw video
MarkTechPost

Induction Labs has released Photon-1, a 106B-parameter mixture-of-experts model that utilizes next-latent-token prediction to learn task policies from raw computer-use video without action labels. The research demonstrates significant compression gains using finite scalar quantization and improved efficiency compared to standard multimodal baselines for task simulation.

Key Takeaways

  • Photon-1 utilizes finite scalar quantization (FSQ) to compress video frames into 960 tokens (~2.2 KB), achieving a reported 100x efficiency gain over OCR for state detection.
  • The model required 30,000 H200 GPU-hours for pretraining, roughly 27x less compute than the estimated requirements for Google's Gemini 3.1 Flash-Lite.
  • Induction Labs sustains a 40% end-to-end Model Flops Utilization (MFU) using custom PyTorch fused kernels and a differential latent encoder.
  • Despite training only on desktop video, Photon-1 outperformed LLM baselines in checkers and billiard physics simulation after specific downstream finetuning.

Why It Matters

The removal of the 'action label' requirement solves the primary data bottleneck for training autonomous video agents, enabling models to learn complex logic from passive internet-scale content. By predicting future latent states rather than raw pixels, Photon-1 achieves a 3x reduction in serving costs, making high-parameter digital agents more economically viable for B2B workflow automation. This shifts the technical frontier from multimodal supervised learning to pure self-supervised video world-modeling. Watch for whether Induction Labs transitions from research results to an open-weight release or a commercial API to challenge existing agentic frameworks.

Additional Context

Induction Labs emerged from stealth in 2025 as part of a growing cohort of San Francisco-based startups focused on autonomous 'computer-use' agents. Per PitchBook, August 2025, the firm received early backing from Y Combinator, which reportedly acquired its first dedicated GPU cluster specifically to support the intensive compute requirements of the 'imagination model' architecture. The team, led by 19-year-old researcher Jonathan Li, is betting that passive observation of human input can mirror the pretraining success seen in large language models while avoiding the manual labeling costs that have hampered previous robotics-focused video models. Technically, Photon-1 builds on a surge of research into next-latent-token prediction (NLTP) as a way to inject a 'recurrent inductive bias' into standard transformers. Recent academic work, such as the NextLat framework presented at NeurIPS in late 2025, demonstrated that predicting future latent states helps models form coherent 'belief states' about their environment. This approach allows transformers to plan multi-step actions without the myopic bias often seen in simple next-token prediction. Industry interest in this method has intensified as companies like Google and DeepMind explore Finite Scalar Quantization (FSQ) as a drop-in replacement for older vector quantization (VQ) methods due to its resistance to codebook collapse. Competitive activity in the world-model space is also heating up. In July 2026, the Reactor research team released 'Open Dreamer,' an open-source world-model pipeline for video tokenization and action-conditioned dynamics. Similarly, Black Forest Labs recently announced FLUX 3, which integrates multimodal flow models for robot action prediction. Induction Labs’ claim of beating Gemini 3.1 Flash-Lite on internal benchmarks indicates that specialized, efficient MoE architectures are reaching parity with generalized frontier models in narrow tasks like desktop simulation and physics modeling.


Read full article at marktechpost.com

Get this in your inbox → Subscribe

Enjoy our coverage?

Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.

Add as preferred source

Related Articles

MarkTechPost: Reactor releases 1.6B parameter open-source Dreamer 4 world-model implementation
Digital Journal: Northwestern’s Spider-Inspired 3D Camera Curbs Machine Vision Power Drain
BigGo: YouTube Ads engineers detail staged evaluation framework for LLM agents

Newest

1 day ago
Barchart: Cerebras and AMD partner on low-latency AI inference architecture
1 day ago
Light Reading: Charter sidesteps Starlink partnership rumors as Q2 broadband losses widen
1 day ago
GuruFocus: Fastly joins Experian to secure autonomous commerce at the edge
1 day ago
The BIG Newsletter: Nexstar and TEGNA Accused of Violating Judicial Order in $6.2 Billion Merger
1 day ago
AI Rights Brief: Google and Disney integrate AI provenance directly into programmatic ad workflows
1 day ago
Front Office Sports: World Cup afternoon ratings spark shift toward earlier U.S. game windows
1 day ago
Futurism: Meta and TikTok face backlash over deceptive AI-generated health ads
1 day ago
Wccftech: Qualcomm Adreno 850 GPU to debut AI Frame Fusion technology
1 day ago
Beet.TV: Brands must re-describe catalogs for AI agents to maintain discoverability
1 day ago
daily.dev: AVIF achieves universal browser support as Edge and Safari close gaps
1 day ago
Los Angeles Times: Disney, Netflix, and Amazon recruit AI talent to automate production workflows
1 day ago
Callaba: Callaba standardizes remote production workflows via SRT and NDI integration
1 day ago
Lib.rs: New zero-dependency Rust decoder vp9dec achieves bit-exact VP9 conformance
1 day ago
IT Brief UK: Fetch.ai and RedSquid TV launch first agentic AI television platform
1 day ago
Vocal: TeqBlaze challenges Epom with modular full-stack white-label ad tech suite
1 day ago
Euronews: EU Expert Panel Backs Age Restrictions and Addictive Feature Bans
1 day ago
Cord Cutters News: FCC chair signals scrutiny for potential streaming-exclusive 2030 World Cup rights
1 day ago
Audio Chocolate: Merging Technologies debuts Anubis Premium SPS for mission-critical broadcast audio
1 day ago
Cord Cutters News: Linear contraction accelerates as 14 cable networks vanish in five years
1 day ago
IPWatchdog: EC mandates Google share search data and Android features under DMA

Upcoming Events

Jul
29–30
Buffer-Free VideoSeattle
Aug
17–20
SET EXPOSao Paulo
Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
View all events →

Top Sources

  1. 1.Sports Video Group104
  2. 2.SiliconANGLE91
  3. 3.YouTube63
  4. 4.Tech Times60
  5. 5.AdExchanger57
  6. 6.TechCrunch55
  7. 7.arXiv50
  8. 8.PPC Land48
Full leaderboards →

Newest

1 day ago
Barchart: Cerebras and AMD partner on low-latency AI inference architecture
1 day ago
Light Reading: Charter sidesteps Starlink partnership rumors as Q2 broadband losses widen
1 day ago
GuruFocus: Fastly joins Experian to secure autonomous commerce at the edge
1 day ago
The BIG Newsletter: Nexstar and TEGNA Accused of Violating Judicial Order in $6.2 Billion Merger
1 day ago
AI Rights Brief: Google and Disney integrate AI provenance directly into programmatic ad workflows
1 day ago
Front Office Sports: World Cup afternoon ratings spark shift toward earlier U.S. game windows
1 day ago
Futurism: Meta and TikTok face backlash over deceptive AI-generated health ads
1 day ago
Wccftech: Qualcomm Adreno 850 GPU to debut AI Frame Fusion technology
1 day ago
Beet.TV: Brands must re-describe catalogs for AI agents to maintain discoverability
1 day ago
daily.dev: AVIF achieves universal browser support as Edge and Safari close gaps
1 day ago
Los Angeles Times: Disney, Netflix, and Amazon recruit AI talent to automate production workflows
1 day ago
Callaba: Callaba standardizes remote production workflows via SRT and NDI integration
1 day ago
Lib.rs: New zero-dependency Rust decoder vp9dec achieves bit-exact VP9 conformance
1 day ago
IT Brief UK: Fetch.ai and RedSquid TV launch first agentic AI television platform
1 day ago
Vocal: TeqBlaze challenges Epom with modular full-stack white-label ad tech suite
1 day ago
Euronews: EU Expert Panel Backs Age Restrictions and Addictive Feature Bans
1 day ago
Cord Cutters News: FCC chair signals scrutiny for potential streaming-exclusive 2030 World Cup rights
1 day ago
Audio Chocolate: Merging Technologies debuts Anubis Premium SPS for mission-critical broadcast audio
1 day ago
Cord Cutters News: Linear contraction accelerates as 14 cable networks vanish in five years
1 day ago
IPWatchdog: EC mandates Google share search data and Android features under DMA

Upcoming Events

Jul
29–30
Buffer-Free VideoSeattle
Aug
17–20
SET EXPOSao Paulo
Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
View all events →

Top Sources

  1. 1.Sports Video Group104
  2. 2.SiliconANGLE91
  3. 3.YouTube63
  4. 4.Tech Times60
  5. 5.AdExchanger57
  6. 6.TechCrunch55
  7. 7.arXiv50
  8. 8.PPC Land48
Full leaderboards →