StreamingMemeStreamingMemeBuyers Guide
AboutLeaderboardsEventsSubmit News
Subscribe

Daily Brief

The streaming industry in your inbox every morning.

Daily Brief

The streaming industry in your inbox every morning.

StreamingMemeStreamingMeme

The independent buyers guide and news aggregator for the streaming technology industry.

Explore

Buyers GuideLeaderboardsEventsSubmit News

Stay updated

Weekly digest of new companies and streaming news.

Categories

Encoding & SoftwareVideo Delivery & CDNStreaming PlatformsAI for VideoProduction HardwareBusiness NewsMonetization & Ad TechRegulatory & Policy

© 2026 StreamingMeme. All rights reserved.

AboutPrivacy PolicyTermsContact
EncodingCDNPlatformsAI & VideoHardwareBusinessAd TechPolicy
← AI for Video
AI & VideoTechnical DevelopmentJune 8, 2026

Meta's VLM³ Achieves 0.9 Depth Estimation Accuracy, Unifying 3D Vision Tasks

Meta's VLM³ Achieves 0.9 Depth Estimation Accuracy, Unifying 3D Vision Tasks
36kr

Meta and Princeton University have introduced VLM³, a framework that unifies four 3D vision tasks through a standard vision-language model, demonstrating that visual models can naturally learn 3D perception. This research significantly improves depth estimation, pixel matching, and camera pose estimation, outperforming previous specialized models by leveraging unified data organization and training methods.

Key Takeaways

  • VLM³ unifies four 3D vision tasks: 3D object understanding, metric depth estimation, pixel matching, and camera pose calculation.
  • The VLM³-4B model achieved an average depth estimation accuracy (δ₁) of 0.90, improving on DepthLM-7B's 0.84.
  • VLM³ significantly reduced the endpoint error (EPE) for pixel matching and outperformed DKM and RoMa.
  • It raised the AUC₃₀° index for camera pose estimation from 5% to 94%, matching DA3-Giant's performance.
  • The framework utilizes Qwen3-VL-4B as its base and relies on unified data organization, image standardization, and text-based spatial localization.

Why It Matters

This development indicates that standard vision-language models can handle complex 3D perception tasks without specialized architectures or task-specific modules. This has immediate implications for fields such as autonomous driving, robotics, and 3D reconstruction, where accurate spatial inference is critical. The ability to unify diverse 3D tasks under a single model simplifies development and could accelerate advancements in AI for real-world applications. Moving forward, continued progress in data mixing strategies and text-based spatial localization will be a crucial signal to watch for further generalization and performance gains in AI-driven 3D understanding.

Additional Context

VLM³ builds on prior work demonstrating that vision-language models can learn pixel-level depth estimation. As noted by arXiv (May 2026), these models historically struggled with fine-grained 3D tasks, often relying on specialized-expert models. However, Meta's research suggests that standard VLMs can achieve competitive or superior performance across diverse 3D tasks. A related paper from Meta (June 2024) provides an overview of vision-language modeling, emphasizing the challenges and advancements in mapping vision to language and extending VLMs to video. Separately, from CVPR 2024, the 'SpatialVLM' framework also addressed enhancing VLMs' spatial reasoning capabilities by generating internet-scale spatial data. While SpatialVLM focused on improving quantitative and qualitative spatial reasoning and robotics applications, VLM³ specifically aims to prove that standard VLMs are 'native 3D learners' with minimal architectural changes, focusing instead on data organization and input representation, as highlighted on Meta's GitHub (May 2026). This ongoing research underscores a broader industry trend toward simplifying AI models for complex tasks through data and training optimization rather than just increasing model complexity.


Read full article at eu.36kr.com

Get this in your inbox → Subscribe

Enjoy our coverage?

Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.

Add as preferred source

Related Articles

MarkTechPost: Induction Labs Photon-1 trains on 18 years of raw video
Digital Journal: Northwestern’s Spider-Inspired 3D Camera Curbs Machine Vision Power Drain
YouTube: NTT's LLMlet enables distributed LLM inference across browsers via WebRTC

Newest

1 day ago
Barchart: Cerebras and AMD partner on low-latency AI inference architecture
1 day ago
Light Reading: Charter sidesteps Starlink partnership rumors as Q2 broadband losses widen
1 day ago
GuruFocus: Fastly joins Experian to secure autonomous commerce at the edge
1 day ago
The BIG Newsletter: Nexstar and TEGNA Accused of Violating Judicial Order in $6.2 Billion Merger
1 day ago
Front Office Sports: World Cup afternoon ratings spark shift toward earlier U.S. game windows
1 day ago
Cord Cutters News: FCC chair signals scrutiny for potential streaming-exclusive 2030 World Cup rights
1 day ago
Cord Cutters News: Linear contraction accelerates as 14 cable networks vanish in five years
1 day ago
Los Angeles Times: Disney, Netflix, and Amazon recruit AI talent to automate production workflows
1 day ago
Callaba: Callaba standardizes remote production workflows via SRT and NDI integration
1 day ago
Wccftech: Qualcomm Adreno 850 GPU to debut AI Frame Fusion technology
1 day ago
AI Rights Brief: Google and Disney integrate AI provenance directly into programmatic ad workflows
1 day ago
Lib.rs: New zero-dependency Rust decoder vp9dec achieves bit-exact VP9 conformance
1 day ago
Futurism: Meta and TikTok face backlash over deceptive AI-generated health ads
1 day ago
Euronews: EU Expert Panel Backs Age Restrictions and Addictive Feature Bans
1 day ago
Audio Chocolate: Merging Technologies debuts Anubis Premium SPS for mission-critical broadcast audio
1 day ago
Vocal: TeqBlaze challenges Epom with modular full-stack white-label ad tech suite
1 day ago
IPWatchdog: EC mandates Google share search data and Android features under DMA
1 day ago
TechRadar: Weka's new WEKApod 3 uses Micron 245TB SSDs for exabyte-scale storage
1 day ago
Lib.rs: Moq-relay 0.3.1 adds mTLS and admission policies for production-grade QUIC streaming
1 day ago
Yahoo: LG mandates removal of residential proxy SDKs from webOS apps

Upcoming Events

Jul
29–30
Buffer-Free VideoSeattle
Aug
17–20
SET EXPOSao Paulo
Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
View all events →

Top Sources

  1. 1.Sports Video Group104
  2. 2.SiliconANGLE91
  3. 3.YouTube63
  4. 4.Tech Times60
  5. 5.AdExchanger57
  6. 6.TechCrunch55
  7. 7.arXiv50
  8. 8.PPC Land48
Full leaderboards →

Newest

1 day ago
Barchart: Cerebras and AMD partner on low-latency AI inference architecture
1 day ago
Light Reading: Charter sidesteps Starlink partnership rumors as Q2 broadband losses widen
1 day ago
GuruFocus: Fastly joins Experian to secure autonomous commerce at the edge
1 day ago
The BIG Newsletter: Nexstar and TEGNA Accused of Violating Judicial Order in $6.2 Billion Merger
1 day ago
Front Office Sports: World Cup afternoon ratings spark shift toward earlier U.S. game windows
1 day ago
Cord Cutters News: FCC chair signals scrutiny for potential streaming-exclusive 2030 World Cup rights
1 day ago
Cord Cutters News: Linear contraction accelerates as 14 cable networks vanish in five years
1 day ago
Los Angeles Times: Disney, Netflix, and Amazon recruit AI talent to automate production workflows
1 day ago
Callaba: Callaba standardizes remote production workflows via SRT and NDI integration
1 day ago
Wccftech: Qualcomm Adreno 850 GPU to debut AI Frame Fusion technology
1 day ago
AI Rights Brief: Google and Disney integrate AI provenance directly into programmatic ad workflows
1 day ago
Lib.rs: New zero-dependency Rust decoder vp9dec achieves bit-exact VP9 conformance
1 day ago
Futurism: Meta and TikTok face backlash over deceptive AI-generated health ads
1 day ago
Euronews: EU Expert Panel Backs Age Restrictions and Addictive Feature Bans
1 day ago
Audio Chocolate: Merging Technologies debuts Anubis Premium SPS for mission-critical broadcast audio
1 day ago
Vocal: TeqBlaze challenges Epom with modular full-stack white-label ad tech suite
1 day ago
IPWatchdog: EC mandates Google share search data and Android features under DMA
1 day ago
TechRadar: Weka's new WEKApod 3 uses Micron 245TB SSDs for exabyte-scale storage
1 day ago
Lib.rs: Moq-relay 0.3.1 adds mTLS and admission policies for production-grade QUIC streaming
1 day ago
Yahoo: LG mandates removal of residential proxy SDKs from webOS apps

Upcoming Events

Jul
29–30
Buffer-Free VideoSeattle
Aug
17–20
SET EXPOSao Paulo
Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
View all events →

Top Sources

  1. 1.Sports Video Group104
  2. 2.SiliconANGLE91
  3. 3.YouTube63
  4. 4.Tech Times60
  5. 5.AdExchanger57
  6. 6.TechCrunch55
  7. 7.arXiv50
  8. 8.PPC Land48
Full leaderboards →