StreamingMemeStreamingMemeBuyers Guide
AboutLeaderboardsEventsSubmit News
Subscribe

Daily Brief

The streaming industry in your inbox every morning.

Daily Brief

The streaming industry in your inbox every morning.

StreamingMemeStreamingMeme

The independent buyers guide and news aggregator for the streaming technology industry.

Explore

Buyers GuideLeaderboardsEventsSubmit News

Stay updated

Weekly digest of new companies and streaming news.

Categories

Encoding & SoftwareVideo Delivery & CDNStreaming PlatformsAI for VideoProduction HardwareBusiness NewsMonetization & Ad TechRegulatory & Policy

© 2026 StreamingMeme. All rights reserved.

AboutPrivacy PolicyTermsContact
EncodingCDNPlatformsAI & VideoHardwareBusinessAd TechPolicy
← AI for Video
AI & VideoTechnical DevelopmentJune 26, 2026

Lightweight CAST Adapter Nearly Doubles Video Retrieval Accuracy on Frozen Backbones

OpenReview

Researchers propose CAST (Context-Aware State Transition), a lightweight adapter for frozen vision-language models that models procedural state transitions for consistent video retrieval. The method outperforms zero-shot baselines across several foundation model backbones and provides a reranking signal that improves coherence of video generation outputs from models like Veo. A new Consistent Video Retrieval benchmark is introduced to diagnose state and identity consistency failures beyond standard semantic matching.

Key Takeaways

  • CAST improves InternVideo2 retrieval accuracy from 36.75% to 71.68% on YouCook2 and from 20.61% to 64.36% on CrossTask, operating entirely in the frozen backbone's native embedding space.
  • The adapter transfers across five frozen backbones — CLIP, InternVideo2, VideoPrism, GME-Qwen2-VL-2B, and Qwen3-VL-Embedding-2B — with identity accuracy rising from roughly 30% to 69–78% across all backbones.
  • A new Consistent Video Retrieval (CVR) benchmark introduces state negatives (temporally misaligned clips from the same video) and identity negatives (appearance-misaligned clips from different videos) across YouCook2, COIN, and CrossTask, using a fixed 1-vs-9 ranking protocol.
  • In a blind human study on 300 YouCook2 prompts, CAST-reranked Veo outputs were preferred over standard text matching across overall preference (55.1% vs 38.6%), physical plausibility (50.6% vs 39.9%), and temporal logic (52.5% vs 38.6%).
  • Residual transition modeling (predicting Δ rather than the target directly) improved state accuracy from 38.92% to 51.03% in ablation, confirming that anchoring prediction around the prior clip's embedding is critical for procedural consistency.

Why It Matters

CAST shows that context-aware state transitions can be modeled as a lightweight query-side adapter without re-indexing video galleries or fine-tuning backbone encoders — a practical advantage for platforms managing large content libraries. The CVR benchmark's hard negatives expose a failure mode that standard retrieval metrics like MSR-VTT miss: clips that are semantically relevant but violate procedural causality or identity continuity. The Veo reranking result, though preliminary with 300 prompts, suggests the same transition signal could guide black-box generation pipelines toward more coherent multi-step narratives. Watch whether the CVR benchmark format — state and identity negatives applied to procedural datasets — gets adopted in broader video retrieval evaluations beyond cooking and task instruction domains.

Additional Context

CAST was accepted as a poster at ICML 2026, scheduled for July 9, 2026 in Seoul, per the ICML virtual site. The project page confirms author affiliations with Google, UC Santa Cruz, and MIT, with lead author Yanqing Liu having completed the work as a research intern at Google. The ICML poster abstract notably broadens the generation-guidance claim to mention both Sora and Veo as target black-box models, whereas the paper itself only evaluates Veo. The CVR benchmark arrives amid a wave of new video retrieval evaluation frameworks targeting gaps that legacy benchmarks like MSR-VTT and DiDeMo leave unaddressed. LoVR, introduced on arXiv in May 2025, provides a long-video retrieval benchmark with 467 videos averaging 26 minutes and over 40,000 fine-grained clips, revealing that even strong multimodal embedding models like GME-Qwen2-VL suffer substantial accuracy drops on long-form content versus short clips. MUVR, also from 2025, evaluates untrimmed video retrieval using 53,000 Bilibili videos with multi-modal queries and finds that EVA-CLIP achieves only 58% mAP, with current MLLMs proving unreliable for reranking tasks. FLARE, posted in 2026, introduces a full-modality audiovisual retrieval benchmark showing that audio-language alignment remains a persistent bottleneck — the best contrastive audio model achieves under 1% Recall@1 on clip-level text-to-clip retrieval. On the generation side, temporal consistency in long-form video output remains an active frontier. LoL (Longer than Longer), presented at CVPR 2026, identifies and addresses a failure mode called "sink-collapse" in autoregressive video generation, where RoPE periodicity causes frames to revert abruptly to initial context frames; the method enables streaming generation up to 12 hours with minimal quality degradation. MilliVid, posted on arXiv in 2026, proposes hierarchical latent tokenization to preserve long-range scene geometry while spending less compute on low-saliency detail. A survey published in ACM Computing Surveys in February 2026 (Yin et al.) catalogs spatiotemporal consistency mechanisms across diffusion-based video generation, noting that even state-of-the-art models struggle to maintain character identity and scene layout beyond 16 seconds without specialized architectural components.


Read full article at openreview.net

Get this in your inbox → Subscribe

Enjoy our coverage?

Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.

Add as preferred source

Related Articles

MarkTechPost: Induction Labs Photon-1 trains on 18 years of raw video
BigGo: YouTube Ads engineers detail staged evaluation framework for LLM agents
Digital Journal: Northwestern’s Spider-Inspired 3D Camera Curbs Machine Vision Power Drain

Newest

1 day ago
Barchart: Cerebras and AMD partner on low-latency AI inference architecture
1 day ago
Light Reading: Charter sidesteps Starlink partnership rumors as Q2 broadband losses widen
1 day ago
GuruFocus: Fastly joins Experian to secure autonomous commerce at the edge
1 day ago
The BIG Newsletter: Nexstar and TEGNA Accused of Violating Judicial Order in $6.2 Billion Merger
1 day ago
AI Rights Brief: Google and Disney integrate AI provenance directly into programmatic ad workflows
1 day ago
Front Office Sports: World Cup afternoon ratings spark shift toward earlier U.S. game windows
1 day ago
Futurism: Meta and TikTok face backlash over deceptive AI-generated health ads
1 day ago
Wccftech: Qualcomm Adreno 850 GPU to debut AI Frame Fusion technology
1 day ago
Beet.TV: Brands must re-describe catalogs for AI agents to maintain discoverability
1 day ago
daily.dev: AVIF achieves universal browser support as Edge and Safari close gaps
1 day ago
Los Angeles Times: Disney, Netflix, and Amazon recruit AI talent to automate production workflows
1 day ago
Callaba: Callaba standardizes remote production workflows via SRT and NDI integration
1 day ago
Lib.rs: New zero-dependency Rust decoder vp9dec achieves bit-exact VP9 conformance
1 day ago
IT Brief UK: Fetch.ai and RedSquid TV launch first agentic AI television platform
1 day ago
Vocal: TeqBlaze challenges Epom with modular full-stack white-label ad tech suite
1 day ago
Euronews: EU Expert Panel Backs Age Restrictions and Addictive Feature Bans
1 day ago
Cord Cutters News: FCC chair signals scrutiny for potential streaming-exclusive 2030 World Cup rights
1 day ago
Audio Chocolate: Merging Technologies debuts Anubis Premium SPS for mission-critical broadcast audio
1 day ago
Cord Cutters News: Linear contraction accelerates as 14 cable networks vanish in five years
1 day ago
IPWatchdog: EC mandates Google share search data and Android features under DMA

Upcoming Events

Jul
29–30
Buffer-Free VideoSeattle
Aug
17–20
SET EXPOSao Paulo
Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
View all events →

Top Sources

  1. 1.Sports Video Group104
  2. 2.SiliconANGLE91
  3. 3.YouTube63
  4. 4.Tech Times60
  5. 5.AdExchanger57
  6. 6.TechCrunch55
  7. 7.arXiv50
  8. 8.PPC Land48
Full leaderboards →

Newest

1 day ago
Barchart: Cerebras and AMD partner on low-latency AI inference architecture
1 day ago
Light Reading: Charter sidesteps Starlink partnership rumors as Q2 broadband losses widen
1 day ago
GuruFocus: Fastly joins Experian to secure autonomous commerce at the edge
1 day ago
The BIG Newsletter: Nexstar and TEGNA Accused of Violating Judicial Order in $6.2 Billion Merger
1 day ago
AI Rights Brief: Google and Disney integrate AI provenance directly into programmatic ad workflows
1 day ago
Front Office Sports: World Cup afternoon ratings spark shift toward earlier U.S. game windows
1 day ago
Futurism: Meta and TikTok face backlash over deceptive AI-generated health ads
1 day ago
Wccftech: Qualcomm Adreno 850 GPU to debut AI Frame Fusion technology
1 day ago
Beet.TV: Brands must re-describe catalogs for AI agents to maintain discoverability
1 day ago
daily.dev: AVIF achieves universal browser support as Edge and Safari close gaps
1 day ago
Los Angeles Times: Disney, Netflix, and Amazon recruit AI talent to automate production workflows
1 day ago
Callaba: Callaba standardizes remote production workflows via SRT and NDI integration
1 day ago
Lib.rs: New zero-dependency Rust decoder vp9dec achieves bit-exact VP9 conformance
1 day ago
IT Brief UK: Fetch.ai and RedSquid TV launch first agentic AI television platform
1 day ago
Vocal: TeqBlaze challenges Epom with modular full-stack white-label ad tech suite
1 day ago
Euronews: EU Expert Panel Backs Age Restrictions and Addictive Feature Bans
1 day ago
Cord Cutters News: FCC chair signals scrutiny for potential streaming-exclusive 2030 World Cup rights
1 day ago
Audio Chocolate: Merging Technologies debuts Anubis Premium SPS for mission-critical broadcast audio
1 day ago
Cord Cutters News: Linear contraction accelerates as 14 cable networks vanish in five years
1 day ago
IPWatchdog: EC mandates Google share search data and Android features under DMA

Upcoming Events

Jul
29–30
Buffer-Free VideoSeattle
Aug
17–20
SET EXPOSao Paulo
Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
View all events →

Top Sources

  1. 1.Sports Video Group104
  2. 2.SiliconANGLE91
  3. 3.YouTube63
  4. 4.Tech Times60
  5. 5.AdExchanger57
  6. 6.TechCrunch55
  7. 7.arXiv50
  8. 8.PPC Land48
Full leaderboards →