Lightweight CAST Adapter Nearly Doubles Video Retrieval Accuracy on Frozen Backbones
Researchers propose CAST (Context-Aware State Transition), a lightweight adapter for frozen vision-language models that models procedural state transitions for consistent video retrieval. The method outperforms zero-shot baselines across several foundation model backbones and provides a reranking signal that improves coherence of video generation outputs from models like Veo. A new Consistent Video Retrieval benchmark is introduced to diagnose state and identity consistency failures beyond standard semantic matching.
Key Takeaways
- CAST improves InternVideo2 retrieval accuracy from 36.75% to 71.68% on YouCook2 and from 20.61% to 64.36% on CrossTask, operating entirely in the frozen backbone's native embedding space.
- The adapter transfers across five frozen backbones — CLIP, InternVideo2, VideoPrism, GME-Qwen2-VL-2B, and Qwen3-VL-Embedding-2B — with identity accuracy rising from roughly 30% to 69–78% across all backbones.
- A new Consistent Video Retrieval (CVR) benchmark introduces state negatives (temporally misaligned clips from the same video) and identity negatives (appearance-misaligned clips from different videos) across YouCook2, COIN, and CrossTask, using a fixed 1-vs-9 ranking protocol.
- In a blind human study on 300 YouCook2 prompts, CAST-reranked Veo outputs were preferred over standard text matching across overall preference (55.1% vs 38.6%), physical plausibility (50.6% vs 39.9%), and temporal logic (52.5% vs 38.6%).
- Residual transition modeling (predicting Δ rather than the target directly) improved state accuracy from 38.92% to 51.03% in ablation, confirming that anchoring prediction around the prior clip's embedding is critical for procedural consistency.
Why It Matters
CAST shows that context-aware state transitions can be modeled as a lightweight query-side adapter without re-indexing video galleries or fine-tuning backbone encoders — a practical advantage for platforms managing large content libraries. The CVR benchmark's hard negatives expose a failure mode that standard retrieval metrics like MSR-VTT miss: clips that are semantically relevant but violate procedural causality or identity continuity. The Veo reranking result, though preliminary with 300 prompts, suggests the same transition signal could guide black-box generation pipelines toward more coherent multi-step narratives. Watch whether the CVR benchmark format — state and identity negatives applied to procedural datasets — gets adopted in broader video retrieval evaluations beyond cooking and task instruction domains.
Additional Context
CAST was accepted as a poster at ICML 2026, scheduled for July 9, 2026 in Seoul, per the ICML virtual site. The project page confirms author affiliations with Google, UC Santa Cruz, and MIT, with lead author Yanqing Liu having completed the work as a research intern at Google. The ICML poster abstract notably broadens the generation-guidance claim to mention both Sora and Veo as target black-box models, whereas the paper itself only evaluates Veo. The CVR benchmark arrives amid a wave of new video retrieval evaluation frameworks targeting gaps that legacy benchmarks like MSR-VTT and DiDeMo leave unaddressed. LoVR, introduced on arXiv in May 2025, provides a long-video retrieval benchmark with 467 videos averaging 26 minutes and over 40,000 fine-grained clips, revealing that even strong multimodal embedding models like GME-Qwen2-VL suffer substantial accuracy drops on long-form content versus short clips. MUVR, also from 2025, evaluates untrimmed video retrieval using 53,000 Bilibili videos with multi-modal queries and finds that EVA-CLIP achieves only 58% mAP, with current MLLMs proving unreliable for reranking tasks. FLARE, posted in 2026, introduces a full-modality audiovisual retrieval benchmark showing that audio-language alignment remains a persistent bottleneck — the best contrastive audio model achieves under 1% Recall@1 on clip-level text-to-clip retrieval. On the generation side, temporal consistency in long-form video output remains an active frontier. LoL (Longer than Longer), presented at CVPR 2026, identifies and addresses a failure mode called "sink-collapse" in autoregressive video generation, where RoPE periodicity causes frames to revert abruptly to initial context frames; the method enables streaming generation up to 12 hours with minimal quality degradation. MilliVid, posted on arXiv in 2026, proposes hierarchical latent tokenization to preserve long-range scene geometry while spending less compute on low-saliency detail. A survey published in ACM Computing Surveys in February 2026 (Yin et al.) catalogs spatiotemporal consistency mechanisms across diffusion-based video generation, noting that even state-of-the-art models struggle to maintain character identity and scene layout beyond 16 seconds without specialized architectural components.
Read full article at openreview.net
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source