StreamingMemeStreamingMemeBuyers Guide
AboutLeaderboardsEventsSubmit News
Subscribe

Daily Brief

The streaming industry in your inbox every morning.

Daily Brief

The streaming industry in your inbox every morning.

StreamingMemeStreamingMeme

The independent buyers guide and news aggregator for the streaming technology industry.

Explore

Buyers GuideLeaderboardsEventsSubmit News

Stay updated

Weekly digest of new companies and streaming news.

Categories

Encoding & SoftwareVideo Delivery & CDNStreaming PlatformsAI for VideoProduction HardwareBusiness NewsMonetization & Ad TechRegulatory & Policy

© 2026 StreamingMeme. All rights reserved.

AboutPrivacy PolicyTermsContact
EncodingCDNPlatformsAI & VideoHardwareBusinessAd TechPolicy
← AI for Video
AI & VideoTechnical Development

NVIDIA’s VideoITG targets better frame selection for Video-LLMs

NVIDIA’s VideoITG targets better frame selection for Video-LLMs
Github

NVIDIA Research has introduced VideoITG, a new framework using "Instructed Temporal Grounding" to improve video understanding for Video Large Language Models (Video-LLMs). The framework includes the VidThinker pipeline for automated annotation, creating the VideoITG-40K dataset with 40K videos and 500K temporal grounding annotations to enhance frame sampling strategies based on user instructions.

Key Takeaways

  • VideoITG is designed to adapt frame sampling to user instructions rather than rely only on redundancy reduction or unsupervised event localization.
  • VidThinker automates annotation through instruction-conditioned captioning, relevant clip retrieval, and fine-grained frame localization.
  • VideoITG-40K includes 40,000 videos and 500,000 temporal grounding annotations.
  • NVIDIA says the plug-and-play VideoITG model uses Video-LLMs’ visual-language alignment and reasoning for discriminative frame selection.
  • The company says VideoITG improves results on multiple multimodal video understanding benchmarks.

Why It Matters

VideoITG addresses a specific bottleneck for Video-LLMs: picking the most informative frames from long videos when instructions and temporal cues matter. The broader point is practical, not abstract — NVIDIA is packaging a dataset, annotation pipeline, and model design around that problem, with VidThinker doing the labeling work at scale. For teams building video understanding systems, the relevant signal is whether this approach holds up across more Video-LLMs and model sizes. NVIDIA already says it extends to multiple benchmarks, so the next thing to watch is how VideoITG performs on other long-video tasks and sampling setups.


Read full article at nvlabs.github.io

Get this in your inbox → Subscribe

Enjoy our coverage?

Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.

Add as preferred source

Related Articles

Qiang Zhang: DeltaToken cuts video tokens from 180K to under 1,000
ayushchat: Whisper runs locally on Apple Silicon with no network access
Medium: Computer vision workflows optimize American football video annotation using automated propagation
Tech Times: Microsoft Mirage cuts AI video memory use 55x via latent caching

Newest

about 10 hours ago
AOL: UK Government considers complete Freeview switch-off between 2034 and 2044
about 10 hours ago
Broadcast: Games of the Future 2026 secures global streaming and broadcast distribution
about 10 hours ago
Investing.com: Alphabet upgraded as Google Cloud revenue surges 82% on AI demand
about 10 hours ago
The Desk: Phynd launches ad-supported cloud gaming beta on LG webOS
about 11 hours ago
Kalkine Media: Adveritas hits A$16.3M recurring revenue, shifts toward cash flow breakeven
about 16 hours ago
Hyper.ai: Google DeepMind and UC Riverside launch framework to trace synthetic video
about 16 hours ago
Exame: Brazil launches TV 3.0 with 4K VVC and interactive IP layers
about 16 hours ago
Digiday: IAB Redefining Media Types Standard targets automated video ad transparency
about 16 hours ago
VentureBeat: Moonshot AI releases Kimi K3 weights with $20M revenue licensing threshold
about 16 hours ago
AdExchanger: Streaming ad tech consolidation turns independent platforms into proprietary gardens
about 16 hours ago
The Fast Mode: AMD and South Korea Partner to Build Heterogeneous Sovereign AI Infrastructure
about 16 hours ago
MediaPost: Microsoft launches Project Perception to defend programmatic supply chains from AI-driven fraud
about 16 hours ago
AdExchanger: Google mandates biometric passkeys for Ads API as AI costs reshape agency deals
about 16 hours ago
Advanced Television: Roku and Fire TV solidify gatekeeper status as OS influence grows
1 day ago
Hackernoon: Production voice pipeline solves African language latency and hallucination problems
1 day ago
Startup Fortune: Higgsfield AI integrates third-party models as revenue run rate hits $300M
1 day ago
Design & Reuse: Stricter ETSI secure boot standards mandate hardware-level chain of trust
1 day ago
UK Parliament: UK Parliament launches investigation into Ofcom's Online Safety Act enforcement
1 day ago
SVG Europe: WBD streams 600 hours of Glasgow 2026 via remote-first infrastructure
1 day ago
TradingView: Amazon settles FTC Prime suit for $2.5B amid AI-focused redesign

Upcoming Events

Jul
29–30
Buffer-Free VideoSeattle
Aug
17–20
SET EXPOSao Paulo
Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
View all events →

Top Sources

  1. 1.Sports Video Group105
  2. 2.SiliconANGLE93
  3. 3.AdExchanger66
  4. 4.Tech Times65
  5. 5.YouTube62
  6. 6.TechCrunch56
  7. 7.PPC Land51
  8. 8.arXiv50
Full leaderboards →

Newest

about 10 hours ago
AOL: UK Government considers complete Freeview switch-off between 2034 and 2044
about 10 hours ago
Broadcast: Games of the Future 2026 secures global streaming and broadcast distribution
about 10 hours ago
Investing.com: Alphabet upgraded as Google Cloud revenue surges 82% on AI demand
about 10 hours ago
The Desk: Phynd launches ad-supported cloud gaming beta on LG webOS
about 11 hours ago
Kalkine Media: Adveritas hits A$16.3M recurring revenue, shifts toward cash flow breakeven
about 16 hours ago
Hyper.ai: Google DeepMind and UC Riverside launch framework to trace synthetic video
about 16 hours ago
Exame: Brazil launches TV 3.0 with 4K VVC and interactive IP layers
about 16 hours ago
Digiday: IAB Redefining Media Types Standard targets automated video ad transparency
about 16 hours ago
VentureBeat: Moonshot AI releases Kimi K3 weights with $20M revenue licensing threshold
about 16 hours ago
AdExchanger: Streaming ad tech consolidation turns independent platforms into proprietary gardens
about 16 hours ago
The Fast Mode: AMD and South Korea Partner to Build Heterogeneous Sovereign AI Infrastructure
about 16 hours ago
MediaPost: Microsoft launches Project Perception to defend programmatic supply chains from AI-driven fraud
about 16 hours ago
AdExchanger: Google mandates biometric passkeys for Ads API as AI costs reshape agency deals
about 16 hours ago
Advanced Television: Roku and Fire TV solidify gatekeeper status as OS influence grows
1 day ago
Hackernoon: Production voice pipeline solves African language latency and hallucination problems
1 day ago
Startup Fortune: Higgsfield AI integrates third-party models as revenue run rate hits $300M
1 day ago
Design & Reuse: Stricter ETSI secure boot standards mandate hardware-level chain of trust
1 day ago
UK Parliament: UK Parliament launches investigation into Ofcom's Online Safety Act enforcement
1 day ago
SVG Europe: WBD streams 600 hours of Glasgow 2026 via remote-first infrastructure
1 day ago
TradingView: Amazon settles FTC Prime suit for $2.5B amid AI-focused redesign

Upcoming Events

Jul
29–30
Buffer-Free VideoSeattle
Aug
17–20
SET EXPOSao Paulo
Sep
11–14
IBCAmsterdam
Sep
13
SportsPro Streamtime Sports LiveAmsterdam
Sep
16–18
RTC.ONKrakow
View all events →

Top Sources

  1. 1.Sports Video Group105
  2. 2.SiliconANGLE93
  3. 3.AdExchanger66
  4. 4.Tech Times65
  5. 5.YouTube62
  6. 6.TechCrunch56
  7. 7.PPC Land51
  8. 8.arXiv50
Full leaderboards →