V-Nova promotes hierarchical data structures to close AI's data-to-tensor gap
V-Nova is advocating for hierarchical, compute-aware data formats to significantly improve the efficiency of Vision AI processing. These formats, designed to be AI-native, aim to reduce compute power, energy consumption, and latency by allowing AI to access visual data with purpose, rather than processing everything.
Key Takeaways
- Up to 70% of GPU time in visual AI is currently spent waiting for data rather than processing it.
- Data movement and preparation account for 50% of visual AI latency and 80% of total processing time.
- V-Nova's hierarchical format enables 'selective decoding,' allowing AI to process low-resolution context before querying high-detail regions only when necessary.
- The format is designed for massive parallelism on GPUs and tensor cores, supporting 'AI-native' real-time perception for multimodal world models.
Why It Matters
The streaming industry is pivoting from video for human consumption to video for machine analysis, yet legacy codecs like H.264 remain a fundamental bottleneck for AI efficiency. V-Nova's push for compute-aware formats addresses the critical 'data-to-tensor gap,' where expensive hardware sits underutilized due to unstructured data pipelines. For operators, this offers a path to scale computer vision and real-time metadata generation without a linear increase in cloud compute costs or energy footprints. Watch for the integration of these hierarchical structures into the upcoming ISO/IEC 23888 (MPEG-AI) standards for Video Coding for Machines.
Additional Context
The push for AI-optimized data formats coincides with a massive shift in streaming engineering. According to Netint’s 2026 State of Video Encoding Report, 70% of industry professionals plan to expand AI capabilities within their encoding workflows this year, marking a transition from experimental R&D to core infrastructure migration. This move is driven by the need to manage the intelligence required for tasks like automated scene segmentation and 4K upscaling, which previously relied on external metadata or secondary compute passes. In April 2026, V-Nova demonstrated this evolution at NAB, showing how multiple AI inference services could leverage a single hierarchical stream to reduce I/O by up to 20x. Simultaneously, the competitive landscape is seeing the rise of end-to-end neural codecs. Per reports from the Streaming Learning Center in May 2026, neural video codecs (NVCs) are beginning to challenge traditional block-based transforms by using trained neural networks for motion prediction and reconstruction. While these neural formats offer superior quality-per-bit, they often require extreme compute resources for decoding. V-Nova’s approach, centering on the SMPTE VC-6 (ST-2117) standard, seeks a middle ground by maintaining compatibility with existing GPU acceleration while redesigning the data hierarchy for selective retrieval. Recent benchmarks released by NVIDIA in April 2026 further validate this strategy, showing that CUDA-accelerated VC-6 implementations achieved up to 85% lower per-image decode times in large-batch inference workloads. By shifting to a 'levels of quality' (LoQ) structure, developers are beginning to report sub-millisecond decode times for 4K context frames. This speed is becoming essential as the global market for streamed content is projected by CDNetworks to exceed $670 billion by the end of 2026, with an increasing percentage of that traffic being analyzed by automated moderation and personalization engines.
Read full article at v-nova.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source