JD.com's AdaCodec cuts video MLLM tokens 84.7%, boosts long-video performance
Researchers from JD.com and Shanghai Jiao Tong University developed AdaCodec, a predictive visual code for video Multimodal Large Language Models (MLLMs). This new approach encodes inter-frame changes as compact P-tokens, reducing visual token usage by 84.7% and significantly lowering inference latency. AdaCodec demonstrates improved performance across various long-video and general video-understanding benchmarks compared to traditional per-frame RGB encoding.
Key Takeaways
- AdaCodec reduces visual token usage by 84.7% by encoding inter-frame changes as compact P-tokens.
- The new code significantly lowers inference latency, cutting time-to-first-token from 9.26s to 1.62s on general video-understanding benchmarks.
- AdaCodec with 32k visual tokens surpassed a 224k baseline on all long-video benchmarks.
- It outperforms traditional per-frame RGB encoding across eleven benchmarks, including long-video and temporal assessments.
Why It Matters
The efficiency gains from AdaCodec represent a notable step forward for video MLLMs, addressing the core challenge of processing increasing volumes of video data more effectively. By drastically reducing token count and latency, models can tackle longer video content and complex temporal reasoning tasks with greater speed and accuracy. This development could accelerate the integration of advanced AI capabilities into video platforms, impacting areas from automated content moderation to enhanced viewer engagement analytics. What to watch next is how quickly these predictive coding principles are adopted by other major AI research labs and integrated into commercial video MLLM offerings.
Read full article at arxiv.org
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source