Bitmovin and AAU Klagenfurt unveil AI-native streaming and XR frameworks
Researchers from Alpen-Adria-Universität Klagenfurt (AAU) and Bitmovin presented several innovations for AI-native media at ACM Multimedia 2026. The research includes the TIGAS remote rendering framework for XR, the LMM-10K multimodal dataset, and LumaID for illumination-aware video head editing.
Key Takeaways
- TIGAS framework offloads 3D Gaussian Splatting rasterization to backends, achieving sub-10ms rendering and 0.88 SSIM over HTTP/3.
- LumaID uses an Omni-Disentangled Diffusion Transformer to isolate identity from environmental lighting, preventing identity leakage in video editing.
- LMM-10K dataset provides 10,000 high-fidelity 4K/60fps sequences with multimodal annotations for training vision-language models.
- Consist-GRPO reinforcement learning mechanism steers generative video processes using multi-dimensional reward signals for pose and lighting alignment.
Why It Matters
This research signals a pivot from pixel-focused codecs to semantic and systems-layer optimization to solve the 'three pressures' of efficiency, low-latency live, and AI-native media. By offloading complex 3D rendering to the backend via TIGAS and providing structured datasets like LMM-10K, Bitmovin and AAU are building the infrastructure required for hardware-agnostic immersive media. For the broader ecosystem, these tools bridge the gap between high-fidelity generative AI and the bandwidth constraints of mobile distribution. Watch for these techniques to integrate into the MPEG Systems roadmap and Bitmovin’s commercial encoding stack to reduce 'identity drift' in AI-generated content.
Additional Context
The ATHENA lab, a joint venture between Bitmovin and AAU Klagenfurt established in 2018, has a documented history of transitioning academic research into commercial B2B products. Per Bitmovin reporting in August 2024, the lab had already filed 16 invention disclosures and 13 patent applications, including granted patents for per-title encoding and AI-driven rate control. These foundational technologies have been deployed at scale, notably supporting events like Super Bowl LIX, which drew over 15.5 million concurrent viewers on Tubi using Bitmovin's player and telemetry stack in May 2025. The shift toward 'thin-client' remote rendering seen in the TIGAS framework mirrors a wider industry trend in 2026 to bypass the computational limits of standalone XR headsets. Per NVIDIA in March 2026, the launch of CloudXR.js enabled similar GPU-rendered VR/AR streaming via standard web browsers using WebRTC. This movement coincides with the rapid maturation of WebGPU, which became an official W3C Baseline in January 2026. This standard provides near-native rendering performance across Chrome, Safari, and Edge, effectively ending the 15-year dominance of WebGL for 3D web content (per VR.org, May 2026). Furthermore, the release of curated datasets like LMM-10K addresses critical bottlenecks in multimodal video intelligence. As of April 2026, industry analysts noted a 40% rise in WebXR adoption as developers sought to avoid the 'walled gardens' and heavy installation friction of native app stores (per TechBullion, January 2026). By combining high-fidelity 4K source material with LLM-generated semantic descriptors, the LMM-10K dataset facilitates the training of long-context vision models that can process hour-long footage at lower token costs, a task that previously demanded prohibitive GPU resources in earlier generative cycles.
Read full article at athena.itec.aau.at
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source