Google launches LiteRT with on-device agentic skills and hardware acceleration
Google presented its new LiteRT LLM APIs and Agent Skills format, part of the Google AI Edge stack designed to facilitate on-device generative AI deployment. The release enables developers to run optimized models across mobile, web, and IoT platforms to address latency, privacy, and offline functionality requirements.
Key Takeaways
- LiteRT-LM APIs provide a dedicated abstraction layer for LLM-specific tasks including KV caching, context window management, and text generation loops.
- The new Agent Skills format uses a directory of metadata, instructions, and executable code to give smaller models domain-specific expertise without re-prompting.
- The AI Edge Portal allows developers to benchmark on-device models across a fleet of physical hardware to identify performance bottlenecks before production.
- LiteRT supports hardware acceleration across CPU, GPU, and NPU for Android, iOS, web, and IoT platforms via the core execution engine and high-level Kotlin/Swift bindings.
Why It Matters
The shift toward on-device generative AI directly addresses the streaming industry's critical priorities of latency and data privacy by eliminating network delays and keeping sensitive user information off external servers. For media companies, LiteRT’s ability to run optimized models locally reduces significant data center costs and enables offline features for applications on mobile and embedded devices. This marks a strategic push to decentralize AI processing, moving the computational burden from the cloud to edge hardware. Watch for how this impacts the adoption of proactive, privacy-centric AI assistants in mobile-first streaming experiences as hardware-specific NPU optimizations become the standard for on-device inference.
Additional Context
Google has significantly matured its edge AI ecosystem since transitioning from TensorFlow Lite to LiteRT. Per a January 2026 update from the Google AI Edge team, the production-ready LiteRT stack now delivers 1.4x faster GPU performance than its predecessor and integrates best-in-class NPU acceleration for chipsets from MediaTek and Qualcomm. This performance boost is aimed at enabling real-time, complex multimodal tasks on-device, such as background video removal and live translation, which were previously reliant on cloud processing. To solve the fragmentation hurdle, Google launched the AI Edge Portal in May 2025, providing developers with an interactive dashboard to benchmark LiteRT models across more than 100 representative mobile device configurations. The competitive landscape for on-device AI is tightening as silicon capabilities reach parity. Per GSM Arena in June 2026, the flagship performance gap between Qualcomm’s Snapdragon 8 Elite Gen 5, MediaTek’s Dimensity 9500, and Apple’s A19 Pro has effectively compressed into a single ultra-high-end tier. While Google’s own Tensor G5 chip tracks slightly behind these benchmarks in raw GPU power, Google’s strategy emphasizes software-layer integration through AICore. Per Android documentation in April 2026, AICore serves as the system-level interface that manages the distribution and security of the Gemini Nano model, ensuring that on-device AI remains consistent across various hardware generations while adhering to privacy principles. Industry analysts note that the proliferation of dedicated Neural Processing Units (NPUs) is resetting the device upgrade cycle. Per GadgetsFocus in October 2025, Gen AI-enabled smartphones were forecast to exceed 30% of total shipments by the end of 2025. This hardware substrate is essential for the localized execution that Google’s LiteRT and Apple’s Intelligence framework require. As Apple continues to leverage its vertically integrated A-series silicon for deep, privacy-led iOS automation, Google’s LiteRT offers a broader, cross-platform alternative that extends from Android and iOS to web browsers and IoT devices like Raspberry Pi 5.
Read full article at youtube.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source