Qualcomm Hexagon NPU architecture adds Element Accelerator for agentic AI
Qualcomm has announced a new Hexagon NPU architecture featuring an Element Accelerator and 50% larger shared memory to support on-device agentic AI. The platform is designed to improve performance for Mixture-of-Experts models and accelerate INT4 prefill tasks for mobile devices.
Key Takeaways
- Element Accelerator purpose-built to optimize transformer operations and multimodal reasoning tasks
- Shared memory capacity increased by 50% to minimize data transfers to external DDR
- Prefill performance for INT4-based models improved by up to 50% for faster decoding
- Support for Mixture-of-Experts models allows 30B-parameter class experiences with only 3B active parameters
Why It Matters
This hardware evolution signals a transition from monolithic AI models to specialized, routed systems that operate locally on mobile chipsets. By integrating the Element Accelerator and expanding on-chip memory, Qualcomm reduces the latency and bandwidth bottlenecks that previously limited complex agentic reasoning in mobile environments. For the streaming and media ecosystem, this enables more sophisticated on-device metadata processing and personalized content discovery without relying on cloud compute. Watch for the first commercial devices featuring this architecture to debut at the Snapdragon Summit 2026 to see real-world token generation benchmarks.
Additional Context
Nokia has built its agentic AI strategy around a tightly integrated hardware and software stack anchored by Nvidia GPUs. In June 2026, Nokia announced an agentic AI framework built into its Network Services Platform for IP network operations, with a root-cause troubleshooting agent as the first commercial deployment expected by end of 2026. The framework operates within operator-defined safety boundaries, allowing AI agents to act on real-time network data without human intervention for routine tasks. This positions Nokia as a direct competitor to Qualcomm's on-device agentic approach, but at the network infrastructure layer rather than the endpoint.
The competitive landscape for agentic AI in telecoms intensified sharply in June 2026. Ericsson launched its AI in RAN commercial software subscription on June 11, claiming up to 20% higher downlink throughput across more than 15 live deployments, while Verizon disclosed that its 60,000-site vRAN network is now applying agentic AI to configuration changes and service assurance. Verizon publicly called for industry-wide interoperability standards for agentic systems, highlighting a gap that mirrors early Open RAN interoperability challenges. The TM Forum's Autonomous Networks L4/5 roadmap and 3GPP 6G standardization process will need to incorporate agentic AI interoperability as a core requirement to prevent vendor lock-in.
Nokia's broader architecture play extends beyond individual agents into a full-stack autonomous network vision. Nokia announced partnerships with AWS and Databricks at DTW Ignite to build data, cloud, and control layers for its Autonomous Network Fabric, claiming operators are already achieving automation rates above 90%, service delivery times under four hours, and up to 85% reduction in slice rollout time. Meanwhile, the fundamental hardware divergence between Nokia and Ericsson on AI-RAN is widening. Light Reading reported that Nokia is designing its entire Layer 1 RAN to run on Nvidia CUDA platforms and GPUs, while Ericsson uses GPUs only for forward error correction and runs all other L1 software on CPUs. This architectural split means the two vendors are building incompatible acceleration paths, a dynamic that approach sidesteps by targeting endpoint inference rather than network-side processing.
Read full article at qualcomm.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source