Expedera's packet-based architecture hits 20x throughput for edge-native AI
Expedera's packet-based NPU architecture optimizes hardware utilization and reduces DDR memory accesses for AI models, achieving up to 20x throughput improvement and 60% power reduction in smartphone implementations. This technology enables more efficient edge AI deployment, which is crucial for real-time processing in devices that handle streaming content.
Key Takeaways
- Packet-based NPU architecture achieves 60-80% typical hardware utilization, rising to 90% peak in production environments.
- Reduces DDR memory accesses by up to 79% for generative models including Llama 3.2 and Qwen2.
- Smartphone OEM implementations reported 11.6 to 16 TOPS/W, cutting power consumption by 50-60%.
- Tailored Origin Evolution platform allows customization of on-chip SRAM and specific attention or vector blocks.
Why It Matters
Immediate gains in hardware efficiency are critical for moving compute-heavy streaming AI, such as real-time upscaling and video analysis, from the cloud to the device. By solving the memory bottleneck—a primary hurdle for edge devices with battery and thermal limits—Expedera enables sub-millisecond AI responses independent of network stability. This shifts the ecosystem's competitive focus from raw TOPS (tera-operations per second) to sustained performance-per-watt. Watch for smartphone OEMs to prioritize integrated NPU IP that minimizes die size while supporting the 128K token context lengths now standard in mobile-first models like Llama 3.2.
Additional Context
The push for dedicated AI hardware reflects a broader industry pivot toward domain-specific ASICs. Per Mordor Intelligence (August 2025), ASIC devices captured 47.2% of the edge AI accelerator market in 2024, as manufacturers moved away from general-purpose GPUs to meet the 5-10W power envelope required for fanless systems. This transition is particularly visible in the smartphone sector, which accounted for over 80% of edge AI hardware volume in 2024 according to MarketsandMarkets research. These devices are increasingly tasked with running advanced generative sets; for instance, testers using Llama 3.2 on iPhone 16 Pro (April 2026) achieved 18 tokens per second, reaching the speed of cloud-based assistants for on-device chat. Beyond mobile, the NPU's role is expanding into high-bandwidth media applications. Per Giznova (February 2026), smart TV NPUs are now essential for real-time 8K upscaling at 60–120Hz, requiring sustained throughput of 50–150 INT8 TOPS to maintain 15ms latency thresholds. In early 2026, MediaTek unveiled its Dimensity 9500 and a 5G-Advanced CPE with Wi-Fi 8, demonstrating how integrated NPUs can reduce application latency by up to 10x through AI-driven Quality of Service (QoS) engines. These advances align with predictions from Precedence Research (November 2025) that the total AI processor market will reach $73 billion in 2026, driven by a 21.5% CAGR in the NPU segment as on-device inference becomes the default for privacy-sensitive and real-time streaming workloads.
Read full article at semiengineering.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source