Arm Neoverse CSS N4 targets agentic AI with 128-core custom silicon
Arm has launched the Neoverse CSS N4 compute subsystem and AGI CPU, designed to support agentic AI workloads in data centers. These hardware solutions offer cloud providers increased memory bandwidth and configurable silicon options to handle the processing demands of AI agents.
Key Takeaways
- Neoverse CSS N4 supports up to 128 cores per die with LPDDR6 memory and PCIe Gen 7 connectivity.
- Arm AGI CPU is gaining adoption from OpenAI, Meta, and Cloudflare for responsive agentic workloads.
- ByteDance division Volcano Engine is developing the first agentic sandboxes powered by the AGI CPU.
- Google Cloud and Microsoft are already utilizing Arm-based Axion and Cobalt 200 processors for AI sandboxing.
Why It Matters
The launch signals a shift in data center requirements where CPU performance per watt and memory bandwidth are becoming as critical as AI accelerators for agentic workflows. For streaming platforms and cloud providers, this provides a path to differentiate infrastructure through custom silicon while reducing development timelines via pre-validated IP. As AI agents move from simple chat to complex tool execution, the reliance on heterogeneous compute environments involving Nvidia, Google, and Microsoft will intensify. Watch for the performance benchmarks of Volcano Engine’s agentic sandboxes to determine if this hardware significantly reduces latency in real-world AI agent deployments.
Additional Context
Arm's Neoverse CSS N4 enters a crowded field of custom silicon designed for AI inference and agentic workloads. Google Cloud announced its Axion processor in April 2024 as the first custom Arm-based CPU for general-purpose cloud workloads, and the chip has since been deployed across Google's internal services and made available to select customers. Microsoft followed with its Cobalt 200 Arm-based processor, which began powering Azure workloads in late 2024, targeting cloud-native and AI inference tasks. Nvidia's Vera CPU, announced as part of its Grace successor roadmap, targets AI agent orchestration workloads with a custom Arm core design, further intensifying competition at the CPU layer for agentic AI. These moves collectively signal that hyperscalers and chip vendors now view the CPU as a strategic differentiator for AI agent scheduling, tool execution, and memory-intensive orchestration rather than a commodity component.
On the business side, Arm Holdings agentic AI demand has been restructuring its licensing model to capture more value from custom silicon programs. The company reported record royalty revenue in fiscal year 2025, driven by increased adoption of Neoverse-based designs in cloud data centers, with cloud infrastructure royalties growing faster than any other segment. Arm's Compute Subsystems (CSS) program, which bundles pre-validated IP blocks to shorten customer design cycles, has attracted more than 20 licensees since its 2023 introduction, including partnerships with cloud providers seeking to reduce time-to-silicon. The CSS N4 specifically targets the agentic AI opportunity by offering configurable core counts up to 128 and enhanced memory bandwidth, positioning Arm to compete directly with Nvidia's integrated CPU-GPU approach for inference-heavy agent workloads.
Technical benchmarks and early deployment signals suggest the agentic AI CPU market is maturing rapidly. Cloudflare announced in early 2025 that it was evaluating Arm Neoverse-based servers for AI inference at the edge, citing power efficiency gains of 30-40% over comparable x86 configurations for token-generation workloads. ByteDance's Volcano Engine platform, which Arm cited as an early CSS N4 design partner, , requiring high memory bandwidth for tool-calling and retrieval-augmented generation pipelines. OpenAI's infrastructure team , noting that agent coordination overhead can consume 40% of total inference compute when tool execution involves database queries and API calls. These data points underscore why Arm is betting that agentic AI will drive a structural increase in CPU attach rates alongside GPU accelerators.
Read full article at eenewseurope.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source