Astera Labs has expanded its Leo Smart Memory Controller portfolio with the X-Series, E-Series, and P-Series to address memory bottlenecks in agentic AI and cloud infrastructure. The new controllers utilize CXL 3.2 and PCIe 6 to enable memory expansion, pooling, and fabric-attached GPU memory for data-intensive workloads.
The shift toward agentic AI requires massive context windows that frequently outpace current hardware memory capacity. By decoupling memory from the CPU and placing it closer to the accelerator fabric, these controllers prevent expensive GPUs from sitting idle while waiting for data. For the streaming and cloud ecosystem, this architecture enables more sophisticated real-time metadata processing and personalized AI agents without a linear increase in hardware costs. As hyperscalers like AMD, Intel, and Samsung align on CXL standards, the focus shifts from raw compute power to efficient rack-scale resource distribution. Watch for adoption rates among neocloud providers as they attempt to optimize inference costs against traditional hyperscale competitors.
New AI agent storage infrastructure is becoming critical as enterprises look to manage the data overhead associated with these memory-intensive workloads. Developers are also exploring hardware-level isolation for AI agent sandboxes to ensure secure execution environments, while others look to open source AI agent frameworks to accelerate deployment.
Astera Labs has launched its new Leo X-Series memory controllers, designed to address memory bottlenecks in agentic AI. By utilizing CXL 3.2 and PCIe 6, these controllers reduce time to first token by up to 62%, allowing AI accelerators to process large context windows more efficiently without increasing hardware costs.
The Leo X-Series creates a dedicated memory tier for AI accelerators, allowing them to store large KV caches and agent context to prevent memory bottlenecks.
The Leo 2 P-Series enables dynamic memory pooling across multiple hosts, which helps reduce stranded DRAM capacity in data centers.
Internal testing using Intel Granite Rapids and Qwen2.5-32B demonstrated a 22% increase in tokens per second and a reduction in time to first token by up to 62%.
The controllers utilize CXL 3.2 and PCIe 6 to decouple memory from the CPU, placing it closer to the accelerator fabric to improve resource distribution.
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source