Majestic Labs debuts 128TB memory-centric server to bypass GPU bottlenecks
Hardware startup Axion Ray has announced a new architecture utilizing Arm-based CPUs and up to 128TB of LPDDR6 memory to handle AI inference tasks. This system aims to provide a cost-effective alternative to traditional HBM-based GPU clusters by increasing memory capacity and reducing power consumption.
Key Takeaways
- Prometheus server utilizes up to 128TB of LPDDR6 memory in a single coherent pool, offering 1,000x more memory per processor than typical GPU setups
- System replaces Nvidia GPUs with Ignite AI Processing Units (AIUs) that combine Arm cores with RISC-V vector and tensor engines
- Hardware design focuses on reducing total cost of ownership by 10x to 50x compared to equivalent H100 clusters while lowering power consumption
- Architecture supports industry-standard frameworks including PyTorch, vLLM, and OpenAI’s Triton to allow existing models to run without modification
Why It Matters
The 'memory wall' is the primary operational bottleneck for large language model inference, where weights must be constantly re-read for every token generated. By prioritizing massive LPDDR6 capacity over raw GPU compute, Majestic Labs provides an alternative for organizations priced out of the current Nvidia H100/H200 market. This shift suggests a growing niche for 'inference-first' hardware that sacrifices high-bandwidth memory (HBM) for sheer volume, making frontier-scale models deployable at a fraction of the power density. Watch for whether this architecture can maintain competitive token-per-second throughput compared to AMD’s high-capacity MI300X or Nvidia’s Blackwell systems.
Additional Context
The trend toward memory-centric AI hardware is accelerating as the cost and supply of High Bandwidth Memory (HBM) remain tight. Per InElectronics (April 2026), startup Positron recently secured a deployment with Oracle for its Asimov processor, which similarly sidesteps HBM in favor of LPDDR-based memory on organic substrates. These designs target the 'inference economics' of air-cooled data centers that cannot support the 120 kW rack densities required by top-tier GPU clusters. Meanwhile, AMD has responded to the memory-capacity gap with its MI355X, which features 288GB of HBM3e to allow 70B+ parameter models to run on a single chip, per Vamsi Talks Tech (June 2026).
This shift occurs as LLM inference is increasingly recognized as a memory-bound rather than compute-bound task. According to TrendForce (January 2026), the reliance on HBM has sparked a supercycle in server-grade DRAM prices, pushing competitors to look for alternatives like LPDDR6. While LPDDR6 was initially targeted at mobile devices, its entry into the data center as a 'middle tier' memory is seen as a way to trade raw bandwidth for improved energy efficiency. Other emerging technologies, such as NEO Semiconductor's 3D X-DRAM, also passed proof-of-concept in early 2026 as potential HBM replacements, underscoring the industry's focus on automated Kubernetes cost optimization over vertical stacking.
Read full article at techradar.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source