AMD and Cerebras Systems have introduced a disaggregated hardware architecture designed to separate prompt processing from token generation for AI inference workloads. The partners claim this approach provides significant improvements in energy efficiency and latency compared to monolithic GPU architectures, with deployments targeting enterprise data centers including Microsoft Azure.