Physical AI adoption shifts focus to hardware-level orchestration and token economics
Industry leaders from Rafay Systems, Positron AI, and Axiado discussed the infrastructure and economic challenges of deploying agentic and physical AI, highlighting the need for hardware-level optimizations. The discussion centered on managing the exponential increase in token consumption and improving inference efficiency through specialized silicon and serverless orchestration.
Key Takeaways
- Agentic workloads consume tokens exponentially faster than chat apps due to multi-turn context reprocessing.
- Rafay Systems is deploying serverless GPU orchestration to allow multi-tenancy without dedicated hardware per customer.
- Positron AI's Archer ASIC systems optimize for tokens per watt to support air-cooled enterprise data centers.
- Axiado's silicon-based platform management offloads thermal and frequency scaling from GPUs to dedicated hardware.
- The inference market is bifurcating into prefill-heavy compute and decode-heavy memory capacity segments.
Why It Matters
The transition from digital chat to physical and agentic AI breaks the standard cloud economic model. As token consumption scales exponentially, general-purpose GPU clusters become cost-prohibitive for enterprise edge and robotics applications. Hardware providers are now forced to integrate intelligence at the silicon level to manage cooling, security, and power consumption dynamically. For the streaming and video ecosystem, this infrastructure evolution is critical for low-latency spatial awareness and real-time processing of high-bandwidth sensor data. Watch for the emergence of 'tokens-per-watt' as the dominant procurement metric for B2B AI infrastructure over raw compute benchmarks.
Additional Context
The pressure on physical AI infrastructure coincides with a massive projected surge in token demand. Per Goldman Sachs Research in May 2026, agentic AI is expected to drive a 24-fold increase in token consumption by 2030, reaching 120 quadrillion tokens per month. This explosion is largely due to the stochastic nature of agentic tasks; according to research from arXiv in April 2026, multi-agent AI token costs can consume 1,000 times more tokens than standard coding chat, with input tokens driving the bulk of the cost. Yale University economists also noted in July 2026 that by this year, more than half of all AI tokens globally now involve agentic systems rather than simple one-off queries.
Technological solutions are rapidly moving toward specialized silicon to mitigate these costs. Axiado’s Trusted Control/Compute Unit (TCU) has already demonstrated hardware-level operational savings of up to $20,000 per rack annually in data center cooling and management, per Electronics Buzz in March 2026. Simultaneously, Positron AI’s Series B funding of $230 million in February 2026 highlights investor interest in its Atlas and Asimov platforms, which claim to deliver 280 tokens per second within a 2000W envelope. Per Business Wire, these systems aim to achieve five times higher efficiency than upcoming general-purpose GPUs like NVIDIA’s Rubin by focusing on memory-centric architectures optimized for the long-context windows required by autonomous robotics and advanced video analytics.
Read full article at siliconangle.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source