AI agent inference costs to surge fivefold by 2028 says Gartner
Gartner projects that inference costs per agentic workflow will increase more than fivefold by 2028 due to the computational demands of multistep reasoning. The report suggests that streaming and tech leaders must shift focus from individual token prices to architectural optimization, model routing, and workflow design to maintain margins.
Key Takeaways
- Agentic reasoning models can increase provider costs by at least five times compared to basic chatbot interactions.
- Gartner identifies an 'Inference Paradox' where cheaper unit costs enable more complex, token-heavy applications.
- Multimodel ecosystems using lightweight models for routine tasks and expensive models for reasoning are recommended to maintain margins.
- Optimization through inference tiering, routing, and orchestration is now critical to product design and ROI.
Why It Matters
The shift from simple chatbots to autonomous agents fundamentally changes the unit economics of AI deployment for streaming platforms. As workflows require repeated rounds of reasoning, the cost of intelligence becomes a variable architectural expense rather than a fixed commodity price. This forces a transition away from tracking token costs toward measuring the specific economic return of each automated task within the ecosystem. Organizations must now prioritize model routing and orchestration to prevent generic autonomous intelligence from ballooning operational budgets. Watch for streaming engineering teams to adopt tiered inference strategies that reserve high-cost reasoning models exclusively for high-value user interactions or complex content metadata processing.
Additional Context
Gartner's projection that agentic inference costs will surge fivefold by 2028 arrives as telecom operators are already deploying AI models directly into network infrastructure to offset rising computational demands. Ericsson's AI-native Scheduler with Link Adaptation, which uses neural networks to predict radio conditions in real time, achieved close to 10 percent improvement in spectral efficiency and up to 15 percent higher downlink throughput in large-scale commercial trials on T-Mobile's live 5G Advanced network across approximately 43 sites in Los Angeles, New York, New Jersey, and Salt Lake City. T-Mobile is targeting commercialization of the technology in the third quarter of 2026, signaling that operators are willing to invest in AI inference at the network edge when the return is measurable and immediate. The economics of AI inference in telecom are being shaped by hardware choices that either amplify or constrain cost growth. Ericsson and Intel jointly benchmarked AI-native Link Adaptation on Intel Xeon 6 processors within AT&T's Cloud RAN stack, achieving up to a 20 percent throughput increase compared to legacy rule-based methods, demonstrating that commercial off-the-shelf silicon can handle inference workloads without dedicated GPU clusters. This matters for the broader inference cost conversation because it shows one path to containing expenses: running purpose-built models on existing or commodity hardware rather than scaling expensive accelerator fleets. Ericsson's approach delivers AI features as software updates on hardware operators already own, avoiding the capital expenditure that would otherwise compound the inference cost problem Gartner identifies. The contrast between Gartner's cost warning and these telecom deployments highlights a key architectural lesson for streaming platforms facing similar agentic AI adoption. Ericsson's Per Narvinger noted at MWC 2026 that the company's AI models improved a link adaptation algorithm that had been deterministically optimized for 30 years by an additional 10 percent, a gain he framed as enormously valuable given that spectrum is one of the largest line items in an operator's budget. The implication for streaming and video infrastructure teams is that inference cost management will depend less on waiting for token prices to fall and more on selecting narrowly scoped models, deploying them on cost-appropriate hardware, and measuring the specific operational return of each agentic AI workflows against the workflow it serves.
Read full article at nationalcioreview.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source