Human oversight drives 75 percent of agentic AI workflows costs
McKinsey analysis indicates that human oversight accounts for up to 75% of variable costs in agentic AI workflows, challenging the common focus on token consumption. The report advises streaming and enterprise leaders to prioritize workflow redesign and agent reuse to achieve sustainable ROI at scale.
Key Takeaways
- Human oversight accounts for 70% to 75% of variable run costs for AI agents in sectors like banking.
- OpenAI GPT-5.5 costs approximately $5 per million input tokens and $30 per million output tokens as of July 2026.
- Single-agent workflows can cost up to $30,000 to run, while multiagent teams reach $200,000 per implementation.
- Fixed costs including public cloud containers, memory, and data scientist capacity remain constant per agent regardless of volume.
Why It Matters
The high cost of human intervention suggests that the streaming industry's focus on model efficiency and token pricing is misplaced. For media companies deploying agentic AI workflows for content moderation or customer support, the real economic hurdle is the 'total cost of ownership' involving expensive functional experts. As the ecosystem shifts toward multiagent teams, the ability to reuse agents across different high-value workflows will determine which platforms achieve sustainable ROI. The industry must now pivot from simple model selection to radical workflow redesign that minimizes manual exceptions. Watch for a rise in 'AgentOps' disciplines to manage these shifting unit economics during quarterly business reviews.
Additional Context
McKinsey & Company has been expanding its agentic AI advisory practice aggressively throughout 2025 and 2026. In early 2025, McKinsey launched a dedicated AI agent deployment framework called Lilli, its internal knowledge platform, to accelerate enterprise agentic workflows across its own consulting operations, signaling that the firm views agent orchestration as a core service line rather than a research exercise. The firm's broader "Superagency" report, published in January 2025, surveyed over 3,400 executives and found that only 1% of respondents said their organizations had deployed agentic AI at scale, underscoring the gap between pilot enthusiasm and production economics that the new cost analysis addresses.
OpenAI's pricing moves directly shape the token-cost side of the equation McKinsey examines. OpenAI announced GPT-5.5 in June 2026 with a 40% reduction in per-token pricing compared to GPT-5, a step that narrows the raw inference gap between frontier and lightweight models. Meanwhile, OpenAI's enterprise revenue surpassed $12 billion annualized in the first half of 2026, driven in part by agentic workflow deployments in customer service and content operations at media companies. This revenue trajectory suggests that despite McKinsey's finding that tokens are a minority cost, enterprises continue to pay premium rates for frontier-model reasoning in high-stakes workflows where error tolerance is low.
Competing consultancies and technology vendors are publishing their own agentic AI cost frameworks, creating a crowded advisory landscape. Gartner projected in April 2026 that 40% of agentic AI projects would be canceled by end of 2027 due to underestimated total cost of ownership, a forecast that aligns with McKinsey's emphasis on hidden labor costs. Deloitte published a companion analysis in May 2026 estimating that human-in-the-loop review adds $4 to $11 per transaction in media content moderation workflows, providing a concrete per-unit benchmark that streaming operators can compare against their own moderation spend. These converging estimates from multiple firms suggest the industry is coalescing around a shared understanding that agent economics cannot be evaluated on model pricing alone.
Read full article at mckinsey.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source