KAIST researchers have identified that AI agents utilizing 70-billion-parameter models can consume up to 136.5 times more energy than standard generative AI, largely due to repeated model invocations and orchestration overhead. The findings highlight significant operational challenges for enterprise streaming operators, noting that these agentic workflows lead to increased latency and high GPU standby times that necessitate improved architectural co-design.