Enterprises overspend on idle GPUs as AI infrastructure visibility lags
A VentureBeat Pulse survey of 107 enterprise decision-makers reveals a disconnect between aggressive AI infrastructure investment and financial visibility, with 83% of respondents reporting GPU utilization at 50% or below. While hyperscalers currently dominate, 64% of enterprises plan to add or switch to specialized AI cloud providers within a year to optimize integration and total cost of ownership.
Key Takeaways
- Only 21% of enterprises currently run AI in production at scale, yet 45% plan to evaluate specialized AI clouds like CoreWeave or Lambda.
- Fewer than 44% of surveyed organizations rigorously track their AI compute costs, leading to unmanaged spending growth.
- Hardware selection is shifting from headline token pricing toward stack integration and total cost of ownership, which influence 41% and 35% of decisions respectively.
- Approximately 32% of enterprises intend to evaluate non-Nvidia accelerators, including AWS Trainium and Google TPUs, within the next 12 months.
Why It Matters
The streaming and video industry faces a structural disconnect where aggressive capacity procurement outpaces operational maturity. For platforms scaling high-density workloads like real-time transcoding or video-generative AI, poor GPU utilization represents a significant margin drain that hyperscalers have yet to solve via traditional general-purpose clouds. As enterprises pivot toward specialized 'neocloud' providers for better integration, the competitive advantage will shift toward teams that can master unit economics rather than those with the largest hardware reserves. Watch for the emergence of 'inference engineering' roles tasked with bridging this 50% utilization gap as production volumes grow.
Additional Context
The 'compute gap' is compounded by extreme levels of resource waste in existing cloud environments. Per Business Insider in April 2026, data from 23,000 Kubernetes clusters suggests average GPU utilization may actually sit as low as 5%, as companies over-provision hardware due to a 'fear of missing out' on limited chip supply. This over-buying is driven by long-term contract requirements that prevent organizations from scaling resources down as easily as they do with legacy CPU compute. Simultaneously, the financial burden of AI is shifting from one-time model training to indefinite inference costs. Per industry analysis from Spheron in April 2026, inference now accounts for up to 80% of enterprise AI spend. As organizations move these workloads into production, the cost of serving tokens becomes a permanent and often unpredictable operational expense. This has fueled the rise of specialized providers; Synergy Research Group reported in May 2026 that 'neoclouds' like CoreWeave and Crusoe now hold a combined 5% share of the total cloud market, with five such providers entering the top thirty globally. To address efficiency, hardware manufacturers are pivoting toward higher-bandwidth architectures. Nvidia's Blackwell GB300, which began shipping in volume in mid-2026 per VRLA Tech, specifically targets the memory bandwidth bottlenecks that frequently leave GPU cores idle during large-scale inference. At a price point of roughly $90,000 to $115,000 per deskside unit, these systems attempt to centralize high-performance compute to reduce the latency and egress costs associated with distributed cloud environments.
Read full article at venturebeat.com
Get this in your inbox → Subscribe
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source