Enterprise AI infrastructure outpaces cost visibility, widening compute gap

Enterprise AI infrastructure spending is outrunning the measurement needed to control it, according to VentureBeat's Q2 2026 Pulse survey of 107 enterprises. The survey uses the term 'compute gap' for the distance between aggressive investment and weak visibility: only 21% of respondents run AI in production at scale, yet 45% plan to evaluate specialized AI clouds in the next year, and 64% intend to switch or add an infrastructure provider within twelve months. The buying pattern looks rational in isolation: integration with the existing stack (41%) and total cost of ownership (35%) dominate selection, while cost per million tokens finishes last at 8%. The problem sits in the second criterion. Fewer than half of enterprises (44%) rigorously track what their AI compute costs, and 83% of GPU-operating respondents report utilization at or below 50%. Buyers are optimizing for an economic dimension most cannot yet measure, on hardware most do not yet use efficiently. The next round of infrastructure decisions will inherit that combination, not escape it.

The survey's deployment curve is the load-bearing fact for everything that follows. Only 21% of respondents describe AI as in production at scale across the organization, while 38% are still running proofs of concept and 37% have some workloads in production but not across the business. A further 4% are not running AI workloads at all. Three-quarters of the sample sits in the building stage, not the operating stage. That matters because the infrastructure decisions the report captures are decisions being made by organizations whose compute footprint is about to expand, not by operators who have already settled into a working pattern. The investment intentions captured here are the leading edge of a build-out, and the source is explicit that the report reads them as such.

The current stack is the familiar one. Google Cloud leads platform usage at 48%, with Microsoft Azure at 29%, AWS at 22%, and Oracle Cloud at 22% rounding out the major hyperscalers. On the model side, Google's Gemini leads at 41%, OpenAI follows at 40%, and Anthropic at 12%. Only 6% of enterprises run their own on-prem or co-located GPU clusters, and 4% operate a custom open-source self-managed stack. The specialized AI clouds that dominate infrastructure headlines, including CoreWeave, Lambda, Crusoe, Nebius, Together, and Fireworks, each register at under 2% of respondents. The methodology note attached to this finding is important: the question allowed multiple selections, respondents averaged 2.1 providers, and the sample skews mid-market, so these figures measure presence in the stack rather than spend or primary status. Read that way, the picture is not that the neoclouds are small. It is that mid-market enterprises are still buying AI from the vendors they already buy cloud from.

Against that base, the intended evaluation pattern is striking. The single largest planned area for evaluation over the next twelve months is AI-specialized clouds at 45%, the same category that sits below 2% today. Behind it, 32% plan to evaluate non-Nvidia accelerators, including AWS Trainium, Google TPU, AMD Instinct, Intel Gaudi, and in-house ASICs, while 28% plan to evaluate next-generation Nvidia silicon, specifically the GB300 generation. Decentralized or distributed compute networks draw 16%, and sovereign or region-specific compute draws 11%. Net momentum, calculated as evaluations minus current usage, puts specialized AI clouds at +24, edging out the hyperscalers themselves at +22. The source itself flags this as the report's sharpest tension, and the consistency check against an earlier April-May survey wave reinforces it: the same category has been the most-cited evaluation target across two differently worded questions.

The switching data sharpens the picture. A clear majority of 64% plan to switch or add an infrastructure provider within twelve months, and 38% within the next quarter alone. For a category as foundational as compute, that level of churn intent is unusual. The providers drawing the most switching consideration are again the incumbents, with Microsoft Azure and Google Cloud at 33% each and OpenAI at 30%, and Gemini at 22%. The neocloud interest registered in the evaluation question is a twelve-month thesis; the near-term switching is mostly reshuffling within the major hyperscalers and consolidating spend around the same set of providers enterprises already use. The two findings read together: the next quarter is about which incumbent wins the reshuffle, and the next year is about whether a non-incumbent layer gains a foothold before or after the hyperscalers extend their stack.

The buying criteria the survey reports are coherent, and the source frames them as such. Integration with the existing cloud and data stack leads at 41%, total cost of ownership follows at 35%, performance, covering latency and throughput, at 24%, and security, compliance, autoscaling, and GPU access cluster at 19% each. Cost per million tokens, the metric that dominates vendor marketing and pricing pages, is the deciding factor for only 8% of respondents, dead last among the choices offered. The pattern is the rational one for a buyer integrating a new category into an existing footprint: it is cheaper and safer to optimize for fit and lifecycle cost than to chase the lowest unit rate on a workload whose total economics are not yet legible. The buyers here are not making a mistake. They are making a reasonable choice on a metric they have not yet learned to measure.

The measurement gap is where the report's claim sharpens into something more useful. Only 44% of respondents say they rigorously track the cost and return of their AI compute. A further 39% track it only partially, 20% cannot quantify it yet, and 6% say it is not a priority. The majority, in other words, are selecting on TCO at the same time as they cannot yet compute it. The source also reports a separate, related, and similarly unflattering number on the asset side: 83% of GPU-operating respondents report utilization at or below 50%, with 49% at or below 25%, and a further 8% not measuring utilization at all. That is the operational expression of the same gap. The infrastructure already in place is not yet producing enough load to amortize its own cost, and most respondents do not have the telemetry to know how close the gap is.

Asked which approach they would rely on as the binding constraint in inference shifts from GPU compute to memory bandwidth, specifically KV-cache capacity, respondents scatter: Dell at 31%, Nvidia at 16%, Hammerspace at 10%, DDN at 9%, with the rest split across open-source KV-cache tooling, model-level efficiency, VAST Data, WEKA, and others. The most telling number is the 18% who say they are not aware of the constraint or have not addressed it yet. For a shift the source describes as reshaping inference cost and architecture, the absence of a coherent answer is itself a finding. Most enterprises will be making near-term memory and storage decisions without a settled view of where the constraint actually lands. The infrastructure cycle ahead is therefore not only bigger than the measurement infrastructure that supports it; it is also running into a different technical bottleneck the same respondents have not yet started to plan for.

The survey is small, with 107 respondents, is self-selected, mid-market skewed, and limited to a single Q2 2026 wave, and the source is explicit about treating the results as directional rather than precise. Read with that caveat, the central claim survives: enterprises are evaluating and switching infrastructure at a pace that outruns their ability to see the unit economics, and the gap between evaluation velocity and measurement maturity is widening faster than the spend. The deployment case for the next wave of specialized clouds, non-Nvidia accelerators, and re-platformed inference depends on a measurement layer most respondents have not yet built, and the source does not show whether that gap closes before the spend does.

Subscribe to AI Enthusiast Log

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe