Enterprise AI agents: Most are chatbots, and the controls lag deployment
VentureBeat Research fielded five parallel surveys in June 2026, reaching 573 qualified respondents at organizations with 100 or more employees, and the central finding cuts through the agentic AI hype directly: enterprises deployed AI agents knowing the controls were not in place. That is not a discovery. It is a confession. Across five measured control layers, 57 to 68 percent of enterprises plan to switch vendors or add new ones within 12 months, and roughly a third plan to move within the quarter. The retrofit is already underway, and it is budgeted for.
The most basic mislabeling in the survey is the one worth starting with. Seventy-one percent of enterprises said a quarter or fewer of their deployed agents can complete multi-step work independently. Only 10 percent said true multi-step agents are the majority of what they run. A single-prompt chatbot with a human reading every answer needs none of the five control layers the survey measured. A true multi-step agent needs all of them. Most enterprises cannot say which category their deployment falls into. That ambiguity is not a technical problem. It is a governance one, and it makes the switching intent numbers harder to interpret: organizations may be planning to add controls to deployments they have already miscategorized.
The evaluation finding is the sharpest measure of a gap between internal confidence and production reality. Sixty-seven percent of enterprises either already allow an agent to push a code or system change to production on automated evaluation results alone, with no human review, or are engineering toward that within 12 months. Only 5 percent fully trust the evaluations that would make that call. Half of all respondents shipped an agent that passed internal evaluations and then caused a customer-facing failure in the past year. The internal evaluation system is not the production environment. Passing a benchmark under controlled conditions does not validate production behavior, and the survey data confirms that organizations know the evaluation system is weaker than the claim implies while still deploying into automated pipelines anyway.
The security data quantifies a risk that is well understood in principle but rarely measured this cleanly. Sixty-nine percent of companies let at least some of their agents share credentials, operating multiple agents under one API key or service account. Organizations that allow credential sharing anywhere experienced a security incident or near-miss at a 63.5 percent rate. Organizations where every agent has its own scoped identity experienced incidents at 40.9 percent. The 22.6-point gap is not a theoretical risk. It is a measured difference in observed outcomes. Scoped identity per agent, starting with production-touching systems, is the recommended control. The survey names it; the adoption gap between 69 percent sharing credentials and the lower incident rate at scoped-identity shops tells the story of where the work is.
The context layer finding connects the data governance problem directly to agent output quality. Fifty-seven percent of enterprises traced a confident, wrong agent answer in the past six months to their own missing or inconsistent business context: wrong metrics, stale definitions, absent documents. Most saw it happen more than once. Agents answer from data nobody has governed, and they do it confidently. The problem is not the agent. It is the definitions and metrics the agent draws on. Governing those definitions before scaling the agent layer is the sequencing issue the survey surfaces, and it is the one that cuts against the common push to deploy first and govern later.
The hardware utilization number belongs in the same category of evidence for where enterprises are versus where they think they are. More than 82 percent of enterprises running their own GPUs reported utilization at 50 percent or less. Only 44 percent rigorously track what their AI compute actually costs and returns. The survey does not establish what a well-run compute operation looks like as a baseline, but the gap between 50 percent utilization and whatever the ceiling should be, combined with weak cost tracking, suggests the capacity planning problem is not solved. The number worth chasing first is not more GPU allocations. It is the utilization and per-workload cost of what is already running.
No layer has an entrenched incumbent. The defaults today are the built-in tools that ship with the large AI platforms enterprises already use. Switching intent runs highest in orchestration, where 68 percent plan to adopt, add, or replace platforms within 12 months, and 34 percent within the quarter. The survey does not ask which direction that money moves: toward the platforms' built-in tools or toward the specialists challenging them. That open question is the market signal the next four quarters will answer. The direction is clear. The destination is not.
The survey samples are self-selected under VentureBeat Research's VB Pulse program. The individual reports carry full methodology notes, and VentureBeat produces both the research and the VB Transform conference where the reports debuted. The directional pattern holds across every survey independently, which carries more weight than any single percentage point. What the five surveys confirm together is not that the agentic transition is failing. It is that enterprises are managing a controlled retrogression: deploying fast, governing slow, and now paying for both at once.