Enterprise AI agents outpace the evals meant to govern them
A VentureBeat Pulse Research survey of 157 enterprises (100+ employees) in June 2026 identifies what it calls an evaluation gap. Enterprises are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less, and they are doing so knowingly. Half of organizations (50%) have shipped an agent that passed internal evaluations and then caused a customer-facing failure; only 5% say they fully trust automated evaluation today; and the most-cited weakness is that evaluations do not align with real-world outcomes. Yet two-thirds (66%) already permit or are engineering toward fully automated, zero-human-in-the-loop deployment. The autonomy is arriving faster than the assurance, and the gap is the distance between the two.
The survey defines the evaluation gap precisely: the distance between how much autonomy enterprises are handing their agents and how far they trust the tests that are supposed to catch the failures before they reach production. What makes the gap consequential is not a coverage shortage but a reality-alignment failure. Enterprises have discovered, often through expensive customer incidents, that a passing eval is not the same as a working agent.
The data is directional at 157 respondents, self-selected, and mid-market-weighted, but the direction is consistent across findings. The sample includes organizations in technology/software (23%), retail/consumer (15%), healthcare/life sciences (12%), and manufacturing (10%), with respondents skewed toward senior and buyer-credible roles. Findings should be read as signals, not precise measurements.
The first finding explains why the gap matters. Half of organizations that run evaluations have deployed an agent or LLM feature that passed internal testing and then failed a customer. A quarter have experienced it more than once. The failure is not hypothetical: the evaluation said the agent was ready, and it was not. Only 36% report no such failure; 8% run no pre-deployment evaluations; and 6% do not track root cause closely enough to know. Enterprises have empirical evidence that evaluation outcomes and production behavior diverge, and that evidence shapes every subsequent decision about how much trust to place in automated gating.
That trust is thin. Only 5% of organizations say they fully trust automated evaluation as it stands. The most common limitation, cited by 29%, is that evaluations align poorly with real-world outcomes. That is the exact mechanism behind the finding that half of organizations have shipped a passing eval that then failed. Bias or inconsistency (21%) and a lack of explainability (18%) follow: enterprises cannot always tell why an evaluation reached its verdict. Data-leakage or privacy concerns in the evaluation process itself affect 17%. The tests meant to certify agents are not trusted to certify them, which makes the autonomy trajectory in the next finding so striking.
Two-thirds of organizations either already allow zero-human-in-the-loop deployment for low-risk agents (34%) or are actively engineering their pipelines to permit it within twelve months (33%). Only 22% rule it out for the foreseeable future. The direction is unambiguous: enterprises are moving to let evaluations gate production autonomously, removing the human check, at the same moment they say those evaluations do not reliably match reality. The autonomy ceiling is rising faster than the assurance beneath it. The mechanism by which the false-confidence failures in Finding 1 occurred will scale rather than shrink unless the evaluation gap closes from the assurance side.
The sample reveals a counter-intuitive split on company size. Larger enterprises (2,500+ employees) are slightly further along the path toward zero-human review than smaller companies (70% versus 64%) and slightly more likely to have shipped an evaluation-passing agent that failed a customer (54% versus 48%). The assumption that large, regulated organizations hold the human in the loop longest is, in this sample, backwards. This is a directional signal given sample sizes of 57 and 100 across those groups, but it suggests that organizational scale does not automatically produce evaluation maturity or deployment caution.
The evaluation stack is fragmented and provider-led. The most common primary tools are provider-native evals, tied with no dedicated tooling at all (17% each). OpenAI's native evals and traces lead at 17%, followed by Anthropic's Claude Console evals at 13%. Specialist evaluation vendors—DeepEval (12%), Braintrust (8%), LangSmith, Weave, Promptfoo, Langfuse, Arize—are scattered across single to low double digits. An additional 11% have built their own. No independent platform has become the category standard, which leaves most enterprises evaluating agents with provider-native tools, homegrown scripts, or nothing. This matters because provider-native tooling measures performance within the provider's own framework, which may not reflect the full range of behaviors agents exhibit when integrated into production workflows outside that framework.
Production monitoring rarely watches output quality. When asked what live production monitoring is built for, 51% of organizations monitor only whether the agent is functioning—uptime, response time, error rate, cost. Only 23% monitor whether its answers are correct. The distinction matters because a confidently wrong answer is invisible to the first kind of monitoring: the request completes, the response is fast, no error is thrown, and every functioning-metric reads healthy. Roughly three-quarters of organizations run no automated, real-time evaluation of output correctness in production. They can see that the system is up and what it costs; they are taking the correctness of its answers on faith. That blind spot is the runtime counterpart to the pre-deployment gap: the same organizations engineering the human out of the deployment decision mostly cannot see, in real time, when the deployed agent starts getting things wrong.
Enterprises buy evaluation tooling on economics and trust it on consistency. Cost of evaluations (28%) leads selection narrowly, ahead of ease of integration (27%) and evaluation accuracy (24%). More than a third (36%) name evaluation consistency—getting the same verdict on the same behavior every time—as their primary measure of success. The emphasis on consistency is telling: before an evaluation's verdict can be trusted, it needs to be stable. That is precisely the property whose absence (bias and inconsistency) ranked among the top trust limitations. Satisfaction with current tooling averages 3.8 on a five-point scale across overall satisfaction, ease of implementation, and value for money—moderate, not confident.
The investment direction complicates the zero-human trajectory. When asked which reliability and evaluation investment will grow most over the next year, the second-largest planned investment after production observability is human review workflows, at 26%. That is the report's quietest contradiction: at the same moment two-thirds of enterprises are engineering the human out of the deployment decision, more of them plan to grow spending on human reviewers than on the automated evaluation pipelines that would replace them. The zero-human trajectory and the human-review budget are rising in the same companies at the same time. Only 8% report that their budget is not increasing. Enterprises are hedging—building toward autonomy while spending to watch agents more closely and keep humans available for the calls that automated evaluation cannot yet be trusted to make.
The evaluation market is in early consolidation. Nearly two-thirds (64%) intend to adopt a new, additional, or replacement platform within twelve months, and 31% within the next quarter. The consideration set points where current usage is thinnest: Confident AI's DeepEval leads what enterprises are evaluating (20%), ahead of OpenAI's native evals (13%) and Braintrust (9%). Given that so many enterprises today rely on provider-native tools or nothing at all, this is less a defection than a first real wave of tooling adoption—the moment the evaluation layer starts to consolidate. Which platforms earn that trust, in a market where almost no one trusts automated evaluation yet, is the question later survey waves will track.
The evaluation gap is not a coverage problem that more tests alone will close. It is a problem of evaluations that reflect reality and can be trusted to gate it. The open question for later waves is whether assurance catches up to autonomy—or whether the false-confidence failures move from customer incidents into changes that deploy themselves without a human in the loop to notice.