Enterprise AI agent evaluation gap leaves half of passed agents failing
A VentureBeat Pulse survey of 157 enterprises found that half have shipped an AI agent or LLM feature that cleared their internal evaluations and then failed a customer in production. Only 5% say they fully trust automated evaluation. Yet two-thirds already allow, or are actively engineering toward, allowing agents