Amazon AGI lab says AI agent pilot to production gap is measurement failure

Amazon's AGI autonomy lead used a VB Transform 2026 stage to argue that the enterprise agent deployment gap is a measurement failure, not a capability one. Bryan Silverthorn, who joined Amazon through the Adept AI acquisition, said teams are still scoring agents on benchmarks that mask the dimensions that decide production viability. Cisco data cited on the same stage puts 85% of enterprises in pilot mode and only 5% in production, a gap that the source frames as a measurement problem rather than a model problem. The framing matters because it shifts the bottleneck from research to operations: the agent that passes an internal eval and fails a customer is not a worse model, it is an unmeasured failure mode.

Silverthorn's prescription borrows a four-part taxonomy from Princeton research: consistency, robustness, predictability, and safety. The taxonomy matters because it separates failure modes that most evals collapse into a single score. Consistency asks whether the agent can do the same task a thousand times in a row. Robustness asks whether it tolerates input variation. Predictability asks whether the team can forecast what the agent will do. Safety asks what it does when it goes wrong. Most enterprise benchmarks report a single number that hides all four. The source's claim is that this compression is exactly why agents pass internal evals and fail customers: the dimensions that decide production are the ones the eval was never designed to test.

The customer example Silverthorn described shows the cost of that compression. A deployed agent for software QA extracted serial numbers from screens for two months without incident. Then it began intermittently reading wrong numbers. The failure traced to a vision encoder whose behavior shifted depending on where the serial number appeared on screen, a sensitivity invisible to a human reviewer. The source does not name the customer, the encoder, or the magnitude of the failure rate. What it does establish is that the agent passed whatever eval came before deployment, then failed in production along a dimension the eval never measured. The gap is not in the model, the source argues; it is in the rubric.

The 85%-to-5% ratio is the framing the source leans on hardest. Silverthorn did not present the Cisco data himself; it was the room's frame for why his reliability argument matters at all. The ratio implies that for every company piloting an agent, 80 percentage points of progress evaporates between pilot and production. The source does not test whether that ratio is causal or merely correlational, but the implication it draws is direct: the agents that escape pilot purgatory will not be the smartest ones, they will be the ones whose teams have built measurement rigs that match the application's stakes. That is a different selection pressure than the one model benchmarks reward.

VentureBeat's own research, presented before Silverthorn's session, complicates the picture further. Half of surveyed companies reported shipping agents that passed internal evals but failed real customers. Enterprises overwhelmingly track uptime while ignoring accuracy. Most default to the model makers' own evaluations as their testing strategy. The source frames that pattern as a coin flip between trusting the vendor and trusting nothing. The implication is direct: enterprises are running a measurement regime that optimizes for one dimension (the system stays up) while ignoring the dimension (does it produce the right answer) that decides whether the deployment is a product. The asymmetry survives because uptime is easy to log and accuracy is hard to define.

The 'intern' framing in Amazon's AGI lab gets quoted often, but it carries an argument that is easy to misread. Silverthorn's argument is not that agents are like interns in capability; it is that managing them requires the same operational habits. Interns are powerful but occasionally clueless, capable of great work and spectacular derailment. The management response is to ask what could go wrong, build in undo and backup paths, and decide what risk to accept. The source describes an Amazon agent that ran experiments around the clock on its own high-level research plan, with the team accepting occasional wrong experiments in exchange for research velocity. That trade is explicit. The intern framing makes it culturally legible. The evaluation regime that decides whether to use the same approach in a customer-facing product is a different question the source does not address.

Silverthorn was candid about the parts of the agent stack that remain unsolved. Self-improving AI is, in his words, a loaded term: Amazon uses AI to improve its models constantly, but fully autonomous self-improvement is distant. Computer use is a core focus, with a commercial trucking customer using browser automation to stitch together warranty claims across fragmented systems. The source stresses that no future agent will rely on computer use alone, it will work alongside MCP, APIs, and other tools. LLM-as-judge techniques are useful but only one of several alignment strategies. None of these caveats undermine the reliability argument. They do locate the deployment ceiling in operational maturity rather than in any one model decision.

Most enterprise AI coverage treats pilot-to-production gaps as evidence that the models are not ready. The source instead relocates the bottleneck to measurement discipline, evaluation design, and management practice. That relocation has a cost. It implies that the work of getting from 5% to a higher production rate is unglamorous: building test suites that match application stakes, instrumenting accuracy rather than just uptime, and treating the agent as an operational system rather than a demo. The source does not show what an enterprise that has crossed the 85%-to-5% gap looks like in practice, only that the gap is closing requires a different discipline than picking a better model.

The argument is persuasive because the data is consistent across sources. Cisco's pilot ratio, VentureBeat's customer-failure rate, and Silverthorn's taxonomy all point in the same direction: agents that pass internal evals and fail customers are the predictable result of evals that test the wrong dimensions. Whether the four-dimension framework becomes a standard rubric or stays a useful mental model depends on whether enterprises invest in the operational work the framework implies. The source does not show that investment taking place at scale, only that Amazon's AGI lab has decided to treat it as the binding constraint.

Subscribe to AI Enthusiast Log

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe