Evals are the new PRD, Expedia's AI chief says, while agent incidents climb

Expedia Group's first chief AI and data officer Xavi Amatriain used his VB Transform 2026 appearance to argue that product requirements documents for AI systems should be replaced by evaluation suites. "The new PRD are the evals," he told the audience, framing evaluation design as the place where product intent, security requirements, and red-teaming get encoded before any code is written. The pitch lands at a moment when most enterprises are not yet equipped to act on it: VentureBeat's VB Pulse research, drawn from 157 enterprises, found that 66% already permit or are building toward production deployment without human review within 12 months, while only 5% fully trust the automated evaluations that would make that decision.

What Amatriain describes is governance by design, not by post-hoc policing. Expedia's model layers principles, then processes and tools, then automation on top. The "agent release toll gates" he describes are risk-calibrated checkpoints: low-risk agents get lighter review, high-risk agents get red-teaming, security review, and multi-round evals that shift from recommended to required as stakes climb. Amatriain says the goal is to keep governance proportional to the action's blast radius rather than uniform across the agent portfolio. He also pushed back on guardrails as a primary mechanism, calling them "a necessary evil" that should shrink over time, since they bias user feedback and degrade the eval loop. Other Transform speakers pushed back, arguing that the highest-risk actions still demand firm guardrails. Both camps are responding to the same audit pressure, and Expedia's bet is that early principle-encoding reduces the need for late-stage rule scaffolding.

The survey data points to a scaling problem in eval fidelity, not a tooling gap. Half of the 157 enterprises surveyed have shipped an agent that passed internal evals but then failed with a real customer. The 50% customer-failure rate is the operational risk the PRD-as-evals idea has to absorb, given that evaluations are meant to be the specification. If evaluations are the specification, the spec is failing in roughly half of enterprise cases before the customer ever sees the output. Amatriain's claim that evals can absorb security requirements and product intent assumes the eval suite is high-fidelity; the survey data suggests the opposite. The architecture he described, with tools composing into skills, skills into sub-agents, sub-agents into orchestrated systems, is structured precisely so each layer can be evaluated in isolation. Whether that decomposition actually improves eval fidelity is something the source does not test, and the 50% customer-failure rate is the implicit baseline any serious test would have to beat.

Amatriain's examples from Expedia's own deployment make the design point concrete. Travel pricing changes minute to minute, so the architecture blends retrieval-augmented generation with direct API tool calls, choosing between them based on latency: cached answers return instantly, while complex queries like a pet-friendly four-star near Lake Michigan with a pool get up to 30 seconds of reasoning. The system also cross-references supplier self-reports against Expedia's own review corpus, surfacing cases where travelers contradicted what the hotel claimed. Amatriain said that Expedia doesn't want the agent to book a hotel or buy a plane ticket. "That's something that the user has to have the agency," he said, calling the user-final-click constraint non-negotiable and framing it as a security decision. The constraint removes a class of irreversible actions from the agent's authority surface, which reduces the audit scope. He tied that back to the governance model: design-time principles make many guardrails unnecessary, because the design itself forecloses the failure path.

The agent security Pulse survey, drawn from 107 enterprises, shows the enterprise-level problem Amatriain is positioning Expedia against. More than half, 54%, have already had an agent security incident or near-miss. Incident rates climb with organization size, reaching 63% among enterprises with more than 1,000 employees versus 49% for companies with 101 to 1,000. Sandbox isolation, the one post-breach control that limits damage, drops from 35% adoption at smaller companies to just 20% at the largest. The data shows that as agents scale inside an organization, the controls that would limit blast radius erode rather than strengthen. 59% of those surveyed plan to adopt, add, or replace agent security tooling within 12 months, and 29% plan to do so this quarter. The procurement urgency suggests that most enterprises are buying their way into controls Expedia is trying to encode in design.

Amatriain warned that the next attackers will be other AI systems, not just humans. External agentic systems, he argued, are going to be "poking at everything you're doing," and once something is detected, the time to fix becomes the limiting factor. He described a feedback loop where monitoring signals from production flow back into the eval suite, with the goal of automating the full cycle. The source does not specify how Expedia measures that cycle's latency or what counts as a fixable signal. The threat model he describes, where adversarial agents probe continuously rather than human attackers opportunistically, compresses the detection-to-fix window to a regime where human-in-the-loop review may not keep up. The eval-as-PRD pitch is the only path he sees to keep ahead, but it requires the eval suite itself to hold up under adversarial probing, which the source does not test.

The PRD-as-evals framing pushes the hardest part of AI product development into the eval suite, which is the part enterprises currently trust least. If Amatriain's model works, the eval suite becomes both the product spec and the security boundary, and governance effort collapses into eval design rather than being layered on top. The 5% trust rate and the 50% post-eval customer failure rate suggest that most enterprises are not yet at the eval fidelity required for that collapse to be safe. Whether Expedia's toll-gate approach produces eval suites with fidelity high enough to absorb product intent, security requirements, and adversarial probing at the same time is the validation question the source leaves open.

Subscribe to AI Enthusiast Log

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe