Rubrik's SAGE pitches AI to police AI agents, but accuracy is unmeasured

Rubrik's SAGE replaces human-in-the-loop agent approval with an AI judge that scores every agent action against natural-language policy in real time. Dev Rishi, Rubrik's GM of AI, made the case at VB Transform 2026, framing agent autonomy as a settled capability question. The problem he did not address is the accuracy of the judge itself. SAGE is a non-deterministic model policing other non-deterministic models, and the company has not published a false-positive or false-negative rate for it. The product is positioned as a governance control; the only evidence of its reliability is a trail of audit receipts rather than a benchmark.

The architecture Rubrik described at the session gives the judge both reach and a reason to exist. SAGE sits inside Rubrik Agent Cloud, which reached general availability in February, and operates as an aggregation of judges built through parameter-efficient fine-tuning. Each judge takes on a task-specific variant of a base model that retains shared organizational context. One judge watches for tool-use hallucinations while another suppresses PII before it can leave the boundary. The economics rest on a small language model that Rishi said operates at an order of magnitude lower cost and latency than a frontier LLM, which is what makes running one call per agent action viable. That same non-frontier positioning is what makes the missing accuracy benchmark the load-bearing question for the design.

The problem Rubrik is selling to is the one Rishi sketched in a CISO roundtable anecdote hosted by Anthropic's CISO. Roughly 14 attendees raised hands when asked if their AI governance and security policies were written down; the same room chuckled when asked how anyone enforces them. Rubrik Zero Labs' April "State of the Agent" report, based on more than 1,600 IT and security leaders, found that 80% of respondents spend more time monitoring and approving agent actions than the agents save. VentureBeat Pulse research presented earlier on the Transform stage placed 66% of enterprises at "allow or actively building toward" production deployment with zero human review, while only 5% said they fully trust the automated evaluations that would make that decision. Rubrik's framing places itself at the intersection of the two percentages.

YOLO mode, the term Rubrik's founder and CTO pushed internally, strips the permission prompt out of agent workflows. In Rubrik's implementation, a second AI judges every action in place of a human clicking approve. The company first ran the experiment on its own Claude Code and Cowork pilots in ask mode, which produced a 120-message Slack thread of developer pushback. "There's no way that I can actually read through this. And it becomes security theater," Rishi said, quoting the complaint. The phrase captures the gap between written policy and effective governance that SAGE, in Rubrik's pitch, is designed to close.

The closing question the fireside did not answer is the one a skeptical CISO would ask. A judge built from a non-deterministic model cannot guarantee that it correctly classifies every agent action, and Rishi offered no false-positive or false-negative rate for SAGE itself. The closest evidence the architecture supplies is auditability. Backtesting, which Rishi said is just starting to roll out, replays an organization's historical agent actions and tool calls against a new policy, marking where the policy would have stepped in and where an action would have sailed through. Batch analysis runs across entire session traces every hour or every day and surfaces what Rubrik calls insights, the problems no individual guardrail caught. The audit trail of each call, plus whatever got past it, is the substitute for a published score.

The attacks that justify the design are the ones no single rule catches. Rishi pointed at what security researcher Simon Willison called the "lethal trifecta" in June 2025: an agent that holds private data while consuming content nobody vetted and has a channel to send what it finds to the outside world. The danger is the stack, not the parts. An agent with both Salesforce access and email access has done nothing wrong yet; the destructive case is the lookup followed by an outbound send that the policy review of either permission in isolation never anticipated. A VentureBeat June Pulse survey of 107 qualified enterprise respondents found that 69% of companies run credential sharing somewhere in their agent fleet. Companies with shared credentials anywhere reported a security incident or near-miss at a 63.5% rate (47 of 74), against 40.9% (9 of 22) where every agent carried its own scoped identity.

The market context for SAGE is what the same VentureBeat research reported: 82% of enterprises name their primary AI provider's built-in guardrails and cloud controls as their main agent security layer. Fifty-nine percent plan to adopt, add, or replace agent security tooling within the next 12 months. Only 12% include an agent-identity product in what they are considering, even with credential sharing still the norm. A separate finding from Rubrik Zero Labs: 88% of respondents say they lack the ability to roll back agent actions without system disruption, a recovery gap that sits squarely inside Rubrik's original backup and restore line of business.

The deployment case for SAGE depends on whether backtest evidence and audit-trail visibility can substitute for a published accuracy number. The same governance policies the product promises to make enforceable are typically the ones that already require an accuracy metric for new controls, which makes the missing scorecard the procurement blocker rather than the engineering one.

Subscribe to AI Enthusiast Log

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe