Multi-turn AI attacks succeed 88% of time — single-turn red teaming misses the pattern

When Cisco ran 6,986 multi-turn attacks against 15 flagship models, attackers who adapted across the conversation broke through as often as 88.3% of the time. That number, presented by Amy Chang, Cisco's head of AI threat intelligence and security research, at the VB Transform 2026 agentic security panel, exposes the gap that single-turn red teaming programs leave open. The 88.3% is not a model failure rate in abstract. It is the rate at which a human adversary who extends and iterates a prompt across multiple conversation turns succeeds against models that pass single-turn evaluations. Single-turn red teaming, Chang explained, tests the one-shot malicious prompt. Multi-turn testing tests what happens when an attacker extends that attack across a conversation, surfaces harmful outputs and misaligned behaviors that a snapshot never catches, and ranks models differently than single-turn testing does.

The survey data from VentureBeat's June 2026 Pulse, covering 107 enterprise respondents, explains why the room was paying attention. More than half, 54%, have already had a confirmed agent security incident or a near-miss caught before harm. Only 32% give every agent its own scoped, managed identity, and 30% isolate their highest-risk agents in sandboxes. Provider-native and hyperscaler controls remain the primary agent security layer at 82% of companies surveyed. Those numbers tell a coherent story: most enterprises have deployed agentic systems using controls designed for a simpler threat model, and the attacks those controls address are not the attacks that are working.

The acquisitions confirm the assessment. Palo Alto Networks closed its $25 billion acquisition of CyberArk in February, CrowdStrike agreed to pay $740 million for SGNL in January, and Cisco announced its intent to acquire Astrix Security for a reported $400 million. Each deal targets the identity and isolation layer that the survey data shows most enterprises have not finished building. The world's largest security vendors have reached the same conclusion about where the risk is concentrated.

Chang described a framework where agents assess a deployment scenario, develop relevant attacks, judge whether they are worth pursuing, execute them, and evaluate their own success. The sophistication of the attack framework is not incidental. It reflects the reality that attackers who target agentic systems are themselves iterating and adapting, which means static, one-shot evaluation understates the actual exposure. What surprised her after building that infrastructure was how the defensive answer stayed simple. "You don't have to get super creative. You just need to think about truly what are the fundamentals and basics of what I'm trying to secure in my organization," she said. Her starting point for CISOs beginning agentic deployments is Cisco's Integrated AI Security and Safety Framework, which catalogs compromise vectors across the AI lifecycle from modality through supply chain. From there, teams work backward from real incidents, trace how each attack succeeded, and build coverage with the right mitigations.

Heather Ceylan, CISO of Box, described a three-layer approach. Permissioning comes first so the agent never accesses more content than the human who invoked it. Ephemeral sandbox environments contain the blast radius if an agent is hijacked. Runtime execution controls restrict the agent's tool calls to only those relevant to the task. "If you want an agent to summarize a doc for you, and you have a prompt injection that says forward this to maliciousattacker at domain.com, it can't do that," Ceylan said. "That action in that tool call is not even in its vocabulary." She classifies agent actions into three oversight categories: actions that are not sensitive need no human in the loop, moderately sensitive actions skip approval but get logged and monitored, and destructive actions like mass deletion always require a human. The categories shift as models and use cases evolve, she acknowledged, but setting them up front gives teams a principled framework rather than ad-hoc decisions.

Rajesh Parekh, VP of AI and ML at Intuit, described why the red teaming surface has expanded. Agents have skills that become vulnerabilities, access to tools that could contain threats, and expanded blast radius when malicious code executes. Intuit has automated common vulnerability patterns from manual red teaming back into its GenOS platform so future agents inherit protection, while red teamers stay focused on new threat vectors. Runtime scanning of prompts and responses adds a final layer that can stop a suspect response and escalate to a human expert.

Ceylan was direct about what development velocity means for security review. "The days of secure code reviews where a human's looking at the code and we're looking at security architecture reviews, design docs, those are done," she said. Box is building toward a fully agentic development lifecycle where agents review design documents, apply security requirements, and scan code for vulnerabilities. "I'm very optimistic that we will get to a point where we will write code without security vulnerabilities because agents and the models are going to get so good at writing code without vulnerabilities," she said. "We're still a long way away from that." Her advice returns to basics that predate agents: "It comes down to very basic least privilege access. If you start giving your agents overly broad permissions at the beginning, it's really hard to comb that back."

The sharpest exchange came on intent detection. When Box's own agent operates, the system controls the prompt, so guardrails and tool-call restrictions can be engineered around known intent. The harder problem, which Ceylan said Box is still working through, arrives when external agents connect and the context behind the request is opaque. Mastercard, speaking earlier at the event, is building an open-source framework to quantify and propagate intent as a standard because complex B2B procurement cannot work without that trust. Endpoint security CTOs, in separate briefings, have taken the opposite position: they are betting on probability rather than intent inference for production workloads. Chang explained that models as trained today cannot reliably derive intent from a prompt, which is why deterministic controls and behavioral proxies remain necessary. Ceylan agreed that both are required. "If you're not doing anything deterministic, you're really relying heavily on that intent, and I haven't seen programs that are there yet," she said.

Ceylan's most direct observation came from Box's own deployment experience. The company placed agents in its security operations center about a year ago, starting with human approval required for every action. Trust built quickly enough that analysts shifted into monitoring mode. Then the agent made one mistake, and every bit of accumulated trust vanished. "They had to start all over again," she said. "So I think that that monitoring piece is so important. Even if you're not gonna have a human in the loop, things change, models change, and we can't control how the models change and interpret things." The story lands because enterprise agentic security is not a problem that gets solved and stays solved. Models change, permissions drift, and adversaries adapt across multi-turn conversations that snapshot tests never capture.

The evidence boundary is worth stating explicitly. The 88.3% figure comes from Cisco's adversarial evaluation against 15 closed and proprietary flagship models. The survey data covers 107 enterprise respondents who self-selected into a technology-focused publication. Neither sample is random, and neither guarantees that the same failure rates appear in production deployments with different threat profiles, access patterns, or control layers. What the numbers do establish is that multi-turn attack success rates differ meaningfully from single-turn results and that models which pass single-turn evaluations fail in extended conversations. For the 82% of enterprises relying on provider-native controls as their primary security layer, that gap is the operational risk. Provider-native controls were not designed to catch multi-turn attack patterns, and the Cisco data suggests they do not.

Subscribe to AI Enthusiast Log

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe