GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Cisco Tests 15 Frontier Models: Multi-Turn Attack Success Reaches 88% — Every Model Fails

Single-turn safety benchmarks have always been a floor, not a ceiling. Cisco’s AI Threat Research team published evidence on May 27 that the gap between floor and ceiling is wider than procurement teams assume — and consistent across every major frontier lab.

The study evaluated 15 closed proprietary flagship models from OpenAI, Anthropic, Google, Amazon, and xAI in a paired-regime evaluation: 30,090 single-turn adversarial prompts and 6,986 multi-turn attacks against the same cohort.

The headline numbers:

  • Single-turn attack success rate: 2.19% to 64.91% across 15 models
  • Multi-turn attack success rate: 7.89% to 88.30% across the same 15 models
  • Models with zero multi-turn success: 0

What Changes in Multi-Turn

Real adversaries do not send one adversarial prompt and give up. They reframe refusals, decompose requests across turns, adopt personas, and escalate gradually. Single-turn benchmarks cannot observe any of that behaviour. Cisco’s design explicitly modelled the iterative attacker.

The two evaluation regimes produce different model rankings, different failure maps, and different tail-risk pictures. A model that scores well on a single-turn benchmark may rank materially worse under multi-turn pressure — and vice versa. The orderings do not correlate cleanly.

Cisco’s earlier study of eight open-weight LLMs found multi-turn ASR 2x to 10x higher than single-turn baselines, with Mistral Large-2 reaching 92.78% under multi-turn attack. The new proprietary study extends that finding: multi-turn vulnerability is a structural property of the frontier, not an artefact of open-weight alignment choices. Whether weights are public or proprietary, the iterative attack surface remains open.

Three Practical Recommendations

Cisco set out three concrete asks for labs and enterprise buyers:

  1. Labs: Publish attack success rates broken down by strategy family on every model release, not aggregate pass/fail rates
  2. Enterprises: Gate deployment decisions on regressions in specific content categories, using a 3-percentage-point threshold as the trigger for manual review
  3. Flagging: Any model with a cross-regime gap larger than 15 percentage points — single-turn to multi-turn — warrants manual security review before production deployment

The underlying conclusion: “If no base model is iteratively safe, the security perimeter has to move outside the model.” Runtime guardrails, monitoring, red-teaming, and application-layer policies become load-bearing for any deployment at risk.

Why This Matters for Enterprise Deployments

Most enterprise AI procurement decisions reference published safety evaluations. Those evaluations overwhelmingly use single-turn methodology. Cisco’s data indicates the resulting risk profile is systematically optimistic — not by a small margin, but by the difference between a 2% ASR and an 88% ASR on the same model.

The research feeds Cisco’s AI Defense product and its LLM Security Leaderboard, which publishes adversarial evaluation signals against leading models. The full study covers all 15 models individually; the published version identifies labs without disclosing per-model attack success rates to avoid providing a targeting map.