Public Leaderboard

Gauntlet

Frontier models run the gauntlet of prompt-injection attacks — ranked by how well they resist direct jailbreaks and indirect, tool-based injection. Higher robustness is better.

Robustness vs efficiencyHigher and further right is better. Each line links one model across reasoning efforts (L/M/H).
robust & efficient ↗
ModelRobustness
1anthropic
claude-fable-5medium
anthropic
100%
2anthropic
claude-fable-5low
anthropic
99%
3anthropic
claude-fable-5high
anthropic
99%
4anthropic
claude-opus-4.8high
anthropic
94%
5anthropic
claude-opus-5low
anthropic
94%
6anthropic
claude-opus-4.8low
anthropic
93%
7anthropic
claude-opus-4.8medium
anthropic
93%
8anthropic
claude-opus-5medium
anthropic
91%
9anthropic
claude-opus-5high
anthropic
89%
10openai
gpt-5.5high
openai
75%
11x-ai
grok-4.3
x-ai
73%
12openai
gpt-5.5medium
openai
70%
13anthropic
claude-sonnet-4.6high
anthropic
68%
14openai
gpt-5.5low
openai
58%
How it works
Two tracks, one robustness score

Every model faces the same frozen attack suite across two escalating threat models. Robustness is 100 − attack-success-rate, so higher means harder to break.

Track 0197 scenarios
Indirect injection
Hand-built agentic scenarios where a malicious instruction hidden in tool output — a document, email, or API response — tries to hijack the model while it works a legitimate task.
Track 0264 probes
Agentic tool-hijacking
ZeroLeaks' Sandbox attack corpus replayed against mock agent tools — tool-poisoning, authority spoofing, and protocol exploits — counting any dangerous tool call as a breach.

Robustness = 100 − attack-success-rate (ASR). Each response is scored by an LLM judge backed by success/failure indicators, so a model is only marked breached on a genuine rule violation — not on a refusal or an error. The chart plots robustness (%) against efficiency (cost per run or output tokens, reversed so more efficient is to the right); each line connects one model across reasoning efforts. Cost uses standard list prices. GPT-5.5 and Grok 4.3 are served via Azure. Latest run: suite indirect-v6+agentic, July 27, 2026.

Ready to secure your
AI infrastructure?

Comprehensive vulnerability assessments backed by evidence and powered by our red team of specialized agents.