AI Model Risk Index
The Industry Benchmark for AI Agent Security
The b³ Benchmark, built by Lakera's research team, is the most comprehensive independent evaluation of how backbone LLMs perform under real-world adversarial attack. Powered by hundreds of thousands of crowdsourced attacks across today's leading models, it gives security and AI leaders the data they need to make informed model selection decisions.
AI agents inherit the security properties of their backbone LLM, and the model you choose directly impacts your risk posture. The b³ Benchmark isolates and measures backbone LLM security using threat snapshots: a framework that captures real-world attack scenarios across agentic applications.
The rankings below reflect aggregated vulnerability scores across all threat categories and defense levels.
Model Rankings
| Rank | Model | Risk Score |
|---|---|---|
| 1 | Claude Sonnet 4 | 0.30717 |
| 2 | Claude 3.7 Sonnet | 0.395582 |
| 3 | GPT-4o | 0.575943 |
| 4 | Gemini 1.5 Pro | 0.671179 |
| 5 | Gemini 1.5 Flash | 0.737874 |
| 6 | GPT-4.1 | 0.748453 |
| 7 | Claude 3 Haiku | 0.774942 |
| 8 | Meta Llama 3.3 70B Instruct | 0.801155 |
| 9 | Meta Llama 3.1 8B Instruct | 0.80151 |
| 10 | Meta Llama 4 Scout | 0.811721 |
| 11 | Gemma 3 12B | 0.818806 |
| 12 | DeepSeek-V3 | 0.888484 |
| 13 | Gemini 2.0 Flash | 0.900635 |
| 14 | Meta Llama 4 Maverick | 0.913013 |
| 15 | Claude 4 Sonnet | 23.86 |
| 16 | Claude 3.7 Sonnet | 31.54 |
| 17 | GPT-4o | 60.04 |
| 18 | GPT-4o-mini | 64.23 |
| 19 | GPT-4.1 | 71.62 |
| 20 | Gemini 1.5 Pro | 72.64 |
Why the b³ Benchmark Matters
AI agents inherit the security properties of their backbone LLM. The b³ Benchmark is designed for security teams and AI leaders who need real-world visibility.
Why This Benchmark Matters:
- Highlights comparative resilience to inform model selection and risk management decisions
- Benchmarks against real world attack techniques like prompt injections, jailbreaks, data exfiltration, and indirect attack vectors
- Quantifies exploitability across key threat categories
- Provides up-to-date, independent benchmarks for security and AI leaders
Agentic Threat Coverage
Evaluates models across realistic agentic threat scenarios, covering the full spectrum of how LLMs are actually deployed today.
Real Attacks, Not Synthetic Prompts
Attacks were collected through large-scale gamified crowdsourcing, where hundreds of participants competed to break AI agents.
Fine-Grained Risk Insights
Gain insights by attack type (direct vs. indirect), task type (instruction override, tool invocation, context extraction), and defense level to find the right model for your specific use case.
What Sets it Apart
Full Attack Categorization
Comprehensive attack coverage covering six major attack task types spanning direct and indirect attacks, tool manipulation, data exfiltration, and denial of service.
Multi-Level Defense Evaluation
Tests models across defense configurations. Every model is evaluated under three defense levels: minimal system prompt constraints, hardened system prompts with extended context, and LLM-as-judge self-defense.
Crowdsourced Attack Quality
The benchmark attacks were selected from hundreds of thousands of human-generated attempts, representing less than 1% of total attack data.
Go Beyond the Benchmark with AI Red Teaming
The b³ Benchmark tells you which models are most resilient. The AI Red Teaming platform tells you whether your AI system is secure.