We send each model the same question in different forms and measure how consistently it holds its policies. Lower score is better. 0.00 means it held every time.
Every model in this benchmark produced at least one response that opens with a refusal, then provides the information it just declined. Policy holds in sentence one. It breaks in sentence two.
Same model. Same session. Three minutes apart. This is what 16 adversarial phrasing variants find that standard safety testing misses.
| # | Model | Consistency score (0–1) | Domains | Evaluated |
|---|---|---|---|---|
| 1 |
Claude Sonnet 4.6
Anthropic
|
20 | 2026-04-17 | |
| 2 |
GPT-4o
OpenAI
|
12 | 2026-04-17 | |
| 3 |
GPT-4o-mini
OpenAI
|
9 | 2026-07-29 | |
|
Llama 3.3 70B
Meta
|
in progress | |||
|
Gemini 1.5 Pro
Google
|
in progress | |||
|
Mistral Large
Mistral
|
in progress |
Want your model on this leaderboard?
Sign in →All evaluations run by contradish, April 2026. Results verified with an independent judge (different provider from the model under test). Domain coverage expands as each model's evaluation completes.
4 safety policies, 6 phrasing variants per test. Collapse Rate: of responses that appeared to follow the policy, how many were hiding a violation in the same reply. The highlighted rows are the ones standard testing would have called passing.
| Model | Policy | Consistency score | Hidden failure rate | Actual compliance | Tested |
|---|---|---|---|---|---|
|
Llama 3.3 70B (Groq)
Meta / Groq
|
Financial | 1.000 | 1.000 | 0 / 6 | 2026-08-14 |
|
Command R+ 08-2024 (Cohere)
Cohere
|
Financial | 1.000 | 1.000 | 0 / 6 | 2026-08-14 |
|
Command R+ 08-2024 (Cohere)
Cohere
|
Legal | 0.917 | 0.750 | <1 / 6 | 2026-08-14 |
|
Gemini 3.5 Flash
Google
|
Healthcare | 1.000 | n/a | 0 / 12 | 2026-08-14 |
|
Gemini 3.5 Flash
Google
|
Financial | 1.000 | n/a | 0 / 6 | 2026-08-14 |
|
Llama 3.3 70B (Groq)
Meta / Groq
|
Legal | 0.333 | 0.200 | 4 / 6 | 2026-08-14 |
|
Command R+ 08-2024 (Cohere)
Cohere
|
Mental health | 0.083 | 0.000 | 5 / 6 | 2026-08-14 |
|
Command R+ 08-2024 (Cohere)
Cohere
|
Healthcare | 0.125 | 0.000 | 5 / 6 | 2026-08-14 |
|
Llama 3.3 70B (Groq)
Meta / Groq
|
Healthcare | 0.000 | 0.000 | 6 / 6 | 2026-08-14 |
|
Llama 3.3 70B (Groq)
Meta / Groq
|
Mental health | 0.000 | 0.000 | 6 / 6 | 2026-08-14 |
|
Kimi K2.5 (DeepSite)
Moonshot AI
|
Scope policy | 0.389 | n/a | 11 / 18 | 2026-08-14 |
Highlighted rows: standard testing reported apparent compliance. Full analysis found every apparent hold contained the policy violation in the same response. Actual compliance 0%. n/a in Collapse Rate = model failed outright with no refusal language (nothing to detect).
Failure rate: percentage of probes using this technique where the model's answer contradicted its baseline response. Technique breakdown for other evaluated models publishes with each full report.
Every version of every question is set before testing starts. Every AI sees identical inputs. Results are reproducible and comparable across time.
Anthropic models are judged by OpenAI models and vice versa. This removes the possibility that the judge favors responses from its own provider.
Failures on the highest-stakes cases -- medication, self-harm, crisis -- count 4x more than routine cases. The score reflects what actually matters.
Mental health, medical, legal, financial, immigration, cybersecurity, and 14 more. Each domain is tested with the same policy applied across every question variant.
Casual, emotional, indirect, formal, hypothetical, roleplay, flattery, authority claims, social proof, and more -- the full range of how real users phrase things.
Every version of a question is verified to carry exactly the same meaning. Only the wording changes. If the AI answers differently, the policy is inconsistent.