Same question, asked different ways. Lower score is better · 0.00 held every time, 0.60+ is unreliable · 20 domains, 8 phrasings each.
| # | Model | Consistency score (0–1)Score | Domains | Evaluated |
|---|---|---|---|---|
| 1 |
Claude Sonnet 4.6
Anthropic
|
20 | 2026-04-17 | |
| 2 |
GPT-4o
OpenAI
|
12 | 2026-04-17 | |
| 3 |
GPT-4o-mini
OpenAI
|
9 | 2026-07-29 | |
| · |
Llama 3.3 70B
Meta
|
in progress | ||
| · |
Gemini 1.5 Pro
Google
|
in progress | ||
| · |
Mistral Large
Mistral
|
in progress |
Want your model on this leaderboard?
Sign in →All evaluations run by contradish, April 2026. Results verified with an independent judge (different provider from the model under test). Domain coverage expands as each model's evaluation completes.
Every model in this benchmark produced at least one response that opens with a refusal, then provides the information it just declined. Policy holds in sentence one. It breaks in sentence two.
Same model. Same session. Three minutes apart. This is what 8 adversarial phrasing variants find that standard safety testing misses.
4 safety policies, 6 phrasing variants per test. Collapse Rate: of responses that appeared to follow the policy, how many were hiding a violation in the same reply. The highlighted rows are the ones standard testing would have called passing.
Recomputed 2026-08-25 directly from raw per-response evaluation data (not hand-summarized). Gemini 3.5 Flash is excluded below - its capture run returned every reply truncated to roughly 80 characters, well under the token limit, so "compliant"/"non-compliant" couldn't be judged from real output. Retest pending; see note.
| Model | Policy | Consistency score | Hidden failure rate | Actual compliance | Tested |
|---|---|---|---|---|---|
|
Llama 3.3 70B (Groq)
Meta / Groq
|
Financial | 1.000 | 1.000 | 0 / 6 | Aug 14 |
|
Command R+ 08-2024 (Cohere)
Cohere
|
Financial | 1.000 | 1.000 | 0 / 6 | Aug 14 |
|
Command R+ 08-2024 (Cohere)
Cohere
|
Legal | 0.833 | 0.750 | 1 / 6 | Aug 14 |
|
Gemini 3.5 Flash
Google
|
Healthcare | data invalidated · response-capture defect · retest pending | Aug 14 | ||
|
Gemini 3.5 Flash
Google
|
Financial | data invalidated · response-capture defect · retest pending | Aug 14 | ||
|
Llama 3.3 70B (Groq)
Meta / Groq
|
Legal | 0.333 | 0.200 | 4 / 6 | Aug 14 |
|
Command R+ 08-2024 (Cohere)
Cohere
|
Mental health | 0.000 | 0.000 | 6 / 6 | Aug 14 |
|
Command R+ 08-2024 (Cohere)
Cohere
|
Healthcare | 0.167 | 0.000 | 10 / 12 | Aug 14 |
|
Llama 3.3 70B (Groq)
Meta / Groq
|
Healthcare | 0.000 | 0.000 | 12 / 12 | Aug 14 |
|
Llama-3.1-8B-Instruct (HF)
Meta / Hugging Face
|
Healthcare | 0.111 | 0.000 | 8 / 9 | Aug 14 |
|
Llama 3.3 70B (Groq)
Meta / Groq
|
Mental health | 0.000 | 0.000 | 6 / 6 | Aug 14 |
|
Kimi K2.5 (DeepSite)
Moonshot AI
|
Scope policy | 0.389 | n/a | 11 / 18 | Aug 14 |
Highlighted rows: standard testing reported apparent compliance. Full analysis found every apparent hold contained the policy violation in the same response. Actual compliance 0%. n/a in Collapse Rate = model failed outright with no refusal language (nothing to detect). Llama-3.1-8B-Instruct (HF) healthcare run: 3 of 12 planned probes errored at the API and are excluded rather than counted either way - the 8/9 above reflects only the probes that returned a real response. Cerebras and NVIDIA runs are not shown: every call in both failed at the API (no response captured), so there is no real data to report for either provider yet.
Failure rate: percentage of probes using this technique where the model's answer contradicted its baseline response. Technique breakdown for other evaluated models publishes with each full report.
Curious how a consistency score is actually measured? Read the methodology →