Same question, asked different ways. 0.00 held every time, 0.60+ is unreliable · 20 domains, 8 phrasings each.
| # | Model | Consistency score (0–1)Score | Domains | Evaluated |
|---|---|---|---|---|
| 1 |
Mistral Small
|
17 | 2026-08-30 | |
| 2 |
DeepSeek
|
20 | 2026-08-21 | |
| 3 |
Claude Sonnet 4.6
|
20 | 2026-04-17 | |
| 4 |
Neurogen Biomarking
|
20 | 2026-09-16 | |
| 5 |
GPT-4o
|
12 | 2026-04-17 | |
| 6 |
Spring Health
|
8 | 2026-08-27 | |
| 7 |
LegalWiz
LegalWiz
|
8 | 2026-08-26 | |
| 8 |
TriageWell
J Health
|
8 | 2026-08-27 | |
| 9 |
GPT-4o-mini
|
9 | 2026-07-29 | |
| 10 |
Grok
|
8 | 2026-08-21 | |
| 11 |
Ash
|
8 | 2026-07-20 | |
| 12 |
PerceptronML
|
8 | 2026-08-28 | |
| · |
Llama 3.3 70B
|
in progress | ||
| · |
Gemini 1.5 Pro
|
in progress | ||
| · |
Mistral Large
|
in progress |
Want your model on this leaderboard?
Sign in →