live demo

Same question. Different phrasing.
Watch the policy drop.

Pick a scenario below. Same question, asked two ways. Watch the answer change.

Neutral question
Asked plainly
same policy
Asked differently
Why it matters:
See all 6 phrasings tested, the system prompt, and the fix
System prompt  CLASH  ESCAPE
CAI Strain
·
▸ Finding
Repair patch
The scenarios above are simulated, canned examples. The section below is real: it calls your actual endpoint.
Zero-install trial · no account

Pick your AI provider, paste your key. We'll ask it one question 4 different ways and show you exactly where the answer changes.

Sent straight to your endpoint for this one request. Never stored.
Advanced options (usually not needed)

1 of 13 test cases · 4 of 8 phrasings · up to 5 runs/day. Sign in for the full 104-call audit.

What's your model's score?

The trial above ran 4 phrasings of one test case against your endpoint. The full audit covers all 104 tests across 13 real-world scenarios in 6 domains, with independent LLM judging and cross-variant consistency scoring, and delivers a full report with every failure case, every phrasing that caused it, and a fix for each one.

Get your CDR → View sample report