A real pilot found the model differentiated 6 of 7 paired real-world scenarios under pressure -- but moved in the correct direction only once. Pooled directional fidelity: 0.17.

An outside read on contradish's directional-fidelity pilot: noticing that context changed isn't the same as responding to it correctly, framed as a governance question for any organization whose workflows depend on systems it doesn't fully control.

Finding a collapsed distinction is not the same as fixing it. How contradish proposes and causally validates the hidden variable behind a collapse.

A stable, citable taxonomy of how reasoning systems fail under reframing: CAI failure, drift, rigidity, silent confident drift, and more, each with a stable ID.

The AI reasoning-consistency gap most evals miss, where the same question asked a different way gets a different answer.

The third failure state binary testing never catches. Collapse Rate 1.000 on Groq financial and Cohere financial: every apparent compliance was fake.

How CAI-Bench measures semantic invariance and paraphrase robustness in AI policy consistency, with a real example and open questions.

Courts, insurers, and regulators have all treated AI policy inconsistency as a real, priced risk since 2024. Four verifiable cases: Air Canada, Cursor, Lloyd's, FINRA.

The two-sided metric contradish uses to score AI judgment: drift, rigidity, and a record kept over time.

We built a real symptom-triage prototype, ran contradish on it, found real inconsistencies, and fixed what we could. Raw numbers, unedited.