contradish research · September 4, 2026

Finding the collapse
is not the same as fixing it

contradish distinguish already tells you when a model collapses two genuinely different situations into the same answer (a healthy adult's ibuprofen dose and a renal patient's are not the same question), and a model that answers them identically has lost the distinction. Knowing that isn't the same as knowing why, or knowing whether a fix actually holds once the world stops handing the model a clean, fully-stated fact. Two new modules close that gap.

The problem these fix

A distinction that collapses looks like agreement.

Ask a model about a healthy adult's ibuprofen dose, then ask about a renal patient's. If it gives the same answer to both, it has collapsed a distinction that should have held: two situations that require different handling produced identical output. contradish distinguish probes a set of these pairs under pressure framing and reports which distinctions hold and which collapse, using the same probing methodology CAI Strain uses, aimed at a different failure.

That report used to be the end of the story. It told you a distinction collapsed. It did not tell you why, and it gave you no way to check whether a proposed fix actually worked, or only looked like it worked.

contradish/resolution.py

The resolution operator finds the hidden variable, then proves it.

A collapsed distinction almost always has a disambiguating variable behind it: a fact the model isn't being told, or isn't weighing, that would make both sides correct at once. Prior work on hidden context in RLHF shows this kind of missing variable is a known cause of inconsistent model behavior. The resolution operator doesn't stop at naming that as a hypothesis. It runs a four-step, falsifiable procedure and only reports success when the fix is measured to work.

1
Propose
An LLM proposes candidate hidden variables that could explain why both sides of the distinction look identical to the model right now, and states, for each, what the situation looks like under either pole.
2
Probe
Each candidate gets a real causal flip test: assert one pole directly, check the model gives the first commitment; assert the opposite pole, check the answer flips to the second commitment. This produces direct_match_rate and flip_rate, not a plausibility guess.
3
Score
The winning candidate is the one with the highest causal_effect_size: the mean of direct-match and flip rate. A candidate that sounds right but doesn't move the model's answer scores low and loses.
4
Validate
The winning candidate becomes a one-line system-prompt patch. The distinction is re-probed with fresh, unconditioned questions and the patch in place. Only if the hold rate measurably rises is the distinction reported resolved.
Why step 4 exists
causal_effect_size = mean(direct_match_rate, flip_rate)
resolved = causal_effect_size ≥ threshold AND validated_hold_rate > baseline_hold_rate
A candidate can pass the flip test perfectly (the model answers correctly whenever the pole is stated directly in the question) and still fail step 4, if surfacing that same fact as system-prompt guidance does nothing. That case is reported honestly as not resolved, not shipped as a guess dressed up as a fix. This is the specific failure mode the operator exists to catch, and it's covered by its own test in the suite.
$ contradish distinguish --domain medication --app mymodule:my_app --resolve

contradish/rate_distortion.py

A resolved distinction can still be brittle.

Real deployments almost never hand a model the disambiguating variable as a plain, fully-confirmed fact. They hand it a chart note that "suggests" something, a support ticket where the customer "believes" they're covered, a form field left blank. The question that matters in production isn't whether a fix works when the model is told everything. It's whether accuracy rises smoothly as information about that variable goes from nothing to certain, or stays stuck at the collapsed baseline until the last possible sentence and then jumps.

The rate-distortion curve answers that directly. It takes a resolution operator's winning candidate and re-probes the distinction across five rungs of stated certainty, from no information at all to the pole asserted as plain fact, and measures accuracy at each rung.

Graded
Trustworthy under partial information
Accuracy climbs with certainty and correlates with it. A hedge helps a little; a stronger hedge helps more. This is the model you can hand a chart note that only "suggests" something.
Threshold
Brittle
Accuracy stays at baseline through every hedge and only recovers once the fact is stated with full certainty. Reliable with a confirmed fact, worthless with anything less, and most real inputs are something less.
Insensitive
Doesn't actually help
Accuracy never meaningfully improves, even at full certainty. The honest negative result: the rate-distortion sibling of "not resolved."
How the shape is decided
correlation = Spearman(certainty rung, mean accuracy)  (computed from scratch, average-rank tie handling, no external dependency)
"insensitive" if (full − baseline accuracy) < threshold; else "graded" if correlation ≥ threshold; else "threshold"
The report always includes the correlation coefficient, not just the label, so the classification is checkable rather than eyeballed, the same standard the underlying research held itself to.
$ contradish distinguish --domain medication --app mymodule:my_app --resolve --rate-distortion

Positioning

Most tools stop at detection. This proposes and grades the fix.

Pass/fail eval & guardrail tools
Report that a case failed
Score a single run
Leave finding the cause to you
Treat "with more context it would work" as untested
contradish
Proposes the disambiguating variable
Proves it's causal with a real flip test
Only ships a fix that's measured to work
Reports whether that fix degrades gracefully or falls off a cliff under partial information

As far as our own research into current eval and guardrail tooling found, none of the tools in this space (pass/fail evals, prompt-compliance checkers, or manual-scoring guardrails) report a graded information curve for constraint resolution. They report a score or a verdict, not a curve you can check.

Honest scope

What this does not establish yet.

The rate-distortion curve is a black-box behavioral analog, not a re-proof

It's the behavioral sibling of an internal, weight-level finding: on a small hand-built transformer, collateral damage to an unrelated constraint during narrow fine-tuning scaled monotonically with the bits of missing information about a hidden disambiguating variable, across a graded noisy channel (pooled Spearman r=+0.96 vs. noise level, over 20 seeds per setting). That result measured real damage under gradient descent on weights we built and could inspect directly.

The rate-distortion curve does not re-run that experiment and does not prove it generalizes to how frontier models are actually fine-tuned. What transfers is the shape of the claim (graded degradation versus a cliff) tested behaviorally, through outputs alone, on a real model, using linguistic hedging in place of an injected noisy channel. We think that's still a useful and currently unclaimed measurement. We don't think it's the same experiment, and we say so in the module's own docstring, not just here.

Find the collapse.
Then check whether the fix actually holds.
Run contradish distinguish --resolve →

pip3 install contradish  ·  source on GitHub