contradish research · September 4, 2026
contradish distinguish already tells you when a model collapses two genuinely different situations into the same answer (a healthy adult's ibuprofen dose and a renal patient's are not the same question), and a model that answers them identically has lost the distinction. Knowing that isn't the same as knowing why, or knowing whether a fix actually holds once the world stops handing the model a clean, fully-stated fact. Two new modules close that gap.
The problem these fix
Ask a model about a healthy adult's ibuprofen dose, then ask about a renal patient's. If it gives the same answer to both, it has collapsed a distinction that should have held: two situations that require different handling produced identical output. contradish distinguish probes a set of these pairs under pressure framing and reports which distinctions hold and which collapse, using the same probing methodology CAI Strain uses, aimed at a different failure.
That report used to be the end of the story. It told you a distinction collapsed. It did not tell you why, and it gave you no way to check whether a proposed fix actually worked, or only looked like it worked.
contradish/resolution.py
A collapsed distinction almost always has a disambiguating variable behind it: a fact the model isn't being told, or isn't weighing, that would make both sides correct at once. Prior work on hidden context in RLHF shows this kind of missing variable is a known cause of inconsistent model behavior. The resolution operator doesn't stop at naming that as a hypothesis. It runs a four-step, falsifiable procedure and only reports success when the fix is measured to work.
direct_match_rate and flip_rate, not a plausibility guess.causal_effect_size: the mean of direct-match and flip rate. A candidate that sounds right but doesn't move the model's answer scores low and loses.contradish/rate_distortion.py
Real deployments almost never hand a model the disambiguating variable as a plain, fully-confirmed fact. They hand it a chart note that "suggests" something, a support ticket where the customer "believes" they're covered, a form field left blank. The question that matters in production isn't whether a fix works when the model is told everything. It's whether accuracy rises smoothly as information about that variable goes from nothing to certain, or stays stuck at the collapsed baseline until the last possible sentence and then jumps.
The rate-distortion curve answers that directly. It takes a resolution operator's winning candidate and re-probes the distinction across five rungs of stated certainty, from no information at all to the pole asserted as plain fact, and measures accuracy at each rung.
Positioning
As far as our own research into current eval and guardrail tooling found, none of the tools in this space (pass/fail evals, prompt-compliance checkers, or manual-scoring guardrails) report a graded information curve for constraint resolution. They report a score or a verdict, not a curve you can check.
Honest scope
It's the behavioral sibling of an internal, weight-level finding: on a small hand-built transformer, collateral damage to an unrelated constraint during narrow fine-tuning scaled monotonically with the bits of missing information about a hidden disambiguating variable, across a graded noisy channel (pooled Spearman r=+0.96 vs. noise level, over 20 seeds per setting). That result measured real damage under gradient descent on weights we built and could inspect directly.
The rate-distortion curve does not re-run that experiment and does not prove it generalizes to how frontier models are actually fine-tuned. What transfers is the shape of the claim (graded degradation versus a cliff) tested behaviorally, through outputs alone, on a real model, using linguistic hedging in place of an injected noisy channel. We think that's still a useful and currently unclaimed measurement. We don't think it's the same experiment, and we say so in the module's own docstring, not just here.