pip3 install contradish
Needs one API key for whichever provider you use: ANTHROPIC_API_KEY, OPENAI_API_KEY, or your LiteLLM provider's own key.
export ANTHROPIC_API_KEY=sk-ant-...
contradish # demo policy pack, no config
contradish benchmark --model claude-sonnet-4-6 # full 20-domain benchmark
The bare command runs twelve cases against a demo policy in about thirty seconds and prints one CAI Strain number. No app code, no test file.
All 18 subcommands. Run contradish <command> --help for full flags.
| Core loop | |
| benchmark | Run CAI-Bench against any model. No app code needed. |
| diagnose | Turn drift cases into a repair package: counterfactuals, prompt fixes, fine-tuning JSONL. |
| improve | End-to-end loop: benchmark, diagnose, rewrite the prompt, re-verify, in one command. |
| monitor | Find drift hotspots in real production conversation logs. |
| Deeper analysis | |
| distinguish | Measure the complementary failure: does the model collapse two genuinely different situations into the same answer? |
| fairness | Audit for disparate treatment across disclosed protected attributes. |
| prompt | Static analysis of a system prompt for internal contradictions. No model call. |
| judge-floor | Measure the judge model's own consistency, the floor every Strain score is bounded by. |
| analyze | Contradiction-forced truth extraction. No API key needed to test your own model. |
| calibrate | Combine saved results into one Calibration Score. Reads existing files, no API calls. |
| findings | Re-mine a saved result for root causes, rigidity, and severity skew. No API calls. |
| Production | |
| replay | Replay logged conversation transcripts for cross-turn self-contradictions. |
| reconcile | Grade a benchmark report against a replay report: surface the validity gap. |
| ledger | Inspect, verify, or anchor the hash-chained ledger that monitor writes to. |
| Setup & utility | |
| init | Interactive setup. Writes .contradish.yaml and an optional GitHub Actions workflow. |
| run | Run manual test cases from a YAML or JSON file. |
| compare | Compare baseline vs. candidate for CAI regression. CI/CD gate. |
| schema | Inspect and validate the published JSON Schema interchange formats. |
Your app is any callable with the shape str -> str: a chatbot, a RAG pipeline, an agent.
from contradish import Suite, TestCase
suite = Suite(app=my_llm_function)
suite.add(TestCase(input="Can I get a refund after 45 days?"))
report = suite.run()
print(report.judgment_strain) # 0.0-1.0, lower is better
Or build a suite straight from a system prompt with Suite.from_prompt(...), or from a built-in policy pack with Suite.from_policy("ecommerce", app=my_app).