Why now · July 18, 2026

The inconsistency risk is already priced, and most teams haven't tested for it

“Semantic invariance testing” is a new term. The failure it measures is not new, and in the last two years it has become a legal, product, insurance, and regulatory line item, not a hypothetical. Four independent, verifiable cases, in order.

Courts
Feb 2024
Air Canada found liable for its own chatbot's words.
Insurers
May 2025
Lloyd's-backed policy launches to cover chatbot hallucination claims.
Regulators
2026
FINRA names hallucination a compliance risk for member firms.
LegalNovember 2022 · decided February 19, 2024
Moffatt v. Air Canada

Jake Moffatt asked Air Canada's website chatbot about bereavement fares. It told him he could apply for the discount retroactively, after booking. Air Canada's actual policy didn't allow that. When Moffatt was refused the refund, he took the airline to the British Columbia Civil Resolution Tribunal.

Air Canada argued the chatbot was “a separate legal entity” responsible for its own words. The tribunal rejected that outright, found negligent misrepresentation, and ordered Air Canada to pay $812.02 CAD. The finding that matters beyond the dollar amount: a company is legally bound by what its AI tells a customer, even when that answer contradicts the company's own written policy.

ProductApril 2025
Cursor's own support bot

This one didn't happen to an airline bolting on a chatbot. It happened to an AI-native developer tools company. Cursor's AI support bot, signing its replies “Sam,” told a user that Cursor was “designed to work with one device per subscription as a core security feature.” No such policy existed.

Users saw the answer, believed it, and cancelled. It took roughly three hours for a human at Cursor to correct the record on Reddit: “this is an incorrect response from a front-line AI support bot” and users are “free to use Cursor on multiple machines.” Three hours is a long time for a confidently wrong answer to sit in public.

InsuranceMay 2025
Lloyd's-backed hallucination coverage

Armilla, a Y Combinator-backed insurer, launched a Lloyd's-backed policy covering legal fees and compensation when a company's AI chatbot hallucinates or underperforms badly enough to trigger a claim. Coverage press explicitly cited Air Canada as the case the policy is built to protect against.

An insurance market only prices a risk it believes is real, recurring, and measurable. This one already is.

Regulatory2026
FINRA's Regulatory Oversight Report

FINRA's 2026 Regulatory Oversight Report names generative AI hallucination as a specific concern for member firms, flagging the risk of AI-generated content that sounds authoritative but doesn't hold up. For any firm using AI in a regulated, customer-facing capacity, this is no longer an engineering footnote. It's a named line item in a federal regulator's guidance.

What this means for testing

Every case above was found out reactively: a customer got burned, a screenshot went viral, a claim got filed. None of these companies knew they had a problem until it became one in public.

Asking the same policy question eight semantically equivalent ways, before a customer does, is a proactive version of exactly what surfaced each of these. Some of the failures above are precisely what this atlas calls CAI failure: a confident answer that would have looked different if asked another way. Others are closer to plain fabrication, the kind that paraphrase testing is also unusually good at catching, because a model that's inventing an answer rarely invents the exact same fabrication twice.

The category name is new. The cost of not testing for it is already on the record, four times over, in two years.

contradish.com / why now