09 July 2026

Outcomes improved after we deployed AI. That is not proof the AI improved them

It is a familiar claim: after we deployed the AI, diagnosis got faster, or mortality fell, or clinicians got more efficient. The claim may even be true. But “things improved after we deployed it” is not the same as “it improved things,” and the gap between those two sentences is where governance has to live.

When a system goes live, many things change at once. The workflow is often redesigned. Clinicians become more practised. The patient mix shifts. Attention and resources flow to the new initiative. Any of these can move the outcome, and a before-versus-after comparison cannot tell you which one did.

This matters because credit drives decisions. If an AI is credited with an improvement it did not cause, it gets scaled, funded, and trusted on false evidence, while the real driver goes unmanaged. The harm is not just an overstated result; it is a misallocation of trust.

So a governance-minded reader treats “improved after deployment” as a hypothesis, not a conclusion. The question is not whether the number moved, but whether the AI moved it, rather than the reorganisation, the learning curve, or the changing population that arrived at the same time.

The questions to ask before accepting a causal claim:

Evidence, not narrative, is what lets a system be trusted with patients. “Better after” is a story. “Better because” is a claim that has to be earned.