09 July 2026
Outcomes improved after we deployed AI. That is not proof the AI improved them
It is a familiar claim: after we deployed the AI, diagnosis got faster, or mortality fell, or clinicians got more efficient. The claim may even be true. But “things improved after we deployed it” is not the same as “it improved things,” and the gap between those two sentences is where governance has to live.
When a system goes live, many things change at once. The workflow is often redesigned. Clinicians become more practised. The patient mix shifts. Attention and resources flow to the new initiative. Any of these can move the outcome, and a before-versus-after comparison cannot tell you which one did.
This matters because credit drives decisions. If an AI is credited with an improvement it did not cause, it gets scaled, funded, and trusted on false evidence, while the real driver goes unmanaged. The harm is not just an overstated result; it is a misallocation of trust.
So a governance-minded reader treats “improved after deployment” as a hypothesis, not a conclusion. The question is not whether the number moved, but whether the AI moved it, rather than the reorganisation, the learning curve, or the changing population that arrived at the same time.
The questions to ask before accepting a causal claim:
- What else changed at the same time as the AI?
- Is there a comparison that isolates the AI’s effect, rather than a simple before-and-after?
- Who benefits from the causal story being true, and did that shape how the evaluation was framed?
Evidence, not narrative, is what lets a system be trusted with patients. “Better after” is a story. “Better because” is a claim that has to be earned.