07 August 2026

Self-certification is not enough when the deployer holds public power: reading the OECD's AI incident monitor from the clinical side

The OECD runs an AI Incidents and Hazards Monitor that documents reported AI harms and sorts each one by sector, and government, security, and defence is a category that keeps recurring rather than sitting at the margin. The monitor’s recent public-sector entries are concrete: a city government’s AI-generated promotional materials included the wrong national symbols and had to be withdrawn after public distrust, and a wave of AI-generated legal advice, unsound because the tools lacked expertise in social law, strained a country’s social courts with frivolous claims. These are not frontier-model doomsday cases. They are ordinary public administration using AI and getting a governance question wrong.

Read at the scale of a national government, this is a public-administration and oversight question. Read from where I actually work, on patient-facing clinical AI, it sharpens something I have to keep honest about in my own practice: who is allowed to certify that a system’s controls work, and when is that self-certification no longer enough.

Why this lands on my own practice, uncomfortably. In the clinical AI I build and govern, I am frequently both the builder and a large part of the accountability. I decide the validated scope, I write the refusal and escalation rules, and I am close to the evidence that says the control held. That concentration of roles is workable in my setting for a specific reason: a clinician sits between the system and the patient, and the patient is present, identifiable, and able to question, refuse, or escalate. Self-certification by the people who built the system is partially checked by a human who owes the patient a duty and by the patient’s own presence in the room. My from-promise-to-evidence note already argues that a control is only governance when there is a test that shows it works and a record that shows it held. What I had not stated is who that record has to convince.

The public-power case removes the checks I had been quietly relying on. When AI is used to help decide a welfare eligibility, a licence, a benefit, or an enforcement action, the person on the receiving end is usually not in the room, did not consent, and often cannot tell that AI was involved at all. There is frequently no clinician-equivalent mediating the decision and owing them a duty. And the same authority that deploys the system is also the one certifying that its controls are adequate. Every safeguard I lean on in the clinical setting, a present user, a mediating professional, a duty of care, is absent, and the one remaining check, self-certification by the deployer, is now being asked to carry all the weight. That is the setup the OECD’s public-sector incidents keep illustrating: a public body deployed an AI tool, judged its own controls sufficient, and the people affected found out only after it went wrong. Consider a benefits agency that uses a model to flag applications for review and certifies internally that a caseworker checks every adverse flag. If that check quietly degrades into rubber-stamping, nobody outside the agency can see it until denied applicants notice a pattern, and by then the harm has landed on people who never knew the model existed. That is a constructed illustration built on the mechanism these incidents show, not a reported case.

What transfers, and the axis the public-power case adds. The discipline transfers cleanly: decide what the system may do, name who is accountable, and require a test and a record before it is trusted. What changes is who the record must satisfy. In my clinical world, the deployer can hold much of the accountability because other checks are present. When the deployer holds coercive public power over absent citizens, the evidence that controls work has to be produced for, and testable by, someone outside the deploying authority, and the affected person needs a route to contest the outcome. Stated as the instrument I would reach for, public-power AI governance needs a layer my clinical checklist never had to specify:

What I am not claiming. I do not govern public administration, and I will not pretend a clinical risk register scales to a welfare system. What I can say precisely is where my own frame was incomplete: it assumed a present user and a mediating professional, and quietly let the builder be much of the accountability. The OECD’s public-sector incidents are a reminder that when neither the user’s presence nor a mediating duty exists, and the deployer holds public power, self-certification is the weakest possible control, and external, independent accountability is not a nice-to-have but the thing that makes the governance real. The same from-promise-to-evidence discipline, pointed at public power, has to answer one more question than it does at the bedside: proven to whom.

Source: OECD, AI Incidents and Hazards Monitor (AIM), which documents reported AI incidents and hazards and classifies them by sector, including government, security, and defence; the public-sector examples referenced here are drawn from its logged entries. The monitor notes its contents should not be reported as the official views of the OECD or its member countries. This note is a governance reflection and not legal or public-policy advice.