Ask the people responsible for AI deployment in any institution one question: when this system is wrong — and it will be wrong — who is accountable?
Not “what is your escalation protocol.”
Not “which vendor do you call.”
Not “where is the incident response runbook.”
Who. Is. Accountable.
In most institutions that have deployed AI in consequential roles, this question produces one of three responses. Some say “the system” — as if a system can be held accountable. Some say “the vendor” — as if accountability can be outsourced with a contract. And some say nothing at all, which is, in its own way, the most honest answer.
None of them are acceptable. And the fact that we have built an entire global AI governance industry without resolving this question tells us something important: we have been solving the wrong problem.
The Question the Industry Keeps Avoiding
The AI governance debate — from the EU AI Act to the proliferating frameworks emerging from governments, think tanks, and standard bodies across the world — has organized itself around a single concern: is the output good enough?
Accuracy rates. Hallucination benchmarks. Bias audits. Fairness metrics. Explainability scores. These are the instruments of AI governance as it is currently practiced, and they represent a genuine attempt to make AI systems more reliable.
But output quality is not governance. It is quality assurance. Governance is a different question entirely.
Governance asks: when this output is used to make a decision that affects a human being, who holds accountability for that decision — and is that accountability traceable, demonstrable, and enforceable?
A system that produces highly accurate outputs but deploys them without a traceable chain of human accountability is not a governed system. It is a trusted system — and trust without the ability to verify is precisely the condition that produces institutional failure, as I argued in an earlier piece on the cost of misplaced trust. The pattern is identical: we have delegated accountability outward — to the model, to the vendor, to the framework — rather than building it as internal architecture.
The EU AI Act’s Article 14, which mandates human oversight for high-risk AI systems, moves in the right direction. But as researchers have noted, there is no clear guidance on what constitutes meaningful human oversight — the standard exists on paper without a mechanism to verify that it is actually happening in practice. Requiring oversight and ensuring accountability are not the same thing.
Three Cases, One Pattern
Consider what actually happens when AI accountability fails in practice. The domain varies. The pattern does not.
In healthcare, AI systems are now embedded in clinical workflows — generating medical notes, supporting diagnostic decisions, recommending treatments. The efficiency gains are real. So are the risks. Malpractice claims involving AI tools rose 14% between 2022 and 2024, with the majority concentrated in diagnostic AI across radiology, cardiology, and oncology. In 2025, insufficient governance of AI was identified as the second most significant patient safety threat in healthcare settings globally.
But here is the question that these statistics do not answer: in each case where AI contributed to patient harm, was there a specific human being who had explicitly reviewed and approved the AI output before it was acted upon — and is there a verifiable record of that review? In most institutions, the honest answer is no. The physician may have technically “approved” the record by signing it. But signing is not reviewing. And reviewing without a traceable record is not accountability.
In financial services, AI systems make or heavily influence credit decisions affecting millions of people. When those decisions are discriminatory — and documented cases confirm they can be — the accountability diffuses immediately. The algorithm was technically compliant. The model was properly validated. The output followed established parameters. And yet the customer was harmed. Courts across multiple jurisdictions are beginning to establish that technical compliance is insufficient: AI systems must be transparent, fair, and accountable. But accountability to whom, through what mechanism, verified how?
In government and public services, AI increasingly determines who receives priority attention, who qualifies for services, who gets flagged for scrutiny. When someone is wrongly excluded or wrongly targeted, they often have no clear avenue to understand why the decision was made, let alone who made it. Researchers studying algorithmic decision-making in public services have identified what they call an “attribution gap” — the distribution of responsibility between developers, deploying organizations, and end users becomes so diffuse that no single party can be meaningfully held accountable.
Three domains. Three manifestations of the same absence: a human being who explicitly, traceably, verifiably holds accountability for how AI output was used.
The Problem Is Usage, Not Output
This is the insight that the governance debate keeps circling without landing on directly: the risk of AI is not primarily in its output. It is in its usage.
An AI system that produces correct output 94% of the time is not dangerous because of the 6%. It is dangerous when the 6% occurs inside an institutional context where no human being has been clearly designated as responsible for reviewing and approving what the AI produced before it was acted upon.
There are two ways institutions fail at this, and both are architectural rather than incidental.
The first failure mode is surrender. The human being is nominally present in the loop but substantively absent from it. The AI recommends; the human executes. The signature goes on the document, but the review did not happen in any meaningful sense. The human is a formality — present enough to satisfy a checklist, absent enough to provide no actual accountability.
The second failure mode is diffusion. Responsibility is distributed across so many parties — the model developer, the platform vendor, the deploying institution, the individual user — that it effectively belongs to no one. When something goes wrong, the diffusion becomes a defense: “we only provided the model,” “we only deployed the platform,” “we only followed the system’s recommendation.” Each statement may be technically accurate. Collectively, they describe an accountability vacuum.
Both failure modes share a root cause: accountability was not designed into the system from the beginning. It was assumed. Someone assumed that the physician would review carefully. Someone assumed that the credit officer would exercise judgment. Someone assumed that the procurement officer would override the system when necessary. Assumptions are not architecture.
Accountability as a Design Requirement
The right response to this is not more regulation — though better regulation would help. The right response is a design principle: chain of accountability must be embedded in AI systems from the moment of deployment, not added as an audit layer after the fact.
What does this mean concretely? Every AI output that will be used to make a decision affecting a human being needs to carry, as part of its institutional lifecycle, a traceable record of who reviewed it, who approved its use, in what context, with what information available at the time of decision. Not as metadata that can be detached. As a chain that follows the output through every stage of its use — verifiable by anyone who later needs to understand what happened and why.
This is not a new concept in institutional governance. It is how medical records are supposed to work. It is how financial audit trails are supposed to work. It is how legal chain of custody is supposed to work. What is new is the necessity of applying this logic explicitly to AI output — because AI creates a specific temptation that prior technologies did not: the temptation to treat the system as the decision-maker, and thereby dissolve the human accountability that civilization requires.
The moment AI output is treated as a decision rather than as an input to a human decision, the chain of accountability breaks. And once it breaks, it cannot be reconstructed after the fact. The audit trail does not exist. The reviewer cannot be identified. The approver cannot be named. What remains is a harm, a harmed person, and an institution that cannot explain what happened.
The Standard We Need to Set
Every institution deploying AI in consequential roles needs to be able to answer one question, for every deployment: when this output contributes to a decision that affects a human being, which specific human being reviewed it, approved its use, and holds accountability for the outcome — and can that accountability be independently verified?
If the answer is “the system,” the institution does not have AI governance. It has AI deployment with liability exposure it has not yet encountered.
If the answer is “the vendor,” the institution has outsourced accountability that cannot legally or ethically be outsourced.
If the answer is silence, the institution has a more serious problem than any AI system can create — a trust architecture that was never built.
AI governance is not a question about output quality. It is a question about whether the institutions deploying AI have built the chain of accountability that makes human responsibility traceable, demonstrable, and real.
That chain does not emerge from the model. It does not come from the vendor. It does not appear in the framework document.
It has to be designed. From the beginning. By humans who understand that they — not the system — are the ones accountable.
Leave a Reply