ETH-06 · Risk, Ethics & Accountability

Accountability for AI Decisions: Who Answers When It Fails

When an automated decision harms someone, accountability can't rest with 'the algorithm'. How responsibility flows from data to deployer under GDPR Article 22.

Accountability chain: data, model and automated decision are links that lead back to a named person who answers for the outcome (GDPR Article 22)dataINPUTSmodelSYSTEMdecisionOUTPUTanswers for itART. 22 · A NAMED PERSON ANSWERS FOR THE DECISION

Who is actually accountable when an AI decision goes wrong?

The people and organisations who chose to build, sell, deploy and rely on the system are accountable — not the system itself. An algorithm cannot be answerable, because answerability is a relationship between humans: someone owes an explanation, a remedy, and sometimes a consequence to someone else. When a loan is wrongly refused, a benefit wrongly stopped, or a candidate wrongly screened out, the question is never “what did the model do”. It is “who decided to let the model decide this, under what conditions, with what oversight, and who carries the duty to put it right”. Accountability for AI decisions is the work of keeping those answers attached to named human roles before anything goes wrong, so that afterwards there is no gap to fall into.

That last clause is the whole problem. The phrase “the algorithm did it” is not a technical observation; it is an accountability move. It takes a decision that humans designed, trained, approved, deployed and depended on, and relocates responsibility into a place where no one can be held to it. The honest version of the same sentence is longer: a data team chose the inputs, a provider built and validated the model, a deployer decided which decisions to automate and how to oversee them, and a manager signed off the level of human review. Each of those is a person who could be asked to account. The job of an accountability framework is to make sure they can be — and that they expected to be all along.

This pillar maps how responsibility for an automated decision actually flows, from the data it was trained on to the organisation that deployed it. It looks at what the law already requires — particularly the right under GDPR Article 22 not to be subject to a purely automated decision — at where the limits of autonomy lie, at what separates real human review from a signature on a screen, and at the concrete ways organisations assign responsibility so it does not dissolve. It is the anchor of our Risk, Ethics & Accountability cluster; the supporting articles go deeper on each piece.

Why does accountability “evaporate into the algorithm”?

Accountability evaporates because automated decisions are produced by many hands, and the more hands involved, the easier it is for each to point at another. Philosophers call this the problem of many hands: when an outcome is the joint product of a long chain of contributors, no single contributor feels — or can easily be shown to be — responsible for the whole. AI systems stretch that chain further than almost any other technology, because the people who build a model, the people who supply its data, the people who deploy it, and the people who act on its outputs are usually in different teams, different companies, and sometimes different countries.

There is a second, subtler reason, and it is about how machines change our sense of agency. When a person makes a judgement, we naturally treat them as its author. When a system produces a score, a flag, or a recommendation, the human who acts on it can feel like a relay rather than a decider — “I only did what it told me” — even when they had full authority to disagree. The output arrives wrapped in an aura of objectivity: it looks like a measurement, not an opinion, and measurements seem to belong to no one. So the human downstream defers, the designers upstream say they only built a tool, and the responsibility that should sit somewhere ends up sitting nowhere.

The third reason is structural neglect. Many organisations simply never decide, in advance, who owns a given automated decision. They procure a tool, switch it on, and discover only after a harm that no role was ever assigned the duty to monitor, question, or override it. The vacuum is not created at the moment of failure; it was there from the day of deployment, waiting. Most “the algorithm did it” failures are not failures of malice or even of technology. They are failures to allocate responsibility while it was still cheap to do so.

The point of taking accountability seriously is to close all three gaps deliberately: to name the hands, to remind the human downstream that they remain a decider and not a relay, and to assign ownership before launch rather than after harm. Everything that follows is a way of doing that.

What does GDPR Article 22 actually say about automated decisions?

Article 22 of the GDPR gives a person the right not to be subject to a decision based solely on automated processing — including profiling — where that decision produces legal effects concerning them or similarly significantly affects them. It is one of the few places in European law that speaks directly to the moment an automated decision lands on an individual, and it has been the backbone of accountability for such decisions since 2018. Understanding its three load-bearing words — solely, significantly, and decision — is the difference between treating it as a slogan and treating it as a control.

When does Article 22 apply to a decision?

Three conditions have to hold together. First, there must be a decision — an outcome that does something to the person, not merely an internal analysis. Second, that decision must be based solely on automated processing, meaning no meaningful human involvement in reaching it. Third, it must have a legal effect or similarly significant effect: refusing credit, terminating a contract, rejecting a job application, cutting off a benefit, or denying access to a service all qualify; a product recommendation usually does not.

When all three hold, the default position of the law is prohibition: the controller may not run the decision this way unless one of three gateways applies — the decision is necessary for entering or performing a contract, it is authorised by Union or member-state law, or it is based on the person’s explicit consent. Article 22 is therefore not a balancing test you can argue your way through case by case. It starts from “not allowed”, and the burden is on the organisation to show it sits inside one of the narrow exceptions.

It is also worth knowing that Europe’s highest court has read the scope generously in favour of individuals. In its 2023 ruling in the SCHUFA case (C-634/21), the Court of Justice held that even producing a credit-scoring value can itself be an Article 22 “decision” where a third party draws strongly on that score to grant or refuse a contract. The practical lesson for accountability: you cannot necessarily escape the duty by saying “we only generate a number, someone else decides”. If your output effectively determines the outcome, the law may treat generating it as the decision.

What safeguards does Article 22 give the affected person?

Where an automated decision is permitted through the contract or explicit-consent gateways, the controller must put safeguards in place, and at minimum these include the right to obtain human intervention, the right to express one’s point of view, and the right to contest the decision. This is the legal hinge on which human-in-the-loop accountability turns: the law does not merely ask that a system be accurate; it gives the affected person a route back to a human who can reconsider.

Read carefully, these safeguards are demanding. “Human intervention” means a person with the authority and competence to actually change the outcome — not a help-desk agent who can only explain why the system said no. The right to “express a point of view” implies the person can introduce information the system never had. And the right to “contest” implies a real reconsideration, not a re-run of the same model on the same inputs. An appeals process that simply asks the algorithm twice satisfies the form of Article 22 while defeating its purpose. We unpack what a genuine reconsideration looks like in meaningful human oversight.

What does “solely” automated really mean — and where is the rubber-stamp trap?

“Solely” is where most organisations think they are safe and often are not. The instinct is to insert a human into the workflow — a click, a sign-off, an “approved” button — and conclude that the decision is no longer based solely on automation, so Article 22 does not bite. European regulators have been explicit that this does not work. To take the decision out of “solely automated” territory, the human involvement must be meaningful: carried out by someone with the authority and competence to change the decision, who actually considers the relevant data rather than routinely applying the system’s output.

That is the rubber-stamp trap, and it is the precise mechanism by which accountability evaporates. A human who approves three hundred algorithmic outputs an hour, without the time, information, or authority to do otherwise, does not convert an automated decision into a human one. They merely lend it a signature — and, dangerously, they may shift the legal appearance of responsibility onto themselves while having none of the real power to exercise it. Token human involvement is worse than honest automation, because it manufactures the look of accountability without the substance, and it does so at exactly the point where a regulator or a court will later go looking for it. Whether human review is meaningful or cosmetic is not a paperwork question; it is the question.

What is the accountability chain, from data to deployer?

The accountability chain is the sequence of human roles, each with its own duties, through which an automated decision passes on its way from raw data to real-world effect. Mapping it is the single most useful exercise an organisation can do, because a harm can originate at any link, and a defensible answer to “who is responsible” usually involves several links at once rather than one scapegoat. The chain is not a way of spreading blame thinly; it is a way of making sure every contribution has an owner.

Link in the chain Who it is What they are accountable for Typical failure
Data sources & curation Data engineers, data owners, upstream providers Lawful basis, representativeness, quality, documented provenance and known gaps Biased, stale or unlawfully sourced training data baked into every later decision
Model development Provider / developer (in-house or vendor) Design choices, validation, documented intended purpose, known limitations and error profile A model used outside the conditions it was validated for, or limitations never disclosed
Procurement & deployment Deployer (the organisation running it) Choosing which decisions to automate, configuring the system, defining oversight, monitoring in production Switching on a tool with no owner, no monitoring and no override path
Human oversight Named reviewers, case handlers, managers Actually reviewing, questioning and overriding within their authority and competence Rubber-stamping under throughput pressure; no real power to say no
Governance & sign-off Accountable executive / DPO / risk owner Approving the level of autonomy, ensuring remedy and contestability exist, carrying the buck No one ever decided the risk was acceptable, or who would answer if it was not

Two distinctions make this chain legally concrete. Under data-protection law, the split between controller (who decides the purposes and means of processing) and processor (who acts on the controller’s instructions) determines who carries the heaviest duties for an automated decision — usually the deployer acting as controller, not the vendor. Under the EU AI Act, a parallel split between provider (who develops or markets the system) and deployer (who uses it under their own authority) allocates obligations along similar lines. The reason both regimes draw the line in roughly the same place is that the same intuition underlies both: the actor who decides to use a system on real people, in a real context, carries responsibility that a component supplier cannot discharge for them.

The practical upshot is that buying an AI system does not outsource accountability for its decisions. A vendor can be responsible for a defective or misrepresented model, and you should hold them to that contractually. But the choice to point that model at your customers, your candidates, or your claimants — and the duty to oversee it once you have — stays with you as the deployer. Responsibility travels with the decision to deploy, not with the code.

Where are the limits of autonomy?

The limit of autonomy is the point beyond which a decision should not be made without a human, no matter how good the system is — and that point is set by the stakes and the reversibility of the decision, not by the model’s accuracy. A system can be right ninety-nine times in a hundred and still be the wrong thing to fully automate, if the hundredth case ends a livelihood, a treatment, or a liberty, and cannot be undone. Accuracy tells you how often the system is right; it tells you nothing about what happens when it is wrong, and accountability lives entirely in that second question.

A workable way to find the limit is to ask three questions of any decision before automating it. How severe and how reversible is the worst plausible outcome for the affected person? How well can the system’s reasoning be explained and contested after the fact? And how legitimate is it, in this domain, for a machine to decide alone — would the affected person, or society, accept “the system decided” as an answer? Where the harm is severe and irreversible, the reasoning opaque, or the legitimacy low, full autonomy is the wrong design even where it is technically permitted. These are the decisions that belong inside a human-in-the-loop or human-in-command pattern; the difference between those control models is the subject of human-in-the-loop vs on-the-loop vs in-command.

There is a deeper principle underneath the practical test. Autonomy and accountability pull in opposite directions: the more a system decides on its own, the harder it becomes to locate a human who genuinely owned the decision. Beyond a certain level of delegation, the human in the workflow can no longer meaningfully understand or override what is happening fast enough to matter, and at that point the appearance of oversight outlasts the reality of it. The limit of autonomy is therefore not only an ethical line but an accountability one: it is the level of delegation beyond which no human can honestly say “I was responsible for this”, and so the level beyond which the decision should not be made alone. Deciding deliberately which decisions sit on each side of that line is the design work covered in designing the decision point.

What makes human review meaningful rather than cosmetic?

Human review is meaningful when the reviewer has the authority to change the outcome, the competence to know when it is wrong, the information to see why, and the time to use all three — and is genuinely expected to disagree when the case warrants it. Strip away any one of those and review degrades into ritual. A reviewer with authority but no competence approves what they cannot evaluate. One with competence but no time approves what they cannot examine. One with both but no power to act produces an opinion the system ignores. Cosmetic review is not a single failure; it is the absence of any one of these conditions, dressed up as oversight.

The most reliable signal that review is real is the disagreement rate. In a healthy process, humans override the system sometimes, those overrides are recorded, supported rather than punished, and they actually change outcomes. When a reviewer almost never disagrees, that is not reassurance that the system is excellent; it is the symptom to investigate, because it is exactly what you would also see if the human had quietly become a rubber stamp. Accountability frameworks that measure oversight by counting sign-offs are measuring the wrong thing — a hundred-percent review rate can sit on top of zero-percent scrutiny. What you want to measure is what happened when the human said no.

Meaningful review also has to be designed against automation bias — the well-documented human tendency to over-trust a system that is usually right, and to stop checking precisely the confident output that most needs checking. The corrosive thing about automation bias for accountability is that it attacks the one capability oversight depends on: the willingness to disagree. The better the system performs on average, the stronger the habit of deference, and the more the rare confident error slips through with a human signature attached. This is why “a human reviewed it” is never, on its own, an adequate answer to “who was accountable”. The honest follow-up is always: could that human have stopped it, would they have known to, and would they have been backed if they did? We treat that whole question in depth in meaningful human oversight.

How do you assign responsibility so it does not evaporate?

You assign responsibility by naming it before deployment, attaching it to a person rather than a process, and building the records and routes that let that person actually be held to it. Accountability is not a value you can declare in a policy; it is an arrangement you have to construct, link by link, and it has to exist on paper before the harm exists in the world. Six practices, taken together, do most of the work — and crucially, all six are completed before a system goes live, not assembled afterwards in response to a complaint.

Name an accountable owner for each automated decision. Every consequential decision a system makes should map to a single named human role that answers for it — not a committee, not “the AI team”, but a person who can be asked, and expected, to explain it. The owner does not make each decision by hand; they own the choice to automate it, the conditions under which it runs, and the duty to fix it when it fails. A decision with no named owner is a decision with no accountability, however sophisticated the governance documents around it look.

Separate the roles along the chain and write them down. Use the controller/processor and provider/deployer distinctions deliberately: be explicit, in contracts and in internal RACI maps, about which duties sit with the vendor and which stay with you. Vendors should warrant the model’s documented behaviour and limitations; you retain responsibility for choosing to deploy it and for overseeing it. Ambiguity here is not neutral — when no one is clearly responsible, in practice everyone is unaccountable.

Keep decision logs that survive scrutiny. For each consequential decision, record what the inputs were, what the system produced, who reviewed it, what they did, and why. Logs are what convert a vague “a human was involved” into a defensible, specific account — and what let an organisation distinguish, after the fact, between a reviewer who genuinely engaged and one who clicked approve. You cannot demonstrate accountability you did not record.

Build a real route to contest and remedy. Article 22’s safeguards are not only a legal duty; they are an accountability mechanism. An affected person needs a usable path to reach a competent human, present information the system lacked, and have the decision genuinely reconsidered. A contest route that re-runs the same model is theatre. Pair it with a defined remedy: if the decision was wrong, what gets put right, by whom, and how fast.

Set and enforce the limit of autonomy explicitly. Decide, per decision type, whether the system may decide alone, must recommend to a human, or must defer to one entirely — and tie that choice to the stakes and reversibility, not to vendor enthusiasm. Then enforce it in the system, not just in a policy document, so the limit is a property of how the software behaves and not an aspiration nobody checks.

Monitor in production and review the limits over time. Accountability is not discharged at launch. Models drift, populations change, and a decision that was safe to automate last year may not be this year. Someone — the named owner — has to watch the disagreement rate, the error profile, and the complaints, and to move decisions back under tighter human control when the evidence says so. A framework that is never revisited is a framework that is slowly going out of date.

None of these practices is exotic, and that is the point. Accountability does not evaporate because the problem is unsolvable; it evaporates because the cheap, dull, pre-launch work of naming owners and writing down roles gets skipped in the rush to ship. The organisations that answer cleanly when an automated decision goes wrong are simply the ones that did that work while it was still easy.

How does the EU AI Act change who answers for an automated decision?

The AI Act adds a second, system-design layer of accountability on top of data-protection law, and it allocates duties along the same provider/deployer line. Where GDPR Article 22 governs the moment an automated decision affects an individual, the AI Act governs how the system that makes such decisions must be built and overseen — most directly through its human-oversight requirement for high-risk systems, Article 14, which obliges that such systems be designed so that a competent person can effectively oversee them while they are in use. The two regimes overlap but are not the same: Article 22 is about a fully automated decision and the rights of the person it lands on; Article 14 is about whether the system was built to be overseeable in the first place, regardless of whether any single decision is fully automated.

A note on timing, because the dates are easy to get wrong and worth stating precisely. The AI Act’s transparency obligations — for example, telling people when they are interacting with an AI system or seeing AI-generated content under Article 50 — apply from 2 August 2026. The heavier high-risk obligations, including the Article 14 oversight duties, sit on a later timeline that the European Commission’s Digital Omnibus simplification package has proposed to push toward December 2027. These dates are still being finalised through that legislative process, so the responsible thing is to design for the principle and confirm the current deadline before relying on a specific one. It is also worth resisting the temptation to frame a weak-oversight gap today as a “seven-percent-of-turnover fine” risk: the Act’s heaviest penalties attach to specific prohibited practices, and the high-risk regime where the oversight duties live is on that later, shifting schedule. The case for getting accountability right does not need an inflated number to stand up. The deeper treatment of the oversight duty lives in our AI Act & human oversight cluster.

What the AI Act ultimately changes about “who answers” is that it makes the deployer’s oversight an explicit, named duty rather than an implied good practice. It does not let the vendor off — providers carry substantial obligations for how the system is built and documented — but it confirms, in statute, the intuition this whole article rests on: the organisation that points an AI system at real people is responsible for making sure a competent human can oversee it, and for the consequences when that oversight fails. The law is catching up to the principle. Organisations that have already built their accountability chain deliberately will find compliance is mostly a matter of writing down what they already do; those who left responsibility to evaporate into the algorithm will find the gap was always there, and that the law has simply made it visible. If you are starting from the beginning, the foundations are laid out in what human-in-the-loop actually means.

Frequently asked

Who is legally responsible when an AI system makes a wrong decision?

Responsibility is shared along a chain, but it concentrates on the organisation that deployed the system to make decisions about real people. Under data-protection law that organisation is usually the controller, which carries the heaviest duties. The vendor or provider can be responsible for a defective or misrepresented model, and you should hold them to that contractually, but the choice to use the system on your customers, candidates or claimants — and the duty to oversee it — stays with the deployer. The system itself is never the responsible party, because accountability is a relationship between humans.

What does GDPR Article 22 say about automated decisions?

Article 22 gives a person the right not to be subject to a decision based solely on automated processing, including profiling, where it produces legal effects or similarly significantly affects them — such as refusing credit, a job, or a benefit. The default is prohibition unless one of three gateways applies: the decision is necessary for a contract, authorised by law, or based on explicit consent. Where it is permitted via contract or consent, the person must get safeguards including the right to human intervention, to express their point of view, and to contest the decision.

Does adding a human approval step take a decision out of Article 22?

Only if the human involvement is genuinely meaningful. European regulators have made clear that token sign-off does not convert an automated decision into a human one. To fall outside 'solely automated', the human must have the authority and competence to actually change the outcome and must consider the relevant data, not routinely apply whatever the system outputs. A reviewer rubber-stamping hundreds of outputs an hour, with no real power to disagree, does not satisfy Article 22 — and may shift the appearance of responsibility onto themselves without any of the power to exercise it.

What is the 'accountability chain' for an AI decision?

It is the sequence of human roles a decision passes through: the people who source and curate the data, the provider who develops and validates the model, the deployer who decides which decisions to automate and how to oversee them, the named humans who review and can override, and the governance owner who approved the level of autonomy. Each link has its own duties and its own possible failure mode. Mapping the chain matters because a harm can start at any link, and assigning a named owner to each one is what stops responsibility from evaporating.

Where is the limit of how autonomous an AI decision should be?

The limit is set by stakes and reversibility, not by the model's accuracy. A system can be highly accurate and still be the wrong thing to fully automate if its rare errors are severe and irreversible. A practical test asks three questions: how bad and how reversible is the worst plausible outcome, how explainable and contestable is the reasoning, and how legitimate is it for a machine to decide alone in this domain. Where harm is severe, reasoning is opaque, or legitimacy is low, the decision should keep a human in the loop or in command, even where full automation is technically allowed.

How is the EU AI Act different from GDPR Article 22 on accountability?

They overlap but address different things. Article 22 governs the moment a fully automated decision affects an individual and gives that person rights to intervention and contestation. The AI Act governs how high-risk systems must be built and overseen — its Article 14 requires that such systems be designed so a competent person can effectively oversee them in use, whether or not any single decision is fully automated. The AI Act also allocates duties along a provider/deployer line, confirming in statute that the organisation deploying a system on real people is responsible for ensuring it can be overseen.

When do the AI Act's human oversight rules actually apply?

The timelines differ by obligation and are still being finalised. The transparency duties under Article 50 — such as disclosing AI interactions and AI-generated content — apply from 2 August 2026. The heavier high-risk obligations, including the Article 14 human oversight duties, sit on a later timeline that the Commission's Digital Omnibus simplification package has proposed to push toward December 2027. Because these dates are moving through the legislative process, design for the principle now and confirm the current deadline before relying on a specific one. Avoid framing a present oversight gap as a fixed large-percentage fine; the heaviest penalties attach to specific prohibited practices, not to the high-risk regime on that later schedule.

How do you assign responsibility so it does not evaporate into 'the algorithm did it'?

By doing the work before deployment, not after harm. Name a single accountable owner for each consequential automated decision; separate vendor and deployer duties explicitly in contracts and RACI maps; keep decision logs that record inputs, outputs, who reviewed and what they did; build a real route to contest and remedy that reaches a competent human rather than re-running the same model; set and enforce the limit of autonomy in the system itself; and monitor in production so decisions move back under tighter control when the evidence requires it. Responsibility evaporates because this cheap pre-launch work gets skipped, not because the problem is unsolvable.