The Question “Should AI Make This Decision” Is Broken. Here’s a Better One.
Most debates about AI ethics get stuck treating “should AI make this decision” as a single yes-or-no question, as if every decision belonged in one of two bins: fine to automate, or too important to touch. That framing collapses the moment you apply it to anything real. A far more precise answer already exists, built nearly fifty years ago for an entirely different technology, and it turns out to translate onto AI delegation with almost no modification required.
Why a 1979 Bioethics Document Is the Right Starting Point
In 1979, following a government investigation into serious ethical abuses in medical research, a federal commission published a short document that became the foundation for how the United States regulates any research involving human subjects. The Belmont Report, published by what is now the U.S. Department of Health and Human Services, established three basic ethical principles for situations where researchers hold significant power over people who can’t fully verify or control the process affecting them: respect for persons, which centers on protecting individual autonomy and requiring informed consent; beneficence, the obligation to maximize benefit and minimize harm; and justice, concerning who fairly bears the risks and receives the benefits of a given decision or intervention. These three principles were built for a specific structural problem: how do you ethically evaluate a situation where one party holds power and expertise the other party doesn’t fully share, and where the second party’s wellbeing depends on decisions made largely outside their direct control?
That’s an almost exact structural description of AI-assisted decision-making, which is precisely why this fifty-year-old framework transfers with so little modification, and why professional bodies working specifically on AI ethics have converged on nearly the same three ideas independently.
The Modern Confirmation This Framework Still Holds
It’s worth noting this isn’t just a clever historical analogy — a major international health body has already applied essentially the same framework directly to AI. The World Health Organization’s guidance on the ethics and governance of artificial intelligence for health established six core principles, with protecting human autonomy listed first, explicitly requiring that humans remain in control of significant decisions, that providers have the information necessary to understand how an AI system is being used, and that informed consent and privacy protections apply specifically to AI-assisted decisions rather than being treated as optional extras. The WHO’s remaining principles map closely onto beneficence (promoting wellbeing and safety) and justice (ensuring inclusiveness and equity) — meaning two independent frameworks, developed decades apart for different purposes, arrived at essentially the same three-part ethical structure once applied to a situation involving delegated, high-stakes decision-making.
Principle One: Autonomy — Where Consent Actually Has to Mean Something
Applied to AI delegation, autonomy asks a specific, testable question: does the person affected by this decision know AI is involved, and do they understand enough about that involvement to make a meaningful choice about it? This is a considerably higher bar than a buried disclosure in a terms-of-service document nobody reads. Genuine informed consent, in the Belmont Report’s own framing, requires information presented in a way the affected person can actually process and act on, not merely technically disclosed somewhere in fine print.
This principle draws a clear, testable line: an AI system quietly screening job applications without any disclosure fails this test outright, regardless of how accurate the system turns out to be, because the affected person never had a genuine opportunity to understand or contest the process shaping an outcome that affects them directly. An AI system clearly disclosed, explained in plain language, with a genuine path to request human review, meets this bar even if the underlying technology is identical — the ethical distinction lives entirely in the transparency and consent surrounding the decision, not in the AI’s technical sophistication.
Principle Two: Beneficence — The Actual Harm-and-Benefit Test
Beneficence asks whether a given use of AI genuinely produces more benefit than harm, and — this is the part often skipped — who specifically bears the harm if the system gets it wrong. An AI tool that speeds up a low-stakes internal task, where an error costs a few minutes of correction, easily clears this bar. An AI tool making or heavily influencing a high-stakes decision, where an error could mean someone loses a job opportunity, a loan, or access to a benefit they were otherwise entitled to, faces a much higher bar — and the ethical question isn’t simply “is this system usually accurate,” it’s “who absorbs the cost on the occasions it isn’t.”
This connects directly to a related point worth drawing out here. Our piece on what humans should never delegate to AI covers the accountability side of this same question — situations requiring a specific person who can be held responsible for an outcome. Beneficence asks a related but distinct question: even setting accountability aside, does the decision genuinely produce more good than harm once the realistic error rate and its consequences are honestly factored in, not just in the aggregate but for the specific individuals who happen to fall on the wrong side of that error rate.
Principle Three: Justice — Who Actually Bears the Cost of Being Wrong
Justice, in the Belmont Report’s framing, asks whether the risks and benefits of a decision are distributed fairly, rather than concentrated disproportionately on people least able to contest or absorb them. Applied to AI delegation, this becomes a question about who ends up bearing the cost when an AI-assisted system makes an error — and whether that cost falls evenly, or falls disproportionately on people with the least practical ability to appeal, understand, or correct the mistake affecting them.
This is where a lot of well-intentioned AI deployment quietly fails the ethical test even when it passes the accuracy and even the consent tests individually. A system that’s 95% accurate sounds impressive in the aggregate, but if the 5% error rate concentrates disproportionately among people already facing structural disadvantages in contesting a wrong decision — limited time, limited resources, limited familiarity with how to request a human review — the justice principle is violated even though the raw accuracy number looks fine on a dashboard. This is precisely the concern regulators like the EEOC have raised about AI hiring tools, where a facially neutral algorithm can still produce unjust outcomes concentrated on specific groups.
Applying the Three-Part Test to an Actual Decision
It helps to see the framework applied concretely rather than described abstractly. Consider a company using AI to help draft initial performance review language for managers to edit and finalize themselves. Autonomy: employees are told AI assists in drafting language, and managers retain full editorial control and final say — this clears the bar, provided the disclosure is genuine rather than buried. Beneficence: the AI-assisted draft speeds up a tedious task without determining the actual substance of the review, and any drafting error is caught by the manager before it reaches the employee — a low-harm, reasonably high-benefit use. Justice: since a human manager reviews and finalizes every case identically regardless of role or seniority, the risk of an AI drafting error isn’t concentrated on any particular group of employees.
Now compare that to an AI system autonomously scoring candidates and auto-rejecting the bottom percentage without human review. Autonomy fails if rejected candidates were never told AI played a determinative role. Beneficence is weaker, since an error directly costs someone a genuine opportunity with no correction mechanism. And justice is genuinely uncertain until someone actually audits whether the error rate falls evenly across demographic groups or concentrates unevenly — which is exactly the kind of test the same underlying framework, applied honestly, forces you to actually run rather than assume.
Why This Framework Beats Vague AI Ethics Guidance
A meaningful share of AI ethics guidance currently circulating stays at the level of aspiration — “use AI responsibly,” “prioritize fairness” — without giving anyone a concrete test to apply to a specific decision. The Belmont framework’s real strength is that each principle generates a specific, checkable question: was genuine informed consent obtained, does the realistic harm-benefit balance hold up once error rates are honestly considered, and does the burden of error fall evenly or concentrate on people least able to contest it. Those three questions, applied consistently, do more real ethical work than most AI-specific guidance currently offers, precisely because they were forged under the pressure of a real historical failure rather than written in the abstract.
Why This Framework Doesn’t Give You an Easy Answer, and That’s the Point
It’s worth being honest that applying these three questions won’t always produce a clean, comfortable verdict, and that’s actually a feature rather than a shortcoming. The Belmont Report itself was explicit that its principles function more like a compass than a formula — they force a genuine weighing of competing considerations rather than mechanically outputting a yes-or-no answer. A hiring tool might pass the autonomy test cleanly, show a reasonable beneficence balance, and still raise a genuine, unresolved justice concern that requires an actual audit rather than a confident assumption either way. That ambiguity is honest, and it’s a meaningful improvement over both extremes common in casual AI ethics discussion — the confident dismissal that treats any AI involvement as automatically fine, and the equally confident alarm that treats any AI involvement as automatically suspect.
The value of a framework like this isn’t that it eliminates hard judgment calls. It’s that it makes the specific dimensions of the judgment call visible and nameable, which is precisely what turns a vague unease about “AI ethics” into an actual, structured decision someone can defend, revisit, and improve over time as more evidence about a given system’s real-world error patterns accumulates.
A Quick Audit Using the Three-Part Test
A useful, honest exercise: take one AI-assisted decision process currently used in your own work or organization, and run it through all three questions specifically. Do the people affected know AI is involved, in a way they could actually understand and act on, not just a buried disclosure? Does the realistic harm-benefit balance hold up once you account for who bears the cost of an actual error, not just the aggregate accuracy rate? And has anyone actually checked whether that error rate concentrates unevenly on any specific group, rather than simply assuming it doesn’t? Most organizations, being honest, have never explicitly run this exercise on their own AI-assisted processes, which is itself a meaningful gap the framework makes visible the moment it’s actually applied.
Frequently Asked Question
What is the Belmont Report and why does it apply to AI ethics?
The Belmont Report is a 1979 federal bioethics document establishing three principles for situations where one party holds power over decisions affecting another: respect for persons and autonomy, beneficence, and justice. These principles apply directly to AI delegation because they were built for the same structural problem — evaluating decisions made largely outside the affected person’s direct control.
What does the World Health Organization say about AI and human autonomy?
The WHO’s guidance on ethics and governance of AI for health lists protecting human autonomy as its first core principle, explicitly requiring that humans remain in control of significant decisions, that people understand how AI is being used, and that informed consent and privacy protections apply specifically to AI-assisted decisions.
What makes informed consent for AI different from a standard terms-of-service disclosure?
Genuine informed consent requires information presented in a way the affected person can actually understand and act on, not merely technically disclosed somewhere in fine print. An AI system quietly influencing a decision without meaningful disclosure fails this test regardless of the system’s accuracy.
How does the beneficence principle apply to AI decision-making?
Beneficence asks whether a given use of AI produces more benefit than harm, specifically accounting for who bears the cost when the system makes an error. A low-stakes task where mistakes are cheap to correct clears this bar easily; a high-stakes decision with a real, uncorrectable cost to the affected person faces a much higher bar.
Why can an AI system with high accuracy still be ethically problematic?
Under the justice principle, even a highly accurate system can fail ethically if its error rate concentrates disproportionately on people with the least practical ability to contest or correct a wrong decision. A system’s aggregate accuracy doesn’t capture whether the burden of its mistakes is distributed fairly.
How can I apply this framework to an actual AI decision in my own work?
Ask three specific questions: do affected people genuinely understand AI’s role and have they meaningfully consented, does the realistic harm-benefit balance hold up once actual error costs are honestly counted, and has anyone verified whether errors concentrate unevenly on any specific group rather than assuming they don’t.
Conclusion
The ethics of delegating a decision to AI were never really a novel philosophical puzzle requiring an entirely new framework — they’re a specific instance of a much older question about power, consent, and fair distribution of risk that a 1979 federal bioethics commission already answered with unusual precision, and that the World Health Organization has since applied directly to AI with almost no modification required.
Autonomy, beneficence, and justice, applied as three specific, checkable questions rather than vague aspirations, draw the actual line far more precisely than “is AI too risky for this” ever could. Any AI-assisted decision that fails genuine informed consent, produces more harm than benefit once real error costs are honestly counted, or concentrates that harm on people least able to contest it, has crossed the line — regardless of how accurate or impressive the underlying technology happens to be.
