RSS Amplifier

Beyond the Slide · Apr 3, 2026

The Illusion of Control: Why "Human in the Loop" Is Clinical AI's Biggest Regulatory Fiction

0
Sign in to vote or save

Dr. Luis Cano · Beyond the Slide

This is a detour from my series on endpoints and digital pathology but this topic keeps generating questions I can't ignore: where are we actually going, and how? Two parts to work through the idea. Here's the first.

It’s 4 a.m. in the emergency department. The on-call physician has been on shift for nine hours. His hands feel heavy when he writes. His speech has slowed. Patients keep coming.

A woman walks in with chest pain. The AI-assisted triage system generates an automatic note: mechanical cause, likely diagnosis costochondritis, probability 90%. The doctor looks at her. She doesn’t seem distressed. The score is high. He prescribes anti-inflammatories and signs the discharge.

At 8 a.m., as the shift is winding down, a middle-aged man bursts through the ER doors. He’s screaming. Crying. Threatening everyone in sight. Paramedics arrive behind him with a stretcher. Cardiac arrest, they announce. Thirty minutes of resuscitation, CPR, defibrillation, epinephrine, nothing worked.

The doctor finishing his shift recognizes the patient. It’s the same woman he saw at 4 a.m.

Who failed here? The AI that generated the diagnosis? The doctor who accepted it without pushing back? The system that presented a 90% probability as though it were certainty? The protocol that called any of this “human in the loop”?

It’s tempting to read that story as a middle-of-the-night anomaly. An exhausted doctor. An unfortunate case. The exception, not the rule.

But that reading is exactly the problem.

4 a.m. in the ER is not an exception. It’s the operating condition for roughly 40% of emergency medicine. And the tired doctor who accepts the algorithmic suggestion without questioning it is not a negligent professional, he’s a cognitive system responding adaptively to an overload it was never designed to handle. The failure is not in the individual. It’s in the architecture that put him there and called it supervision.

“Human in the loop” appears in FDA regulatory documents, EMA ethics guidelines, medical AI company pitch decks, and hospital implementation protocols. It gets repeated as though it were enough. As though the mere presence of a physician somewhere in the process guarantees that everything else is under control.

It doesn’t. And there’s enough evidence to show why.

What follows isn’t an argument against human-AI collaboration. It’s the opposite, it’s an argument for taking that collaboration seriously. Because as long as the medical field keeps repeating HITL as a mantra, it will keep confusing the physician’s presence with the physician’s actual participation. And that confusion has clinical, legal, and ethical consequences we haven’t finished measuring.

Share

To understand why the HITL model is structurally fragile in clinical settings, you have to follow the term back to where it came from, and the answer is uncomfortable. It didn’t emerge from medical ethics or from any science of clinical decision-making. It came from military systems engineering in the mid-twentieth century.

In that original context, the “loop” was an operational fix: autonomous systems of the era couldn’t handle environmental variability, so a human operator was inserted to close the control cycle. The intervention was binary, technical, and measurable. The human fired or didn’t fire. Approved or aborted. There was no epistemic nuance.

When medicine adopted this language, it did so under similar pressure: the need to manage complex systems without surrendering moral accountability. But nobody stopped to ask whether the metaphor actually held. In a missile system, the loop is operational. In medicine, it’s epistemic. That difference changes everything.

The history of medicine is, at its core, the history of technologically mediated diagnosis. Every era has had its instrument and its form of authority:

The fundamental difference between all those eras and the current one is this: the stethoscope doesn’t generate a hypothesis. AI does. And as deep learning systems move from simple pattern detection toward complex clinical reasoning, the physician in the loop is increasingly reduced to the role of notary someone who signs off on decisions that were already configured by the algorithm.

One of the most persistent errors in medical AI discussions is treating HITL as if it were the only interaction model that exists. It isn’t. The European Commission’s Ethics Guidelines describe at least four distinct models, each with its own logic, its own risk profile, and, as we’ll see, its own failure points.

The most restrictive model. The AI generates a recommendation, a suggested diagnosis, a grading score, a risk alert, but the system blocks any further action until a professional explicitly reviews and approves it. In digital pathology, this means systems that flag regions of interest on a slide, while the pathologist performs and signs the final diagnosis.

The theoretical advantage is real: AI processing power combined with the human capacity to grasp contextual nuance, patient history, the full clinical picture. The practical disadvantage is that in high-volume workflows, this model generates extreme cognitive load and slows down processes exactly where speed has diagnostic value.

Here the AI has autonomy to execute processes or surface information directly. The human doesn’t validate each individual decision; they act as a supervisor who steps in only when they detect an anomaly or when the system fires a high-risk alert. Continuous monitoring systems in ICUs are the clearest example: the AI tracks hundreds of physiological parameters in real time and flags imminent sepsis risk. The clinician doesn’t review each data point, they oversee the system’s overall behavior.

The efficiency gain is real, but the risk profile shifts: automation bias operates quietly. If the system fails in a subtle, sustained way, not with an obvious alert, but through gradual drift, the human supervisor may not catch it until the error already has consequences.

This model operates at a more strategic level. The human doesn’t supervise each decision or each alert; they decide when to deploy the tool, for which patient population, under what ethical and legal conditions. It’s multidisciplinary governance: the committee that approves an AI triage protocol, the lab director who defines in which cases the platform can act more autonomously and in which it requires mandatory review.

The risk here is different and, in some ways, deeper: the disconnect between high-level policy and day-to-day clinical practice. A decision made in a conference room can have consequences that only show up weeks later, at the microscope.

The system operates fully autonomously during execution. The human sets the rules at the start and reviews results at the end, but doesn’t intervene in between. Today this is limited to administrative tasks, automated billing, note transcription, scheduling, but the regulatory and commercial pressure to expand this model toward lower-complexity diagnostic tasks is real and growing.

The primary risk is loss of traceability: when something goes wrong, errors can become systemic and invisible before anyone notices.

The full picture, with risk profiles:

What this map reveals is uncomfortable: no interaction model is free of risk. The question isn’t which one eliminates human failure, none of them do. The question is where in the spectrum failure occurs, under what conditions, and with what consequences.

It would be dishonest to ignore the cases where HITL genuinely works. In digital pathology, combining algorithmic detection of regions of interest with expert review has been shown to reduce interobserver variability and improve detection of mitotic foci in high-density tumors. In radiology, the human-AI combination outperforms either one alone on specific tasks. In the ICU, early warning systems have reduced sepsis mortality in centers with well-implemented protocols.

Well-structured collaboration has real advantages. The human filters artifacts the AI can’t contextualize. They catch biases in algorithmic recommendations that only become visible when you hold them against the patient’s actual reality. They anchor legal accountability in an agent capable of empathy and moral judgment. And they feed the model with corrections that make it better over time.

The problem isn’t the model in its ideal form. The problem is the distance between that ideal and what actually happens in clinical practice.

Share Beyond the Slide

The ER doctor from the opening story wasn’t negligent. He was a cognitive system responding the only way it could to nine hours of overload. That has a precise name: automation bias. The tendency to defer to automated system suggestions even when contradictory evidence is available. It’s not an individual weakness it’s an adaptive response, and the current design of HITL amplifies it rather than compensating for it.

A study on clinical decision support systems for prescription documented that when the system incorrectly flagged a medication as inappropriate, physician prescribing errors increased by 56.9%. The human in the loop didn’t prevent the error. The AI’s “assistance” actively led the clinician to make mistakes they wouldn’t have made working alone.

Then there’s what researchers call ersatz understanding, false comprehension. Explainable AI tools (XAI) generate visual or textual justifications that look logical, sound clinical, and carry the syntax of evidence. The physician reads them, finds them reasonable, and trusts the system more. What they can’t know is that the explanation may be a post-hoc rationalization of a spurious correlation: the model flagged pneumonia not because of the lungs, but because of the brand stamp on the X-ray machine.

And then there are hallucinations. In digital pathology, virtual staining models can generate cellular structures that look perfectly real but don’t exist in the original tissue. Morphologically plausible, histologically coherent, completely fabricated and by their very statistical nature, designed to pass through the filter of the human eye.

HITL has become the mechanism that allows clinical decision support tools to avoid classification as high-risk medical devices. The regulatory argument is simple: if the physician can override the recommendation, the system isn’t deciding it’s assisting. Responsibility stays with the human. The loop is closed.

But that legal architecture rests on an assumption the evidence directly challenges: that the clinician, under real working conditions, actually possesses the cognitive resilience needed to override the algorithm when it matters.

The result is a lose-lose situation for the professional. Follow the AI’s recommendation and it turns out to be wrong potential negligence. Ignore the recommendation and it turns out to be right failure to follow the standard of care. In that scenario, deferring to the algorithm becomes the path of least legal resistance, regardless of clinical intuition. The loop doesn’t protect the physician. It traps them.

The opening story happens in a specific context: an emergency department, the middle of the night, an exhausted physician. It’s easy to assume that other specialties with their more controlled rhythms, their quiet laboratories, their digital slides reviewed without time pressure are somehow insulated from this kind of failure. They aren’t.

A pathologist receives a digital slide. The AI system has flagged the regions of highest mitotic density and generated a grading score: grade 2, moderate proliferative activity. The interface displays the marked areas with a visual overlay. The preliminary report is pre-loaded in the system.

The pathologist reviews the flagged regions. They’re consistent with what you’d expect to see. The cellular morphology holds. The score fits the patient’s clinical context, a 52-year-old woman with a 1.8 cm breast nodule, no documented lymph node involvement.

What the system didn’t flag, because it fell outside the highest-density regions, because its distribution was peripheral, because statistically it read as noise, were atypical mitotic foci scattered along the tumor margin. Foci that, integrated with the rest of the picture, shifted the grade from 2 to 3. That changed the protocol. That changed the conversation with the patient.

The pathologist didn’t look for them because the system had already defined where to look. The loop was closed. The validation was logged. The report went out as grade 2.

No overnight shift. No nine hours of accumulated fatigue. No visible stress. Just a well-designed workflow, a clean interface, a competent professional and the same structural failure as in the ER, running quietly beneath conditions nobody would have called high-risk.

Who failed here? The algorithm that missed the peripheral foci? The pathologist who trusted the flagged regions? The system that presented a pre-loaded report as a starting point rather than a suggestion? The regulator who approved the tool under the assumption that human supervision was sufficient safeguard?

This is the question the HITL paradigm cannot answer. And it’s the question we need to learn to ask before the next system gets deployed in the next laboratory, with the next pathologist, in front of the next slide.

In part two, we’ll explore what should replace the loop and why the answer isn’t more supervision, but a fundamentally different architecture for the relationship between the clinician and the machine.

References:

  1. Human-in-the-loop. In: Wikipedia [Internet]. 2025 [cited 2026 Apr 3]. Available from: https://en.wikipedia.org/w/index.php?title=Human-in-the-loop&oldid=1327052927

  2. Li Zheng E, Jin W, Hamarneh G, Lee SSJ. From Human-in-the-loop to Human-in-power. Am J Bioeth. 2024 Sep;24(9):84–6. doi:10.1080/15265161.2024.2377139 PubMed PMID: 39226019; PubMed Central PMCID: PMC11384285.

  3. Ingram L. A brief history of medical diagnosis and the birth of the clinical laboratory.

  4. Google Cloud [Internet]. [cited 2026 Apr 3]. Qué es la intervención humana. Available from: https://cloud.google.com/discover/human-in-the-loop

  5. EMA and FDA set common principles for AI in medicine development | European Medicines Agency (EMA) [Internet]. 2026 [cited 2026 Apr 3]. Available from: https://www.ema.europa.eu/en/news/ema-fda-set-common-principles-ai-medicine-development-0

  6. FDA Oversight: Understanding the Regulation of Health AI Tools • Bipartisan Policy Center. Bipartisan Policy Center [Internet]. [cited 2026 Apr 3]. Available from: https://bipartisanpolicy.org/issue-brief/fda-oversight-understanding-the-regulation-of-health-ai-tools/

  7. Human in the loop requirement and AI healthcare applications in low-resource settings: A narrative review [Internet]. [cited 2026 Apr 3]. Available from: https://scielo.org.za/scielo.php?script=sci_arttext&pid=S1999-76392024000200007

  8. Human-In-The-Loop: What, How and Why. Devoteam [Internet]. [cited 2026 Apr 3]. Available from: https://www.devoteam.com/expert-view/human-in-the-loop-what-how-and-why/

  9. Guide to Optimizing Human AI Collaboration Systems [Internet]. [cited 2026 Apr 3]. Available from: https://www.deepscribe.ai/resources/optimizing-human-ai-collaboration-a-guide-to-hitl-hotl-and-hic-systems

  10. Rasmussen M. Michael Rasmussen. GRC 20/20 Research, LLC [Internet]. 2026 Feb 10 [cited 2026 Apr 3]. Available from: http://www.GRC2020.com

  11. What the EMA–FDA AI Principles Really Mean for Clinical Development & Regulatory Affairs [Internet]. 2026 [cited 2026 Apr 3]. Available from: https://www.precisionformedicine.com/blog/what-the-ema-fda-ai-principles-really-mean-for-clinical-development-regulatory-affairs

  12. Abd-Alrazaq A, Solaiman B, Mekki YM, Al-Thani D, Farooq F, Alkubeyyer M, et al. Hype vs Reality in the Integration of Artificial Intelligence in Clinical Workflows. JMIR Form Res. 2025 Dec 12;9:e70921. doi:10.2196/70921 PubMed PMID: 41385778; PubMed Central PMCID: PMC12700513.

  13. Roy SS. AI-Instigated Human Oversight: Rethinking Human-in-the-Loop Safety in Clinical AI.

  14. JD Supra [Internet]. [cited 2026 Apr 3]. AI in Healthcare: Legal and Ethical Considerations at the New Frontier. Available from: https://www.jdsupra.com/legalnews/ai-in-healthcare-legal-and-ethical-6634372/

  15. AI in healthcare: legal and ethical considerations in this new frontier [Internet]. [cited 2026 Apr 3]. Available from: https://www.ibanet.org/ai-healthcare-legal-ethical

  16. Barkley Z. Legal Challenges and Patient Protections in AI-Driven Healthcare – Fordham Undergraduate Law Review [Internet]. [cited 2026 Apr 3]. Available from: https://undergradlawreview.blog.fordham.edu/healthcare/legal-challenges-and-patient-protections-in-ai-driven-healthcare/

  17. Diaz-Asper C, Dagne M, Terhune E, Staker E, Heyn P. Artificial Intelligence (AI) as a Cognitive Function Digital Biomarker: Analyzing Speech in Older Individuals. Innovation in Aging. 2025 Dec 31;9. doi:10.1093/geroni/igaf122.1757

  18. Howard C, Johnson A, Baratono S, Faust K, Peedicail J, Ng M. Machine Learning–Based Cognitive Assessment With The Autonomous Cognitive Examination: Randomized Controlled Trial. Journal of Medical Internet Research. 2025 Jul 30;27(1):e67446. doi:10.2196/67446

  19. Sokol K, Fackler J, Vogt JE. Artificial intelligence should genuinely support clinical reasoning and decision making to bridge the translational gap. NPJ Digit Med. 2025 Jun 10;8:345. doi:10.1038/s41746-025-01725-9PubMed PMID: 40494886; PubMed Central PMCID: PMC12152152.

  20. Beck J, Eckman S, Kern C, Kreuter F. arXiv.org [Internet]. 2025 [cited 2026 Apr 3]. Bias in the Loop: How Humans Evaluate AI-Generated Suggestions. Available from: https://arxiv.org/abs/2509.08514v1

  21. Agudo U, Liberal KG, Arrese M, Matute H. The impact of AI errors in a human-in-the-loop process. Cogn Res Princ Implic. 2024 Jan 7;9:1. doi:10.1186/s41235-023-00529-3 PubMed PMID: 38185767; PubMed Central PMCID: PMC10772030.

  22. iatroX [Internet]. 2026 [cited 2026 Apr 3]. AI Hallucination in Medicine: Real Examples, Real Risks, and How to Protect Yourself | iatroX Clinical AI Insights. Available from: https://www.iatrox.com/blog/ai-hallucination-medicine-real-examples-risks-how-to-protect-yourself-2026

  23. UCLA [Internet]. [cited 2026 Apr 3]. AI watching AI: Dangerous errors in digital pathology caught by UCLA system. Available from: https://newsroom.ucla.edu/releases/dangerous-AI-errors-digital-pathology-caught-ucla-artificial-intelligence

  24. The Myth of the Human-in-the-Loop and the Reality of Cognitive Offloading. Perry World House [Internet]. [cited 2026 Apr 3]. Available from: https://perryworldhouse.upenn.edu/news-and-insight/the-myth-of-the-human-in-the-loop-and-the-reality-of-cognitive-offloading/

Read the original on beyondtheslide.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.