RSS Amplifier

Beyond the Slide · Apr 10, 2026

Beyond the Loop: AI as an Extension of Clinical Reasoning

0
Sign in to vote or save

This page did not load. You can still read it on the original site — the toolbar below keeps your place in the directory.

If the problem wasn’t a lack of supervision, the solution can’t be more supervision either.

If the problem wasn’t a lack of supervision, the solution can’t be more supervision either. A proposal for rethinking from the ground up the relationship between the clinician and the machine.

Second and final part of the series on human in the loop. If you missed part one, here’s the short version: no human-AI interaction model is free of risk, and HITL has structural failures that no supervision protocol can fix. This part proposes what should replace it.

Part one ended with two questions that had no answers. The ER doctor who accepted the AI’s diagnosis at 4 a.m. The pathologist who didn’t look beyond the flagged regions. Who failed? How do you prevent it?

The answer the system usually gives is predictable: more supervision. Stricter protocols. Double verification. More humans in the loop.

It’s the wrong answer. And it’s wrong because it starts from a misdiagnosis of the problem.

The failure didn’t happen because supervision was lacking. It happened because the system’s design put the physician in a position where supervising was, cognitively, nearly impossible. Nine hours into a shift. Hundreds of alerts. A pre-loaded report with the answer already filled in. Asking for more supervision in that context is like telling someone swimming against the current to swim harder. The problem isn’t the effort. It’s the current.

If the problem is one of cognitive architecture, the solution has to be one of cognitive architecture too. Not more loops. A different model entirely.

How a physician actually reasons and why that matters

To understand why HITL fails structurally, you first have to understand how clinical reasoning actually works. Not the textbook version. The version that happens in real practice, under real conditions.

A physician facing a patient doesn’t run through a linear algorithm. What they do is build, in real time, a probability distribution. Every symptom, every finding on physical exam, every lab result acts as a Bayesian modifier it updates the likelihood of each diagnostic hypothesis based on what was already known. Experienced clinicians do this almost automatically, through what cognitive researchers call illness scripts: mental structures that organize knowledge about specific conditions and allow rapid pattern recognition.

That speed has a cost.

  • System 1: the intuitive, fast, pattern-based mode of reasoning, is efficient but fragile. It works well when the case is prototypical, when the presentation matches what the physician has seen before, when the data are complete and coherent. When any of those conditions breaks down, System 1 doesn’t catch it on its own. It needs…

  • System 2: the slow, deliberate, analytical mode, to step in and challenge the initial hypothesis.

The problem is that System 2 is expensive, cognitively expensive. Under fatigue, time pressure, or information overload, the brain suppresses it. Not out of negligence, out of economy. The result is reasoning that closes too early, anchors on the first plausible hypothesis, and stops looking for alternatives. Researchers call this premature closure, and it’s the most consistently documented source of diagnostic error in medicine.

The physician doesn’t fail because they don’t know enough. They fail because the cognitive system that allows them to know has real biological limits and the clinical environment routinely ignores them.

The blind spot nobody sees, because it’s in the map, not the terrain

There’s a particular kind of diagnostic error that’s especially hard to detect and to correct: the epistemic blind spot. Not the error that happens when a physician has the information and misreads it. The error that happens when the right diagnosis never enters the search space at all.

Physicians build differential diagnoses from what they know. That seems obvious and it is. But the implication doesn’t always get stated clearly: if a disease doesn’t exist on the clinician’s mental map, no amount of clinical evidence can surface it. From a Bayesian standpoint, if the perceived prior probability is zero, no data point will move it. The physician doesn’t actively rule it out. They simply never consider it.

This isn’t a training problem. It’s a structural one. Every clinician’s hypothesis space is shaped by their specialty, their geography, their case volume, and above all by the diagnoses they’ve seen frequently enough for those conditions to occupy real estate in memory. Rare diseases, atypical presentations, conditions a physician might encounter once in a career, all of those fall outside the map. Not because the physician is careless, but because the map is built from experience, and individual experience has limits that no supervision protocol can compensate for.

Confirmation bias compounds this. Once System 1 proposes a hypothesis, the brain tends to seek out evidence that confirms it and to downweight what contradicts it. That’s not irrationality, it’s cognitive efficiency. But in an environment where the AI has already proposed its own diagnosis, that bias gets amplified. The physician isn’t just closing down their own reasoning. They’re closing it down around the same answer the system already suggested.

Epistemic blind spots aren’t fixed by more supervision. They’re fixed by a system that actively searches for what the physician isn’t searching for and surfaces it before the reasoning closes.

The stethoscope didn’t diagnose. It amplified.

In 1816, René Laennec invented the stethoscope because he needed to hear the heart sounds of a patient more clearly. The solution was simple: a tube that amplified what the ear couldn’t catch on its own.

What matters isn’t the instrument. It’s what it didn’t do: the stethoscope didn’t diagnose. It didn’t generate a hypothesis. It didn’t tell the physician what to think. It just gave access to information that was otherwise unreachable, and left clinical reasoning to do the rest.

That’s exactly the distinction missing from the current debate on medical AI. Most systems are designed to give answers. The problem is that when AI gives an answer, the physician stops looking. The search space closes, and whatever falls outside it (the mitotic foci at the margin) becomes invisible.

What AI should be doing is not closing the search space. It should be expanding it. Acting directly on the limits we just described: availability bias, premature closure, epistemic blind spots. Not replacing the physician’s reasoning, but extending it into the territory that the individual map doesn’t cover.

The difference between a system that gives answers and a system that amplifies reasoning isn’t technical. It’s philosophical. And it has direct clinical consequences.

The cognitive prosthetic: what it means in practice

The concept I’m proposing to replace HITL is the cognitive prosthetic, an AI that doesn’t supervise or validate, but extends the clinician’s reasoning capacity in the same way the stethoscope extended their perceptual capacity.

The difference from current clinical decision support systems isn’t minor. It’s structural:

The most important difference is in that last row. A traditional decision support system generates a response with the same apparent confidence regardless of whether the case falls inside or outside its training range. A cognitive prosthetic knows when it doesn’t know and it says so.

That changes the entire dynamic of the interaction. Instead of an answer that closes down reasoning, the system offers an uncertainty map that keeps it open. The clinician is still the one deciding. But they decide with more hypotheses visible, with a clearer sense of where the limits of available knowledge lie and without the system having closed the search space before they had a chance to explore it.

The case that shows the difference

A young patient arrives in the ER with periumbilical pain migrating to the right lower quadrant, low-grade fever, and a positive McBurney’s sign. The diagnosis is almost automatic: acute appendicitis. The clinician orders a CT to confirm and starts preparing for surgery.

In a standard HITL system, the AI confirms the suspicion. The hypothesis space closes. The patient goes to the OR.

In a cognitive prosthetic system, something different happens. The AI analyzes the volumetric CT data not just to confirm appendicitis, but to actively search for predictors of alternative diagnoses that the human eye doesn’t prioritize in an acute presentation. In this case, the system flags two findings the clinician wasn’t looking for:

Appendiceal neuroendocrine tumors are rare, they appear in fewer than 1% of surgical specimens. A clinician who operates on ten appendicitis cases a month may not see one in years. By definition, they’re not on the mental map. They’re exactly the kind of diagnosis that the epistemic blind spot makes invisible.

The AI doesn’t change the decision to operate. That remains clinical. What changes is the surgeon’s situational awareness walking into the OR. Alerted to the possibility of an aNET at the base of the appendix, they can adjust the margins, inspect the mesoappendix more carefully, and, if the tumor measures between 1 and 2 cm or there’s invasion, avoid a second major surgery that nobody would have anticipated otherwise.

The AI didn’t replace the physician’s judgment. It gave them access to a diagnosis that was outside their usual search space, not because of any lack of competence, but because no individual can accumulate in one career the equivalent of thousands of simultaneous cases.

That’s exactly what a cognitive prosthetic should do: not answer, but expand. Not close, but open.

What this means for design, regulation, and accountability

If we accept this model, several things need to change at once and not all of them are technical.

In system design, the most immediate implication is that AI shouldn’t offer a single diagnosis with a confidence percentage. It should offer a probability distribution with explicit uncertainty and it should be able to say, when a case doesn’t resemble anything in its training data, that its reasoning in that specific context isn’t reliable. A system that always responds with the same apparent certainty, whether it knows or not, isn’t an amplifier of reasoning. It’s a source of false confidence.

In regulation, the paradigm shift means stopping to ask whether there’s a human in the loop, and starting to ask what quality of reasoning the human-AI combination actually produces. Does the system keep the hypothesis space open or close it? Does it communicate its uncertainty or hide it? Does it push the clinician to think, or allow them to stop thinking? Those are the relevant questions. None of them are part of any current regulatory approval framework.

In accountability, the model has uncomfortable implications. If AI is an extension of clinical reasoning, not a separate tool the physician watches over, responsibility remains human, but the standard of care has to evolve. A physician who doesn’t use available amplification tools in a case where they would have made a difference could find themselves, in the not-too-distant future, in a position legally comparable to a physician who today operates without preoperative imaging when there was a clear indication for it.

We’re not talking about replacing the clinician. We’re talking about giving them, for the first time, access to the collective knowledge of thousands of cases that no individual can accumulate in a single career.

The stethoscope didn’t make physicians less of physicians. It made them better physicians, because it gave them access to something their unaided senses couldn’t reach. Clinical AI has the same potential, but only if we stop designing it as a system that gives answers and start designing it as a system that amplifies questions.

The loop was a necessary starting point. It was how a cautious society learned to coexist with a technology it didn’t fully understand. But the loop isn’t enough anymore. Not because AI is good enough to replace the physician, it isn’t, and in many respects it never will be. But because the loop was designed to protect the legal system, not to improve clinical reasoning.

And those are different things.

Beyond the Slide is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

References:

  1. Bouabida K, Chaves BG, Anane E. The augmented physician: AI and the future of clinical cognition. Front Artif Intell. 2026 Feb 23;9. doi:10.3389/frai.2026.1744544

  2. Greengrass CJ. Transforming clinical reasoning—the role of AI in supporting human cognitive limitations. Front Digit Health. 2026 Jan 5;7. doi:10.3389/fdgth.2025.1715440

  3. Primer 3: The Role of Clinical Reasoning in Diagnostic Excellence | Coordinating Center for Diagnostic Excellence [Internet]. [cited 2026 Apr 10]. Available from: https://codex.ucsf.edu/primer-3-role-clinical-reasoning-diagnostic-excellence

  4. Scribd [Internet]. [cited 2026 Apr 10]. Clinical Reasoning in Differential Diagnosis | PDF | Medical Diagnosis | Sensitivity And Specificity. Available from: https://www.scribd.com/document/669425490/CLINICAL-REASONING

  5. Mann SE. Clinical reasoning.

  6. Weber P, Binder K, Krauss S. Why Can Only 24% Solve Bayesian Reasoning Problems in Natural Frequencies: Frequency Phobia in Spite of Probability Blindness. Front Psychol. 2018 Oct 12;9. doi:10.3389/fpsyg.2018.01833

  7. Changkui Li. The Double Helix Model for Epistemic Reconstruction in Contemporary TCM Education: Structured Co-Existence, AI Mediation, and a Five-Year Roadmap. MRHK. 2025 Dec 15;7(4):27–40. doi:10.6913/mrhk.070404

  8. Greengrass CJ. Transforming clinical reasoning—the role of AI in supporting human cognitive limitations. Front Digit Health. 7:1715440. doi:10.3389/fdgth.2025.1715440 PubMed PMID: 41561162; PubMed Central PMCID: PMC12813117.

  9. Exploring the feasibility of conversational diagnostic AI in a real-world clinical study [Internet]. [cited 2026 Apr 10]. Available from: https://research.google/blog/exploring-the-feasibility-of-conversational-diagnostic-ai-in-a-real-world-clinical-study/

  10. HealthManagement.org. Radiology Management, ICU Management, Healthcare IT, Cardiology Management, Executive Management [Internet]. [cited 2026 Apr 10]. Available from: https://healthmanagement.org/c/healthmanagement/issuearticle/are-we-ready-for-ai-mental-prosthesis

  11. Davies S. Neurological Scaffolding: How Artificial Intelligence Can Support Cognitive Recovery. Medium [Internet]. 2026 Mar 13 [cited 2026 Apr 10]. Available from: https://medium.com/@driversrepublic/neurological-scaffolding-how-artificial-intelligence-can-support-cognitive-recovery-e96c7eae5f17

  12. Şimşek O, Şirolu S, Irmak YÖ, Hamid R, Ergun S, Kepil N, et al. Challenges and predictive radiological findings in the diagnosis of neuroendocrine tumors in patients with acute appendicitis. Ulus Travma Acil Cerrahi Derg. 2024 Nov 4;30(11):780–5. doi:10.14744/tjtes.2024.70392 PubMed PMID: 39498711; PubMed Central PMCID: PMC11843387.

  13. Andrini E, Lamberti G, Alberici L, Ricci C, Campana D. An Update on Appendiceal Neuroendocrine Tumors. Curr Treat Options Oncol. 2023;24(7):742–56. doi:10.1007/s11864-023-01093-0 PubMed PMID: 37140773; PubMed Central PMCID: PMC10271885.

  14. Moris D, Tsilimigras DI, Vagios S, Ntanasis-Stathopoulos I, Karachaliou GS, Papalampros A, et al. Neuroendocrine Neoplasms of the Appendix: A Review of the Literature. Anticancer Research. 2018 Feb 1;38(2):601–11. PubMed PMID: 29374682.

  15. Shanthi I, Na S. When AI intervene Clinical Decision-Making: The influence of Organisational Support, Cognitive Load, and Perceived Autonomy. Malaysia Journal of Invention and Innovation. 2025 Feb 12;4:33–9. doi:10.64382/mjii.v4i3.112

  16. As’ad M, Faran N, Joharji H. AI-Supported Shared Decision-Making (AI-SDM): Conceptual Framework. JMIR AI. 2025 Aug 7;4:e75866. doi:10.2196/75866 PubMed PMID: 40773762; PubMed Central PMCID: PMC12331219.

  17. Swinton M. The Sentience Halo: The Risk of Unopposed Mirroring and Perceived Awareness in AI Therapy. 2025. doi:10.31234/osf.io/g6axf_v1

  18. Hsu JYC. Tonal Isomorphism: A Methodology for Cross-Domain Mapping in the Generative Age. Philosophies. 2025 Nov 5;10(6). doi:10.3390/philosophies10060122

  19. Position Paper: Integrating Explainability and Uncertainty Estimation in Medical AI [Internet]. [cited 2026 Apr 10]. Available from: https://arxiv.org/html/2509.18132v1

  20. Lindenmeyer A, Blattmann M, Franke S, Neumuth T, Schneider D. Towards Trustworthy AI in Healthcare: Epistemic Uncertainty Estimation for Clinical Decision Support. J Pers Med. 2025 Jan 31;15(2):58. doi:10.3390/jpm15020058 PubMed PMID: 39997335; PubMed Central PMCID: PMC11856777.

  21. Out of Distribution Detection: Knowing When AI Doesn’t Know [Internet]. 2025 [cited 2026 Apr 10]. Available from: https://www.sei.cmu.edu/blog/out-of-distribution-detection-knowing-when-ai-doesnt-know/

  22. Lotfi D, Mahani MAN, Koohi-Moghadam M, Bae KT. Safeguarding AI in Medical Imaging: Post-Hoc Out-of-Distribution Detection with Normalizing Flows. IEEE J Biomed Health Inform. 2026 Feb 19;PP. doi:10.1109/JBHI.2026.3666215 PubMed PMID: 41712399.

  23. Parmar M, Silpasuwanchai C. From Overload to Convergence: Supporting Multi-Issue Human-AI Negotiation with Bayesian Visualization [Internet]. 2026 [cited 2026 Apr 10]. Available from: http://arxiv.org/abs/2603.22766doi:10.1145/3772318.3790358

  24. Sharma S, Singh M, McDaid L, Bhattacharyya S. XAI-based Data Visualization in Multimodal Medical Data [Internet]. bioRxiv; 2025 [cited 2026 Apr 10]. p. 2025.07.11.664302. Available from: https://www.biorxiv.org/content/10.1101/2025.07.11.664302v1 doi:10.1101/2025.07.11.664302

  25. Iseko A. Diversity as Ethical Infrastructure: Reimagining AI Governance for Justice and Accountability. Int J Sci Technol Soc. 2025 Sep;13(5):190–204. doi:10.11648/j.ijsts.20251305.13

Read on beyondtheslide.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.