RSS Amplifier

A&Ox2 · Jul 28, 2026

Tales from the Wards 2.0

0
Sign in to vote or save

Anil Makam · A&Ox2

Practicing medicine is the systole to the diastole of research and evidence appraisal. Hospital medicine, in particular, is a fast-paced, intellectually demanding exercise in solving complex, deeply human problems—often with near-immediate feedback.

Every stretch on service teaches me something new (usually several somethings new) despite nearly two decades in practice caring for thousands of patients. Clinical presentations are endlessly variable. Knowledge evolves. There are thousands of ways to be sick enough to require hospitalization. Few careers offer that kind of breadth and pace.

Tales from the Wards is a recurring series where I share what I learned with others.

Below is a rundown of the medical inaccuracies, controversies, and errors I encountered in the wild during my last stretch on the consult service. I was planning to include cases from from my own direct care service, including two of my own errors, but the list became unwieldy. I’ll save those for a Tales from the Wards 2.5.

The throughline of these cases is whether today’s AI, or a version that exists in the next 3 years, could help.

For clinicians, please play along, and chime in the comments yes/no for each of the 8 scenarios represented by the 6 cases below.

1

I staffed a new consult for a guy disengaged with medical care with a gnarly gangrenous toe. I was told by team members who reviewed his chart that his hemoglobin A1c was 6.7% and his foot X-ray showed acute osteomyelitis. Neither of those were true. His A1c has been consistently over 15% and his X-ray said no such thing. Although after discussion with the surgeon, he did have strong clinical evidence of osteomyelitis.

Could AI be more factually accurate? Yes. This is low hanging fruit.

2

A middle-aged man presented with subacute left hand numbness and found to have cervical myelopathy with cord compression. I consulted on him postoperatively after decompression surgery. Unrelated to the reason for consult, I learned that he also developed subacute dysarthria at the same time as his hand numbness began. His brain MRI upon admission showed a “chronic” right thalamic stroke. At that time, the patient was unable to provide an accurate history and his closest family members who he lived with were not present. His history morphed into chronic dysarthria, and he was considered ‘at baseline’—perhaps the two most dangerous words in the hospital.

After a multidisciplinary discussion, I re-diagnosed his presentation as a subacute right thalamic stroke one month prior to his admission with incidental cord compression. Whether some component of his hand numbness was attributable to the cord compression cannot be teased out at the bedside. But, probabilistically, developing slurred speech from a stroke and hand numbness from spinal cord compression simultaneously would be more rare than getting struck by lightning twice.

Could AI have made the diagnosis? No. A correct diagnosis required careful history-taking skills to recognize when a patient’s history is incomplete or inaccurate and the need to corroborate symptom onset with his family who saw him everyday. An AI would be able to sift through the MRI images or the report itself to identify a possible lesion. But the radiographic chronicity combined with documented ‘chronic dysarthria’ would have stumped its compute. Perhaps a future version with ambient sight and sound capabilities would recognize the need for better history taking and bring to bear the clinical pearl that I learned that strokes on MRI can evolve from the acute to the ‘chronic’ phase as early as a few weeks out.

3

An older gentleman with mild COPD, compensated cirrhosis, and recurrent episodes of self-limited hemoptysis (typically while receiving perioperative heparin infusions) underwent salvage bypass surgery for severe peripheral arterial disease. Unsurprisingly, he developed another bout of hemoptysis while on heparin. His bleeding gradually improved, and he was discharged with stable, persistent small-volume hemoptysis after tolerating rechallenge with his oral blood thinner. The thought was that his hemoptysis was a combination of friable airways and mild coagulopathy from cirrhosis that was unmasked by heparin, but ultimately no one was certain.

The dilemma in this case wasn’t the diagnostic uncertainty. Nothing was amiss. Practicing medicine is hard. The controversy was the interpretation of his CT scan findings of his lungs. It showed new large patchy groundglass opacities (GGOs) and confluent consolidation in dependent portions of his right lower lobe (RLL), “consistent with aspiration pneumonia.” It certainly looked impressive. But, he never had shortness of breath, chest pain, chills, sweats, headaches, trouble swallowing, delirium, or loss of consciousness. He never spiked a temperature nor had leukocytosis during his hospital stay. My colleague, who first saw him at night on our consult service after the hemoptysis began and his CT resulted, showed remarkable restraint by not starting antibiotics. I saw the patient the following morning, and concurred. The patient’s pre-test probability for pneumonia was incredibly low. We transferred his care to a medicine service now that hemoptysis, and not his peripheral vascular disease, was his most pressing concern.

The new medicine attending disagreed and prescribed 5 days of broad spectrum antibiotics. Differences in management are common. Out of curiosity I reached out wondering if the patient’s clinical course evolved. It hadn’t. The decision to treat rested entirely on how “nasty” the CT appeared.

Would this patient’s clinical course have been similar without antibiotics? I would wager so, but I have no crystal ball to ascertain the counterfactual. While he had no adverse effects from antibiotics, these decisions matter because you don’t know ex ante who gets kidney injury, a drug rash, or C. difficle.

Could an AI reconcile differences in Bayesian clinical decision making? No. The patient’s pre-test probability for aspiration pneumonia was undoubtedly very low. I am doubtful that the three attendings involved would disagree. The difference in clinical management rested on how much weight to lend to the CT findings. Does ‘nasty’ GGOs in the RLL have such a strong positive likelihood ratio to push the post-test probability past the treatment threshold (say 50% perhaps)? No study tells us the likelihood ratio of increasingly “nasty” GGOs. In my experience, infiltrates without corresponding signs or symptoms are rarely diagnostic of pneumonia, and these CT findings alone were insufficient to cross my treatment threshold. No AI in the foreseeable future can practice the art and science of medicine.

4

An older man with a liver transplant for hepatocellular cancer from now cured hepatitis C was hospitalized for surgical repair of his fractured leg after he was pinned by a car against a wall. I was consulted the following day for new kidney injury after his tibia was repaired. When my team was presenting his case to me, his hemodynamics during the operating room were unknown. His alcohol history, relevant given his liver transplant, was also underreported. We reviewed his chart together and learned that he was hypotensive during the operating room and was on a vasopressor for several hours. At the bedside, I also learned that he drank so often bartenders knew him by name and from a later chart review, that his PeTH levels from another hospital (which are stored in a separate section of our EHR) were consistently through the roof, meaning that he drank like a fish. Consistent with an aphorism I learned as a wards resident (thanks Dr. Ellwood Jones!), people lie about sex, drugs, and taking their iron pills (really, all pills) With more complete information gathering, we diagnosed him with prerenal acute kidney injury from transient post-operative hypotension due to sedation and surfaced his very severe AUD, which makes liver transplant failure much higher.

Could an AI diagnose his acute kidney injury? Yes. An AI would have no difficulty reviewing hour-by-hour intraoperative vitals and identifying vasopressor use, and could reason sufficiently well enough to connect his intraoperative hypotension to his AKI.

Could an AI recognize severe alcohol use disorder and its relevance? No. An AI could locate relevant laboratory results buried elsewhere in the EHR that showed excessive alcohol consumption. My hesitation is whether an AI, without prompting, can identify the unknown knowns. Would it independently recognize severe alcohol use disorder from chart review and appreciate its relevance in a liver transplant recipient? I am doubtful.

5

An older gentleman with Parkinson’s disease was hospitalized for a spinal vertebral fracture after falling. He was nonoperatively managed with a brace, pain control, and physical therapy. He was euvolemic. Eating and drinking, albeit less than normal. With physical therapy the day prior he was noted to have asymptomatic orthostatic hypotension. This was attributed to volume depletion, and to a lesser extent, his amlodipine and finasteride. He was given intravenous fluid boluses which transiently improved his orthostasis, but was short lived. I saw him the following day, again with asymptomatic orthostasis. However, I noted that while his systolic blood pressure dropped > 20 points after standing up, his heart rate did not change with position, which was consistent with his prior orthostatic vitals. Given his Parkinson’s disease, I re-diagnosed him as neurogenic orthostatic hypotension and recommended harm reduction measures (eg. having something sturdy to grab) in the advent he ever became dizzy when standing up too quickly).

Could an AI diagnose his orthostatic hypotension correctly? Yes. A chart review recognized the the lack of heart rate change despite a sizeable drop in blood pressure (∆HR /∆SBP < 0.5 has a negative LR of 0.1 for hypovolemia) should indicate a neurogenic etiology given his euvolemic state and his Parkinson’s disease which can cause autonomic dysfunction. This requires manipulation of his vital sign changes into a ratio, an understanding of the diagnostic test characteristic from literature, and chart review of his physical exam and comorbidities. This seems solvable.

6

A young woman with uncontrolled diabetes (A1c 13%) was admitted for left upper thigh cellulitis with a phlegmon from folliculitis that evolved over several days in the hospital into an abscess requiring incision & drainage in the operating room. We were asked by the surgical service to see her for hyperglycemia the following day. She was billed to me by my team as “straightforward” with a correspondingly pithy presentation of standard hyperglycemia inpatient management – tweak her insulin, continue sliding scale, and connect to her primary care doctor. Reviewing her labs on my phone, I immediately noticed a pattern in her basic metabolic panel that I’m highly attune to in patients with poorly controlled diabetes: a high blood glucose, low bicarbonate, and an elevated anion gap. She also had ketones on her urine analysis. After brief hallway teaching on diagnosis and management of early diabetic ketoacidosis (no infusions needed, just dose subcutaneous regular insulin semi- frequently and ask your patient to drink more water), we entered her room.

During my history, I learned this was not her first rodeo. She described prior inguinal abscesses that began as “pimples”. She was obese, had scarring in her opposite groin, and suddently the diagnosis shifted from uncomplicated soft skin tissue infection to suspected hidradenitis suppurativa (HS), a chronic inflammatory skin disorder that can lead to disabling deep-seated inflammatory nodules, abscesses, and tunneling in the intertriginous regions like the groin. I am not particularly savvy in skin. Dermatologists confirmed stage 1 disease after identifying open comedones of photographs of her dressed wound that I had overlooked. In addition to the planned course of oral antibiotics, she was discharged with topical therapies.

Could AI recognize early DKA? Yes. The laboratory pattern is classic

Could an AI differentiate early hidradenitis from garden variety SSTI? No. Even perfect chart review cannot recover history that was never elicitied. Nor am I convinced that an AI with ambient vision and audio could independently synthesize recurrent inguinal infections, groin scarring, subtle comedones, and the broader clinical context than the internists and surgeons who saw her.

I surmised that AI could help in half of the 8 discrete issues represented by 6 cases I saw in my practice.

Now it’s your turn. For each of the 8 scenarios, what’s your vote—yes or no? Which ones did I get wrong? Which cases do you think AI could solve, and which still fundamentally require human judgment? Leave your scorecard in the comments. I’ll be interested to see where the consensus emerges.

No posts

Read the original on aaox2.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.