Every healthcare AI pitch I sit through leads with the same slide. A risk score, a curve, an AUC somewhere north of 0.85, and a confident promise that the model will catch deterioration or readmission or sepsis before a clinician would. The demo always lands. The model is almost never the thing that decides whether any of it actually works.
I went back to that instinct after reading a systematic review of predictive models embedded in electronic health record systems, not models that were built and validated and written up, but models that were wired into live clinical use and then measured against patient outcomes. Of the 32 studies that reported clinical outcomes, 22 showed improvement after the model went live. A 69 percent hit rate is encouraging on its own. What’s more useful is what separated the wins from the misses, because it had very little to do with how good the model was.
The pattern that runs through the successful implementations is that the prediction arrived where a human was positioned to do something about it. The review’s authors highlight a heart-failure readmission model described by Amarasingham that was deliberately not run on weekends or holidays. It fired on weekdays only, specifically because that was when a heart-failure case manager was available to coordinate follow-up for the patients it flagged.
Sit with that for a second. The team intentionally narrowed when their model spoke, not to improve its accuracy, but to match it to the hours when its output could be converted into action. A flagged patient on a Saturday is a number with nowhere to go. The same flag on a Tuesday, handed to a case manager whose entire job is the next step, is a workflow. The model didn’t change between those two days. Everything around it did.
This is the part the demo slide never shows. Anticipating a patient’s needs before they arise is only half a sentence. The other half (and then a specific person does a specific thing in time to matter) is where the value lives, and it’s the half that almost never makes it into the procurement conversation.
The clearest cautionary tale here is the Epic Sepsis Model, a proprietary tool that had been switched on at hundreds of US hospitals before anyone outside the vendor had rigorously checked whether it worked. When Andrew Wong, Karandeep Singh, and colleagues at Michigan Medicine externally validated it across more than 38,000 hospitalizations, the results were rough: an AUC of 0.63, well below the 0.76 to 0.83 Epic had reported to its customers, and a sensitivity of 33 percent. The model missed roughly two-thirds of actual sepsis cases while flagging enough patients to bury the clinicians who had to chase every alert.
The most instructive finding wasn’t the miss rate. It was why the alerts that did fire were so often useless. The model used antibiotic administration as one of its inputs, and antibiotics get ordered when a clinician already suspects sepsis. So the score frequently climbed after the bedside decision had been made. The “early warning” was, a lot of the time, an echo of a judgment call the team had already reached on their own. The validation found the model surfaced only about 7 percent of the sepsis cases clinicians hadn’t already caught.
A prediction that mostly confirms what you already decided isn’t anticipation. It’s latency with a dashboard. And it carries a real cost: every false alert spends a clinician’s attention, and attention is the most rationed resource in the building. Epic eventually overhauled the model and started telling hospitals to retrain it on their own patient data before going live, a quiet admission that a model shipped as one-size-fits-all wasn’t one.
Thanks for reading Digital Evolutionary! This post is public so feel free to share it.
Put those two stories next to each other and the lesson is hard to miss. The heart-failure model worked because someone designed the conditions under which its prediction would turn into care. The sepsis model struggled because it was bolted onto a workflow that hadn’t been redesigned to receive it. It just fired into the same stream of alerts everyone was already learning to ignore.
The review names this directly. Alert fatigue, lack of training, and added work for the care team show up again and again as the reasons implementations stall, and the authors note that plugging in a previously validated model without local customization tended to make those problems worse, not better. The off-the-shelf score that performed beautifully somewhere else is not a head start. It’s an unvalidated assumption about your population, your staffing, and your existing alert load.
The implication for anyone building or buying one of these is less glamorous than the model work. The hard, valuable design happens around the prediction: deciding who owns the response, whether that person actually has the capacity to own it, what the alert threshold should be for this population rather than someone else’s, and (the step everyone skips) what gets pulled out of the alert stream to make room for the new signal. That’s operational design, not data science, and on most projects it belongs to nobody.
So the first question I want to ask about any predictive tool (ours included) is no longer “what’s the AUC.” It’s narrower and more annoying: when this thing is right, who acts, what do they do, and does the signal reach them early enough to change the outcome? If there’s no clean answer to that, a better model won’t save the project. It’ll just be a more accurate way to generate noise. The prediction was the easy part. The workflow it lives inside is the whole job.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.