Artificial intelligence in health care has a problem that I continue to emphasize in these articles. Every week brings a press release or an article about a new company or model that reads a scan, drafts a note, or flags a diagnosis with magical speed and accuracy. Medical technology seems to captivate readers because it provides hope (and sometimes actual solutions) that we can solve some of our most horrible afflictions.
The articles, marketing messaging, and even the spoken beliefs of policy makers seem to romanticize AI even more than is usual. What little evidence we do have for most technologies comes from retrospective studies and internal validation sets, where a model is developed and then turned loose on a historical set of data and then evaluated on how well it predicts things or otherwise functions. Those studies are useful. They tell us that a model can do something under controlled conditions. Said differently, the model can transform raw data into information. However, the models almost never tell us whether the model changes anything once it is dropped into a real clinic, with real clinicians who are busy, skeptical, and free to ignore it, and real patients who may or may not show up for the follow-up test.
This is why I keep coming back to randomized trials. A randomized controlled trial isolates the effect of the technology from everything else swirling around it. With AI specifically, that isolation matters more than usual, because the effect of an AI tool is almost never just the model and the systems that create the data on which the AI operates are dynamic and complex. Said differently, it is the model plus the clinicians that use it, plus the workflow it lives in, plus the other external considerations related to the patient’s condition and the context of where care is being rendered.
The single most important question in any of these studies looking at the effects of AI is simple: compared to what? To what do we compare the effects of the AI technology? In drug trials, we use placebos, or pills that look and feel very similar to the experimental drug. Or, we compare a drug to another drug to see how they compare.
In the case of AI, what do we compare the AI to? Is it compared to the current non-AI process? Is it compared to an expert physician? Is it compared to a new graduate physician? Is it compared to another medical device? These are material questions for how we understand the value of AI in health care. Unfortunately, these questions are increasingly missed in the current discussion on the value of AI in health care.
Three recently published clinical trials give us evidence across three very different use cases. I mention the use cases because AI is not a monolith, but rather each technology or algorithm or set of algorithms has an intended use. In this case, one AI technology uses AI to find hidden disease. One uses AI to support mental well-being. One uses AI to change how clinical work gets done.
Each study also illustrates a different randomized study design suited to the key question being asked and the comparator of interest. Taken together, they are a useful reminder of both what specific AI technologies can and cannot deliver based on the current state of the evidence.
The first study, the DULCE trial published in Nature Medicine, tackles a gap in primary care. Advanced chronic liver disease affects an estimated 2 to 5 percent of the general population, and it is frequently caught only after the liver has already failed.
The AI technology, a machine learning model, reads a standard 12-lead electrocardiogram and flags patients at higher risk of advanced liver disease. The key data element here is that cirrhosis produces measurable cardiovascular changes, so an ECG carries a faint signal of liver health. Because ECGs are cheap, fast, and already performed constantly in at-risk populations, this turns a routine cardiac test into an opportunistic liver screen. The underlying model was trained on 77,507 individuals and posted strong discrimination in validation, with a C-statistic of 0.858 (it performs much better than random chance).
The study design here was a pragmatic, cluster-randomized trial run across 45 primary care practices in southern Minnesota and Wisconsin. Ninety-eight care teams were randomized, not individual patients. When a positive result came back, clinicians in the intervention arm were notified and prompted to consider confirmatory liver testing. The comparator was usual care: clinicians who received no alert and practiced exactly as they otherwise would.
That comparator, paired with the pragmatic design, is what makes the results believable. This is not a test of whether the model is accurate in a set of testing data from the ECGs, it is a test of the entire care pathway, alert included, in the messy reality of primary care. Importantly, it mirrors how the technology would actually be adopted at the care team level. Patients do not adopt technologies, so randomizing at the care team level (as opposed to the patient level) tests the technology in a more realistic scenario.
As for the results, the intervention doubled new diagnoses of advanced liver disease, from 0.5 percent in usual care to 1.0 percent (odds ratio 2.09). But the yields still landed below what epidemiology would predict (the expected prevalence), and the authors note that clinician follow-through was only moderate. The alert works only when someone acts on it, and plenty of clinicians did not order the next test. The AI found the signal. The care system did not complete the loop to allow the effects of the AI to be realized.
This is a critical point to note because most clinical AI technologies are just cool tools and not magical entities that solve all our problems. Once a technology is validated the more daunting task in health care is integration into daily use where the AI is subject to humans and all the complexities, carelessness, and biases that come with them.
The second study focuses on prevention and is delivered via a smartphone as opposed to a clinician mediated use case. A multi-institutional, preregistered, longitudinal RCT evaluated Flourish, a consumer generative-AI app built around a chatbot named Sunnie, powered by OpenAI’s GPT-4o. Most clinical chatbots are “deficit-based” in that they wait for symptoms of depression or anxiety and then try to treat them. Flourish is “strengths-based” and proactive, nudging users toward gratitude, connection, and mindfulness before distress escalates.
In the study, researchers enrolled 486 undergraduates across three campuses and followed them for six weeks in the fall of 2024, spanning the natural stress of a real semester. (Side note: most of our psychology studies are based on undergraduate students, which is a big weakness because they are not representative of the general population).
The study practiced good science and preregistered their hypotheses, used an intent-to-treat approach that kept every participant in the analysis regardless of how much they used the app or if they dropped out, and measured outcomes at four points in time rather than one.
The comparator, though, is where many of these mental health app studies are inhibited. The authors note it too.
Flourish was tested against a waitlist control: students with their existing campus supports but no app. That tells us whether the app produces effects compared to doing nothing. It cannot tell us whether the benefit comes from the app’s specific design or if merely paying attention to well-being drives the observed results. If we wanted to contextualize the effects of the app in the greater well-being tool kit, we would need an active comparator such as psychotherapy, SSRI medications, another app, or something else with existing effects.
The results suggest that the app is mildly effective. Students using the app reported higher positive affect, more resilience, stronger belonging, and less loneliness, and they were buffered against the declines in mindfulness and flourishing that the control group experienced. Yet there was no significant effect on depression, anxiety, or stress and effect sizes were modest.
There are many mental health apps on the market and the addition of LLMs will make them more personalized and engaging. Given the massive prevalence of anxiety and depression, we should pay attention to this space and keep looking for apps and study designs that show bigger and more significant effects.
The third study is not about diagnosis or treatment effects. It is about workflow, which is one of the “hottest” areas of AI in care delivery at the current time. At least according to the hype articles.
Published in the Journal of the American Heart Association, the AI-Echo trial evaluated US2.ai (which is both a website and the technology name apparently, which is very cool), an FDA-cleared platform that automatically measures echocardiographic parameters from ultrasound images. The sonographer’s job, in theory, shifts from measuring everything by hand to verifying what the AI produced, with expert echocardiologists finalizing every report.
The study design in this study is a single-center randomized crossover trial. Four sonographers at a Tokyo hospital were randomly assigned to work with AI on some days and without it on others, across 585 patients over 38 study days. Because each sonographer serves as their own control, the design cancels out the individual differences in skill and speed that would otherwise affect the results, although the use of 4 sonographers is very small sample, so keep that in mind as we evaluate the results. The comparator is the same operators doing the same work, their normal manual workflow, on their non-AI days.
The results suggest that examination time dropped from 14.3 to 13.0 minutes, daily exams per sonographer rose from 14.1 to 16.7, and the number of parameters analyzed per study jumped 3.4-fold. The sonographers reported lower mental fatigue on AI days. Image quality also improved. The AI’s measurements agreed with the final expert-endorsed values within acceptable limits for roughly 90 percent of parameters.
Many policymakers and hospital operators are actively exploring whether AI can enhance the productivity of care. Whether that productivity trickles down to lower costs is another question entirely, but evidence is starting to support the idea that bumps in productivity are possible.
In every case, the AI was rarely the whole story. DULCE’s benefit depended on whether clinicians followed up. Flourish’s meaning depends on whether the comparison was fair. AI-Echo’s payoff depended on what sonographers did with the time they saved. The model is one input. The system around it decides how much that input is worth.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.