I read a lot of AI vendor pitches in a given month, and a good share of them boil down to some version of “we’ll fix your clinical documentation.” It’s an easy pitch to make and a hard one to verify, because most of what’s been published on ambient scribes so far has been a small pilot or a handful of enthusiastic early adopters talking to a friendly journalist. So when a real-world evaluation of an ambient AI scribe tracking 2.3 million patient encounters over sixteen months came out of a Spanish hospital network this week, it was the first study in a while with the scale to actually say something.
Alcázar-Peral and colleagues at the Quirónsalud Healthcare Network in Madrid ran a multicenter observational study of their ambient documentation tool, Scribe, across outpatient care from September 2024 to December 2025. Of 11.6 million outpatient visits in that window, 2.3 million were completed using the tool — an overall usage rate above 20 percent, with monthly adoption climbing from 2.7 percent at launch to a peak above 31 percent by the study’s end.
That climb is the number most people will skim past. Digital health tools that plug into the electronic health record typically take years to reach a third of an organization’s visits, if they get there at all. Going from single digits to nearly a third of every outpatient visit in a little over a year puts ambient documentation in different company than almost any other EHR-adjacent technology of the last decade.
Fast adoption alone doesn’t prove a technology works — it proves clinicians didn’t hate it enough to stop using it, which is a lower bar than most vendor decks imply. What makes this study genuinely useful is that it didn’t stop at the adoption curve. It kept measuring transcription accuracy, consultation length, and clinician experience across the entire sixteen months, which is where the honest numbers show up.
Here’s the honest number. Average consultation length barely moved: 15.01 minutes without the tool versus 14.65 with it, with several months favoring Scribe before converging toward parity as workflows matured. Call it twenty seconds, and even that faded as clinicians settled into the tool. Transcription accuracy, measured as semantic agreement between the AI-generated note and the actual encounter, stayed remarkably stable the entire time (87 to 89 percent).
That’s a smaller time win than UCLA’s own trial found for its best-performing scribe, and this one comes from a real deployment ten times the size. The study’s authors offer a more interesting explanation than “the tool doesn’t save time”: their working hypothesis is that whatever minutes got reclaimed weren’t banked as shorter visits — they were reinvested into the visit itself, in patient interaction and the kind of care that doesn’t show up on a stopwatch. Clinician-experience surveys, run in two waves (184 clinicians, then 163), support that read: five of the seven measured domains improved meaningfully as the tool matured, including administrative workload and stress and quality of life, with gains holding steady across heavy and light users alike.
That pattern lines up with the most rigorous controlled evidence available. UCLA Health ran the first randomized trial of ambient scribes in U.S. outpatient practice — 238 physicians, three arms, two months — and found a real but modest time savings for one tool (about 9.5 percent versus control), no significant change for the other, and improvements in burnout and workload scores for both. Different country, different design, same basic shape: genuine gains, smaller than the pitch, uneven across tools.
The UCLA trial also surfaced something a usage log can’t: occasional clinically significant inaccuracies in the AI-generated notes. Senior author John Mafi’s response to that finding is the line worth keeping: “This technology requires active physician oversight, not passive acceptance.”
That’s a governance requirement dressed up as a safety caveat, and it applies regardless of how smoothly a rollout scales. A tool that’s improving on average, at 2.3 million encounters, can still be wrong in one specific chart on one specific day — and the fix for that isn’t a better model. It’s a review step nobody’s tempted to skip.
Thanks for reading Digital Evolutionary! This post is public so feel free to share it.
Zoom out further and the field’s own literature agrees on where this goes next. A recent narrative review of ambient scribe research — Razaghi and colleagues, published this past January — found the same recurring limitations across eighteen studies: omission errors, inconsistent performance, and what they call “note bloat,” notes that get longer without getting more useful. Their conclusion is that today’s scribes, however well adopted, are still fundamentally passive. They wait for a clinician-led encounter, transcribe it, and stop.
That’s the distinction I keep coming back to in my own work. A scribe that transcribes reliably at scale is a genuinely useful tool. A system that turns the same encounter into a structured, reasoned signal (flagging what needs follow-up, catching the pattern a tired physician might miss at 4 p.m. on a Friday) is a different category of thing. The first is documentation done well. The second is documentation that starts pulling weight the note never could.
Two million consultations is real proof that clinicians will adopt this technology, and keep using it. It’s not proof that the technology has moved past capturing what happened toward understanding what it means. That gap is still the frontier, and it’s the one worth watching.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.