RSS Amplifier

The AI Value Gap · Mar 30, 2026

No.20: The Regulation Paradox

0
Sign in to vote or save

Amin Mrini · The AI Value Gap

TL;DR: The industries we expected to resist AI longest - medicine, law, finance - are adopting it fastest. The reason: decades of codified knowledge give AI both a reliable substrate and a measurable ROI surface that most of the economy lacks. Medicine is the clearest proof, with 81% physician adoption in under three years in the US. But we are scaling adoption faster than we are measuring whether it works - and in a profession that codified its knowledge specifically because errors kill people, that gap should worry us more than it does. OpenEvidence is the case study, and its lessons extrapolate directly to enterprise knowledge work.

Two years ago, it was reasonable to assume that regulated industries would be among the slower adopters of AI. The liability exposure, the compliance burden - all pointed to a cautious, multi-year ramp. The thinking was that AI would move fastest where the rules were loose - marketing, software, media - and reach medicine or law once the technology had proven itself elsewhere.

The speed of what actually happened caught nearly everyone off guard. Healthcare now captures nearly half of all vertical AI spend - tripling YoY and outspending the next four verticals combined. Harvey, the legal AI platform, hit an $8bn valuation in December 2025 after a $160M raise led by Andreessen Horowitz, and was reportedly in talks for another $200M at $11bn just two months later - with over 1,000 customers and $190M in annual recurring revenue. Financial services accounts for more than 20% of all AI spending globally, according to IDC. Three of the most regulated sectors in the Western economy moved first, and by a wide margin.

Regulated industries share a feature that has nothing to do with regulation per se: decades of peer-reviewed, hierarchically organised knowledge that has been written down, structured, and made retrievable. Medicine has its clinical guidelines, indexed literature in PubMed, structured pharmacological databases, and formalised diagnostic criteria. Law has centuries of case law, annotated statutes, and regulatory filings. Finance has accounting standards, Basel frameworks, and filing taxonomies so granular that they were practically built for machines to read.

LLMs perform well with structured, authoritative information. They struggle with ambiguity, with problems that require cross-functional judgement, with contexts where the answer depends on organisational or tacit knowledge. The peer-reviewed research backs this up. Dell’Acqua et al. (2026), the Harvard/BCG “jagged frontier” study published in Organization Science this March, tested 758 BCG consultants on GPT-4 and found a clean split: 25% faster and more accurate on tasks inside AI’s capability frontier, 19% worse on tasks outside it. One way to read that frontier is through codification: the more structured and externally representable the task, the better AI tends to perform.

But industrialised knowledge does something else beyond making information machine-readable: it makes value legible. You can measure error rates in clinical diagnosis. You can track time-to-completion in document review. You can quantify compliance risk in financial reporting. These industries didn’t just adopt AI first because the knowledge was structured - they adopted first because the pain was already quantified and priced. Healthcare’s admin cost crisis, law’s billable hour pressure, finance’s compliance cost explosion - these are problems with visible P&Ls attached to them.

The most striking evidence comes from the AMA’s 2026 Physician Survey on Augmented Intelligence, released this March. It surveyed 1,692 physicians between January and February 2026, and the adoption numbers are extraordinary: 81% of US physicians now use AI in clinical practice, up from 38% in 2023 (for reference: The US Census Bureau shows 18% of US businesses using AI for ANY business function).

The highest adoption rates cluster around structured, documentation-heavy tasks: 39% of physicians use AI for summaries of clinical and research information (up from 13% - triple the rate), 30% for patient discharge instructions, 28% for billing documentation, 28% for chart summaries and progress notes. These are tasks where the acceptable output format is known and the cost of a minor error is relatively contained. Assistive diagnosis - the harder application that requires genuine clinical judgement - sits at 17%. Doctors are reaching for AI where the task is well-defined, and they are slower to trust it where judgement is required.

The case study that captures both the promise and the problem is OpenEvidence. The platform claims 40% of US physicians as daily users, 760,000 registered clinicians, and 18 million consultations per month. By the company’s own account, 100 million Americans will be treated this year by a clinician using OpenEvidence. On pure adoption metrics, it’s one of the fastest-scaling AI products in any vertical.

With adoption at that scale, you'd expect the quality-of-care uplift to be well documented. I was surprised by how thin the evidence is. The Mayo Clinic evaluation (Hurt et al., 2025) studied five patients with common chronic conditions - hypertension, diabetes, COPD, heart failure, chronic kidney disease. Four physician raters assessed the responses across multiple dimensions. OpenEvidence scored well on clarity (3.55 out of 4) and relevance (3.75 out of 4). But impact on clinical decision-making - the metric that actually matters for patients - scored 1.95 out of 4, the lowest of all metrics evaluated. The platform “primarily reinforced rather than modified plans.” In five cases, not a single treatment decision changed.

Then there’s the MedXpertQA pilot (a preprint, not yet peer-reviewed): 100 complex subspecialty questions, the kind of cases where AI could theoretically add the most value by surfacing evidence a generalist might miss. OpenEvidence scored 31-34% accuracy, with 75% discordance between its reasoning methods. On hard, real-world clinical scenarios, it was wrong more than two-thirds of the time. The authors' conclusion: "expert oversight remains essential" - pretty much word for word the conclusion I drew on AI coding.

Five patients is not a definitive study. Accuracy on curated exam-style questions doesn’t map directly to clinical impact. OpenEvidence may well be improving outcomes in ways no one has measured yet - reinforcing correct diagnoses, catching drug interactions, surfacing relevant literature faster. But read the Mayo Clinic finding carefully: high marks on clarity and relevance, lowest marks on decision impact. If OE “reinforces rather than changes decisions,” then its real value is confidence amplification and workflow acceleration.

A platform used by 40% of US doctors daily, influencing care for 100 million patients a year, has a published evidence base consisting of five retrospective cases and a pilot showing 31% accuracy on complex questions. We are scaling adoption faster than we are measuring whether it works. In medicine - a profession that normally requires evidence before adoption - we are running the experiment live, at population scale.

The platform is free for physicians. Revenue comes from pharmaceutical advertising - sponsored placements alongside AI-generated clinical guidance, with CPMs reportedly ranging from $70 to over $1,000 and a $150M+ revenue run rate at 90% gross margins. Every time a physician asks for evidence-based recommendations, the response environment is shaped by pharmaceutical sponsorship.

UpToDate, the clinical reference platform that OpenEvidence is displacing, charges roughly $500 per year per seat - a high-friction, subscription-funded model whose independence from pharmaceutical influence has been part of its value proposition for decades. OpenEvidence is free. The trade physicians are making - consciously or otherwise - is exchanging an unbiased, paid reference tool for a faster, ad-funded one. It is the Google playbook applied to clinical decision-making: give the product away, monetise through sponsored access to the decision environment. The line between content and sponsorship becomes harder to trace when the content itself is dynamically generated by an LLM.

AI interfaces are becoming the decision layer for knowledge work. The tool that synthesises information, frames options, and presents recommendations is increasingly the tool that shapes the decision. Whoever funds that interface - whether it’s a pharma company buying CPMs, a vendor subsidising a freemium copilot, or a platform monetising through adjacent services - has structural influence over the decisions made through it. The AMA data suggests physicians already sense this intuitively: 71% express privacy concern about non-institutional AI tools versus 42% for tools provided by their health system. They distinguish between AI they control and AI controlled by someone else’s business model. Enterprise leaders buying AI tools for their teams should be asking the same question.

The most troubling findings in the AMA survey aren’t about adoption - they’re about what comes after. 88% of physicians express concern about AI-related skill loss. 70% are specifically concerned about trainees losing clinical skills. Among early-career physicians - the ones still building the clinical instinct that AI is meant to supplement - 35% rate personal skill loss as a high concern.

Published evidence now supports their worry. A study in The Lancet Gastroenterology & Hepatology (August 2025) tracked adenoma detection rates in colonoscopy before and after AI-assisted practice. Detection rates dropped from 28.4% to 22.4% when physicians subsequently worked without AI assistance. The study’s authors described it as “the first study to suggest a negative impact of regular AI use on healthcare professionals’ ability to complete a patient-relevant task.” Measured, patient-relevant deskilling in a clinical setting.

Springer Nature’s systematic review of AI-induced deskilling in medicine (2025) maps the vulnerability surface more broadly: physical examination, differential diagnosis, clinical judgement, and physician-patient communication. The pattern it describes - physicians shifting from independent diagnosticians to validators of AI-generated recommendations, progressively disengaging from complex cognitive tasks - aligns precisely with what 88% of AMA respondents are worried about.

The fair counter-argument: AI may be redistributing skills rather than purely degrading them. Less memorisation, more supervisory judgement. Less time retrieving literature, more time evaluating and synthesising it. A physician who offloads routine research to AI and spends that time on complex differential diagnosis could end up with a different skill profile, and potentially a better one. The AMA data supports this reading - 76% of physicians see AI as advantageous for patient care, and 74% say it improves their diagnostic ability.

But skill redistribution requires deliberate redesign of how physicians are trained and assessed, and there is very little evidence that this is happening. Medical training still rewards the skills AI is replacing - recall, pattern recognition from volume, independent synthesis - and has barely begun to build curricula around the skills AI demands: critical evaluation of AI outputs, recognition of model limitations, judgement about when to override. Without that redesign, the trajectory defaults to erosion.

Doctors have always relied on tools - guidelines, imaging, lab results. AI is different because it collapses the effort required to form a view. Previous tools supported reasoning. A guideline gives you a framework; you still have to apply it. An AI-generated summary compresses much of the reasoning path into an answer-shaped output. The shift is from effortful cognition to frictionless delegation. As I referenced in a recent piece on “cognitive surrender”, the Wharton research is striking here: people followed incorrect AI answers 80% of the time, with confidence increasing despite wrong results. Effortful cognition is how expertise is built and maintained, and AI is compressing exactly the part of the process that builds it.

The feedback loop is where the danger is: AI handles the structured tasks well, so physicians delegate more. With less practice, their own skills atrophy. As skills atrophy, dependence increases - and as dependence increases, the ability to catch AI errors decreases. The colonoscopy study is the first concrete evidence of this loop operating in clinical practice. My guess is that it won’t be the last.

Medicine is the most data-rich example, but the structural logic applies across regulated verticals. Harvey’s adoption in legal AI follows the same pattern - document review, contract analysis, due diligence, compliance work. The highest-value tasks in law are the most structured, and that’s where AI gains traction first. Allen & Overy’s early trial saw 3,500 lawyers run 40,000 queries across 43 offices and 250 practice areas. The VLAIR benchmark - the first independent study of legal AI tools against a lawyer control group - found AI completed document review up to 80 times faster than lawyers, with Harvey scoring highest on accuracy across most tasks. The efficiency gains are real, but the evidence that legal outcomes improve is - just as in medicine - largely unmeasured.

Financial services, same story again: regulatory compliance, fraud detection, risk modelling, algorithmic trading - all well-defined and backed by decades of formalised data. The Contrary Research vertical AI playbook identifies the pattern explicitly - vertical AI companies are winning where domain-specific expertise creates a defensible moat.

Medicine is not a special case and I think the lessons here extrapolate directly to the enterprise.

The AMA data is a near-perfect map of where AI will land in any organisation: documentation and retrieval first, decision support second, genuine augmentation of expert judgement a distant third.

The enterprise data backs this up. McKinsey’s 2025 State of AI survey tested 25 attributes across nearly 2,000 organisations and found that only 39% see any EBIT impact from AI, despite 88% having deployed it in at least one function. The single strongest predictor of value? Workflow redesign - the high-performing organisations (roughly 6% of the sample) were nearly three times more likely to have fundamentally restructured their processes around AI.

The pattern is the same one medicine reveals, just less visible. Organisations that succeed with AI are the ones whose workflows were already well-defined and whose outcomes were already measurable. The ones that struggle are deploying AI against undocumented processes and knowledge trapped in people’s heads. The constraint is their own operating infrastructure.

What does fixing that constraint actually look like? It means turning tribal knowledge into retrievable, maintained, version-controlled assets - decision trees, process documentation, tagged and structured data. It sounds straightforward, but companies systematically underestimate three things. First, the political dimension: knowledge is power, and people hoard it. Asking a senior analyst to formalise their research process is asking them to hand over their competitive advantage. Second, the maintenance cost: knowledge decays. Documented processes drift from reality within months if no one owns the upkeep. Medicine solves this with peer review and continuous guideline updates; most enterprises have no equivalent discipline. Third, the gap between what’s documented and what people actually do. The process maps say one thing, reality says another, and AI trained on the documentation produces outputs that don’t match how work gets done - which is exactly where trust breaks down and adoption stalls.

Medicine had a century’s head start and still landed at 81% adoption concentrated in documentation tasks. Enterprises without that foundation should expect a much steeper climb.

The deskilling risk travels too. If 88% of physicians - a profession with decades of formal training and continuous assessment - are worried about skill atrophy, enterprise knowledge workers with far less formalised expertise should be more worried. Published evidence already shows measurable deskilling after just three months of AI-assisted practice, and most enterprise roles lack the feedback loops that would even surface equivalent degradation.

The regulation paradox resolves into a market selection principle. Industrialised, written-down expertise is where AI works - and where it justifies real spend. The enterprises that will capture genuine value are the ones that treat their information infrastructure with the same discipline that medicine brought to clinical evidence: continuously maintained, rigorously governed, and measured against outcomes.

About

I analyse AI progress beyond the headlines, focusing on enterprise execution, incentives, and real-world economic impact.

No posts

Read the original on aminmrini.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.