RSS Amplifier

Beyond the Slide · Jul 3, 2026

The Sophistication Trap : Why More Technology Doesn’t Necessarily Create More Value

0
Sign in to vote or save

Dr. Luis Cano · Beyond the Slide

The hospital had done the transition right. On paper.

They bought the high-resolution scanner. Set up cloud storage for gigapixel images. Rolled out digital viewers on every pathologist’s desktop. In the board presentation, the project looked exactly as it should: digital pathology, end to end, ready to scale.

Six months later, the scanner had become the most expensive paperweight the hospital ever bought. It sat idle most of the week. Of every thousand images that were supposed to reach the cloud, barely four hundred did. Technicians weren’t complaining about the technology itself, they were complaining that it added a step instead of removing one. Day to day, in the hallways, people kept doing what they’d always done: looking through the microscope.

One day, a resident had a hard case and wanted a second opinion. He found the attending in the hallway: “The image is ready in PACS, everyone can log in whenever you’re free.” “Do you have the slides?” the attending asked. “Yes, right here, I just finished scanning them,” the resident said. “Let’s go to the five-head scope. Call the others.” Thirty-year-old technology was still the first instinct and it worked better, faster, with less friction than any digital workflow the hospital had built that year.

That scene has repeated, with variations, in more than one hospital I’ve seen up close. And every time, the question left hanging in the air isn’t “why aren’t they using what we bought?” It’s an earlier question, one nobody asked before buying: was the problem we had actually the one this technology solves?

The question almost everyone asks when evaluating health tech is the wrong one. It’s not which is best. It’s what works, what fits, what actually solves the problem.

No hospital buys thinking about the sharpest image resolution or the most advanced algorithm on the market. Or it shouldn’t. It buys what fits the workflow that already exists. What an exhausted technician can operate at six in the evening without extra friction. What solves the real problem — not the problem the vendor described in the demo, under the best light, in the best-case scenario, in front of the committee that approves the budget.

This is a story about that gap. About how the health industry and the pharma industry around it, has learned to confuse two questions that sound alike but aren’t: how sophisticated is this technology, and what problem does it actually solve in production? The evidence, not just the anecdote, suggests that gap is wider, and far more expensive, than most organizations are willing to admit.

For months I thought the problem was artificial intelligence. It wasn’t. It’s a much older pattern that AI just makes more visible, and more expensive. Over time I started plotting the projects I saw (the scanner gathering dust, the five-head scope beating fiber optics) on two very simple axes. I didn’t know yet that it would turn into a matrix with a name of its own. We’ll get to it, with the data behind it.

For a while I treated the five-head-scope scene as a nice anecdote, the kind of story you tell in a conference hallway and forget by the next day. But I went looking for whether it was an outlier, and after going through dozens of reports and studies on AI adoption in health care, I landed on an uncomfortable conclusion: it wasn’t. It’s the local version of a pattern that industry-wide evidence confirms with a consistency that surprises even people who expected it.

RAND CORPORATION, 2024

80% of AI projects fail, double the failure rate of IT projects that don’t involve AI.

RAND’s study isn’t a market estimate or an attention-seeking headline. It’s the result of structured interviews with sixty-five data scientists and engineers, each with at least five years building AI models in industry or academia. Its most uncomfortable finding isn’t technical: the most common cause of failure isn’t a broken algorithm. It’s a problem that was never properly defined in the first place.

The gap between investment and real value shows up in every serious report published on the subject in the last two years. Boston Consulting Group’s October 2024 report, “Where’s the Value in AI?,” surveyed a thousand C-suite executives across fifty-nine countries. Its headline finding: despite $252.3 billion in aggregate spend in 2024, 74% of companies couldn’t demonstrate tangible value from their AI investments.

In health care, the pattern sharpens. Qventus’s 2026 report, built on interviews with more than sixty technology, clinical informatics, and AI leaders at U.S. health systems, documents a tension any clinical leader will recognize instantly: 65% rated board pressure to operationalize AI a 7 or higher on a 10-point scale. 94% said delaying adoption would put them at a direct competitive disadvantage. And yet only 4% of those same health systems have actually scaled an AI solution with measurable clinical or operational results.

Board pressure: 65%. Feeling of competitive disadvantage from waiting: 94%. Systems that actually scaled with measurable results: 4%.

The gap isn’t a problem of insufficient ambition. It’s almost the opposite: there’s plenty of ambition, and what’s missing is rigorous problem selection before budget and reputation get committed to the solution.

The case closest to the five-head-scope scene, though, comes from a study published in July 2025 in the Journal of the American Medical Informatics Association, which surveyed forty-three leading U.S. health systems across thirty-seven distinct AI use cases. The finding is almost a mirror of Act I: imaging and radiology are, by far, the clinical area with the highest AI deployment, 90% of organizations report at least partial deployment, but only 19% of those same organizations report a high degree of real success with those tools. In clinical risk stratification, perceived success drops to 38%. The only category with consistent success, 53%, is automated clinical documentation: precisely the application with the least algorithmic sophistication and the least clinical decision authority of the set.

The pattern, taken as a whole, is clear: the closer an AI tool gets to the center of clinical decision-making, and the more sophisticated its architecture, the wider the gap between deployment and perceived success. This isn’t a case of technology that doesn’t exist or doesn’t work in the lab. It’s a case of an organization that never checked, before buying, whether that was actually where the bottleneck lived.

Share

Health care has been running a natural experiment for more than a decade without calling it that: two bets, nearly opposite in scale, aimed at the same kind of problem (reducing preventable harm to patients) with results that should be required reading in any innovation committee.

IBM Watson for Oncology started from a seductive premise: if an AI system had beaten the best human players on Jeopardy!, why couldn’t it process the world’s oncology literature and generate treatment recommendations more rigorously than any individual oncologist? IBM acquired Truven Health Analytics for $2.6 billion and Merge Healthcare for $1 billion to consolidate the training data it needed. Aggregate internal investment topped $4 billion. The collaborative project with MD Anderson Cancer Center alone cost more than $62 million before it was cancelled.

The system never operated with the reliability it promised. Independent evaluations documented clinically unsafe recommendations, including suggesting bevacizumab for lung cancer patients with severe hemoptysis, a contraindication any first-year oncologist would catch immediately. In concordance studies against tumor boards in South Korea, the system reached just 77% agreement in breast cancer and 54.5% in gastric cancer, far below what’s needed to sustain clinical trust.

The root cause wasn’t a lack of computational power. It was a model trained almost exclusively on Memorial Sloan Kettering’s prescribing patterns, marketed as a universal validation system, with no transparency into how it reached its recommendations, the classic black-box problem, and no real adaptation to the local clinical and epidemiological contexts where it was deployed.

In parallel, a team led by surgeon and researcher Atul Gawande designed an intervention that fit on a single sheet of paper: a nineteen-item surgical safety checklist, organized around three moments (before anesthesia induction, before incision, before the patient left the operating room) read aloud by the surgical team itself.

MULTINATIONAL STUDY, EIGHT HOSPITALS

−36% in major surgical complications after introducing the checklist (from 11.0% to 7.0%).

In-hospital deaths fell more than 40%, from 1.5% to 0.8%. Infrastructure or licensing cost was, in practice, zero. The barrier wasn’t technical, it was cultural: surgeons who initially saw the checklist as a bureaucratic step slowing them down.

The same pattern shows up at a smaller scale inside pathology itself. An improvement program at a histology lab, published in the American Journal of Clinical Pathology, combined a barcode system with a Lean redesign of the microtomy workflow (no new equipment purchased) and achieved a 78.64% reduction in labeling errors, along with a 35.28% drop in reprocessing time.

$4 billion and a black box recommending unsafe treatments, against a sheet of paper read aloud that saved measurable lives. The difference wasn’t the budget. It was whether anyone asked, before starting, where the problem actually lived.

Neither case exists to argue that technological sophistication is never worth it. Digital pathology has its own legitimate high-value quadrant, and we’ll name it explicitly further down. The point is more precise, and more uncomfortable: sophistication and value are independent variables. They can move together. But just as often they split apart, and when they do, almost no organization catches it before the contract is signed.

It would be comfortable and wrong, to close this argument by saying health care simply buys bad technology. That’s not what the evidence shows. Paige Prostate Detect’s algorithms work. Mitosis-detection models work. Three-dimensional radiotherapy dose planning works, and it’s exactly the kind of problem, high mathematical dimensionality, impossible to solve by hand, where sophistication is a necessity, not vanity.

The villain of this story isn’t the algorithm. It’s the expectation built around it: the idea, almost never said out loud but present in every innovation committee, that not having a visible AI strategy is itself a strategic risk, regardless of whether the problem the organization actually faces requires that level of sophistication.

And it’s worth saying plainly: this pattern didn’t start with AI, and it won’t end with it either. Health care and pharma have lived it before, under other names. The ERP wave of the nineties and 2000s promised to integrate the entire hospital operation and, more often than not, ended in underused systems kept alive out of fear of reversing the investment. The Big Data craze promised that more data, on its own, would produce better clinical decisions. Blockchain was going to solve health record interoperability and, in the vast majority of hospital pilots, never made it past the innovation committee. AI is simply the newest, most expensive version of a question health care has spent thirty years avoiding: is the problem we have organizational, or technological? Because only one of those two gets solved by buying something.

We jumped into the pool because everyone around us was jumping. Did anyone bother to check if the pool had water in it?

That expectation has a name in the organizational behavior literature, and it’s worth naming precisely, because a problem without a name is much harder to manage in a board meeting.

The human brain tends to assume that intricate systems (the ones that demand more cognitive effort to understand) are inherently superior to direct approaches. Research on clinical behavior around language models shows that both residents and specialists are vulnerable to this bias: their diagnostic accuracy drops when they discard a simple, correct answer in favor of an elaborate, sophisticated one generated by an intelligent system, just because it sounds more rigorous.

Adopting cutting-edge AI carries enormous external symbolic value. It lets a board project technological leadership to investors, regulators, and high-profile patients. When the primary goal becomes announcing the acquisition of the system, not auditing, months later, the real value it generated for the patient or the technician operating it, innovation stops being a bet and becomes a performance.

Once an institution commits a large initial investment (financial and reputational) to an AI platform, whoever made that call tends to justify further investment down the same path, to avoid the psychological cost and public scrutiny of admitting a strategic loss. The project doesn’t stop when it stops making sense. It stops when there’s no more budget (or career) left to protect.

A CIO’s or CMIO’s professional standing is often tied to the size and technical sophistication of the systems portfolio they run. Approving a Lean Six Sigma workflow redesign doesn’t carry the same weight on a corporate résumé as leading the rollout of “autonomous AI agents”, even when the former has a better shot at moving the number that actually matters.

The system was designed, without anyone planning it that way, to reward announcing the AI project, not auditing, six months later, whether that project actually solved anything.

Naming the pattern isn’t new, the literature on tech solutionism, innovation theater, and escalation of commitment has existed for years, in health care and beyond. What’s still rare is an organization pausing, before approving the budget, to ask which of these mechanisms is driving its own decision. Naming the pattern hasn’t prevented it. What’s missing isn’t more literature about the problem. It’s a filter used before signing, not a diagnosis written after.

Sophistication is not an attribute of the solution. It is a property of the problem.

Put differently: every problem has an optimal level of sophistication, not the maximum available, not the easiest to justify in a boardroom, but whatever the problem actually demands. Below that level, the solution falls short. Above it, the solution gets fragile, expensive, and hard to maintain, without buying any additional value. That’s the principle behind the tool that follows.

Share Beyond the Slide

If the villain is the expectation, the defense can’t be an opinion about whether AI is “worth it” in the abstract. It has to be a sequence of concrete questions, applicable before the budget is committed, not a damage audit after the scanner is already gathering dust. This is the Complexity-Value Matrix: the tool I kept sketching, almost without noticing, every time a project started looking more like the scanner from Act I than the five-head scope.

Before the sequence of questions, Quadrant I deserves its due, it exists, it’s real, and it isn’t a marginal exception. A pathologist with a light microscope can’t calculate, by eye, the spatial distribution of sixteen hundred morphological biomarkers in a single H&E image. A medical physicist can’t manually optimize beam-incidence angles for radiotherapy to minimize damage to a glioblastoma without compromising surrounding healthy tissue, that’s a high-dimensional geometric problem no heuristic shortcut can handle. In both cases, algorithmic sophistication isn’t a signal for the board. It’s the only way to make visible something the human eye structurally cannot see. That’s the standard any other project should be measured against before approval: does the problem genuinely exceed human cognitive capacity, or does it just feel more comfortable to treat it as if it did?

This is the sequence I use with the teams I work with, four questions that capture what separates Quadrant I from Quadrant II on the map below.

Is the bottleneck a pattern the human eye can’t process (mathematical density, data volume, dimensionality) or is it a problem of coordination, communication, or workflow between people? If the root cause is organizational, no amount of algorithmic sophistication will fix it. The surgical checklist and the five-head scope share the exact same diagnosis: the problem was never computational.

Before building or buying a complex model, it’s worth benchmarking it against the simplest clinical heuristic available. In predicting 30-day hospital readmissions, published comparisons show the LACE index (four variables, calculated by hand) performs on par, statistically, with deep neural networks, and in some cases beats models like XGBoost, which tend to overfit the noise in local clinical data. If the simple solution already captures 80% of the value, that’s the starting point — not the finish line.

Not the model’s technical performance metric, sensitivity, specificity, AUC on a curated dataset. The operational or clinical metric you’ll measure in production, against a documented baseline, six months after deployment. If no one in the room can name that metric before approving the project, what’s being purchased isn’t a solution. It’s the feeling of having bought one.

There’s another reason to take Question 3 seriously: the price on the quote is never the real cost of sophistication. Every additional degree of complexity comes with a bill that shows up nowhere in the contract, staff training, integration with the existing clinical workflow, ongoing model maintenance, governance of the data feeding it, periodic revalidation against performance drift, and the almost never budgeted cost of the cultural change required before anyone in the ER trusts what the machine recommends. None of that shows up in the demo slide. All of it shows up, late, in next year’s budget. A project that looks cheap on the quote and turns expensive in operation didn’t fail on price, it failed because nobody added up the cost of complexity before signing.

This is the question almost no organization asks, and it’s the most important of the four. It’s not about defining success, Question 3 already does that. It’s about defining, in advance, the concrete evidence that would trigger walking away. Without that threshold written down before starting, the escalation of commitment from Act IV has a clear runway: any ambiguous result gets read as “needs more time,” never as “this isn’t working.”

These four questions can be plotted as two independent axes, technological sophistication and real problem-solving value, producing four quadrants. Placing a project there, before approving it, forces into the open what otherwise stays implicit in the excitement of a demo.

The Complexity-Value Matrix

The most useful thing about this matrix isn’t classifying finished projects, that’s easy, and it changes nothing. It’s using it in the room where the budget gets approved, before the money is committed, and forcing someone to say out loud which quadrant they think the project will land in, and why. The evidence gathered here suggests a good share of today’s digital health investment ecosystem sits, without anyone deciding it explicitly, in Quadrant II: high sophistication, low value.

Quadrant II projects rarely fail because the algorithm is bad. They fail because they try to solve an organizational problem with a computational tool.

There’s a conceptual framework, originally developed by Clayton Christensen for consumer markets, that translates surprisingly well to an operating room or a pathology lab: Jobs to Be Done. Its premise is that people and organizations don’t buy products for their technical attributes. They “hire” them to do a specific job, in a specific circumstance.

The real job of a pathology chief isn’t operating the most advanced algorithm on the market. It’s cutting diagnostic turnaround time and reducing interobserver variability on hard cases. The real job of a hospital CIO isn’t having an “AI strategy” to present to the board. It’s making sure resources (people, compute, budget) produce better, measurable, sustained clinical decisions per patient.

When the job gets defined independently of the solution, something the “buy the most advanced system” logic systematically hides comes into view: in plenty of cases, a change in physical workflow design (where the five-head scope sits, how the report sign-off rotation is organized, who sees the urgent case first) achieves the same result, at a fraction of the transaction cost, without the adoption curve that quietly kills so many well-intentioned AI projects.

The organization that wins this decade won’t be the one that announces the most ambitious AI project. It’ll be the one that knows, before approving the budget, which quadrant it’s going to land in.

This isn’t an argument against sophistication. It’s an argument for earning it. Quadrant I exists, it’s real, and the organizations that occupy it (in computational pathology, precision radiotherapy, biomarker discovery) didn’t get there by buying the most advanced thing available. They got there after ruling out, with evidence, that a simpler solution could solve the same problem. Sophistication, when it’s earned, stops being a signal for the board and becomes what it was always supposed to be: the right answer to a question someone actually bothered to ask first.

Next time a vendor walks into a committee room with a flawless demo of an autonomous agent system, or a foundation model promising to solve the hospital’s diagnostic bottleneck once and for all, there’s one question worth more than any sensitivity or specificity figure on the slide: if the problem is organizational, why are we evaluating a computational solution?

You don’t need to reject the technology to ask that question. You just need to refuse to jump into the pool before checking there’s water in it. Every organization has its own five-head scope waiting in some corner, its own thirty-year-old solution that still outperforms last year’s sophisticated project. The question was never whether it exists. It was whether anyone bothered to look for it before buying something new.

Sophistication is not an attribute of the solution. It is a property of the problem.

This is the short version of everything above, built to print and bring to the room.

Before approving your next AI project, it's worth spending an hour answering these four questions. It's usually a lot cheaper than finding out, six months later, that the project landed in the wrong quadrant. If you want to run the Complexity-Value Matrix against your next investment, or one you've already made that isn't delivering, I'd be glad to help, just send me an email to luiscanoayestas@gmail.com.

References:

  1. Ryseff J, De Bruhl BF, Newberry SJ. The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed: Avoiding the Anti-Patterns of AI. Santa Monica (CA): RAND Corporation; 2024. Report No.: RR-A2680-1.

  2. Boston Consulting Group. Where’s the Value in AI? Boston: BCG; 2024 Oct.

  3. Qventus. Beyond the Pilot: How CIOs Are Operationalizing AI Across Health Systems in 2026. San Francisco: Qventus; 2026 Apr.

  4. Poon EG, Lemak CH, Rojas JC, Guptill J, Classen D. Adoption of artificial intelligence in healthcare: survey of health system priorities, successes, and challenges. J Am Med Inform Assoc. 2025 Jul;32(7):1093-1100.

  5. Flach RN, van Dooijeweert C, van Diest PJ, et al. Prospective Clinical Implementation of Paige Prostate Detect Artificial Intelligence Assistance in the Detection of Prostate Cancer in Prostate Biopsies: CONFIDENT-P Trial. JCO Clin Cancer Inform. 2025.

  6. Strickland E. IBM Watson, heal thyself: how IBM overpromised and underdelivered on AI health care. IEEE Spectr. 2019;56(4):24-31.

  7. Haynes AB, Weiser TG, Berry WR, et al. A surgical safety checklist to reduce morbidity and mortality in a global population. N Engl J Med. 2009;360(5):491-499.

  8. Heher YK, Chen Y, Pyatibrat S, Yoon E, Goldsmith JD, Sands KE. Achieving high reliability in histology: an improvement series to reduce errors. Am J Clin Pathol. 2016;146(5):554-560.

  9. Christensen CM, Hall T, Dillon K, Duncan DS. Competing Against Luck: The Story of Innovation and Customer Choice. New York: HarperBusiness; 2016.

  10. Kim H, Ock M, Park JY, et al. Real-world concordance rate of Watson for Oncology with multidisciplinary tumor board decisions in South Korea. BMC Cancer. 2023.

Read the original on beyondtheslide.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.