In 1998, a group of GI pathologists ran an uncomfortable experiment. They took the same histological slides (esophageal, gastric, colorectal biopsies) and sent them to 31 expert pathologists in 12 countries. The question was simple: does everyone see the same thing?
They didn’t. On esophageal lesions, agreement between Western and Japanese pathologists dropped to 14%. Essentially chance. The same slide, under the same microscope, read as invasive carcinoma in Tokyo and high-grade dysplasia in Boston.
No one in that study failed to see what was on the tissue. The problem wasn’t anyone’s eyes. Each pathologist was carrying a different rule for deciding when an abnormal cell crosses the line into cancer. The Western school required demonstrated histological invasion of the lamina propria or submucosa. The Japanese school classified primarily by cytological atypia and nuclear hyperchromasia, without requiring proof of invasion. Two rigorously applied systems, producing incompatible labels on the same human tissue.
The consequence wasn’t academic. With disagreement at that scale, the clinical risk is obvious even without a single documented case: a patient diagnosed with carcinoma in Japan, read under Western criteria, would be classified differently, with the surgery already indicated at home now in question. Or the reverse: a malignancy call that, under the other system, would never have justified a total gastrectomy with lymphadenectomy. You don’t need a specific chart to see that two different labels on the same tissue, crossing a border, cost something.
This wasn’t solved by training anyone to “see better.” It was solved by negotiating an agreement: the 1998–2000 Vienna Classification, which restructured the categories into five uniform clinical-management tiers. Interobserver agreement rose from 14–37% to over 70%. No pathologist’s eye changed. The rule everyone agreed to follow did, and that distinction matters: Vienna didn’t erase the underlying disagreement between the two diagnostic schools. It translated both criteria into comparable management categories. The scientific disagreement is still documented in the literature. What changed is that it stopped automatically producing incompatible clinical decisions.
I open with this story because, at a smaller scale, it’s exactly the problem the medical AI industry keeps refusing to look at.
Everyone is talking about foundation models. The dominant narrative, repeated in Nature Reviews Drug Discovery, Nature Medicine, and nearly every strategic conversation in the sector— is that these models are building digital pathology’s infrastructure layer, the way an operating system once structured personal computing. Since 2024, half a dozen of these models have shipped from academic and corporate labs, trained on banks of millions of whole-slide images:
Not all of these models pitch themselves as replacing the pathologist, many are, by their own technical description, general-purpose feature extractors, meant to power downstream applications someone else will build. But the ambition they share, explicitly or not, is that the core problem is perception: if a system learns to recognize morphological patterns more precisely, at greater scale, over more tissue, most of the clinical value follows from that. It’s a bet that performs well on academic benchmarks, and pulls in capital no standardization consortium ever has.
Worth naming a distinction here that carries the rest of this piece: computational complexity (processing gigapixel images, scaling transformer architectures, inferring genomic correlations from a routine stain) is not the same as institutional complexity getting two medical schools to agree on where a diagnostic line sits, and making that agreement auditable, defensible, and acceptable to a regulator. Foundation models are built to absorb the first. The question this piece is asking is what happens when the real problem is the second.
The largest independent benchmark run on these models to date (32 models, 41 tasks, over 17,500 slides, from Olivier Gevaert’s team at Stanford, in Nature Communications) offers evidence consistent with that reading: neither model size nor pretraining dataset size consistently predicts downstream performance. Differences among the top performers are, in most cases, statistically insignificant. And once models leave their training domain (from TCGA into real clinical data, a different hospital, a different scanner) the rankings shift. No single model dominates every setting. The same study found that ensembling several models does improve aggregate robustness, suggesting these models capture complementary strengths, even if none resolves the underlying problem alone.
In other words: scaling any one model isn’t buying the edge it promises. And that same study, by design, doesn’t evaluate (because no foundation model can, on its own) whether two hospitals in two countries will call the same thing carcinoma.
But there’s another bet. And someone placed it eight years before “foundation model” meant anything in medicine.
In 2016, at the ISBI challenge on breast cancer metastasis detection in lymph nodes, the algorithm from Andy Beck and Aditya Khosla’s team (the founders of PathAI) scored an AUC of 0.925 on whole-slide classification. An expert pathologist, working alone with no time limit, scored 0.966. The algorithm, on its own, still lost to the human.
The finding that actually mattered was different: when the pathologist’s read was combined with the algorithm’s assistance, the combined AUC rose to 0.995, roughly an 85% reduction in the specialist’s error rate. This wasn’t a race between the human eye and the machine. It was a different question: what does the pathologist need in order to be wrong less often?
“A model can in no way replace the work of a pathologist. It should be used as a tool to help the pathologist identify problem areas within a tissue sample.”
Andy Beck, co-founder and CEO of PathAI
That early bet held. PathAI built AIM-NASH, the name under which the FDA and EMA qualified the tool as a Drug Development Tool, built specifically to measure and reduce interobserver variability in grading metabolic dysfunction-associated steatohepatitis within MASH clinical trials. The company was founded, per its own corporate materials, “to apply technology to pathology to reduce error rates and increase accuracy and reproducibility.”
Reproducibility. Not replacement. That was the bet, funded by a $60M Series B led by General Atlantic in 2019, at a moment when the rest of the sector’s dominant ambition was already something else. I don’t offer this as proof the whole industry should follow one strategy (it’s a case, not a law) but as evidence that the instrumentation-and-reproducibility bet is viable, gets funded, and has already cleared a regulator.
The ambition of “the model just diagnoses” isn’t new, and it isn’t unique to pathology. It’s the same narrative pattern we’ve already lived through, with exact dates and names everyone recognizes.
On October 19, 2016, at an official press call, Elon Musk promised that by the end of 2017 a Tesla would drive fully autonomously from Los Angeles to New York, “without the need for a single touch” on the wheel. On April 22, 2019, at Tesla’s Autonomy Investor Day, he raised the stakes: “I feel very confident predicting autonomous robotaxis for Tesla next year… we’ll have more than a million on the road.”
Seven years since that first promise. We still have taxi drivers.
In November 2016, at a Creative Destruction Lab talk in Toronto, Geoffrey Hinton (today a Nobel laureate, then already deep learning’s most influential figure) said of radiologists: “People should stop training radiologists now. It’s just completely obvious that within five years deep learning is going to do better than radiologists.” Nearly ten years later, radiology is one of the highest-demand, best-paid specialties in the U.S., and of the more than one thousand FDA-cleared AI devices for radiology, not one is authorized to diagnose autonomously without a radiologist’s sign-off.
I’m not citing Musk or Hinton to cast them as villains. They’re just the most visible names in a much broader pattern: an entire industry that, faced with a hard problem, would rather promise to eliminate it than ask what’s causing it. “We’re going to replace the pathologist” is a headline. “We’re going to harmonize diagnostic criteria across twelve countries” is not. But only the second one actually fixes what costs the patient time, money, and safety.
This is exactly what my Complexity-Value Matrix predicts: sophistication is a property of the problem, not of the tool used to solve it. I planted this distinction earlier, but it’s worth pausing on, because it’s this piece’s central conceptual contribution.
Computational complexity is processing gigapixels per image, scaling vision transformer architectures, inferring genomic correlations from a routine stain. It’s the complexity foundation models are built to absorb and where they genuinely excel.
Institutional complexity is something else entirely: getting two medical schools to agree on where displasia ends and carcinoma begins, folding that consensus into multidisciplinary tumor boards, making it auditable to a regulator, defensible in a lawsuit. It’s slow, it doesn’t scale with more training data, and there’s no benchmark you can win in a paper.
Applying massive computational complexity to a problem of institutional complexity is a category error. A feature extractor with billions of parameters doesn’t, by itself, resolve decades of disagreement between Japan and the West over the histological definition of a tumor margin. That problem isn’t won with scale. It’s won with governance.
This isn’t just my intuition anymore. The Stanford benchmark offers evidence consistent with this reading: comparing small, base, and large ViT architectures, and pretraining datasets from under 100,000 slides to over a million, scaling benefits are most consistent within TCGA, the same domain most of these models were trained on. On external and out-of-domain data, the benefits are weak, variable, and dependent on domain and task. Scale appears to buy more performance inside the distribution it trained on than outside it. It’s a correlation, not an airtight causal proof (the study itself flags this as a limitation) but it points the same direction as the Complexity-Value Matrix: computational complexity mostly solves computational problems. It doesn’t, by itself, get two hospitals to agree.
I don’t think we got foundation models wrong. I think we got it wrong expecting a better visual representation, on its own, to solve clinical reproducibility. These are related problems (better perception can help standardization, if it’s designed for that) but they’re not the same problem, and treating them as if they were is the choice worth revisiting.
Why did the industry bet so heavily on replacement instead of standardization? I can’t prove it was a purely commercial decision, that would claim more than the evidence supports. What I can say, as a reasoned reading rather than an established fact, is that “the model just diagnoses” is an easier story to fund than “the model makes what the pathologist already decides auditable,” and that asymmetry likely shaped where the capital went first.
Nor do I think perception and standardization are mutually exclusive bets. A well-governed foundation model, (inside a validated system, with regulatory oversight, with an explicit measurement purpose, as with AIM-NASH) does produce real clinical value. The question isn’t foundation model yes or no. It’s whether the model was designed to make a human judgment auditable, or to replace it with nothing auditing anything.
CAP, RCPath, and the ICCR have spent years working on synoptic reporting protocols and harmonizing histopathology terminology across the U.S., Europe, Australasia, and Asia. They receive a fraction of the capital any new foundation model attracts. And yet, by every regulatory and clinical signal available, they’re the ones actually solving the problem that costs the patient real time.
Foundation models don’t have to sit outside that effort, they can become one of its instruments, rather than its alternative. There’s a concrete agenda here, and pieces of it are already scattered across what these models already do well:
Quantify interobserver disagreement instead of just predicting a label, turn variability itself into a visible metric, not noise to discard. Automatically translate between different classification systems (Japan/West, AJCC/UICC, different Gleason versions) instead of demanding the world converge on one. Audit diagnostic decisions already made, flagging where human judgment departs from a reference consensus, without overriding it. Version diagnostic criteria the way code gets versioned, so you can trace when and why a rule changed. And cross-validate across centers, using the model itself as a thermometer for how much a diagnosis varies by who reads it and where. None of these five is “the model just diagnoses.” All of them use the same technology.
The question isn’t whether foundation models will keep improving. They will. The question that actually matters (the one almost no one is asking before the next funding round) isn’t which model has the best AUROC this week. It’s which clinical decision, currently resting on one person’s judgment in one hospital, becomes more reproducible and more auditable because of what’s being built.
Schlemper RJ, Itabashi M, Kato Y, et al. Differences in diagnostic criteria for gastric carcinoma between Japanese and western pathologists. Lancet. 1997;349(9067):1725-1729.
Schlemper RJ, Riddell RH, Kato Y, et al. The Vienna classification of gastrointestinal epithelial neoplasia. Gut. 2000;47(2):251-255.
Chen RJ, Ding T, Lu MY, et al. Foundation models in computational pathology: methods, applications and clinical implications. BMJ Oncology. 2024.
Wang D, Khosla A, Gargeya R, Irshad H, Beck AH. Deep learning for identifying metastatic breast cancer. arXiv:1606.05718 [Internet]. 2016 Jun 18.
PathAI. About PathAI [Internet]. Boston: PathAI; [cited 2026 Jul]. Available from: https://www.pathai.com/about-us
General Atlantic. PathAI Secures $60M in Series B Funding Led by General Atlantic and Existing Investor General Catalyst [press release]. 2019 Apr.
PathAI. PathAI Announces EMA Qualification for AIM-MASH AI Assist (qualified by FDA/EMA under the name AIM-NASH as a Drug Development Tool for MASH clinical trials) [Internet]. PathAI; 2025.
Massachusetts Institute of Technology. 6.S897 Machine Learning for Healthcare, Lecture 12 Notes [Internet]. MIT OpenCourseWare; Spring 2019.
Lawler R. Musk targeting coast-to-coast test drive of fully self-driving Tesla by late 2017. TechCrunch. 2016 Oct 19.
Tesla. Autonomy Investor Day [corporate presentation]. Palo Alto: Tesla; 2019 Apr 22.
Creative Destruction Lab. Machine Learning and the Market for Intelligence [conference]. Toronto: University of Toronto; 2016 Nov.
Bareja R, Carrillo-Perez F, Zheng Y, Pizurica M, Nandi TN, Tian L, et al. A benchmark study of vision and pathology foundation models for computational pathology. Nat Commun. 2026.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.