THE SCENE
Every pathologist has seen a case like this. HER2-positive breast cancer, IHC 3+, the kind of staining that makes a slide worth showing to residents. The pathologist orders in situ hybridization anyway, just to document it properly. Some cases are too clean not to study closely.
In the mid-2010s, a phase II trial enrolled 164 patients with exactly that profile (HER2-positive by immunohistochemistry) to test, experimentally, trastuzumab emtansine combined with pertuzumab as neoadjuvant treatment, before surgeons decided how much tissue to remove. It was not yet standard of care. It was the question the trial existed to answer. Picture two of those patients, enrolled with the same pathology report in hand.
Neither knows yet that the outcome will diverge. Same waiting room, different dates, same question, will this work?, and the same reassuring answer: the tumor expresses the target, the drug was built for exactly this, there’s real reason for optimism.
Their oncologist has no reason to treat them differently. The pathology report says the same thing for both: “HER2 positive, IHC 3+.” Same box checked, same trial protocol. Both walk through the same door.
Months later, one has no residual tumor in the resected tissue, pathologic complete response, the best possible outcome in this setting. The other has substantial residual disease. Same drug. Same dose. Same molecular target, on paper.
What the pathology report didn’t say (because the protocol never asked) is that in one patient HER2 was distributed uniformly across the tumor, and in the other there were entire regions where amplification simply wasn’t present, subclones nobody looked for because nobody had to. The same binary label, “positive,” concealed two completely different tumor architectures.
THE SCENE HAS A REAL ANCHOR, METZGER ET AL., PHASE II NEOADJUVANT TRIAL
55% vs. 0%
Pathologic complete response rate in homogeneous vs. heterogeneous HER2 tumors (n=157 evaluable of 164 enrolled, 16 with intratumoral heterogeneity, p<0.0001, adjusted for hormone receptor status).
No hypothetical patients were needed to make this scenario real: in that trial, none of the 16 patients with intratumoral HER2 heterogeneity achieved a complete response. None. While 55% of those with a homogeneous tumor did. The pathology report, in both groups, said exactly the same thing: HER2 positive.
The word heterogeneity is worth pausing on, because in everyday clinical language it sounds like a footnote. In this trial it had a precise definition: a tumor region with HER2 amplification in fewer than 50% of cells, or an outright HER2-negative area, detected by FISH in biopsies from two different sites of the same tumor. A single biopsy wasn’t enough. The tumor had to be sampled in more than one place for the pattern to show up.
And there, perhaps, is the first hint of where this essay is headed: the problem isn’t only what we measure. It’s how many times, and in how many places, we bother to measure it.
THE QUESTION THAT SHOULDN’T MAKE SENSE
There’s something unsettling about a finding like this, and it isn’t only clinical. It’s epistemological. We’ve spent two decades building precision oncology on a simple premise: find the right molecular target, direct the drug at it, and the outcome should be predictable. HER2 positive. Trastuzumab. Benefit. That causal chain is, in large part, the founding promise of targeted oncology.
But the pathology report (the document that triggers that chain) was never built to capture whether the target was evenly distributed or patchy, whether the tumor had enough stroma to physically block drug delivery, whether the immune microenvironment was hospitable or hostile. The report answers a binary question: is the target there or not?, about a phenomenon that was never binary.
I’ve spent months noticing something similar in a field that, on the surface, has nothing to do with this: the race to build foundation models for digital pathology. The more I look at these two territories side by side, the more the shape of the problem looks the same, even though the content is entirely different.
This is the question that has followed me through most of what I write here: why do innovations that look inevitable take so long to change medicine. Almost every time I dig into that question, the answer isn’t the technology (which usually already exists, somewhere, in some paper) it’s the institutional architecture that decides which question gets solved first.
TWO OBJECTS THAT SHARE NOTHING
A foundation model trained on millions of histology images and a single-marker immunohistochemistry assay share no technology, no budget, no invention date. One is the leading edge of applied AI in medicine. The other is a staining method that has existed, essentially, for decades.
Foundation models chase perception without an equivalent decision infrastructure. Companion diagnostics chase reproducibility by reducing biology to whatever can be standardized. From opposite directions, both reveal the same gap: we can measure more than our institutions know how to use.
The foundation model industry (I’ve written about this before) has bet heavily on perception: models increasingly capable of “seeing” subtle tissue patterns, matching or exceeding the pathologist’s eye. What almost nobody is building with the same intensity is the decision infrastructure that would turn that perception into something clinically actionable, reproducible, auditable, transferable across labs and countries.
The companion diagnostic for an ADC runs the opposite path and arrives at the same place. It doesn’t chase perception (it’s deliberately simple, one marker, one binary cutoff), it chases reproducibility: the quality that lets it work the same way in any hospital, anywhere. The price of that reproducibility is reducing tumor biology to the one variable that’s easy enough to measure that way. Two opposite paths, perception without decision, reproducibility without full biology, and the same result: we measure more than our institutions know how to use.
That’s the blind spot these two objects, so different on the surface, share without knowing it.
It isn’t a coincidence that diagnosis and treatment selection took separate paths in recent precision oncology history. Diagnosis became AI’s territory, with research budgets in the hundreds of millions. Treatment selection for ADCs, meanwhile, stayed anchored to a last-century instrument (immunohistochemistry) because there was never an institutional moment where anyone asked whether both problems were, underneath, the same question asked twice.
WHY TISSUE MATTERS MORE THAN THE REPORT ADMITS
For an antibody-drug conjugate to work, the target has to be there, but that’s not enough. The drug has to physically reach it, crossing a tissue architecture the pathology report never describes in that level of detail.
Start with the intuitive part: tissue as physical obstacle. Dense extracellular matrix and proliferating cancer-associated fibroblasts act as a real barrier against diffusion of antibody-sized molecules. In some tumors, specific fibroblast subpopulations drive immune exclusion and physically constrict vasculature, restricting drug passage from blood vessel to tumor core. Chaotic tumor vascularization also creates perfusion gradients: central regions literally receive less drug than the periphery.
So far, intuition holds: more barrier, less drug. But biology, again, resists simplicity.
In third-generation conjugates like trastuzumab deruxtecan, cathepsin L, a protease overexpressed and secreted into the tumor microenvironment by stromal and inflammatory cells, can cleave the drug’s linker outside the cell, releasing payload directly into the interstitial space. This kills neighboring tumor cells that don’t even express the target (the so-called bystander effect) without requiring receptor-mediated internalization. It’s a finding demonstrated in specific models and still recent; there isn’t enough evidence to assume this mechanism operates with equal force in every patient or every cleavable-linker ADC, but it points at something real: the same stroma that acts as a barrier in one context acts as an activation mechanism in another.
Target heterogeneity, by contrast, doesn’t carry that same duality, it’s close to uniformly bad news, as far as the evidence shows.
There’s a third variable, less discussed than stroma or vasculature, that also decides the drug’s fate: the target’s subcellular location. HER2 retained at the cell membrane, where the antibody can bind easily, is not the same as HER2 displaced into the cytoplasm, partially out of reach. That distinction, membrane versus cytoplasm, is what the Normalized Membrane Ratio quantifies, a metric that resurfaces, with a starring role, in the final act of this essay.
None of these three variables (stroma, vasculature, subcellular localization) appears on a standard pathology report today. Not because they’re invisible to a trained eye: an experienced pathologist can often describe them qualitatively in the margin of a report. They’re absent because there’s still no quantitative vocabulary, reproducible across observers, that turns them into data a clinical trial can use to stratify patients. That, in essence, is the standardization problem running through this entire essay.
THE QUANTITATIVE PAYOFF, NOW IN ITS PLACE
55% vs. 0%
Metzger et al., spatial HER2 heterogeneity, not just its presence, determined whether pathologic complete response occurred. A factor invisible to the standard IHC assay.
WHAT THE COMPANION DIAGNOSTIC CHOSE NOT TO SEE
The regulatory design of ADC companion diagnostics isn’t an accident or a negligence, it’s a reasonable answer to a different problem than the one we now know exists. When these assays were designed, regulators needed to answer a simple, urgent question: does it make sense to expose this patient to an expensive, non-trivial-toxicity drug? A cheap, cross-lab-reproducible IHC assay, backed by decades of infrastructure already installed in hospitals worldwide, answered that question reasonably well.
What that design never incorporated is whether, within the “positive” group, subgroups exist with tissue architectures different enough to change their odds of benefit. Sacituzumab govitecan, a TROP2-directed ADC, doesn’t even require quantifying the target in its approved indication: the trial showed clinical benefit across such a wide expression range that a cutoff would have improperly excluded patients who did benefit.
No clinical guideline, not NCCN, not ESMO, requires or even suggests characterizing stromal, vascular, or immune infiltration variables to decide whether to prescribe an ADC as monotherapy today. Clinical practice still treats broad populations based almost exclusively on the binary presence or absence of the antigen. Not because the biological evidence doesn’t exist (as Act IV showed, it does, and it’s solid) but because there’s still no regulatory vehicle, no commercial incentive, built to carry it through to the clinical decision.
Regulatory review documents justify single-marker assays on concrete grounds: analytical simplicity, high baseline antigen prevalence in the target population, and the risk that too strict a cutoff would improperly exclude patients who’d still benefit clinically. It isn’t an arbitrary choice.
The practical result of that logic is a single-speed regulatory architecture: every new ADC reaches market with, at best, an expression assay for its own target, designed and validated in isolation. There’s still no framework to evaluate, across different ADCs, different targets, different payloads, whether the same set of microenvironment variables consistently predicts response. Every drug reinvents its own diagnostic question from zero. It’s the same pattern from Act III, again: an industry capable of investing in perception hasn’t yet found the institutional vehicle to invest with equal intensity in the decision infrastructure that would make it useful.
THE COUNTERARGUMENT THE THESIS ITSELF DEMANDS
It would be tidy to close here with a clean conclusion: the industry needs to reclassify patients by microenvironment, and digital pathology is the tool that will make it possible. That conclusion would be exactly the kind of overconfidence this essay has spent three acts criticizing elsewhere.
In the ASCENT trial, sacituzumab govitecan showed clinical benefit, in progression-free and overall survival, consistently across every TROP2 expression subgroup analyzed, including low expression. In TROPiCS-02, the finding was even more decisive: “there was no clear level of TROP2 expression at which a better treatment effect was observed,” as the trial’s principal investigator reported. Drug efficacy did not, in practice, depend on how much target was there.
THE EVIDENCE THAT LIMITS THE CENTRAL HYPOTHESIS
Benefit across every subgroup
ASCENT and TROPiCS-02: consistent clinical benefit from sacituzumab govitecan regardless of TROP2 expression level, including very low expression.
How do this finding and Metzger’s coexist in the same thesis? The answer, again, lies in drug architecture, not the microenvironment alone. Third-generation ADCs, cleavable linkers, permeable payloads, high drug-to-antibody ratio, appear potent enough to partially overcome both target heterogeneity and microenvironment barriers. The conjugate’s biological sophistication, in some cases, is already doing the work a microenvironment score (or a foundation model) would otherwise have to do from outside.
This is exactly what the Complexity-Value Matrix predicts: sophistication is a property of the problem, not automatically of the solution built on top of it. Sometimes the problem (tissue heterogeneity) demands an equally sophisticated answer, as with T-DM1. And sometimes the problem is already solved by drug design, and layering additional stratification on top adds no clinical value, only regulatory complexity and cost.
The question that actually matters, then, isn’t “do we need to measure the microenvironment?” in the abstract. It’s: for which ADC family, with which linker type and which payload, would that measurement change a real clinical decision? Most of the time, that question still doesn’t have an answer grounded in prospective evidence.
Put differently: HER2 heterogeneity matters enormously for T-DM1, because that drug depends almost entirely on receptor-mediated internalization inside the cell that expresses it, there’s no bystander effect to compensate for uneven tissue. TROP2 expression seems to matter less for sacituzumab govitecan: its pH-sensitive CL2A linker also enables continuous extracellular payload release in the acidic stroma. That doesn’t mean internalization stops mattering, available mechanistic evidence indicates both pathways contribute, it means the drug has a second mode of action that partially compensates when the first one fails.
But this shouldn’t be mistaken for a free solution. The same mechanism that lets an ADC bypass part of the tissue heterogeneity problem (a cleavable linker, capable of releasing payload outside the target cell) is, per available evidence, the one most associated with systemic toxicity. A meta-analysis of 40 clinical trials and roughly 7,900 patients found significantly more grade ≥3 adverse events in ADCs with cleavable linkers than non-cleavable ones, and higher drug-to-antibody ratio was independently associated with higher probability of severe toxicity.
THE HIDDEN PRICE OF THE “POTENCY WITHOUT CLASSIFICATION” PATH
47% vs. 34%
Grade ≥3 adverse events in cleavable-linker (bystander) ADCs vs. non-cleavable linkers, meta-analysis of 40 clinical trials (n=7,879). The same mechanism that compensates for tissue heterogeneity reduces drug selectivity throughout the body.
“Build a more potent drug” isn’t a neutral alternative to “classify the patient better.” It’s a different bet, with its own cost, except that cost is paid by the patient, in adverse events, not by the pharmaceutical company, in diagnostic research budget.
THE INCOMPLETE MAP
There is, however, one real exception, and it’s worth naming precisely, not inflating it. Based on public sources identified through August 2026, TROPION-Lung17 is the only phase III trial that prospectively randomizes patients based on a computational pathology parameter: TROP2’s Normalized Membrane Ratio, calculated through Roche and AstraZeneca’s Quantitative Continuous Scoring system, which received the FDA’s first Breakthrough Device Designation ever granted to an AI-driven companion diagnostic, in 2025. That designation opens a priority regulatory pathway, it does not, yet, equal demonstrated efficacy or definitive clinical validation.
THE ONLY KNOWN PROSPECTIVE EXCEPTION
1 of 1
TROPION-Lung17: the only identified phase III trial prospectively randomizing by a computational pathology biomarker. Ongoing; no published results to date.
It’s worth being precise about what that exception actually measures. NMR quantifies the ratio between membrane and cytoplasmic target, a subcellular distribution measurement, not a direct demonstration of how much payload ends up internalized in the cell. It showed predictive potential in retrospective analyses and is being validated prospectively; it does not measure stroma, vasculature, or immune infiltration. It’s a real advance in one of the four microenvironment domains biology tells us matter (the target itself) and it still has no published clinical results confirming it.
The map, then, is only beginning to be drawn, and only in one corner. The rest of the territory (stromal architecture, immune competence, vascular accessibility) still has no standardized measurement instrument, no companion diagnostic, no regulatory vehicle carrying it through to daily clinical decisions.
There’s an asymmetry here worth naming, not to judge it but to understand it. Six ADCs are approved today with biomarker selection, each with its own linker, its own payload, its own engineering optimized in isolation. On the other side, there’s a single computational pathology platform with regulatory designation, shared so far between two drugs in the same family. A more potent linker’s engineering doesn’t transfer to a competitor, it’s investment that doesn’t compound. A tissue classification infrastructure, once validated, does transfer from one ADC to the next. It’s worth asking what would happen if a fraction of what’s currently spent making each drug individually more potent and less selective were invested in that shared infrastructure instead, not because that’s, given today’s evidence, the correct answer, but because it’s the question current investment logic doesn’t seem to be asking.
And there’s a more uncomfortable question underneath that one, perhaps the real blind spot of this entire essay. Almost everything we’ve built in precision oncology, the IHC assay, the regulatory approval criterion, even the microenvironment score this essay imagines, is designed to produce a dichotomous answer about a phenomenon that was never dichotomous: expresses or doesn’t, responds or doesn’t, high score or low score. Tissue biology, as Act IV showed, is a continuum of grays, of barriers that are sometimes mechanism and sometimes obstacle, of heterogeneity that matters enormously in one drug and almost nothing in another.
Maybe the question isn’t how well we classify, or how potent we make the drug, but whether we’re building institutions capable of acting on an answer that was never yes or no, but a range of probability.
I don’t claim this essay closes that gap. What I have is the certainty of having seen the same shape of problem in two territories nobody puts side by side: foundation models chasing perception without decision infrastructure, companion diagnostics chasing reproducibility by reducing biology to what can be standardized. From opposite directions, both reveal the same gap, we can measure more than our institutions know how to use. I’m still drawing this map. If someone else is seeing the same pattern from another point in the territory, it probably draws better in company.
Beyond the Slide is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
References:
Metzger Filho O, Viale G, Trippa L, et al. HER2 heterogeneity as a predictor of response to neoadjuvant T-DM1 plus pertuzumab: results from a prospective phase II clinical trial. J Clin Oncol. 2019;37(15_suppl):502.
Predictive Biomarkers of Antibody-Drug Conjugate Efficacy for Solid Tumors: Current Challenges and the Potential Role of Quantitative Proteomics. Clin Cancer Res. 2026;32(4):661.
Roche. Roche granted FDA Breakthrough Device Designation for first AI-driven companion diagnostic for non-small cell lung cancer [press release]. April 29, 2025.
Bardia A, Rugo HS, Tolaney SM, et al. Final results from the randomized phase III ASCENT clinical trial in metastatic triple-negative breast cancer and association of outcomes by HER2 and TROP2 expression. J Clin Oncol. 2024;42(15):1738-1744.
Rugo HS, Bardia A, Marmé F, et al. Sacituzumab govitecan (SG) vs treatment of physician’s choice (TPC): efficacy by Trop-2 expression in the TROPiCS-02 study. Cancer Res. 2023;83(SABCS22-GS1-11).
Daiichi Sankyo / AstraZeneca. TROPION-Lung17 TROP2 Biomarker Directed Phase 3 Trial of DATROWAY Initiated in Patients with Previously Treated Advanced Nonsquamous NSCLC [press release]. January 13, 2026.
Tang SC, Wynn C, Le T, McCandless M, Zhang Y, Patel R, Maihle N, Hillegass W. Influence of antibody-drug conjugate cleavability, drug-to-antibody ratio, and free payload concentration on systemic toxicities: a systematic review and meta-analysis. Cancer Metastasis Rev. 2024;44(1):18.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.