In the previous chapter, we left a question unanswered.
We saw how endpoints are constructed, how they are validated, and why some bring us closer to clinical truth while others pull us further away without us realizing it. We saw that a biomarker can work perfectly while the patient gets worse. We saw that for years we measured how many meters a child with Duchenne could walk in a hospital corridor while his real life was happening somewhere else entirely.
And at the end, we raised something uncomfortable: what happens when all that evidence reaches the clinic and the physician doesn’t know what to do with it?
That’s what we’re going to explore now.
But to understand it properly, we need to start with a problem that goes deeper than endpoints, deeper than biomarkers, deeper even than trial design.
The problem is that we have spent decades building better maps without asking whether the territory we’re measuring is the right one.
The Science Says It Works. The Trial Confirms It. The Patient Doesn’t Improve.
There is a sequence that repeats itself in modern medicine with a frequency that should trouble us more than it does.
A drug enters development with a solid biological hypothesis. Phase II trials show a clear signal in the biomarker of interest. The pivotal trial demonstrates statistical significance on the primary endpoint. The regulatory agency approves it. Clinical guidelines incorporate it.
And then the drug reaches the real world.
And something doesn’t work.
Not always. Not in every patient. But frequently enough that the question becomes unavoidable: are we measuring the right thing, or are we measuring what we can measure?
The architecture of modern medicine rests on a rarely named divide: the difference between a drug that is approvable from a regulatory perspective and a drug that is usable from a clinical perspective. These are two different questions. And for too long we have treated them as if they were the same.
The Surrogate Problem: A Promise of Speed Against Biological Uncertainty
The use of surrogate endpoints (progression-free survival (PFS), objective response rate (ORR), reduction in amyloid burden) responds to a logic of efficiency. Clinical development is slow and expensive, and waiting decades to confirm benefit in overall survival isn’t always viable.
But something gets forgotten with remarkable consistency: the validity of a surrogate is not an intrinsic property of the measure. It is a mathematical relationship that must be rigorously validated for each specific context of use, each disease type, and each drug mechanism of action.
And that validation, more often than the field admits, is insufficient.
The case of immune checkpoint inhibitors in oncology illustrates this with uncomfortable precision. PFS is assumed to be a good predictor of overall survival. But the strength of that association depends critically on how the drug is being used:
Let’s look closely at that third row. R² of 0.01 to 0.22. No correlation. What that means in practice is that in combination ICI therapies, PFS is a statistical mirage that offers no guarantee of additional survival benefit.
For the clinician, this creates a paralyzing ambiguity: should a drug that delays tumor growth by three additional months be prescribed if the probability that the patient lives one more day is statistically null?
There is another problem layered on top of this. A common error in interpreting evidence is equating a p-value < 0.05 with meaningful clinical benefit. Statistical significance measures the probability that a result isn’t due to chance. Clinical relevance evaluates the magnitude of the effect and its impact on the patient’s life. In many oncology trials, PFS improvements are statistically significant but clinically marginal, a delay in progression of just 1.2 or 1.5 months.
The concept of the Surrogate Threshold Effect (STE) attempts to address this by defining the minimum magnitude of effect on the surrogate needed to predict real clinical benefit with confidence. Many drugs approved under accelerated pathways don’t reach that threshold. Which means the observed benefit is, in practical terms, clinical noise a massive cost to the healthcare system without tangible public health return.
Approval based on unvalidated surrogates transforms the pharmaceutical market into a post-commercialization experimentation laboratory, where the patient assumes the biological risk and the system assumes the financial risk.
When Clearing the Brain Is Not the Same as Healing the Mind
In 2021, the FDA approved aducanumab for Alzheimer’s disease. What followed was one of the most revealing regulatory controversies in recent medical history.
The drug worked on the biomarker. That’s not in dispute. It reduced beta-amyloid plaques in the brain in a significant, measurable way through PET imaging. The agency’s premise was that this reduction was “reasonably likely” to predict clinical benefit in cognitive function.
The trial data told a different story:
One positive trial, one negative trial, with a trend toward cognitive worsening in some subgroups. And an advisory panel that voted 10 to 1 against approval. The agency approved the drug anyway.
What followed was telling. The Centers for Medicare and Medicaid Services decided not to cover the drug outside of clinical trials. Neurologists divided. Some centers prescribed it. Others refused. And patients, people with Alzheimer’s and their families, who had waited decades for an answer, were caught in the middle of a debate that the evidence couldn’t resolve.
Beyond this, the treatment was associated with a rate of ARIA (amyloid-related imaging abnormalities), including cerebral edema and hemorrhage of approximately 35-40%, requiring constant monitoring via MRI.
The amyloid was cleaner. The cognition kept deteriorating.
The map said one thing. The territory said another.
When Stopping the Tumor Is Not the Same as Saving the Patient
The aducanumab case is not an anomaly. It’s a pattern.
Bevacizumab entered the market for metastatic breast cancer with data that seemed compelling. The E2100 trial showed a 5.5-month improvement in progression-free survival, with a hazard ratio of 0.48 and impeccable statistical significance (p < 0.0001). The drug looked like a revolution.
Five years later, the FDA withdrew that indication.
The confirmatory trials, AVADO and RIBBON-1, failed to replicate the magnitude of the initial effect. And more critically: none of the four randomized trials demonstrated that patients lived longer. The toxicity, meanwhile, was real: severe hypertension, gastrointestinal perforations, and treatment-related deaths in 0.8-1.2% of cases.
A drug that delays radiological progression without extending life or improving well-being is a victory for the trial protocol. But it is a defeat for the purpose of medicine.
Between approval and withdrawal, during those years when uncertainty was already visible in the data, thousands of patients received that treatment. With its costs. With its side effects. With the hope that an FDA-approved drug generates.
PFS had worked perfectly as a biomarker. It simply didn’t predict what mattered.
The Trial Patient Doesn’t Exist in Your Clinic
There is a structural reason why this happens, and it has nothing to do with bad faith.
Clinical trials are designed to demonstrate efficacy under optimal conditions, what in regulatory language is called efficacy, not effectiveness. To guarantee internal validity, investigators select homogeneous populations, systematically excluding elderly patients, those with multiple comorbidities, or those on concurrent medications. The result is an ideal patient who rarely exists in actual practice.
The data in multiple myeloma say this with a precision that should generate more debate than it does. Population studies have demonstrated that patients treated in real clinical practice have a 51% greater risk of progression or death than patients in pivotal trials for the same regimens (HR = 1.51). The risk of death in the real world is 76% higher (HR = 1.76).
The same drug. The same diagnosis. Double the risk of dying.
The difference is not in the drug. It’s in the fact that the real patient is 74 years old, has diabetes, moderate renal insufficiency, and takes seven medications. The trial patient was 58, had no relevant comorbidities, and was monitored every two weeks by a clinical research team.
Added to this is the heterogeneity that trials fail to capture: the patient’s functional status, tumor burden, organ reserve, life circumstances. When trials don’t adequately report on these subgroups (or when surrogates are used as an average across an entire indication) the physician lacks the information needed to make a precise risk-benefit assessment. The result is variability in clinical adoption that doesn’t reflect physician preference. It reflects insufficient evidence.
The Regulatory Maze: Speed Against Certainty
The acceptance of surrogate endpoints by the FDA and EMA is not a technical error. It’s a political and social decision, designed for serious diseases with unmet medical needs.
And it carries a known price.
The use of surrogates can reduce drug development time by a median of 11 to 19 months in oncology. But approximately 15% of oncology indications approved under accelerated pathways have subsequently been withdrawn because confirmatory trials failed or were never completed. This phenomenon of “dangling approvals” creates a situation where physicians and patients use treatments whose real efficacy is, at best, uncertain for periods of 3 to 10 years.
The differences between agencies add another layer of complexity:
Between 2019 and 2024, the EMA approved the same indications as the FDA with a median delay of 181 days and with labels more restricted to specific subgroups. This “European caution” reflects a greater emphasis on long-term safety, while the FDA prioritizes early access and exploratory innovation. Neither position is wrong. But the divergence creates a real problem for global development programs that must function across both jurisdictions simultaneously.
Clinical Practice Under the Shadow of Uncertainty
When regulatory evidence is ambiguous, the burden of decision shifts to clinical guidelines and the individual ethics of the physician.
And there, too, consensus is absent.
The NCCN in the U.S. tends to rapidly incorporate drugs with accelerated approval into its recommendations, often without highlighting the underlying surrogate uncertainty. ESMO, in contrast, has taken a more critical stance. Its Magnitude of Clinical Benefit Scale (ESMO-MCBS) evaluates drugs not by their legal approval, but by the strength of their survival and quality-of-life data.
When the ESMO scale was applied to 267 oncology trials, only 12% demonstrated substantial clinical benefit. Among genomic targeted therapies approved between 2015 and 2022, fewer than one third met ESMO criteria for significant clinical benefit at the time of approval.
This creates a real ethical dilemma: is it appropriate to offer a patient a treatment that appears in NCCN guidelines but that ESMO rates as low clinical benefit? For the physician, evidence stops being a guide and becomes a source of conflict.
And at the end of that chain is the patient. Who expects certainty. And whom the physician must explain that a drug has been approved but that we don’t know with confidence whether it will help them live longer.
Communicating uncertainty is the final link in this gap. Frameworks exist to address it the SPIKES protocol, the Ask-Tell-Ask method, models of prognostic communication that promote transparency about what is unknown. But the clinical reality is that many physicians fear that honesty about the fragility of the evidence will destroy the patient’s hope. And the lack of specific training in managing biological and statistical uncertainty remains one of the greatest barriers to closing this gap at the bedside.
The Algorithm That Works in the Paper and Collapses in the Hospital
This gap is not exclusive to drugs. It exists everywhere we confuse performance under controlled conditions with performance in real life.
In digital pathology, we see it with particular clarity.
An algorithm reaches the market with impressive metrics. 94% accuracy. AUC of 0.97. Validated in three independent cohorts. The paper passes peer review.
And then the algorithm arrives at the real hospital.
The scanner has three years of deferred maintenance and the images contain artifacts the model never encountered during training. The slides are prepared with a staining protocol slightly different from the center where the model was developed. The computer that must process them takes four minutes per image in a workflow that needs to be agile. The pathologist who must interpret the algorithm’s output received no training on how to calibrate their judgment when the model and their experience disagree.
The 94% accuracy becomes something no one can measure well in that context. And what cannot be measured well cannot be trusted.
The problem wasn’t the model. The model was good. The problem is that we optimized for the benchmark and not for the deployment environment. For the map, not for the territory.
We don’t need the best model. We need the model that works in the real life of that hospital, with that scanner, with those slides, with that pathologist.
That is a different question. And it is the right question.
Three Things That Could Change This
There are no simple solutions to a structural problem. But there are decisions the field could make today.
Mandatory surrogate validation proportional to risk. No drug should receive traditional approval based on a surrogate without a prior meta-regression demonstrating strong correlation (R² > 0.70) with final clinical outcomes in that specific disease. The correlation between a biomarker and the endpoint that actually matters is not an intrinsic property of the marker it is a relationship that must be proven before approval, not after.
Default inclusion of the real patient. Pivotal trials should include representative cohorts of patients with comorbidities and advanced age, or alternatively be immediately followed by mandatory pragmatic trials as a condition for commercialization. A perfectly controlled trial that cannot predict what will happen in clinical practice is not as rigorous as it appears.
Explicit transparency about uncertainty. Clinical guidelines and drug labeling must clearly communicate the nature of the endpoint used (surrogate versus clinical) and the status of confirmatory studies. Today, a drug approved under accelerated pathways based on a reasonably likely surrogate appears in guidelines with the same weight as a drug with fifteen years of overall survival data. That equivalence is false. And building clinical decisions on it has real consequences.
The integration of real-world data, pragmatic trials, electronic health records, evidence generated outside controlled trials, offers a concrete opportunity to close this gap. From 2025 onward, there is a growing trend toward using RWE not only as post-marketing support but as an integral part of trial design. The ICH M14 guideline, adopted in September 2025, establishes harmonized principles for the use of real-world data in safety assessment. Its successor, ICH E23, will seek to extend those principles to effectiveness evaluation, ensuring that RWE is accepted consistently by regulatory agencies and health technology assessment bodies worldwide.
The direction is right. How quickly the field can move in that direction depends on resolving these problems, not ignoring them.
The Question That Part 3b Will Try to Answer
Everything we have explored here lives at a level that is necessary but incomplete.
There is a place where this gap between the map and the territory is especially visible, especially technical, and especially relevant for those of us who work in pathology: histological scoring.
When a trial’s endpoint depends on how a pathologist interprets a biopsy, every problem we have discussed converges at a single point. Interobserver variability. The subjectivity inherent in scoring. The difference between what the criterion defines in the protocol and what the eye sees under the microscope.
How do you build solid evidence when the measurement instrument is, by nature, imprecise?
That’s what we’ll explore in the next installment.
Because before we talk about solutions, we need to understand the problem clearly.
And the histological problem runs deeper than it appears.
Beyond the Slide is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
References:
Gyawali B, Hey SP, Kesselheim AS. Evaluating the evidence behind the surrogate measures included in the FDA’s table of surrogate endpoints as supporting approval of cancer drugs. eClinicalMedicine. 2020 Apr 13;21:100332. doi:10.1016/j.eclinm.2020.100332 PubMed PMID: 32382717; PubMed Central PMCID: PMC7201012.
Sala I, Pagan E, Pala L, Oriecuia C, Musca M, Specchia C, et al. Surrogate endpoints for overall survival in randomized clinical trials testing immune checkpoint inhibitors: a systematic review and meta-analysis. Frontiers in Immunology. 2024 Jan 29;15:1340979. doi:10.3389/fimmu.2024.1340979
Durer S, Fu P, Chen Z, Dowlati A. Meta-analysis of surrogate endpoints for overall survival in extensive-stage small-cell lung cancer. ESMO Open. 2025 Oct 15;10(11):105843. doi:10.1016/j.esmoop.2025.105843 PubMed PMID: 41101157; PubMed Central PMCID: PMC12550574.
Prasad V. Surrogate end points in oncology: the speed–uncertainty trade-off from the patients’ perspective. Nat Rev Clin Oncol. 2025 May;22(5):313–4. doi:10.1038/s41571-025-01007-z
Yuvraj. Real-World Data in Pharmaceutical Efficacy-Effectiveness Gap [Internet]. 2025 Aug 27 [cited 2026 Apr 24]. Available from: https://www.worldpharmatoday.com/it-data-management/real-world-data-in-managing-pharmaceutical-efficacy-effectiveness-gap/
Visram A, Chan KKW, Seow H, Pond G, Gayowsky A, Mohyuddin GR, et al. Comparing the clinical trial efficacy versus real-world effectiveness of treatments for multiple myeloma: a population-based study. Haematologica. 2024 Aug 22;110(1):228–33. doi:10.3324/haematol.2024.285768 PubMed PMID: 39219457; PubMed Central PMCID: PMC11694101.
PhD MS MD. The Importance of Real-World Evidence in Medical Research and Drug Development | Applied Clinical Trials Online [Internet]. 2026 [cited 2026 Apr 24]. Available from: https://www.appliedclinicaltrialsonline.com/view/real-world-evidence-medical-research-drug-development
Berg A, Mayers RP, Richards S. Biogen’s Aduhelm controversy as a case study for accelerated approval biomarkers in Alzheimer’s and related diseases. MIT SPR. 2022 Aug 29;3. doi:10.38105/spr.y9p3sxuqg1
What is the clinical evidence supporting, and controversy surrounding, the use of aducanumab (Aduhelm TM) for Alzheimer disease? | Drug Information Group | University of Illinois Chicago [Internet]. [cited 2026 Apr 24]. Available from: https://dig.pharmacy.uic.edu/faqs/2021-2/october-2021-faqs/what-is-the-clinical-evidence-supporting-and-controversy-surrounding-the-use-of-aducanumab-aduhelm-tm-for-alzheimer-disease/
Rethinking FDA’s Accelerated Approval Pathway: New Draft Guidances and Implications for Drug Companies. McGuireWoods [Internet]. [cited 2026 Apr 24]. Available from: https://www.mcguirewoods.com/client-resources/alerts/2025/1/rethinking-fdas-accelerated-approval-pathway-new-draft-guidances-and-implications-for-drug-companies/
Mehta GU, Pazdur R. Oncology Accelerated Approval Confirmatory Trials: When a Failed Trial Is Not a Failed Drug. J Clin Oncol. 2024 Nov 10;42(32):3778–82. doi:10.1200/JCO-24-01654
Ardito V, Falk L, Gasol Boncompte M, Pontes C, Robinson J, Schreyögg J, et al. Do conditional marketing authorisations actually accelerate patient access? Time-to-access of conditional vs. standard marketing authorisations in Italy, Spain, and Germany. J Pharm Policy Pract. 19(1):2651404. doi:10.1080/20523211.2026.2651404 PubMed PMID: 41983243; PubMed Central PMCID: PMC13072672.
Review of All Solid Tumor Drug Approvals From 2019 to 2024 by US Food and Drug Administration, European Medicines Agency, and Brazilian Health Regulatory Agency. JCO Global Oncology [Internet]. [cited 2026 Apr 24]. Available from: https://ascopubs.org/doi/10.1200/GO-25-00326
Lau F, Seifert R. Comparison of drug approvals of the FDA and EMA between 2013 and 2023. Naunyn Schmiedebergs Arch Pharmacol. 2026;399(1):279–99. doi:10.1007/s00210-025-04412-4 PubMed PMID: 40608116; PubMed Central PMCID: PMC12894151.
Mooghali M, Mitchell AP, Skydel JJ, Ross JS, Wallach JD, Ramachandran R. Characterization of accelerated approval status, trial endpoints and results, and recommendations in guidelines for oncology drug treatments from the National Comprehensive Cancer Network: cross sectional study. BMJ Med. 2024 Apr 5;3(1):e000802. doi:10.1136/bmjmed-2023-000802 PubMed PMID: 38596814; PubMed Central PMCID: PMC11002412.
Tibau A, Hwang TJ, Avorn J, Kesselheim AS. Clinical value of guideline recommended molecular targets and genome targeted cancer therapies: cross sectional study. BMJ. 2024 Aug 20;386:e079126. doi:10.1136/bmj-2023-079126 PubMed PMID: 39164034; PubMed Central PMCID: PMC11333991.
Most Targeted Cancer Drugs Lack Substantial Clinical Benefit | MDedge [Internet]. [cited 2026 Apr 24]. Available from: https://www.mdedge.com/content/most-targeted-cancer-drugs-lack-substantial-clinical-benefit
Gyawali B, de Vries EGE, Dafni U, Amaral T, Barriuso J, Bogaerts J, et al. Biases in study design, implementation, and data analysis that distort the appraisal of clinical benefit and ESMO-Magnitude of Clinical Benefit Scale (ESMO-MCBS) scoring. ESMO Open. 2021 Apr 20;6(3):100117. doi:10.1016/j.esmoop.2021.100117 PubMed PMID: 33887690; PubMed Central PMCID: PMC8086024.
Cox C, Hatfield T, Parry M, Fritz Z. To what extent should doctors communicate diagnostic uncertainty with their patients? An empirical ethics vignette study. J Med Ethics. 2025 Feb 26;51(11):e109932. doi:10.1136/jme-2024-109932 PubMed PMID: 40011039; PubMed Central PMCID: PMC12573412.
Piórek A, Płużański A, Kowalski DM, Krzakowski M. Best Practices and Communication Strategies for Informing Oncology Patients About Treatment Discontinuation and Transition to Palliative Care—A Practical Guide for Oncologists. Cancers (Basel). 2025 Nov 3;17(21):3566. doi:10.3390/cancers17213566 PubMed PMID: 41228358; PubMed Central PMCID: PMC12608836.
Libert Y, Peternelj L, Bragard I, Liénard A, Merckaert I, Reynaert C, et al. Communication about uncertainty and hope: A randomized controlled trial assessing the efficacy of a communication skills training program for physicians caring for cancer patients. BMC Cancer. 2017 Jul 10;17:476. doi:10.1186/s12885-017-3437-8 PubMed PMID: 28693515; PubMed Central PMCID: PMC5504708.
Real-World Data (RWD) & Real-World Evidence (RWE). EUPATI Toolbox [Internet]. [cited 2026 Apr 24]. Available from: https://toolbox.eupati.eu/resources/patient-toolbox/real-world-data-rwd-real-world-evidence-rwe/
CCRPS Clinical Research Taininrg [Internet]. [cited 2026 Apr 24]. Real World Evidence RWE Integration Trends in Clinical Trials 2025 Report. Available from: https://ccrps.org/clinical-research-blog/real-world-evidence-rwe-integration-trends-in-clinical-trials-2025-report
ICH Official web site : ICH [Internet]. [cited 2026 Apr 24]. Available from: https://www.ich.org/news/ich-m14-guideline-reaches-step-4-ich-process
International Council for Harmonisation of Technical Requirements for Pharmaceuticals for HumanUSE. E23: Considerations for the Use of Real-World Evidence (RWE) to Inform Regulatory Decision Making with a focus on Effectiveness of Medicines. Geneva, Switzerland.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.