In the first article of this series, we corrected a misunderstanding: the FDA did not approve an AI-based endpoint. What it did was qualify a more precise measurement instrument for endpoints that already existed.
That seemed like a technical distinction at the time. Now, having traveled this whole road together, I think it was something more.
It was a signal that the field was beginning to ask the right question.
Not what do we measure, but how do we measure it. And underneath that, something more unsettling: are we measuring biology, or are we measuring our capacity to simplify it?
This is the final part of the series. Its purpose is not to catalog promising technologies. It’s to propose something harder, a shift in the question we ask when we design clinical evidence.
Because the future of endpoint science isn’t about measuring the same things more precisely. It’s about measuring something different.
For over a century, clinical medicine has operated under an assumption that was never seriously challenged: that complex biological phenomena can be captured with sufficient fidelity through discrete variables and static categories.
That assumption was useful. Without it, we couldn’t have standardized anything, not trials, not comparisons, not the regulatory infrastructure that gets drugs to patients.
But it has a limit. And we’re reaching it.
The problem isn’t technical. It’s ontological. I remember when I started working in liver cancer, the hardest thing for me was evaluating cirrhosis. How could I be sure that the fibrous bridges actually connected the portal spaces? Was I seeing a single layer of a three-dimensional process that the microscope couldn’t show me in depth? How do you know that what you see is what the disease is actually telling you, and not just what a two-micron section allows you to see? I’m fairly sure every pathologist has asked themselves something like this at least once; and that most of the time, we answered as best we could with the instrument we had.
That discomfort had a name we couldn’t quite articulate then. Disease is not a state. It’s a trajectory, continuous, dynamic, heterogeneous in space and shifting over time. And the measurement systems we use to capture it (ordinal scales, arbitrary cut-points, single biopsies, categorical response criteria) are snapshots of a process that doesn’t pause to be photographed.
RECIST 1.1 in oncology makes this concrete in a way that’s hard to dismiss. A patient with 29% tumor burden reduction gets classified the same way as one with 1% reduction, both as Stable Disease. A further 2% change flips the first patient into Partial Response. That boundary has no biological basis. It’s a line drawn to make statistical analysis manageable.
And when only 23% of surrogate-survival pairs in oncology show a strong correlation, the question that surfaces isn’t how to improve the surrogates. It’s whether we’ve been asking the right question at all.
The tension between biological complexity (redundant networks, nonlinear dynamics) and regulatory simplification (mean comparisons, response rates) has reached a point where continuing to ignore it has real consequences. Drugs discarded that may have worked. Drugs approved that may not have. Evidence built on instruments we know are imprecise, but keep using because validated alternatives don’t yet exist.
That’s changing. Slowly, with friction, with all the resistance that accompanies any paradigm shift in a field where regulatory caution is a genuine virtue.
But it’s changing.
The most fundamental shift in endpoint science isn’t technological. It’s conceptual.
We’re moving from asking how much did the marker change? to asking how did the disease trajectory change?
That difference sounds subtle. It isn’t.
Measuring a state at two points in time, baseline and end, and comparing the values ignores everything that happened in between. The speed of change. The direction. The fluctuations. The moments of regression and recovery. The difference between a patient who improved slowly and steadily and one who improved quickly and then relapsed. Both can have the same delta at the end of a trial. But their biology, their prognosis, and probably their response to future treatments are entirely different.
Trajectory modeling addresses exactly this. Rather than assuming all patients progress the same way, it recognizes that a population is made up of distinct subgroups with divergent progressions (rapid progressors, stable patients, partial recoverers) patterns that get buried in a cross-sectional analysis. Longitudinal biomarkers, repeated measurements of the same parameter over time, capture intra-subject variability and improve the statistical efficiency of studies in ways that a single endpoint cannot.
Multi-state models and Hidden Markov Models add another layer: they allow us to understand how patients move between different levels of disease severity, and to predict the probability of future transitions based on the trajectory observed so far
The key regulatory question is whether these systems can become actionable primary endpoints. The honest answer today is not yet, but the direction is clear. AI is opening the door to a medicine where the endpoint isn’t the resolution of a symptom but the deviation of an imminent risk trajectory. We’re not asking whether the patient got worse. We’re asking whether their trajectory changed direction.
The second structural shift is the transition from qualitative interpretation toward continuous computational quantification.
We saw this in the previous installment with AIM-MASH. But the reach of this transition goes well beyond the liver.
In nephropathology, AI allows precise quantification of interstitial fibrosis and tubular atrophy outperforming pathologist reproducibility and providing an objective, comprehensive pathological profile of the biopsy specimen.
In immuno-oncology, quantifying the spatial organization of immune cells relative to tumor cells (proximity analysis) offers continuous biomarkers that predict immunotherapy response better than simple PD-L1 positive cell counts. It’s not just how many immune cells are present. It’s where they are and how they relate spatially to the tumor.
In tissue heterogeneity, computational models can measure architectural entropy and stromal heterogeneity, factors that often correlate with therapeutic resistance and that no traditional scale captures.
Do these variables improve clinical evidence generation? In terms of sensitivity and reproducibility, yes. In terms of interpretability and stability, the answer is more complicated.
AI models can be sensitive to technical variations in slide preparation and scanner hardware, introducing a new form of instability into results. And the opacity of some deep learning models creates a challenge that is not only technical but ethical: a pathologist might agree with an AI score but still need to understand which specific morphological features are driving that metric in order to validate it biologically.
Replacing human subjectivity with algorithmic opacity isn’t progress unless it comes with explainability. Continuous quantification is the right direction. But the black box problem is real, and the field can’t keep treating it as a footnote.
The third shift, and the most ambitious, is the integration of multiple data modalities into composite measurement systems.
The core idea is this: the state of a disease cannot be faithfully captured through a single variable, however precise. Disease is the result of interactions across multiple systems, molecular, cellular, tissue-level, functional, behavioral. Measuring only one of those levels and extrapolating to the rest is exactly the problem we’ve been documenting throughout this series.
Foundation models in medicine (what the literature calls GMAI) represent the current frontier of this integration. Trained on massive datasets combining radiological images, pathology slides, genomic sequences, and electronic health records, these models are capable of cross-modal reasoning. A model can infer gene expression profiles from cellular morphology, providing a deeper understanding of tumor biology without the need for costly additional molecular assays.
The relationship between multimodal integration and systems biology is intrinsic. By capturing data from the molecular level all the way to functional level through wearables, these systems enable continuous estimation of disease state that is more robust against the noise of any single modality. The weakness of one data source gets compensated by the strength of the others.
But the barriers are real and deserve to be named honestly.
Prospective validation in real clinical environments is still scarce. Most models show exceptional performance in retrospect, but their behavior in real time with new data is much less documented.
Interpretability is the deepest obstacle. Integrating five or six data sources makes it extremely difficult to trace the cause of a prediction. And without that traceability, the confidence of regulators and clinicians has a reasonable ceiling.
The required infrastructure is significant. Platforms capable of managing, harmonizing, and processing petabytes of diverse data while meeting strict governance and privacy standards. That’s not a minor technical problem, it’s an access barrier that could deepen the inequalities between centers that already exist in traditional pathology.
And the regulatory framework for qualifying a “measurement system” as a primary endpoint simply doesn’t exist yet. Regulators tend to prefer isolated, well-understood variables over complex latent constructs. Which is understandable. And which also needs to evolve.
Digital twins in medicine are personalized virtual representations of patients that simulate their molecular and physiological state over time. They’re moving (slowly) from speculative narrative to practical tool in clinical development.
Today they’re used mainly in three concrete areas.
Virtual patient modeling: creating digital replicas using genomic and clinical data to predict individual response to different doses or therapeutic regimens.
Trial simulation: running in silico studies to evaluate the operational characteristics of a trial design before recruiting the first human patient.
Ssynthetic control arms: using digital twins conditioned on the baseline data of active-arm patients to predict how they would have evolved if they’d received placebo, reducing the size of concurrent control groups.
The FDA has discussed cases where sponsors propose using digital twins to predict placebo outcomes in Phase II and III, as long as the model is validated with representative datasets and its influence on the final decision is proportional to its demonstrated risk. The FDA’s ISTAND program provides a concrete pathway for qualifying these tools.
But there’s a risk the field needs to name clearly: conceptual overreach.
A digital twin cannot capture unknown biological interactions or unexpected side effects that only direct observation in humans can reveal. A simulated model is only as good as the data it was built on, and human biology always has the capacity to surprise in ways no model anticipated. Digital twins are a mechanism for structured information borrowing, not a substitute for randomized evidence.
Used with that awareness of their limits, they’re a powerful tool. Used as if they were reality, they become another layer of abstraction removed from the patient.
There’s a paradox at the heart of traditional clinical development: the moment we stop measuring the patient is exactly the moment the drug enters their real life.
The trial closes. Follow-up ends. And what happens after falls outside the record.
Digital biomarkers and continuous monitoring are changing that.
Concrete examples already exist. The stride velocity endpoint in DMD (which we explored earlier in this series) was qualified by the EMA as a primary endpoint. Continuous glucose monitoring is now standard in diabetes trials for measuring time-in-range; a continuous metric that captures glycemic control with a granularity that quarterly HbA1c measurements never could. Actigraphy in COPD allows qualification of physical activity measures to evaluate the real-life impact of respiratory therapies.
But the transition to real-world data is not without serious technical and behavioral challenges that would be naive to ignore.
Sensors can generate noisy signals due to incorrect use, battery problems, or hardware variability. Patient behavior acts as a massive confounding factor. Device fatigue can lead to loss of longitudinal data, compromising the integrity of continuous endpoints. And for a digital endpoint to be regulatorily acceptable, it has to pass through three levels of validation: technical verification, analytical validation of measurement precision, and clinical validation of its relationship to health status.
The fundamental challenge is converting gigabytes of raw sensor data into a clinical outcome measure that is meaningful to both the patient and the regulator. An increase in stride velocity has to translate into a tangible improvement in patient autonomy to count as evidence of efficacy. The data doesn’t speak for itself. It needs to be interpreted in the context of what matters.
In January 2026, the FDA and EMA jointly published ten guiding principles for the use of artificial intelligence across the drug lifecycle. It’s the first joint framework from these two agencies on AI, and its existence signals something important: regulators can no longer ignore this transition. They have to build the rules while the science moves forward.
The distinction between three phases of validation is critical for understanding where we are and where we need to get to.
Analytical validation asks whether the algorithm measures the biomarker precisely and reproducibly across different scanners or technical conditions.
Clinical validation asks whether that biomarker associates with a meaningful health outcome: survival, progression, quality of life.
Clinical utility asks whether using that endpoint in a trial actually leads to better therapeutic decisions or a more accurate evaluation of the drug.
These are three distinct questions with answers that can diverge. An endpoint can have impeccable analytical validation and weak clinical validation. Or it can have both and still lack demonstrated clinical utility because physicians don’t know how to integrate it into practice.
One of the biggest structural regulatory challenges is this: agencies are designed to approve static tools. But AI models may need periodic updates to correct biases or adapt to new data. The FDA’s Predetermined Change Control Plan concept is an attempt to allow algorithms to update in a controlled way without requiring a full new submission each time. It’s a step in the right direction. And it’s insufficient on its own to manage systems that evolve continuously.
The regulatory framework is not prepared for the pace of the science. That gap is real. Closing it will require a conversation the field has not yet had with sufficient honesty.
We’ve reached the end of the series. And there’s one question that can’t go unanswered.
What does it actually mean to measure disease?
For a century, the implicit answer was: assign the patient to a category. Stage 2. Grade 3. Partial response. Stable disease. Those categories gave certainty. They enabled comparison. They made decision-making possible in the middle of uncertainty.
What this series has tried to demonstrate is that the certainty was partially illusory. The categories were models. The models were imprecise. And the imprecision had consequences, in trials, in approvals, in patients.
The direction the data points toward is different. Not static labels but dynamic representations. Not isolated biomarkers but network-based inference. Not endpoint-centric trials but continuous evidence ecosystems where data collection doesn’t end when the trial closes, it integrates into the real lifecycle of the drug.
But that future can’t avoid a hard question: does increasing measurement complexity improve medicine, or does it simply increase abstraction?
If a clinical trial is won or lost based on a subtle shift in a health index derived from fifty latent variables, we need to be sure that shift is genuinely meaningful for the patient’s life. Medicine cannot move from a discipline of clinical observation to one of pure geometric navigation, disconnected from the human experience of disease.
As we move away from isolated biomarkers toward multimodal systems, the most powerful endpoints may become latent variables, mathematical constructs integrating thousands of signals with no single intuitive visual or biological representation. That raises a tension that is not only technical but ethical: can we trust what we can’t directly see? Interpretability becomes not just a scientific need but a moral one. For AI to be a trusted partner in generating clinical evidence, its decisions must be, if not fully transparent, at least consistent with established biological reasoning.
This series started with a sentence someone said in a consulting call.
“PathAI managed to get the FDA to approve an AI-based endpoint.”
That sentence was wrong. But the mistake was understandable, and revealing. It revealed that the field was beginning to sense that something important was changing, even if it didn’t yet have the precise words to describe it.
What has changed is not the definition of endpoints. It’s our understanding of their limits.
We know now that the gold standard was made of clay. That the photograph doesn’t capture movement. That the map simplifies the territory in ways that have real consequences for real patients.
And we know (and this is what this final installment has tried to demonstrate) that the tools to build something better already exist or are emerging. Trajectory modeling. Continuous quantification. Multimodal systems. Digital twins. Endpoints captured where the patient lives.
None of these tools is perfect. All of them carry barriers, scientific, regulatory, philosophical. And the temptation to replace the known imprecision of histological scoring with the unknown opacity of a deep learning model is not progress if it doesn’t come with rigorous validation, explainability, and honesty about limits.
But the direction is right.
The future of endpoint science is not measuring the same things better. It’s asking a different question. Instead of how much did the marker change, asking how did the patient’s disease trajectory move. Instead of did it cross the threshold, asking did it move in the direction that matters for their life.
That is the new science of endpoints.
And building it, with rigor, with honesty about its limits, and without losing sight of the patient on the other side of every data point, is the responsibility of everyone working in this field.
References:
Zhai J, Smith F, Soon G. Comparison of continuous, binary, and ordinal endpoints. J Biopharm Stat. 2025 Oct;35(6):1143–60. doi:10.1080/10543406.2025.2489288 PubMed PMID: 40285714.
An MW, Mandrekar SJ, Branda ME, Hillman SL, Adjei AA, Pitot H, et al. Comparison of continuous versus categorical tumor measurement-based metrics to predict overall survival in cancer treatment trials. Clin Cancer Res. 2011 Oct 15;17(20):6592–9. doi:10.1158/1078-0432.CCR-11-0822 PubMed PMID: 21880789; PubMed Central PMCID: PMC3195893.
Sanyal AJ, Loomba R, Anstee QM, Ratziu V, Kowdley KV, Rinella ME, et al. Utility of pathologist panels for achieving consensus in NASH histologic scoring in clinical trials: Data from a phase 3 study. Hepatol Commun. 2023 Dec 22;8(1):e0325. doi:10.1097/HC9.0000000000000325 PubMed PMID: 38126958; PubMed Central PMCID: PMC10749704.
Zayan TA, Khafagy HA. Artificial intelligence-powered digital pathology: A potential gold standard for kidney biopsy interpretation? World Adv Renal Med. 2025 Dec 20;1(3):72–9. doi:10.25259/WARM_21_2025
Cinar I, Cinar E. Assessment of interobserver variability in Gleason grading for prostate carcinoma. North Clin Istanb. 2025 Jun 23;12(3):337–43. doi:10.14744/nci.2025.11456 PubMed PMID: 40843333; PubMed Central PMCID: PMC12365472.
Weintraub WS, Lüscher TF, Pocock S. The perils of surrogate endpoints. Eur Heart J. 2015 Sep 1;36(33):2212–8. doi:10.1093/eurheartj/ehv164 PubMed PMID: 25975658; PubMed Central PMCID: PMC4554958.
Mittal A, Kim MS, Dunn S, Wright K, Gyawali B. Frequently asked questions on surrogate endpoints in oncology-opportunities, pitfalls, and the way forward. eClinicalMedicine. 2024 Sep 13;76:102824. doi:10.1016/j.eclinm.2024.102824 PubMed PMID: 39764569; PubMed Central PMCID: PMC11701476.
Columbia University Mailman School of Public Health [Internet]. 2016 [cited 2026 May 8]. Trajectory Analysis. Available from: https://www.publichealth.columbia.edu/research/population-health-methods/trajectory-analysis
Kwon BC, Achenbach P, Dunne JL, Hagopian W, Lundgren M, Ng K, et al. Modeling Disease Progression Trajectories from Longitudinal Observational Data. AMIA Annu Symp Proc. 2021 Jan 25;2020:668–76. PubMed PMID: 33936441; PubMed Central PMCID: PMC8075441.
Park KC, Yoo W. Translating multimodal foundation models into oncology: Toward a future where AI directs diagnosis and therapy. Genes Dis. 2025 Nov 27;13(4):101958. doi:10.1016/j.gendis.2025.101958 PubMed PMID: 41847347; PubMed Central PMCID: PMC12989828.
A Multimodal Foundation Model of Spatial Transcriptomics and Histology for Biological Discovery and Clinical Prediction [Internet]. [cited 2026 May 8]. Available from: https://arxiv.org/html/2604.03630v1
Ding R, Yang Z, Yu Q, Zhou J, Ni B, Zheng M, et al. AI-based pathomics in kidney diseases: progress and application. Ren Fail. 47(1):2598080. doi:10.1080/0886022X.2025.2598080 PubMed PMID: 41403092; PubMed Central PMCID: PMC12713211.
Rowsthorn E, Xia Y, Breakspear M, Fripp J, Robinson GA, Ashton N, et al. Multimodal latent composites are associated with cognition and Alzheimer’s disease dementia: a framework for systems-level brain health [Internet]. medRxiv; 2026 [cited 2026 May 8]. p. 2026.02.21.26346745. Available from: https://www.medrxiv.org/content/10.64898/2026.02.21.26346745v1 doi:10.64898/2026.02.21.26346745
Analytics TS Sonrai. The Scientist [Internet]. [cited 2026 May 8]. Accelerating Biomarker Discovery with Multimodal Data and Foundational AI Models. Available from: https://www.the-scientist.com/accelerating-biomarker-discovery-with-multimodal-data-and-foundational-ai-models-73836
Translating Multimodal Foundation Models into Routine Clinical Practice [Internet]. [cited 2026 May 8]. Available from: https://eam.edu.eu/news/Translating%20Multimodal%20Foundation%20Models.html
The Latent Space Hypothesis Toward Universal Medical Representation Learning [Internet]. [cited 2026 May 8]. Available from: https://arxiv.org/html/2506.04515v1
Ph.D SS. FDA’s Bayesian Guidance and the Emerging Role of Digital Twins in Clinical Trials. Medium [Internet]. 2026 Mar 16 [cited 2026 May 8]. Available from: https://medium.com/@subhrajit.samanta/fdas-bayesian-guidance-and-the-emerging-role-of-digital-twins-in-clinical-trials-6bd972c916d3
Silva A, Vale N. Digital Twins in Personalized Medicine: Bridging Innovation and Clinical Reality. J Pers Med. 2025 Oct 22;15(11):503. doi:10.3390/jpm15110503 PubMed PMID: 41295204; PubMed Central PMCID: PMC12653454.
www.hoganlovells.com [Internet]. [cited 2026 May 8]. FDA’s evolving regulatory framework for AI use in drug & device clinical trials and research. Available from: https://www.hoganlovells.com/en/publications/fdas-evolving-regulatory-framework-for-ai-use-in-drug-device-clinical-trials-and-research
Regulatory Engagement Opportunities when Developing Digitally Derived Endpoints.
IntuitionLabs [Internet]. [cited 2026 May 8]. FDA Draft Guidance on AI in Drug Development Explained. Available from: https://intuitionlabs.ai/articles/fda-draft-guidance-ai-drug-development
Digital Twins in Clinical Trials: How AI Virtual Control Arms Are Transforming Study Design in 2026 [Internet]. [cited 2026 May 8]. Available from: https://www.pienomial.com/blog/digital-twins-in-clinical-trials-how-ai-generated-virtual-control-arms-are-rewriting-study-design-in-2026
Tarnanas I, Seixas A, Wyss M, Vlamos P, Çöltekin A. Merging multimodal digital biomarkers into “Digital Neuro Fingerprints” for precision neurology in dementias: the promise of the right treatment for the right patient at the right time in the age of AI. Front Digit Health. 7:1727707. doi:10.3389/fdgth.2025.1727707 PubMed PMID: 41602207; PubMed Central PMCID: PMC12832889.
IntuitionLabs [Internet]. [cited 2026 May 8]. FDA Digital Health Guidance: 2026 Requirements Overview. Available from: https://intuitionlabs.ai/articles/fda-digital-health-technology-guidance-requirements
Katsoulakis E, Wang Q, Wu H, Shahriyari L, Fletcher R, Liu J, et al. Digital twins for health: a scoping review. NPJ Digit Med. 2024 Mar 22;7:77. doi:10.1038/s41746-024-01073-0 PubMed PMID: 38519626; PubMed Central PMCID: PMC10960047.
A new regulatory milestone: what the joint FDA and EMA’s AI principles can mean for clinical trial technology [Internet]. [cited 2026 May 8]. Available from: https://www.suvoda.com/insights/blog/a-new-regulatory-milestone
Shanmugam U, Rajendran MK, Natarajan J, Karri VVSR. Clinical Trial Design and Regulatory Requirements for Artificial Intelligence as a Medical Device: A PRISMA-ScR–Guided Scoping Review of Global Guidance and Evidence (2017–2025). J Clin Med. 2026 Mar 4;15(5):1937. doi:10.3390/jcm15051937 PubMed PMID: 41827353; PubMed Central PMCID: PMC12985890.
Reiss J, Ankeny RA. Philosophy of Medicine. In: Zalta EN, Nodelman U, editors. The Stanford Encyclopedia of Philosophy [Internet]. Winter 2025. Metaphysics Research Lab, Stanford University; 2025 [cited 2026 May 8]. Available from: https://plato.stanford.edu/archives/win2025/entries/medicine/
Allam A, Feuerriegel S, Rebhan M, Krauthammer M. Analyzing Patient Trajectories With Artificial Intelligence. J Med Internet Res. 2021 Dec 3;23(12):e29812. doi:10.2196/29812 PubMed PMID: 34870606; PubMed Central PMCID: PMC8686456.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.