RSS Amplifier

Brain Trials · Aug 4, 2026

Anatomy of a Failed Trial

0
Sign in to vote or save

Jose-Alberto Palma MD PhD · Brain Trials

A couple of weeks ago I wrote about CELIA, Biogen’s Phase 2 trial of the anti-tau ASO diranersen, which missed its primary endpoint and had its largest clinical signal at its lowest dose.

A question I got often was some version of: so did the tau hypothesis just fail?

No, it did not. In fact, it’s the opposite: that at least one dose showed evidence of clinical slowing means the tau hypothesis is well alive.

But… the primary endpoint failed right? Yes it did.

Which brings us to the topic of this post: “failed” has different interpretations, at least eight, and the differences among them are everything.

A trial can miss because the biology is wrong. It can miss because the biology is right and the drug never reached the tissue. It can miss because the drug reached the tissue but the dose was too low, or high enough to cause harm. It can miss because the drug worked but the measuring instrument couldn’t see it. It can miss because nothing went wrong at all but chance had its say.

Each points to a different next step: abandon the target, change the molecule, change the dose, change the patients, change the endpoint, or run the same study again. Putting them together into the same “failed” basket throws away useful information the trial might have produced.

Six checks on a negative trial. Checks 1–3 are Morgan's three pillars of survival (Drug Discovery Today, 2012); checks 4–6 are how the clinical trial was built. Fail any one and your negative trial may have not tested the hypothesis it intended, regardless of what it may have taught you.

In 2012 a team at Pfizer, led by Paul Morgan, went back through 44 of their own Phase 2 programs and asked why each had succeeded or failed.¹ The finding was not the failure rate. It was that in almost half of the programs (43%), it was impossible to tell whether the mechanism had been tested at all.

Not just “the drug didn’t work.” But whether the drug had even engaged the target it was supposed to.

They also found something about the programs that did work: every success had demonstrated that the drug engaged its target. None of the failures had.

From this they proposed three things a program should establish before anyone reads a negative result as meaningful, which they called the three pillars of survival:

  • Exposure: the drug reaches the tissue where the target lives, at adequate concentration, for long enough.

  • Binding: it engages the target once it gets there.

  • Pharmacology: engaging the target produces the biological effect it was supposed to produce.

The framework is now 14 years old and has held up. Pfizer reconfirmed it internally in 2020, and it has become standard in brain drug development, where all three pillars are harder to establish than anywhere else in medicine.

The same author later helped build a broader version at AstraZeneca, the 5Rs: the right target, right tissue, right safety, right patient, and right commercial potential, plus what they called the right culture.² Under it, success rates from candidate nomination through completion of Phase 3 rose from 4% in 2005–2010 to 19% in 2012–2016.

Each pillar has a standard measurement in neurodegeneration clinical trials, and each has a catch.

Exposure is usually cerebrospinal fluid (CSF) drug concentration, or its ratio to blood, which tells us how much of a given drug crosses into the CNS. Of course, for intrathecal or intracerebral administration, this says very little: we know that the drug gets into the CSF. The tricky thing is to determine if the drug gets into the relevant brain area. Those programs would need imaging of labeled drug, or an honest admission that distribution into specific brain tissue is unknown.

Binding is cleanest when the target is a receptor, because we can use PET occupancy studies to measure it directly. Nowadays, most neurodegeneration targets aren’t receptors, so engagement has to be inferred from the nearest biochemical consequence. Which creates a problem: in neurodegeneration, binding and pharmacology are often the same measurement. Tominersen’s CSF huntingtin protein levels, or total tau CSF levels for MAPT-targeted therapies, serve as both.

Pharmacology means the biology moved: tau, alpha-synuclein, or huntingtin falling in CSF, or plaque and tangle burden falling on PET. Clearing this, however, does not guarantee that the clinically-relevant endpoints will move. We have many examples for this. For instance, tominersen, Roche’s anti-HTT ASO, lowered both mutant huntingtin and neurofilament light in CSF, and still produced no clinical benefit in HD patients.⁵ Verubecestat, Merck’s BACE1 inhibitor, lowered Aβ42 by up to 90%, and patients got worse.⁹

There is a third possibility this pillar tends to hide: the biomarker moves, but in the direction opposite to the one the mechanism predicts. Roche’s NLRP3 inhibitor selnoflast suppressed its cytokine targets convincingly — stimulated IL-1β down about 91% against placebo — and then raised the TSPO-PET neuroinflammation signal by roughly 41% in the midbrain, the region an anti-inflammatory drug should be quieting.¹⁷ That was presented as a pharmacodynamic effect supporting further study. I’ve written separately about this pattern, which I’d call hypothesis flipping. A wrong-way biomarker is not a nuance to be explained. It means the pillar was not cleared.

A problem in many brain trials is that there is no good biomarker of target engagement. The temptation here is to skip ahead to clinical endpoints. “If we cannot measure a biomarker, let’s just go ahead and measure clinical endpoints.” That’s precisely what the 43% of Pfizer trials in Morgan’s paper did. So the answer is to build the biomarker first (slow, expensive, and likely to benefit your competitors as well), or to proceed with the clinical endpoint: you may get lucky and hit it (particularly if you expect to see movement in a few weeks, in fact a lot of symptomatic drugs were approved this way over the decades), but if you get a negative result, you will not know what failed.

So here is the sorting I actually use. It’s just two questions, not a list.

Six things have to be true. Morgan’s three pillars cover the pharmacology. Three more cover the trial itself.

Did the drug get where it’s needed? This is the hardest pillar in brain disease and the one most often assumed rather than shown. Antibodies against intracellular targets are one example. First, antibodies given intravenously reach CSF at a fraction of a percent of their blood concentration. Not only is their CNS concentration low, but they are unable to reach the intracellular space, where relevant protein aggregates occur. Cinpanemab, an antibody against alpha-synuclein, was negative across essentially every endpoint in Parkinson’s disease.⁷ Like other antibodies, it targeted alpha-synuclein outside cells. A drug that never reaches the compartment where the pathology sits has not tested whether that pathology matters. We’ll see how the prasinezumab Phase 3 trial turns out.

Did it engage the target? UniQure's AMT-130, an intracerebral AAV therapy for Huntington disease, did not show reductions in mutant huntingtin in CSF — the entire mechanistic point of the drug.⁶ UniQure argued that local delivery might not register in CSF, yet its own animal studies (minipig) showed exactly that reduction. A trial can report some inconclusive, potential benefit and still leave you unable to say whether the mechanism was engaged.

Did engagement do anything? Gantenerumab's twin Phase 3 trials enrolled 1,965 people and produced CDR-SB differences of 0.31 and 0.19, neither significant.⁸ Roche also disclosed that amyloid removal came in lower than expected, at roughly half of projection. Turbocharging gantenerumab by conjugating it to a transferrin-receptor shuttle — trontinemab — did result in dramatically more robust amyloid plaque clearance: at the higher dose, 92% of participants dropped below the amyloid positivity threshold within 28 weeks, with ARIA-E under 5%.¹⁶ Two identical Phase 3 trials are now underway, and it would be surprising if the CDR-SB difference weren't more robust than that with gantenerumab.

Did the patients have the disease? Clinical diagnosis of Alzheimer's is wrong against autopsy in 25 to 30% of cases.¹⁰ In one research cohort, about a quarter of patients with clinically diagnosed early-onset Alzheimer's had no amyloid on PET. For 20 years, anti-amyloid drugs were given to trial populations recruited on clinical grounds, meaning a substantial fraction of participants had no amyloid plaque to remove. We have, unfortunately, too many examples for this.

Could the clinical endpoint see the effect? Some neurodegenerative disorders are slowly progressive. Parkinson's disease, HD, some SCAs — you need several years to see progression on current clinical scales.⁴ In Parkinson's disease, MDS-UPDRS parts 1 and 2 barely move after 1–2 years, and part 3 moves a little more.¹¹ But the harder question is whether that movement is clinically meaningful and reflects how the patient feels and functions, rather than what the physician observes. This is a complicated question that deserves its own post.

Could the design and statistical plan answer the question? CELIA’s primary endpoint was a dose-response relationship on CDR-SB.¹⁵ But Phase 1b had already shown spinal fluid tau lowering flattening out around 60% across doses, so the pharmacology predicted no dose-response. Everybody is wondering whether they should have powered the study to detect individual differences in each dosing arm versus placebo. The answer is, of course, yes. That would have required more patients (likely double) but then Biogen might have had a “positive Phase 2” instead.

Off to one side sits a 7th question that isn’t about pharmacology or design at all: are the data real? Occasionally the problem is unblinding, misconduct, or images that don’t survive scrutiny. I’ve written about one such cases before.

It’s the one category where the answer is not a better trial.

Only three things left.

The idea is wrong. The target isn’t causal, or modifying it doesn’t change the course of the disease. Cardiology has the cleanest example. Epidemiology linked higher HDL cholesterol to fewer cardiac events, so drugs were built to raise it.¹⁴ Dalcetrapib raised HDL by about 30% and did not reduce events at all. Anacetrapib, far more potent, roughly doubled HDL and did cut major vascular events by 9%, but the benefit was due to lowering non-HDL cholesterol, not elevating HDL, and it was never approved. Potent engagement, repeated, adequately powered, nothing at the other end. That is what a refuted hypothesis looks like.

Is there any way to know in advance whether a target is worth the attempt? Partly. Drug mechanisms with human genetic support are roughly 2.6 times more likely to reach approval.³ Tau has that support: mutations in MAPT cause inherited dementia. The tau hypothesis therefore enters this discussion with a real prior in its favour, which is one more reason not to write it off on the strength of a trial that never tested it.

The effect is real but too small to detect here. Common, and hard to distinguish from the above without a bigger or longer trial. It partially overlaps with the question of whether the clinical instrument used was the right one.

Chance. Aducanumab’s twin Phase 3 trials ran to the same protocol in the same populations.¹³ In EMERGE, the high dose slowed functional decline by 40%. In ENGAGE, the high dose produced a CDR-SB difference of +0.03 — nothing, in the wrong direction. Two identical experiments, two opposite answers. The power of chance.

Huntington’s disease has just supplied the model, and it’s worth pausing on because it’s so rare.

Roche’s tominersen, an anti-HTT ASO given by lumbar puncture, was the first drug to lower mutant huntingtin in humans. Its Phase 3 was halted in 2021, with patients on the more frequent schedule doing worse than placebo. Roche then analyzed the wreckage and made a specific claim: lower exposure in younger patients with less disease burden might help, where higher exposure in more advanced patients had harmed. It built a trial to test exactly that, in adults aged 25 to 50 with early disease, at a lower dose given less often: 100 mg three times a year, against 120 mg every eight weeks in the halted study.

Top-line results came in this month. The drug was well tolerated. It lowered mutant huntingtin, and it lowered neurofilament light. And Roche discontinued the program.5

Every one of the four tests was passed. The explanation was named in advance, it predicted something specific, the next trial was built to test it, and there was a stopping rule the company honored.

It also leaves huntingtin lowering in an uncomfortable place. Tominersen is the only program in this space that has clearly shown both target engagement and downstream pharmacology, and with a corrected dose in a better-chosen population it produced no clinical benefit. Whether the drug ever reached the striatum in adequate amounts is still not established, so the first pillar remains open. But nobody should claim the hypothesis has been rescued by another route until a competing program demonstrates the engagement this one did.

When a trial misses, “failed” is the beginning of a finding, not the finding. Ask first whether the experiment happened: did the drug get there, engage the target, do something, in the right patients, measured by an instrument that could see it, in a design that could answer the question. Only if all of that holds does the result tell you anything about the biology.

Then ask who named the explanation, and when.

That changes how a press release reads. A company saying it didn’t meet the primary endpoint but saw encouraging trends has told you nothing, on purpose, because vagueness preserves every option. A company saying it achieved 40% of target exposure and is advancing at triple the dose has told you which pillar failed, made a prediction, and given you something to hold it to.

One of those is a vague statement. The other is a scientific hypothesis. Only the second is worth anything, to a patient deciding whether to enroll, an investor deciding whether to stay, or a field deciding whether twenty years of negative trials meant what everyone said they meant.

Two lists I'd like to build in the comments: your experience on negative trials that genuinely taught the field something, and negative trials that taught us nothing because the experiment never really happened. Please, let me know in the comments:

Leave a comment

If reading trials this way is useful, consider subscribing.

What an endpoint means, why a missed primary isn’t the whole story, how to weigh benefit against harm - that’s what my book is about.

A Patient’s Guide to Clinical Trials: Navigating the Promise and Pitfalls of Experimental Treatments (Bloomsbury) — available now.

This analysis represents my personal views, not necessarily those of my employer, and is based entirely on publicly available information.

References

  1. Morgan P, Van Der Graaf PH, Arrowsmith J, Feltner DE, Drummond KS, Wegner CD, Street SDA. Can the flow of medicines be improved? Fundamental pharmacokinetic and pharmacological principles toward improving Phase II survival. Drug Discovery Today. 2012;17(9–10):419–424. https://doi.org/10.1016/j.drudis.2011.12.020— 44 Pfizer Phase 2 programs; mechanism testing indeterminate in 43%; the three pillars of survival.

  2. Morgan P, Brown DG, Lennard S, et al. Impact of a five-dimensional framework on R&D productivity at AstraZeneca. Nature Reviews Drug Discovery. 2018;17:167–181 — the 5R framework, with success from candidate nomination through Phase 3 completion rising from 4% in 2005–2010 to 19% in 2012–2016. See also Cook D, Brown D, Alexander R, et al. Nature Reviews Drug Discovery. 2014;13:419–431, where the framework was first set out.

  3. Minikel EV, Painter JL, Dong CC, Nelson MR. Refining the impact of genetic evidence on clinical success. Nature. 2024;629(8012):624–629. https://doi.org/10.1038/s41586-024-07316-0 — mechanisms with human genetic support are 2.6 times more likely to reach approval; the advantage increases with confidence in the causal gene and is largely independent of genetic effect size, minor allele frequency, or year of discovery.

  4. Mortberg MA, Vallabh SM, Minikel EV. Disease stages and therapeutic hypotheses in two decades of neurodegenerative disease clinical trials. Scientific Reports. 2022;12:17708.

  5. Tominersen: GENERATION HD1 (NCT03761849), halted March 2021, with the more frequent dosing arm faring worse than placebo. GENERATION HD2 (NCT05686551) was a Phase 2 study of 301 participants across 15 countries, aged 25–50 with early HD, testing 100 mg of tominersen by spinal fluid injection three times a year over a minimum 16-month treatment period, assessed using the composite Unified Huntington’s Disease Rating Scale (cUHDRS) or Total Functional Capacity. Primary source: Roche and Genentech, “Update on Two Clinical Studies: GENERATION HD2 (tominersen) and POINT-HD (RG6496),” global patient community letter, 9 July 2026. The letter reports that the study met its safety and biomarker objectives but not its efficacy objective, that tominersen was well tolerated with no new safety signals, that it “significantly lowered mutant huntingtin protein and Neurofilament Light Chain (NfL)” relative to placebo, and that there was no meaningful impact on clinical efficacy. https://hdsa.org/wp-content/uploads/2026/07/Roche-GENERATION-HD2-_-POINT-HD-global-patient-community-letter-July2026.pdf

  6. AMT-130: uniQure Phase 1/2 three-year topline, September 2025 — approximately 75% slowing on cUHDRS (about 1.0 point) and 60% on Total Functional Capacity (about 0.6 points) versus propensity-matched external controls from ENROLL-HD and TRACK-HD, 12 participants per cohort. Earlier sham-controlled data showed no significant cUHDRS difference at 12 months at either dose, with the sham group declining more slowly than external natural-history data predicted; CSF neurofilament light fell approximately 7% below baseline in sham patients at 12 months, against approximately 8% in treated patients at 36 months. No clear mutant or wild-type huntingtin reduction was reported in CSF, despite reductions being demonstrated in the company’s minipig work (Science Translational Medicine, 2021). See my earlier appraisal for the full argument.

  7. Cinpanemab: SPARK trial, New England Journal of Medicine 2022;387:408–420. Gosuranemab in progressive supranuclear palsy achieved over 95% CSF target engagement with no clinical effect.

  8. Gantenerumab: GRADUATE I and II topline, Roche, November 2022 — CDR-SB differences 0.31 (p=0.095) and 0.19 (p=0.30), n=1,965, amyloid removal below projection.

  9. Verubecestat: APECS, New England Journal of Medicine 2019 — n=1,454, higher dose worsened CDR-SB and daily function, hippocampal volume loss approximately 0.5% greater and evident by week 13, Aβ42 lowered up to 90%. EPOCH (n=1,957) halted for futility.

  10. Diagnostic accuracy in Alzheimer’s: clinical versus autopsy misdiagnosis rates of 25–30%; approximately 25% of clinically diagnosed sporadic early-onset Alzheimer’s patients in the LEADS cohort are amyloid-PET-negative.

  11. Prasinezumab: PASADENA, primary endpoint MDS-UPDRS Parts I+II+III at week 52 not met, less progression on Part III; PPMI natural history shows minimal 52-week change in Parts I and II. The follow-up PADOVA trial, using time to motor progression, also missed its primary endpoint.

  12. Lecanemab: Clarity AD, New England Journal of Medicine 2023 — CDR-SB 1.21 versus 1.66, difference 0.45, 27% less decline.

  13. Aducanumab: EMERGE high dose, functional decline slowed 40%; ENGAGE high dose, CDR-SB difference +0.03 (p=0.833).

  14. CETP inhibitors: dal-OUTCOMES (dalcetrapib) stopped for futility; REVEAL (anacetrapib) — HDL roughly doubled, major vascular events reduced 9%, benefit attributed to non-HDL cholesterol lowering; not submitted for approval.

  15. CELIA/diranersen: AAIC 2026 presentation, data cut 11 March 2026 — primary dose-response endpoint p=0.21.

  16. Trontinemab: an antibody built from gantenerumab’s Aβ-binding regions coupled to a transferrin receptor 1 shuttle module (Roche’s Brainshuttle platform), which raises CNS exposure at low systemic doses. Phase 1b/2a Brainshuttle AD study (NCT04639050), ongoing. At 3.6 mg/kg over 28 weeks: adjusted mean amyloid PET reduction of approximately 99 centiloids, 92% of participants below the 24-centiloid positivity threshold, 71% cleared to 11 centiloids or below, and ARIA-E in under 5% — presented at CTAD, December 2025, in cohorts 3 and 4 (151 participants). An earlier dose-expansion cut presented at AAIC in July 2025 reported 91% (49/54) below 24 centiloids and 72% (39/54) below 11 centiloids, so figures vary slightly by data cut. Accompanied by early reductions in CSF and plasma total tau, pTau181, pTau217 and neurogranin. Note that all of these are pharmacodynamic endpoints: no clinical efficacy data exist yet. The Phase 3 programme comprises two identical trials, TRONTIER 1 and TRONTIER 2, in early symptomatic Alzheimer’s disease, targeting roughly 1,600 participants across 18 countries, with a further Phase 3 study in preclinical disease planned. Roche media release, 28 July 2025: https://www.roche.com/media/releases/med-cor-2025-07-28

  17. Selnoflast (RO7486967): Pagano G, Zinnhardt B, Mracskó EZ, et al. A Phase 1b study to test the safety, pharmacokinetics and pharmacodynamics of selnoflast in early stage Parkinson’s disease. Presented at AD/PD 2026, Copenhagen, 17–21 March — 57 participants, 17 sites, 28 days. Target engagement: ex-vivo LPS-stimulated IL-1β suppressed approximately 91% versus placebo (p=0.003), CSF IL-18 down approximately 30% (p=0.02), plasma hsCRP down approximately 58% (p=0.01). TSPO-PET ([18F]DPA-714, most affected side): midbrain +41% versus placebo (80% CI 24–60%, p=0.0009), putamen +42% (80% CI 26–61%, p=0.0006), against a pre-specified study alpha of 0.20. For the fuller argument, and the Bial pariceract case alongside it, see my earlier piece.

Read the original on braintrials.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.