RSS Amplifier

The Munro Report · May 26, 2026

How much evidence is enough?

0
Sign in to vote or save

Alasdair Munro · The Munro Report

Adam Kucharski recently posted a thoughtful piece reviewing Helen Pearson’s new book Beyond Belief, in which he picks up her concept of “aluminium-standard” evidence. This describes evidence which is faster, cheaper, and often more feasible than the gold standard, but still good enough to inform the decision being made. The argument is that demanding gold-standard evidence before action can become a form of “weaponised certainty”; a way for people opposed to an intervention to keep raising the bar of evidence until no decision is possible.

This is true, but it left me considering another problem. If gold-standard evidence isn’t always available, and isn’t always required, how do you know when the evidence you have is enough? There is a different problem where “Aluminium-standard” risks becoming weaponised in its own right, or a way of saying “whatever evidence I happen to find convincing, applied to whichever decision I happen to want to make.” We need a way of knowing when good enough really is good enough.

I want to offer a framework for thinking about this that is rarely articulated in public-facing evidence debates, and I think makes the argument easier to navigate.

The first thing to get out of the way is the idea that knowledge is something we either have or we don’t. We essentially never know anything. We hold beliefs with varying degrees of confidence, and evidence is what moves those beliefs.

This probability language used for this is known as Bayesian (after the Reverend Thomas Bayes), but it is intuitive and it does not require any formal training in probability to use. We are all doing it constantly.

You have a prior belief, for example about whether your child is unwell (”she seems off today”), you gather evidence (she has a temperature, she won’t eat, she’s unsettled), and you update your belief accordingly. By the time you decide whether to call the GP, your confidence has moved from thinking “she might have something”, to, “she probably has something”.

Three things matter for any decision of this kind. Your prior belief: how likely the thing was before you had this evidence. The strength of the evidence itself: how much it should move that belief. And the key, often overlooked element - the threshold for action: how confident you need to be before doing whatever it is you are considering doing.

The amount of evidence you need is not a fixed property of the evidence. It is a function of the gap between where your prior sits and where the threshold for decision sits.

That gap is what the evidence has to close.

Medicine has earned its reputation for high evidentiary standards, and rightly so. Kucharski wisely refers to the story of Archie Cochrane and his attempts to overcome “eminence based medicine” with practice based on high quality evidence.

For most novel interventions, the prior is uninformative. We have a mechanistic hypothesis, e.g., this drug should bind that receptor and produce that effect, but mechanism has a long and humiliating history of failing to predict clinical outcome. The graveyard of plausible-sounding interventions that turned out not to work, or to actively harm patients, is enormous.

This is partially a reflection of the fact that humans, including those of us trained in medicine, systematically overestimate the benefits of interventions and underestimate their harms. Plausible mechanisms feel like evidence. Anecdotal successes are remembered; failures are attributed to the disease. This is true of clinicians, of patients, of researchers, and of policymakers.

Meanwhile, the threshold to act is high. Most medical interventions are expensive, and recommended to large numbers of people, many of whom are otherwise well, would have recovered anyway, or had other options available. The downside of recommending the wrong thing is borne by people who did not choose to take on the uncertainty. This is important because an individual choosing to try an experimental treatment for their own incurable illness can apply one threshold; a public health body recommending an intervention to millions applies a much higher one.

Another crucial element is that in medicine (or even public health) we usually can run the trial, and generate the “gold standard”. Unlike Kucharski’s caribou where the breeding programme cannot meaningfully be randomised, most drug interventions, vaccines, devices and procedures can be tested in proper clinical trials. In fact, many things you might imagine to be impossible to randomise could be, if scientists are able to think more flexibly. The feasibility of generating better evidence raises the bar on whether you should act on worse evidence.

When the gold standard is genuinely unavailable, aluminium becomes a defensible substitute. When the gold standard is available but slow, it’s more complicated.

The framework also explains why we sometimes accept lower-quality evidence in medicine and are right to do so.

Severe outcomes raise the cost of inaction, which lowers the threshold to act. Limited alternatives narrow the comparison: when there is no other plausible option, you are no longer asking “is this better than nothing” against a high bar, but against a much lower one. Rare diseases change feasibility: a trial that would resolve uncertainty cannot practically be run, so the alternative to acting on imperfect evidence is acting on no evidence at all. Established safety profiles reduce the potential costs of lowering the threshold.

Each of these is a specific lever being pulled, not a general permission to be more relaxed. That distinction is what separates exceptions with good reasons from special pleading.

Few topics have been argued more poorly than masks during the COVID-19 pandemic. People saying “we know masks do/don’t work” was worse than unhelpful, because almost everyone involved was using the word “masks”, AND the word “work”, to mean different things, and was applying their conclusion to different decisions and different thresholds.

Consider a fit-tested N95/FFP3 worn by a healthcare worker entering a respiratory isolation room. The prior is informative and favourable: regulators have tested these devices extensively, the mechanism by which they work is well understood, and we are confident they remove most particulate matter under fixed conditions. The threshold to act is low: the intervention is brief, the cost to the wearer is minimal, the contact is high-risk, and the alternative is exposure. Even a thin RCT evidence base does not need to do much work to clear that gap. Recommending this is straightforward.

Now consider intermittent community mask use. Cloth or surgical masks, worn variably, removed during the periods of highest transmission risk (at home, eating, socialising with close contacts). The prior is much weaker, because most of the mechanistic confidence applies to a different intervention under different conditions. The evidence is weak: the RCTs that exist are inconclusive, and observational evidence that is heavily confounded and (depending on who generated the study) often produces effect sizes so large they are implausible at face value. The threshold to act depends on what you are deciding to do.

If the decision is whether to recommend mask use to the public, the threshold is much lower. The intervention is relatively cheap, the imposition is small for most people, and a recommendation merely shifts the choice for people who would not have considered it otherwise. Even weak evidence could feasibly clear that bar.

If the decision is whether to mandate mask use across entire populations for years at a time, the threshold is much higher. Mandates impose real social, educational, economic and environmental costs, and risks damage to public trust in institutions. The evidence required to justify a mandate is therefore substantially more than the evidence required to justify a recommendation, even though the intervention itself is nominally the same.

I think this helps some of the mask debate become intelligible:

Much of what was thought of as debates about the evidence were really debates about thresholds; and the thresholds are less governed by science than they are by personal values.

Libertarians will have vastly different evidence thresholds to socialists on public health interventions. This isn’t right or wrong in any testable or scientific sense. I think many people are just resistant to admitting it.

Thanks for reading The Munro Report! This post is public so feel free to share it.

Share

It is absolutely right that weaponised certainty exists and is dangerous. Opponents of an intervention have clearly raised the evidentiary bar as a way of blocking action. Climate change is a current example and vaccine debates often have the same shape. Demanding perfect evidence for things you already oppose, and accepting anecdote for things you already support is dangerous and intellectually dishonest.

Within the realm of medicine and public health however, I think the inverse of this has been more consequential. Call it weaponised urgency: invoking the impossibility of perfect evidence to bypass evidence the field actually could have generated. “We cannot wait for an RCT” when the RCT could have been run in a few months but never was. “Action is needed now” when the burden of disease will still be there in a year, and a year is what would have produced a defensible answer. The pandemic offered repeated examples.

The reason this matters is that despite the efforts of people like Archie Cochrane and Austin Bradford-Hill, medicine’s track record is littered with interventions adopted on insufficient evidence and later withdrawn; not with interventions withheld until certainty arrived and never came.

The textbook examples include hormone replacement therapy being recommended for cardiovascular prevention in healthy postmenopausal women on the strength of observational data, until the Women’s Health Initiative trial showed it caused more harm than benefit. Class I antiarrhythmic drugs were given routinely to patients after heart attacks on the mechanistically obvious basis that suppressing arrhythmias should save lives, until the CAST trial found they tripled mortality. The list is long, the harms in each case were substantial, and in every instance the evidence to know better was either available, or could have been generated, before the intervention was rolled out.

Weaponised certainty is absolutely real, and is a problem, but it is rarely the move that ends up doing the most damage in medicine. The more common pattern is acting too fast, on optimistic priors, with evidence that turns out not to support the confidence we placed in it.

Where two people work from the same evidence and reach different decisions, the disagreement is usually about the threshold, not the evidence. This framework helps us see that, but it does not tell you which threshold is right - or if there truly is a “right”. We must avoid mistakes such as claiming that evidence is insufficient when your real objection is to the threshold.

When a threshold is genuinely contested, and the gold standard is available and feasible to generate, we should err on the side of asking for more or better evidence.

There may be some exceptions where waiting is itself a decision with costs, and those cases need to be argued explicitly. But the prior on “we should act now on what we have” deserves to be more sceptical because historically that is the more common mistake we have made.

Aluminium standard evidence is not lower-quality gold-standard evidence. It is what you reach for when the gold standard is not available, and only when the gap you are trying to close is small enough that aluminium can do the work. The art of evidence based decision making lies in being able to tell the difference. The art of evidence based public debate lies in being able to say, explicitly, which one you think you are doing.

Work out whether you are arguing about the evidence (science) or about the threshold (values) before you dive in.

No posts

Read the original on alasdairmunro.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.