Fourteen randomised trials. 182,880 adults. No reduction in deaths, not overall, not from heart disease, not from cancer. What general health checks reliably produced was more new diagnoses. That’s the Cochrane review by Krogsbøll and colleagues, published in the BMJ in 2012, and the 2019 update did not rescue it. Across four trials and 164,881 people, health checks had little or no effect on fatal and non-fatal heart disease, and the reviewers graded that finding high certainty. Their conclusion ran to one line: general health checks are unlikely to be beneficial.
Now hold that next to your last five years.
If you’re the kind of person who reads this newsletter, you probably have a folder. Panels going back years. A wearable on your wrist since 2021. Maybe a scan. You’ve listened to the podcasts and you can hold your own on ApoB. And if you’re honest about it, the file has grown considerably more than you have.
That’s not a personal failure. It’s the predictable output of a model, and the model has a name.
Camp one says: measure more. The implied theory is that health behaviour is bottlenecked on information, and that once you can see the number, you’ll act on it. It’s a seductive theory because it’s true of almost everything else in a competent adult’s life. It’s mostly false here.
The most-cited evidence in camp one’s favour is the wearable literature, and it’s worth being precise about what it shows. An umbrella review in Lancet Digital Health in 2022, led by Ferguson at the University of South Australia, pooled 39 systematic reviews and meta-analyses. Activity trackers did move behaviour: roughly 1,800 extra steps a day, about 40 minutes more walking, around a kilo of body weight. Those are real effects and I’m not going to pretend otherwise.
Look at what happened downstream, though. In the same review, effects on blood pressure, cholesterol and glycated haemoglobin were typically small and often not statistically significant. The device moved the behaviour it counts. It did not reliably move the physiology you bought it for.
Then there’s the trial nobody in camp one quotes. The IDEA study, published in JAMA in 2016, put 470 young adults with overweight or obesity through six months of intensive diet and exercise counselling, then randomised them: half added a wearable with a web interface, half continued with web-based self-monitoring alone. Over 24 months, the group with the wearable lost less weight. One trial, one population, and I’d caution anyone against building a worldview on it. But it should at minimum retire the assumption that adding measurement to an intervention makes the intervention stronger.
Scans deserve the same honesty. A 2018 BMJ systematic review by Gibson and colleagues pooled 32 studies of brain and body MRI in apparently asymptomatic adults and found potentially serious incidental findings in 3.9% of scans, rising to 12.8% once findings of uncertain seriousness were included. Roughly half of the serious ones were suspected cancers. A separate 2019 systematic review in the Journal of Magnetic Resonance Imaging put the pooled false-positive proportion for whole-body screening MRI at around 16%, with wide uncertainty around it. Both reviews report substantial heterogeneity, so treat the point estimates as ranges rather than facts. The direction is clear enough: in a population without symptoms, a very sensitive test mostly finds things that turn out to be nothing, and each one costs a follow-up, a wait, and a few weeks of a person’s peace.
So camp one gave a generation of careful, well-resourced adults a folder, a monthly subscription, and a low-grade dread. The one thing it didn’t reliably give them was a different body.
Here’s the mechanism, and it’s been measured.
Webb and Sheeran’s 2006 meta-analysis of experimental studies found that a medium-to-large change in people’s intentions (d = 0.66) produced only a small-to-medium change in their behaviour (d = 0.36). Reviews of health behaviour more broadly find intention accounts for somewhere around a fifth of the variance in what people actually do. You can move someone’s intention a long way and move their behaviour barely half as far. Every intervention built on informing, persuading and alerting is fishing in that shallow half.
This is where I have to be careful, because the sloppy version of my argument is that measurement doesn’t work, and that’s not what the evidence says.
Harkin and colleagues, writing in Psychological Bulletin in 2016, meta-analysed 138 experimental studies covering 19,951 people. Prompting someone to monitor their progress towards a goal did promote goal attainment, d = 0.40, and the change in monitoring frequency mediated the effect. Monitoring works.
But read the moderators, because they’re the whole argument. The effect was larger when the outcome was physically recorded, and larger when it was reported or made public. In other words, monitoring works when it’s attached to a goal, written down, and seen by someone else. That is a very specific behaviour. It is not the same activity as a device silently accruing readings you glance at on the train.
Camp one borrowed the credibility of progress monitoring and shipped something structurally different: data collection with no goal fastened to the other end, no accountability, no consequence for the reading being bad. That version has no evidence behind it, and the reason is legible in the mechanism rather than mysterious.
The backlash arrived on schedule. Put the wrist device in a drawer. Switch off. Rest. Nervous system regulation. The unspoken claim is that the missing input was recovery.
For a subset of people that’s exactly right, and I want to be careful here because burnout is real and the people carrying it are not helped by a founder telling them to try harder.
But the evidence on what a restful break does over time is not ambiguous. A more recent meta-analysis in European Psychologist, by Wendsche and colleagues, pooled 13 studies covering 1,428 employees, with an average holiday length of 11 days. Well-being improved, d = 0.25. Then people went back to work, and after the first post-holiday week the differences from pre-holiday levels were no longer significant, all effect sizes at or below 0.12. The earlier meta-analysis by de Bloom in 2009 found the same shape: a positive effect during the break, a fade after work resumed. Longer holidays didn’t protect the gain.
Rest is necessary. Six days of it is not a stimulus, it’s a pause, and a pause does not change a slope. If your problem is that you already know what to do and don’t do it, a week of switching off returns you to exactly the life you left, slightly better rested, with the same architecture waiting.
Camp two isn’t wrong about rest. It’s answering a question most of camp one’s refugees weren’t asking.
The Diabetes Prevention Program is the study I’d put in front of anyone who thinks structured doing is the soft option.
3,234 adults at high risk of type 2 diabetes, randomised to placebo, metformin, or an intensive lifestyle programme, published in the New England Journal of Medicine in 2002. Over a mean 2.8 years, the lifestyle arm cut diabetes incidence by 58%, with a confidence interval of 48 to 66. Metformin cut it by 31%. The trial was stopped early because the answer was that clear.
Now look at what the lifestyle arm actually consisted of, because this is the part that gets lost when people summarise it as “diet and exercise”. Every participant had an individual coach. Frequent contact. A structured 16-session core curriculum teaching behavioural self-management. Supervised activity sessions. A maintenance phase with restarts built in for the people who fell off. The published goals were modest, 7% weight loss and 150 minutes a week of brisk walking, and the participants hit them because the programme was engineered around the doing rather than the knowing.
Nobody in that trial needed new information. They needed the thing built.
The smallest version of the same principle has its own literature. Gollwitzer and Sheeran’s meta-analysis of 94 independent tests found that forming an if-then plan, specifying when, where and how, produced a medium-to-large effect on goal attainment, d = 0.65, over holding the goal alone. One sentence, written in advance, moves behaviour more than most of what camp one sells for four figures.
If you’re going to keep something, keep the number that responds to work.
The strongest candidate is cardiorespiratory fitness. Mandsager and colleagues at the Cleveland Clinic followed 122,007 adults who underwent treadmill testing between 1991 and 2014, publishing in JAMA Network Open in 2018. Over 1.1 million person-years, fitness was inversely associated with all-cause mortality, and the association had no observed ceiling. Elite performers, two standard deviations above the age and sex mean, had roughly an 80% lower adjusted mortality risk than the lowest group.
Be precise about what that is. It’s a cohort study of people referred for stress testing, not a randomised trial, and reverse causation is a live concern: unwell people perform worse on treadmills. So this is a strong association, not a proven cause. But it points the same direction as decades of exercise trials, and unlike almost every other marker in your folder, it’s a number you can personally move within weeks by doing something difficult on purpose.
That’s the test I’d apply to everything you currently measure. Does this number change when I do the work? If the answer is no, or no one has ever checked, it’s not a priority, it’s a subscription.
Assessment still matters. It just isn’t a lifestyle. Measure properly, at a point in time, so the work can be built from where you actually are. Then close the file and go and do the work. That’s the shape of the six days we run: assess on day one, build a protocol from those numbers, then implement it with people beside you and measure again before you leave. It’s also why the week doesn’t end at the gate: your team stays on your protocol for the first month home, which is where most retreat gains historically go to die.
I built Atlas Cove as an outsider to this industry, in Portugal, and the reason we run one small cohort a month with ten rooms rather than something scalable is unglamorous. Coaching density is not a luxury feature. In the DPP it was the intervention.
If you’ve spent five years in camp one, the instinct when it fails is to conclude you need better data. That’s the trap, and the industry is happy to sell you the upgrade.
You don’t have an information problem. You’ve got a folder that proves it. What you have is an energy account running a quiet overdraft, a body that has been described in detail and asked to change nothing, and a set of intentions that have never once been converted into a structure with a coach, a schedule and a witness.
So the question isn’t which panel to run next.
What’s the last thing you measured that changed what you did on a Tuesday? And if the honest answer is nothing, what would it take to spend the next six days building instead of watching?
If this is the argument you’ve been waiting for someone to make, subscribe. The next essays go deeper into what actually moves a baseline: what the training evidence supports, what the recovery science does and doesn’t show, and how to tell a marker worth your attention from one worth ignoring.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.