This is Part 1 of a two-part series on health wearables. Part 1 covers consumer devices: what they actually measure, and what the peer-reviewed literature says about the claims on the box. Part 2 will cover digital biomarkers in clinical trials: as of 2026, wearables have been used in 1,021 drug trials, and the FDA has formally qualified zero.
A Wall Street Journal technology columnist recently wore four fitness trackers at once for three weeks. The Oura Ring 5, the Google Fitbit Air, the Whoop MG embedded in athletic clothing, and an Apple Watch Series 11 — all simultaneously. She then took the results somewhere the manufacturers would probably prefer she hadn’t: a Stanford sleep lab, where a polysomnogram and a chest-strap ECG provided the clinical reference points.
The comparison was well done and the findings are instructive.
But the test has three significant gaps that, once you fill them in with the peer-reviewed literature, change the picture considerably. Importantly, only about 11% of commercially available wearable devices have been independently validated for even one biometric outcome.
But first, the results that actually held up.
Thanks for reading Brain Trials! Subscribe for free to receive new posts.
What held up
Resting heart rate. All four devices came within one beat per minute of the sleep lab reference. The Oura hit it exactly at 59 bpm. This is consistent with what the science shows: a 2024 umbrella review in Sports Medicine covering 24 systematic reviews and more than 430,000 participants found a mean heart rate bias of just ±3% across consumer wearables. Wrist-based optical sensors are reasonably reliable for resting heart rate, particularly during sleep, when motion, the main source of error, is minimal. If your tracker says your resting heart rate is 58 bpm, it is probably 57 or 59 bpm.
Total sleep duration. The Apple Watch matched the lab’s recorded sleep duration to the minute: six hours and fifty-two minutes. The Fitbit Air was close behind. This too is consistent with the literature, which finds that most devices can approximate total sleep time to within a workable margin, even if the same umbrella review notes a tendency to overestimate it by more than 10%. The Apple Watch result was better than the literature average. That is worth noting.
Those are the two things that held up.
Everything else is more complicated, and in one important case, not tested at all.
The finding that should be on the box
All four devices overestimated deep sleep. Considerably.
The polysomnogram, which measures brain wave activity via electrodes placed directly on the scalp, recorded 28 minutes of deep sleep. Every device reported more than 60 minutes.
This is not a calibration problem. It is a physics problem.
Deep sleep, also called slow-wave sleep, is defined by specific low-frequency, high-amplitude electrical patterns called delta waves. The only instrument that can detect them is one measuring the brain’s electrical activity directly. A wrist sensor inferring sleep stages from heart rate fluctuations and movement has no access to that signal. It is making an educated guess, and on deep sleep specifically, the guess is systematically wrong.
This finding is not new. A landmark 2021 study in the journal Sleep (Chinoy et al.) tested 7 consumer sleep-tracking devices against polysomnography and found that deep sleep was consistently overestimated across the board. A 2025 study testing six devices against PSG in 62 adults found that the errors (in duration and staging) ranged from 12% to 180% compared to the gold standard (Schyvens et al.) A 2026 study (Searles et al.) found the problem is even more pronounced in older adults, who are precisely the population for whom sleep quality matters most clinically.
The sleep doctor who reviewed the WSJ columnist’s results offered the right framing:
“A sleep tracker is equivalent to having a bathroom scale.”
The point being that what matters is whether the number trends up or down over time, not what it says on any given night. That is a useful reframe. It is also a significant step down from “precision sleep science,” which is the register most wearable marketing operates in.
What the WSJ piece didn’t test — and why it matters
The comparison tests three things: sleep duration, sleep staging, and heart rate during exercise. It does not test the metric that may be most widely used and most consequentially inaccurate: calorie and energy expenditure estimation.
Millions of people use their fitness trackers to manage weight, plan nutrition, and close their calorie rings. The 2024 umbrella review I mentioned earlier (Doherty et al.) found that energy expenditure error ranges from -21% to +15% depending on the device and the activity. That is the average range. Individual studies have found errors exceeding 100% during high-intensity exercise. Apple Watch has been documented miscalculating energy expenditure by as much as 115% during graded exercise testing in controlled conditions. Only about 9% of the energy expenditure comparisons in the validation literature fell within ±3% of the criterion measure.
Put plainly: the calorie number your watch shows you after a workout is not a reliable figure.
It is a very rough estimate built from your heart rate, your movement, your height and weight, and a proprietary algorithm that has not been validated against direct calorimetry. Using it to decide how much to eat, how hard to train, or whether Tuesday’s workout “earned” a particular meal involves more trust in the number than the evidence warrants.
The WSJ piece also involves one person. One skin tone, one fitness level, one sleep architecture, one set of wrists. The scientific literature is consistent that wearable accuracy varies meaningfully by age (devices perform worse in older adults), by BMI, by sex, and by skin pigmentation. The optical sensors that all four tested devices use work by shining light into the skin and measuring how much blood is pulsing beneath. Darker skin pigmentation absorbs more light, reducing the signal quality for heart rate and, particularly, for blood oxygen saturation estimates. This is a documented limitation with implications.
A third gap: the comparison runs for 3 weeks. Consumer validation studies are increasingly showing calibration drift over time; accuracy that is acceptable in the first month can degrade as the device’s reference model diverges from the individual’s actual physiology. Three weeks is not enough to observe this.
The numbers behind the marketing
Here is a figure worth holding onto: only 11% of commercially available consumer wearable devices have been validated for at least one biometric outcome (Doherty et al). And because each device measures multiple outcomes, the number of actual validation studies conducted represents roughly 3.5% of those needed for a comprehensive assessment of what these devices claim to measure.
The market is projected to reach $186 billion by 2030. The validation literature covers 3.5% of the claims.
The readiness scores, recovery grades, body battery levels, strain indices, and illness predictions that now populate wearable apps sit almost entirely in that unvalidated 96.5%. They are generated by proprietary algorithms, assessed against outcomes the companies themselves define, and not available for independent researchers to audit.
Some of them may be tracking something real. We don’t know yet, because the evidence needed to know hasn’t been generated in a form that independent science can evaluate.
For instance, the Oura ring flagged the WSJ columnist as showing “major signs of strain” the day before she woke up sick, while other devices rated her “peak” or “primed for training.” This is certainly interesting. But it is just one data point, interpreted retrospectively, in one person. Resting heart rate typically rises one to two days before symptomatic infecious (flu-like) illness regardless of which device you use to measure it (Radin et al).
Whether the Oura’s specific score reflects something more sophisticated than an elevated resting heart rate, is not known.
What they are actually good for
None of this is an argument for leaving the device in a drawer.
The behavior-change evidence is real: tracking activity makes people move more, and arbitrary but visible targets work for adherence regardless of their biological precision — a point I’ve made in this newsletter before in the context of step counts and sleep hours. A numerical goal you can check before bedtime is more likely to be pursued than a vague instruction to “be more active.”
For identifying long-term trends rather than delivering nightly verdicts, these devices offer potential value. A resting heart rate that has climbed 5 bpm over 6 weeks of poor sleep is a real signal. Total sleep time trending downward through a stressful month is worth knowing. The bathroom scale analogy is right: used for trend detection, not precision measurement, these tools can do something useful.
The issue is the gap between how they are sold and what they have demonstrated. The same features (deep sleep duration, readiness score, VO2 max estimate, calorie burn) are presented with identical visual authority regardless of how different the underlying evidence is. A resting heart rate reading and a recovery score look the same in the app. They are not the same at all.
The bottom line
Your fitness tracker is likely accurate on your resting heart rate and approximately accurate about how long you slept. That’s it.
It is systematically inaccurate about your sleep stages, particularly deep sleep, in ways that physics rather than software sets the ceiling on.
Its calorie estimates can be off by 20% or more and in some conditions by considerably more than that.
Its readiness scores, recovery metrics, illness predictions, and other proprietary indices have not been independently validated against clinical outcomes in any form that medicine would recognize as evidence.
Only about 11% of consumer wearable devices have been validated for even one of the metrics they report. The rest is trust in the algorithm.
Used as a trend-tracking tool, a motivation device, and a rough monitor of the basics, a fitness tracker has genuine value. The problem is not what these devices measure. It is that the confidence with which they present every number, validated or not, is identical, and most users, reasonably, assume that a number appearing on a screen in that format was earned.
A bathroom scale that also claimed to measure your metabolic age, your immune readiness, and your cardiovascular fitness from the same morning weigh-in would be met with appropriate skepticism. We have not yet extended that skepticism to the wrist.
In the next post, I’ll take this question from the consumer into the clinical trials — where digital biomarkers have been deployed in over 1,000 drug trials over 25 years, and the regulatory record tells a considerably more complicated story than the industry’s confidence would suggest.
I left out two things worth flagging. First, VO2 max estimates (now standard on most devices) which have received even less independent validation and people use to make serious training decisions. And, second, does any of this actually change health outcomes? Owning a tracker and being healthier because of it are not the same claim, and the evidence on the second one is thinner than the industry suggests. Happy to go deeper on any of these, let me know in a comment.
If this kind of analysis, the kind that reads the study behind the headline and tells you what it actually found, is useful to you, consider subscribing. Brain Trials looks critically at the science of the brain: the trials that shape treatment, the studies that shape understanding, and the claims that deserve a closer read.
Thanks for reading Brain Trials! Subscribe for free to receive new posts.
Views expressed here are my own and not necessarily those of my employer.
References:
Doherty C, Baldwin M, Keogh A, Caulfield B, Argent R. Keeping Pace with Wearables: A Living Umbrella Review of Systematic Reviews Evaluating the Accuracy of Consumer Wearable Technologies in Health Measurement. Sports Med. 2024;54(11):2907-2926. doi: 10.1007/s40279-024-02077-2.
Chinoy ED, Cuellar JA, Huwa KE, Jameson JT, Watson CH, Bessman SC, Hirsch DA, Cooper AD, Drummond SPA, Markwald RR. Performance of seven consumer sleep-tracking devices compared with polysomnography. Sleep. 2021;44(5):zsaa291. doi: 10.1093/sleep/zsaa291.
Schyvens AM, Peters B, Van Oost NC, Aerts JM, Masci F, Neven A, Dirix H, Wets
G, Ross V, Verbraecken J. A performance validation of six commercial wrist-worn
wearable sleep-tracking devices for sleep stage scoring compared to
polysomnography. Sleep Adv. 2025;6(2):zpaf021. doi:
Searles ME, Licata A, Cucinotta M, Kainec K, Spencer RMC. Performance
evaluation of consumer sleep-tracking wearables and nearables in healthy young
and older adults. Sleep Adv. 2026;7(1):zpag006. doi:
Radin JM, Wineinger NE, Topol EJ, Steinhubl SR. Harnessing wearable device
data to improve state-level real-time surveillance of influenza-like illness in
the USA: a population-based study. Lancet Digit Health. 2020;2(2):e85-e93.
The WSJ wearable comparison (Nicole Nguyen, June 7, 2026) tested the Oura Ring 5, Google Fitbit Air, Whoop MG, and Apple Watch Series 11 against polysomnography at Stanford Health Care’s Sleep Medicine Center.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.