Part 2 of 2. Part 1 examined what consumer trackers actually measure, and what the peer-reviewed literature says about the claims on the box.
In Part 1, I looked at what happens when you take a consumer fitness tracker to a sleep lab. The verdict: resting heart rate holds up, total sleep duration is approximately right, deep sleep staging is systematically wrong, calorie estimation can be off by more than 20%, and only about 11% of commercially available wearable devices have been validated for even one of the metrics they report.
Now move from the consumer market to the clinical trials, where wearables have been doing their most serious work for 25 years, and where a precise regulatory distinction, consistently misunderstood in the trade press, explains almost everything about why the record looks the way it does.
When people in the digital health industry say a wearable-derived measure has been “accepted by the FDA,” they can mean three very different things, and the differences matter enormously.
The first and highest level is formal DDT qualification — Drug Development Tool qualification under Section 507 of the Federal Food, Drug, and Cosmetic Act. The FDA formally endorses a measure for a defined context of use, following a rigorous, multi-stage review of analytical validity, clinical validity, and context-specific interpretability. The payoff is significant: once a measure is formally qualified, any future drug sponsor can use it without re-justifying it from scratch. It becomes a shared evidentiary asset for the field.1
The second level is trial-specific endpoint acceptance. The FDA agrees, through a regulatory discussion with one sponsor, that a particular measure can serve as a primary or secondary endpoint for a specific trial. This is not qualification. It is a one-time agreement. It does not generalize to other trials, other sponsors, or other therapeutic contexts.
The third level is contributing to an approval as supporting evidence. A wearable-derived measure appears in the regulatory briefing package for a drug and contributes to the evidentiary case alongside primary endpoints. No formal status is conferred. The measure is useful in context. It does not become a reusable regulatory tool.
These three levels are often conflated in industry presentations, conference talks, and science journalism. Keeping them separate is not pedantry. It is the only way to understand what 25 years of wearables in clinical trials has actually produced.
A review published this month in Nature Reviews Drug Discovery catalogued 1,021 interventional drug trials registered between 2001 and 2025 that incorporated wearable-derived data.2 Adoption has accelerated sharply, from fewer than 30 new studies per year through 2014, to 128 new trials in 2024 alone. The infrastructure, the data volumes, and the industry investment are genuinely large.
Against that backdrop, the formal qualification record: as of 2024, not a single sensor-derived digital health technology has received formal DDT qualification through the FDA. This was confirmed in a review in Clinical Pharmacology & Therapeutics (Bakker et al., 2025).3 The FDA’s Biomarker Qualification Program has accepted 61 projects and formally qualified eight biomarkers across its entire history — none of them sensor-derived digital endpoints.
The single formal regulatory qualification in the wearable space belongs to the European Medicines Agency: SV95C, the 95th percentile of stride velocity derived from ankle-worn sensors, endorsed as a primary endpoint for ambulatory patients with Duchenne muscular dystrophy. It took almost a decade of rigorous analytical validation, longitudinal patient data, and careful stakeholder alignment. The FDA has cited it in review materials and it has been incorporated into pivotal US-based gene therapy trials. It is a genuine milestone. It is also the only one, in 25 years, from either agency.
1,021 trials. One formal qualification, from the EMA, in a rare pediatric disease. Zero through the FDA.
The continuous glucose monitor story is the closest analog to regulatory success on the US side, and it illustrates the distinction precisely.
Time in range — TIR, the percentage of time glucose stays within 70 to 180 mg/dL — is explicitly cited in FDA guidance as a valid surrogate endpoint and has been incorporated into regulatory submissions and product labeling. It functions as a de facto accepted endpoint in diabetes trials. But it has not been formally qualified through the DDT program. It occupies a space between trial-specific acceptance and formal qualification — close enough to regulatory legitimacy to be used, not formally enshrined as a reusable tool.
Why has TIR gotten this far when almost nothing else has? Because glucose is the disease. The CGM measures the thing that is being treated. The signal and the clinical outcome are the same entity, captured continuously rather than periodically. There is no inferential gap between what the sensor detects and what the patient experiences as meaningful.
That alignment is rarer than it sounds, and its absence explains most of the failures.
The Bellerophon case is the most instructive recent example of the second level — trial-specific acceptance — and why it is not the same as formal qualification.
In Bellerophon’s Phase 2b trial of pulsed inhaled nitric oxide for pulmonary hypertension, ActiGraph devices measured moderate-to-vigorous physical activity alongside conventional endpoints. Conventional measures showed little change. Actigraphy showed something real: patients on treatment maintained their activity levels while the placebo group declined, yielding a placebo-corrected improvement of roughly twelve minutes per day. On that basis, the FDA endorsed MVPA as the primary endpoint for the pivotal Phase III REBUILD trial — the first time a wearable-derived activity measure had been accepted as the primary endpoint in a pivotal cardiopulmonary study.
The REBUILD trial stopped for futility. No benefit over placebo. The program was discontinued.
The lesson is not that the FDA was wrong to engage, or that wearable-derived endpoints cannot work. The lesson is more obvious: trial-specific acceptance of an endpoint does not validate whether that endpoint reliably tracks drug benefit. Formal qualification exists precisely because of cases like this — because a measure that works in one Phase 2 context may not generalize, and the only way to establish generalizability is through the full evidentiary pathway.
The nemolizumab story from atopic dermatitis trials is the one I find most instructive for how wearables should be designed — but it is important to be precise about what it was and wasn’t.
In trials of nemolizumab (a monoclonal antibody targeting IL-31) researchers strapped wrist actigraphs to patients at night and measured scratching. A statistically significant reduction in the scratch-to-sleep duration ratio of 32 minutes per hour versus a 28-minute-per-hour increase in the placebo group appeared by week one. The measure was carried into pivotal Phase III trials. Actigraphy data was included in regulatory briefing packages. It contributed, as supporting evidence, to nemolizumab’s approval in 2024.
This was the third level: wearable-derived data contributing to an approval alongside primary endpoints. The actigraphy scratching measure was not formally qualified. It did not receive trial-specific acceptance as a primary endpoint. It was well-designed supportive evidence, genuinely useful, and something other sponsors would still need to re-justify in their own submissions.
Why is it still instructive? Because it was designed correctly. Researchers started from what matters to the patient: nocturnal scratching is something a patient with atopic dermatitis knows is happening, measures subjectively, and uses as personal evidence of whether their disease is better or worse — and worked backward to a sensor that could capture it objectively. The digital measure and the clinical reality pointed at the same thing from the start. The path from sensor signal to patient experience required no inference. There was no gap to cross.
This is the design principle that most wearable work in neuroscience has missed.
In neurodegeneration, Verily provides the cautionary tale.
Verily, Google’s life sciences division, built the Virtual Motor Exam for Parkinson’s Disease using the Verily Study Watch. The device digitized elements of the MDS-UPDRS Part III, the standard neurologist-administered motor assessment, capturing finger-tapping, tremor, and gait signals continuously in the real world rather than periodically in a clinic. They followed thousands of patients with Parkinson’s for many years. One of the largest longitudinal studies ever conducted in this disease. The resources behind this effort were essentially unlimited. The regulatory experience Verily brought to it was substantial.
The FDA declined to accept the Letter of Intent for DDT qualification.
The rejection letter (DDT COA, publicly available) makes the reasoning explicit.4 The MDS-UPDRS Part III — and by extension the Verily algorithm attempting to digitize it — measures things that are not directly meaningful to how a patient functions in daily life. A change in rigidity or finger-tapping speed, however precisely captured, cannot be straightforwardly interpreted as a meaningful change in what a patient can actually do. The FDA’s preference was explicitly for measures aligned with MDS-UPDRS Part II: speech, eating, dressing, getting around the house.The things patients experience as the disease taking or returning something from them.
Finger tapping speed, measured with a precision no neurologist could match, is not the same as being able to button a shirt. Ridigity, assessed by a neurologist, is not the same as being able to go to the grocery store. The FDA said so, directly and on the record.
Read that twice, because it is the key to understanding why 25 years of digital biomarker research in neurodegeneration has produced so little regulatory traction. The field’s response to this consistent position has been to build more sensitive instruments. The FDA’s position is that more sensitivity produces precise measurements of things that don’t matter to patients. These two trajectories have been running in parallel for two decades.
There is a catch-22 at the center of digital biomarker development for neurodegeneration that the field discusses in private more often than it publishes.
Specifically for Parkinson’s, the premise is: “wearables detect motor changes too subtle for the human eye to catch; subtle early signals might predict clinical decline; earlier detection enables earlier intervention; therefore wearables capturing subtle signals are valuable clinical endpoints”.
But the FDA’s qualification pathway requires a measure to reflect something the patient experiences as meaningful change in their functioning. And if a change is too subtle for a trained clinician to observe, and not perceptible to the patient as affecting their daily life, it won’t meet that standard.
If the signal is detectable only by algorithm, and not by neurologist or patient, it is, by the FDA’s framework, not evidence of clinical benefit.
This is not a regulatory failure. It is a conceptual one. The two things the field wants to be simultaneously true, that wearables detect signals too subtle for humans, and that those signals are clinically meaningful, are in tension with each other at the level of regulatory logic. The FDA’s consistency on this point, across multiple submissions and multiple years, suggests it is not going to resolve in the field’s favor without a fundamental reorientation of what is being measured and why.
EEG wearables — the form factor most directly relevant to neurological disease — appear in exactly two proof-of-concept studies in the entire 1,021-trial dataset. Two, in 25 years. The most direct window into brain function that wearable technology currently offers has essentially not been tried in drug trials. This is particularly striking because the FDA already accepts EEG-derived endpoints as primary outcomes in sleep medicine: the Maintenance of Wakefulness Test and polysomnography-based measures like latency to persistent sleep have served as primary or co-primary endpoints in narcolepsy and insomnia approvals for decades. The FDA demands it. The regulatory precedent for brain-electrical-activity-based endpoints exists.
Wearable EEG for tracking sleep architecture, sleep continuity, and related endpoints in neurological disease trials is, in my view, the lowest-hanging fruit the field has not picked.
Instead, neuroscience has organized around motor signals in Parkinson’s disease and cognitive signals in Alzheimer’s disease.
To be fair, some of what’s being measured is patient-relevant. Falls are patient-relevant — patients, families, and caregivers know when a fall happens, and the event is objectively detectable and clinically meaningful. Freezing of gait is patient-relevant. Speech changes in Parkinson’s disease are patient-relevant — dysarthria is something patients and families consistently describe as among the most distressing symptoms, and it is tractable by acoustic sensors. These endpoints start from what the patient experiences. They have not been prioritized because they don’t produce the clean signals that look like good data.
In Alzheimer’s disease, the task is harder because patient self-report degrades as the disease progresses. The largest multi-technology validation study in AD to date — RADAR-AD, a European consortium involving Janssen, Novartis, and 6 academic centers — tested 6 remote monitoring technologies across 237 amyloid-confirmed participants (Fröhlich et al., Alzheimer’s Research & Therapy, 2025).5 The wearable-sensor-based machine learning models identified prodromal AD with an AUC of 73% — modestly above chance and well below what a 10-minute paper cognitive screening test achieves routinely. The single best-performing technology in the study was not a sensor at all. It was the Amsterdam iADL questionnaire — a digital version of the kind of functional assessment that asks patients and caregivers whether daily tasks like managing finances or preparing meals have become harder. The most useful remote monitoring technology in the largest AD wearable study turned out to be asking the right questions, not deploying sensors.
But daily activity rhythms, sleep continuity, mobility patterns within the home, and naturalistic conversational speech are all sensor-accessible, genuinely meaningful to families and caregivers, and worth the validation work to demonstrate if they have real value. None of them requires novel technology. All of them require a different starting question.
The nocturnal scratching study from dermatology did not require a Verily-scale massive investment. It required clever researchers willing to start from what the patient already knew mattered and build backward to the sensor. The sophistication was in the question, not the instrument (which had existed for decades).
Twenty-five years. 1,021 drug trials. One formal regulatory qualification in the wearable space — from the EMA, in a rare neuromuscular pediatric disease — and zero through the FDA for sensor-derived digital endpoints.
The trial-specific acceptances are real regulatory engagements that reflect genuine FDA openness. They are not formally qualified endpoints. The field sometimes presents them as though they were, which inflates the apparent progress and obscures the gap between what has been agreed to in a single trial and what has been validated as a reusable clinical tool.
The digital biomarker field has tended to interpret this record as a sign that the regulatory pathway is slow and the evidentiary bar unreasonably high. Both things may contain some truth. But the Verily letter is public, the FDA’s reasoning is clear, and the pattern across therapeutic areas is consistent: measures developed by starting from what a sensor can detect, and hoping the signal turns out to be clinically relevant, have not been meeting the standard. Measures designed by starting from what the patient experiences as meaningful, and building backward to the sensor, have a better record — even if that record is still modest, and even if the formal qualification count remains at zero.
The instruments are more sensitive than they have ever been. Sensitivity pointed at the wrong question is the most expensive way to generate inconclusive data.
After twenty-five years, the question the field needs to answer is not how to make the sensors better. It is whether it is measuring the right things in the first place.
I left out the question of whether Neuron23’s digital biomarker primary endpoint in the NEULARK LRRK2 trial (which I covered in the LUMA piece) will face the same regulatory logic — a sensor-derived motor measure as a primary endpoint in a disease where the FDA has already told one sponsor that motor digitization doesn’t meet the “meaningful to the patient” standard. If there’s interest, I’ll connect those dots in a future piece. Also, let me know if there is any digital biomarker that should get more attention.
If this kind of analysis is useful to you — the kind that reads the FDA rejection letters, not just the press releases — consider subscribing. Brain Trials covers neuroscience through the lens of evidence, mechanism, and what the regulatory record actually says when you look at the primary sources.
The gap between what a technology measures and what a patient experiences is where most digital health promises go to die. Knowing how to see that gap is what clinical trial literacy is about. I wrote a book about that.
A Patient’s Guide to Clinical Trials: Navigating the Promise and Pitfalls of Experimental Treatments (Bloomsbury) — available now.
This analysis is based entirely on publicly available information, including published regulatory documents, and represents my personal views, not necessarily those of my employer.
References
Fayad et al. Wearable Technologies in Clinical Trials for Drug Development: Trends and Emerging Opportunities. Nature Reviews Drug Discovery. June 2026.
Bakker et al. Regulatory Pathways for Qualification and Acceptance of Digital Health Technology-Derived Clinical Trial Endpoints. Clinical Pharmacology & Therapeutics. 2025.
FDA Drug Development Tool Letter of Intent Determination, DDT COA (Verily Study Watch, Parkinson’s disease).
Nemolizumab atopic dermatitis pivotal trials: NCT03181503 and NCT03989349. FDA approval 2024.
Bellerophon REBUILD trial: NCT03267108. Stopped for futility.
EMA Qualification Opinion on SV95C (stride velocity) as primary endpoint in Duchenne muscular dystrophy: EMA/CHMP/SAWP/178058/2019.
Fröhlich H, Vairavan S, Muurling M, et al. RADAR-AD: assessment of multiple remote monitoring technologies for early detection of Alzheimer’s disease. Alzheimer’s Research & Therapy. 2025;17:29. doi:10.1186/s13195-025-01675-0
The DDT qualification pathway is codified under Section 507 of the Federal Food, Drug, and Cosmetic Act (21st Century Cures Act). The full list of qualified biomarkers is publicly available on the FDA’s Biomarker Qualification Program website. As of 2024, eight biomarkers have been formally qualified — all traditional (blood-based, imaging-based, or physiological), none sensor-derived digital.
Fayad et al., “Wearable Technologies in Clinical Trials for Drug Development: Trends and Emerging Opportunities,” Nature Reviews Drug Discovery, June 2026. The review catalogued 1,021 interventional drug trials registered on ClinicalTrials.gov between 2001 and 2025 incorporating wearable-derived data. Adoption grew from <30 new trials/year through 2014 to 128 in 2024. The most common therapeutic areas were cardiopulmonary, metabolic, and neurological disease.
Bakker et al., “Regulatory Pathways for Qualification and Acceptance of Digital Health Technology-Derived Clinical Trial Endpoints,” Clinical Pharmacology & Therapeutics, 2025. This review confirmed zero formal DDT qualifications for sensor-derived digital endpoints through the FDA as of 2024.
FDA Drug Development Tool Letter of Intent Determination, DDT COA
#000142, Verily Study Watch Virtual Motor Exam for Parkinson’s Disease. Publicly available through the FDA’s DDT Qualification Program. The rejection explicitly cited the disconnect between MDS-UPDRS Part III motor assessments (what the device digitizes) and patient-meaningful functional outcomes aligned with MDS-UPDRS Part II (what the FDA considers clinically relevant for qualification).Fröhlich H, Vairavan S, Muurling M, et al. RADAR-AD: assessment of multiple remote monitoring technologies for early detection of Alzheimer’s disease. Alzheimer’s Research & Therapy. 2025;17:29. doi:10.1186/s13195-025-01675-0. 237 participants across four diagnostic groups (healthy controls, preclinical AD, prodromal AD, mild-to-moderate AD), all AD participants amyloid-confirmed. Six remote monitoring technologies. ML-based prodromal AD detection: AUC 73%. The Amsterdam iADL questionnaire — a digital functional assessment, not a sensor — outperformed wearable-derived features.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.