RSS Amplifier

Faces of Digital Health Newsletter · Jan 22, 2026

Can 45 Seconds of Voice Really Detect Disease?

0
Sign in to vote or save

Tjaša Zajc · Faces of Digital Health Newsletter

This is a newsletter of Faces of Digital Health - a podcast that explores the diversity of healthcare systems and healthcare innovation worldwide. Interviews with policymakers, entrepreneurs, and clinicians provide the listeners with insights into market specifics, go-to-market strategies, barriers to success, characteristics of different healthcare systems, challenges of healthcare systems, and access to healthcare. Find out more on the website, tune in on Spotify or iTunes.

Let’s connect on Linkedin!

Ambient speech tech has exploded in healthcare but mostly as “documentation glue”, helping clinicians with note-taking through speech recognition. But what if the real opportunity isn’t the transcript at all?

I recently spoke with Henry O’Connell, Founder and CEO of Canary Speech. Frankly, because I saw the company claims their tech can diagnose conditions such as Alzheimer’s disease from nothing more than a 45 seconds long discussion. Too good to be true? Not necessarily.

As Henry argues, the last 25 years of “voice biomarker” progress stalled for a simple reason: the field focused on what people say (words) instead of how the nervous system produces speech (signal).

O’Connell traces the roots of the space through decades of speech and language work, from early NLP to Nuance and far-field speech systems. For a long time now, the industry could find correlations by analyzing spoken words and their textual representation. But it never produced widely adopted clinical products. The limitation, says O’Connel, is data density: word-based analysis gives you a few hundred-ish elements; analysing voice characteristics however, yield millions.

Share

Canary Speech’s bet was to turn the industry paradigm upside down: ignore words and instead analyze the articulatory system; how the central nervous system coordinates breath, vocal cords, tongue, and timing to generate language. It’s the same human machinery regardless of language, and it’s the layer clinicians often “read” intuitively like when you can hear someone isn’t okay even if they say “I’m fine.”

Surprisingly little audio is needed for diagnosis: 45 seconds of conversational speech captured during a clinician-patient exchange. Canary extracts 2,590 voice features every 10 milliseconds (with a 25ms window sliding by 10ms), which adds up to roughly 13 million data elements for the model to evaluate.

Those features include characteristics tied to vocal cord vibration (e.g., jitter/shimmer), respiratory sounds, resonance, and how those properties change over time; plus derivatives (how fast features change and recover). Many of these features were already known in speech science. Canary focuses on stacking multiple extraction approaches and validating them clinically.

Canary partners with clinical teams (neurology, psychiatry, specialty services), captures patient audio during clinical interviews in medical practice, and uses clinician diagnoses and test batteries as labels. This is then the ground truth. Machine learning then learns correlations between the diagnostic labels and the extracted features.

The result is an additional insight for the clinician. Tech is meant as clinical decision support, not a standalone diagnostics tool.

Progressive neurological diseases like Huntington’s, Parkinson’s, and Alzheimer’s can be detected with 98%+ accuracy, says O’Connell. Behavioral health conditions such as anxiety or depression are more nuanced. The accuracy there is “in the 80s.” A central point: the system can measure multiple dimensions at once e.g., cognitive decline plus depression, fatigue, and anxiety without extra time burden on the clinician. Still, an insight can prompt the clinician to ask questions they otherwise wouldn’t.
One example O’Connell shares: “A doctor is meeting with a young mother six weeks after her birth. The mother is being very much like every mother on the planet Earth is often. Focused on her child, focused on the responsibility of caring for that child, feeling that obligation, a good one, and a responsibility to do so. The doctor says to her, how are you doing? Her focus is all over here and she says, I’m doing fine. Canary’s speeches assessment came back at virtually the same time and said her depression levels were clinically as high as they could get. So the doctor said: ‘Why don’t we explore that a little bit more?’ And by the time they had a few more questions in, the woman related how deeply depressed she was.“

If the results can be strong, why hasn’t this type of tech become routine? O’Connell points to earlier research norms that relied on structured read speech (scripts) borrowed from speech pathology. Since, as claims O’Connell, the brain regions involved in reading vs conversing differ, script-based studies can be counterproductive for building conversational biomarkers.

Canary avoids read speech entirely.

In addition, AI and infrastructure shifts, such as real-time streaming, fast compute, better transcription and ambient tooling are accelerating development outside of text analysis. Canary’s integration with ambient documentation workflows (e.g., a clinical copilot listening anyway) turns voice biomarkers into “no extra workflow” intelligence.

Share

What if a clip of a conversation reveals more than an individual would like to know? Similarly as with genetics tests, where, clinicians can face a dilemma of finding and sharing information outside of the scope of the initial diagnostic purpose. Canary’s framing is that voice biomarkers describe what’s present in the moment, not a future risk. That matters in primary care screening, where objective signals can prompt better follow-up without relying on subjective questionnaires alone.

The most unexpected use case for voice biomarkers may be in-room monitoring for aggression risk. Unfortunately, in U.S. hospitals, violence against nurses is a serious operational and safety issue. Canary can contribute to minimizing incidents by indicating a specific room has aggressive individuals, enabling teams to enter these environment with more caution (for example, two people enter, not a nurse alone) and prepared to de-escalate conflicts.

For pharma and trials, Canary can support:

  • Pre-screening participants (e.g., who is or isn’t demonstrating mild cognitive impairment)

  • Longitudinal tracking to detect changes over time (e.g., whether cognitive function improves with a therapy)

  • More direct outcome measures than proxies like plaque reduction alone

The company is planning to expand the number of diagnoses it covers, adding models for conditions such as PTSD, postpartum depression, ADHD and autism, pain, and COPD. It also plans to expand into additional languages, ensuring each one is properly validated before launch.

Canary is already deploying in multiple regions (U.S., Canada, Japan, UK/Ireland, parts of Europe, and expanding in South America and Arabic-speaking markets). They plan on deploying a wellness product in 2026, which will still be built on clinically developed models rather than scraping uncontrolled public sources.

Bottom line: “Voice in healthcare” isn’t only about transcription. If clinical validation scales, voice biomarkers could become a quiet layer of decision support: always on, non-intrusive, and potentially transformative for early detection, behavioral health, and even workforce safety.

Leave a comment

Share

00:12 Intro: voice biomarkers beyond ambient documentation
01:03 Potential vs reality: what voice can (and can’t yet) prove clinically
01:54 Why earlier voice-biomarker work focused on words—and why it stalled
06:34 The “intuition” problem: we hear mood without words
07:33 Canary’s clinical-first approach and global clinical partnerships
08:39 Do you need long-term data? (Agatha Christie example)
09:41 Method: ~45 seconds of conversational speech, ambient capture
10:36 Scale: 2,590 features every 10ms (~13M data elements)
11:02 Where the features come from (vocal cords, respiration, resonance)
14:36 How models are built: IRB, clinician ground truth, ML correlation
17:30 Accuracy + adoption: why it’s not standard practice yet
18:22 Reported performance: 98%+ neuro, ~80s behavioral health
22:44 Culture/language bias concern: why validation per language matters
24:22 Guardrails: validate every new language and population testing
31:07 Why read-speech scripts misled the field; conversational-only stance
34:16 What changed: AI, compute, real-time streaming, workflow fit
37:10 How it’s used: screening vs suspected disease; incidental findings
39:24 Primary care example: postpartum depression flagged despite “I’m fine”
47:56 Wellness/employee use: de-identified dashboards and ethics
53:42 Technical requirements: device capture, signal-to-noise checks
57:10 In-room monitoring: aggression risk signals for staff safety
1:02:22 Clinical trials: pre-screening + measuring therapeutic impact
1:07:14 Global rollout: regions, languages, and partnerships
1:09:42 Consumer access: wellness product planned in 2026
1:12:32 Wrap-up: why this matters as cognitive health needs grow

No posts

Read the original on fodh.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.