We always overestimate the change that will occur in the next two years and underestimate the change that will occur in the next ten.
— Bill Gates
When Tesla set out to build self-driving cars, they quickly realized that a PDF of traffic laws or a transcript of a driving lesson couldn’t train an AI to navigate a busy intersection. What unlocked autonomy was video—millions of hours of real drivers navigating real roads. Video showed the AI not just the rules, but the reality: when drivers hesitate, when they accelerate, and how they recover from mistakes.
That’s what made it possible for AI to move from memorizing rules to understanding the world in motion.
Healthcare is on the edge of a similar transformation.
Over the past decade, we’ve seen AI in medicine evolve along a familiar path:
Text 📄 — the foundation. Electronic health records, patient portals, and medical literature provided the raw material for algorithms to summarize, predict, and search.
Voice 🎙️ — the next frontier. Ambient clinical documentation, digital scribes, and triage lines gave AI the ability to listen, transcribe, and interact.
But soon we’ll be approaching the video era 📹 of AI in healthcare. Video is the richest signal we have. It combines what patients say, how they look, how they move, and the environment around them. It’s not just a transcript of a visit or a recording of words—it’s the lived reality of care.
And in the spirit of Gates’ reminder, we may be underestimating how transformative this leap will be over the next decade.
It’s tempting to think of AI + video as futuristic. But in many ways, it’s already here:
Telehealth visits 💻: Millions of virtual encounters now happen each year across primary and specialty care.
Tele-ICU programs 🏥: Critical care teams monitor patients remotely through continuous video streams.
Virtual nursing sitters 👩⚕️: Hospitals use remote video monitoring for patients at risk of falls, confusion, or self-harm.
Surgical video libraries 🔪: Archived procedures are being mined for training and quality improvement.
Virtual physical therapy 🏃♂️: Computer vision tracks patient movement, measures progress, and guides recovery—all without the therapist in the room.
These examples show that video is no longer peripheral—it’s becoming a primary input in how healthcare is delivered, monitored, and improved.
Video is not just “longer voice” or “moving text.” It’s fundamentally different in three ways:
Multi-modality. 🎛️ Video brings together words, tone, facial expression, posture, movement, and environment.
Function and behavior. Video reveals not just what patients say, but what they do.
Context. Video situates the patient in their environment—whether that’s the ICU, their living room, or a PT session at home.
For decades, healthcare has been limited to what can be documented in notes or heard in a brief interaction. Video offers a way to capture the fullness of the encounter.
Clinic visits 🩺: AI could analyze speech, facial cues, gait, and subtle changes to flag early disease.
Home health 🏡: Video check-ins could support rehab, detect falls, or assess wound healing.
Behavioral health 💬: Recognizing mood, affect, and cognitive decline through visual signals.
Physical therapy (expanded) 🏋️: Assess posture, muscle strength, and trajectory across sessions.
Emergency & urgent care 🚨: Live video feeds analyzed before a clinician enters the room.
Each of these opportunities could bring more objective, continuous, and actionable data into care.
Privacy and consent. 🔒 Patients must trust how video is used and stored.
Data infrastructure. Video is massive and requires significant resources to process.
Workflow integration. Insights must reduce burden, not add to it.
Ethics and boundaries. Guardrails are needed to prevent surveillance creep.
Healthcare has faced these challenges before with EHRs, telehealth, and genomics. But video adds a layer of intimacy that raises the stakes.
Cars couldn’t learn to drive from text alone. They needed video in the wild—millions of examples of humans navigating unpredictable environments.
Similarly, healthcare AI won’t fully mature until it learns from the rich, messy reality of how care happens.
And just as autonomous vehicles raised new questions about liability, trust, and safety, healthcare will need to wrestle with the same as video-based AI enters the mainstream.
Bill Gates’ quote is worth coming back to:
“We always overestimate the change that will occur in the next two years and underestimate the change that will occur in the next ten.”
Most leaders today might point to text summarization or voice transcription as the big drivers of healthcare AI. Those are real advances, but they’re just the beginning. The underestimated leap is what happens when video becomes the primary fuel.
Text gave AI knowledge 📄.
Voice gave AI interaction 🎙️.
Video can give AI understanding 📹—of function, behavior, and context.
The next ten years may feel incremental at first—better telehealth tools, smarter PT platforms, more efficient ICUs. But in retrospect, we may look back and realize that video was the bridge that moved AI in healthcare from reading and listening to seeing and understanding.
The opportunity is in how we shape it. Will video become just another tool of surveillance? Or will it be designed to enhance care, deepen connection, and extend access?
That choice, like the steering wheel 🚗, is still in our hands.
This post is not an endorsement or investing advice. It is personal opinions and does not reflect the views of my past, present or future employers, clients, or colleagues.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.