For two years the default mental model for healthcare AI has been a large model in someone else’s cloud. That assumption is quietly breaking. Capable open-weight models now run on hardware a clinic can buy off the shelf, and the economics have inverted: a one-time ~$2,000 machine can serve a local model at ~100 tokens/second with no per-token bill and no patient data leaving the building. In this roundtable episode of Signal & Symptoms, Dr. Junaid Kalia — with co-hosts Ed Marx and Dr. Harvey Castro — argues that “where the model runs” (edge, near-edge, or cloud) is now a clinical and privacy decision, not merely an infrastructure one.
The upside is real: data sovereignty, offline resilience, and cost control, with obvious relevance to rural and low-connectivity care. So is the catch. The most capable cheap open models today largely originate in China (DeepSeek, Qwen, Kimi K2), which turns a technical choice into a question of standards, security, and national competitiveness — and collides with a regulatory reality: a neural network is a black box that resists the itemized Software Bill of Materials (SBOM) the FDA now requires, while prompt injection remains the #1 documented LLM security risk. This briefing maps the edge-AI market, the open-model economics, the security and regulatory friction, and what health-system leaders should ask before running a model themselves.
Edge AI — running models on local devices rather than a central cloud — is scaling fast. The broader edge-AI market was estimated at roughly $20.8B in 2024, growing ~21.7% CAGR through 2030, with some forecasts putting it near $105B by 2030 (≈27.6% CAGR) [1]. In healthcare specifically, edge computing was ~$8.16B in 2025, projected to ~$23.2B by 2031 (≈19% CAGR), with healthcare & life sciences among the fastest-growing segments [2].
Demand isn’t only institutional. West Health-Gallup research (survey of 5,500+ U.S. adults, late 2025) found roughly 1 in 4 adults — ~66 million people — have used AI for health information or advice, most often to supplement care before or after a visit [3]. On the show, Ed Marx cited office visits trending down (a figure he attributed to Gallup/West Health) as patients increasingly self-serve with AI [3]. The pull toward AI-at-the-point-of-need is already here; the open question is where that intelligence physically lives.
Open-weight model makers. The “onslaught” the hosts describe is real: capable, freely downloadable model families now span tiers — small router/tool-calling models, mid-size workhorses (e.g., Qwen), and frontier-class releases (e.g., Kimi K2 from Moonshot AI). DeepSeek anchors the low-cost end. (Model quality comparisons cited on-air — e.g., “beats Gemini Flash” — reflect the hosts’ own testing, not standardized benchmarks.)
Silicon & local runtimes. Apple Silicon with MLX and Core ML is a serious local-inference platform; Junaid reported models running ~30% faster on the same GPU after Core ML optimization (his own measurement). NVIDIA’s CUDA/vLLM stack remains the cloud/data-center default.
The state as a player. Governments are now direct participants: in August 2025 the U.S. took a ~10% equity stake in Intel ($8.9B), converting CHIPS-Act grants into ownership — a signal that the compute layer is being treated as national infrastructure [4].
The regulator. The FDA is the gatekeeper for any model used as a device — and its cybersecurity expectations (SBOM, Section 524B) are now mandatory, not aspirational [6].
Cost structure flips from rent to own. Cloud inference is a recurring per-token cost; local inference is a capital cost amortized across unlimited use. Junaid’s framing — a ~$2,000 Mac at ~80–100 tokens/sec versus an ongoing Azure bill — is the core argument. At the low-cost end, DeepSeek’s API pricing runs a fraction of frontier U.S. models — on the order of tens of cents per million tokens vs. several dollars [5], which is also what makes “give it away, own the standard” a viable strategy.
Data-exfiltration cost avoided. Keeping inference local removes an entire class of privacy and compliance exposure — no PHI transiting a third-party cloud.
Resilience ROI. On-device models keep working with bad or no connectivity (the “on a plane / in a rural clinic” case), turning AI from a connectivity-dependent service into a reliable local tool.
The provenance problem (the “catch”). The best cheap open models are largely Chinese; the hosts were explicit that this creates a security risk that makes them unsuitable for production clinical use as-is. Model origin is now part of the risk assessment.
The FDA black box. For a regulated device, manufacturers must submit a Software Bill of Materials listing every software component — mandatory for “cyber devices” under FD&C Act Section 524B since October 2023 [6]. A neural network’s learned weights resist that kind of itemization: you can list the libraries, but not fully account for what a model “knows” or was trained on.
Prompt injection. It is the #1 risk in the OWASP Top 10 for LLM Applications and, by design, cannot be fully patched — LLMs process instructions and data in the same channel [8]. Layered injections are hard to detect and harder to fix, exactly as Junaid described.
Standards capture. If a single cheap model becomes the default substrate, whoever controls it controls the standard — the Jevons-paradox dynamic Harvey raised (cheaper → more demand → lock-in).
The crisis is a squeeze: patients are racing ahead to AI while the average clinical encounter shrinks. U.S. primary-care visits average roughly 18 minutes, with about 1 in 4 under 12 minutes [9] — the “13 minutes” the hosts cited sits squarely in that reality. Meanwhile ~66 million adults are already consulting AI for health, sometimes instead of a visit [3]. Care can’t stretch time; it can only add capacity.
Local AI is one credible pressure-release valve — if it’s deployed with discipline. The episode’s implicit playbook: match the model to the task (don’t send a limousine across the street); keep inference local where privacy or connectivity demands it; treat model provenance as a security input, not a footnote; and keep a clinician in the loop rather than surrendering judgment to a black box. Capability without provenance and oversight isn’t a solution — it’s a new liability.
CIOs / CMIOs: Decide where each AI workload should run (edge / near-edge / private cloud) as a deliberate policy, not by default. Inventory model provenance the way you inventory any supply chain, and require an SBOM posture for anything approaching a regulated use [6].
Clinician leaders: Pilot local models on non-PHI, low-risk tasks first. Treat “the model told me” with the same skepticism as any unvalidated source — prompt injection makes blind trust dangerous [8].
Investors / builders: The durable value is in the governance and fine-tuning layer — provenance, validation, and bring-your-own-data infrastructure — not in the raw model, which is trending toward free [5].
Policymakers: The competition is now national (the Intel stake is the tell [4]). Fund and standardize trustworthy, auditable domestic models before a cheaper foreign standard becomes the default substrate for clinical AI.
The endpoint the hosts point to isn’t cloud or local — it’s the clinician being able to choose, deliberately, for each task and each patient. A model small enough to run on a phone handles the quick question offline; a private on-prem cluster runs the heavier reasoning where PHI can never leave; the cloud is used when, and only when, it’s the right tool. In that world the most advanced clinic isn’t necessarily the one in the biggest city — it might be the rural community that leapfrogged straight to self-contained AI because it had to. The lesson traveling from a $2,000 Mac to a clinic with no road is the same one that keeps recurring on this show: the technology is finally cheap and capable enough that the real questions are governance ones — who owns the model, where the data lives, and who stays accountable for the decision.
[1] Grand View Research / Precedence Research — Edge AI Market (~$20.8B in 2024, ~21.7% CAGR; alt. forecast ~$105B by 2030, ~27.6% CAGR) — https://www.grandviewresearch.com/industry-analysis/edge-ai-market-report · https://www.precedenceresearch.com/edge-ai-market
[2] Mordor Intelligence — Edge Computing in Healthcare Market ($8.16B 2025 → $23.22B 2031, ~19% CAGR) — https://www.mordorintelligence.com/industry-reports/edge-computing-in-healthcare-market
[3] West Health–Gallup — Millions of Americans Now Consult AI Before, After, and Sometimes Instead of Seeing a Doctor (≈1 in 4 adults / ~66M; survey Oct–Dec 2025) — https://westhealth.org/news/millions-of-americans-now-consult-ai-before-after-and-sometimes-instead-of-seeing-a-doctor/ · https://news.gallup.com/poll/707789/americans-turning-supplement-healthcare-visits.aspx
[4] CNBC / NPR — U.S. government takes ~10% stake in Intel ($8.9B), Aug 2025 — https://www.cnbc.com/2025/08/22/intel-goverment-equity-stake.html · https://www.npr.org/2025/08/22/nx-s1-5509673/trump-says-us-government-will-take-stake-intel
[5] CloudZero — DeepSeek Pricing (per-token cost a small fraction of frontier U.S. models) — https://www.cloudzero.com/blog/deepseek-pricing/
[6] FDA / Emergo by UL — Final Guidance on Medical Device Cybersecurity; SBOM mandatory for “cyber devices” under FD&C Act §524B (since Oct 1, 2023) — https://www.emergobyul.com/news/fda-releases-final-guidance-medical-device-cybersecurity
[7] Texas Farm Bureau / Texas Land Trends — ~85% of Texas land is managed by those who farm and ranch; rural land loss to metro growth — https://texasfarmbureau.org/texas-land-trends-tracks-changing-state/
[8] OWASP Gen AI Security Project — LLM01: Prompt Injection — #1 risk in the OWASP Top 10 for LLM Applications — https://genai.owasp.org/llmrisk/llm01-prompt-injection/
[9] Healio / HealthDay (Neprash et al., analysis of 21M+ visits) — Average U.S. primary-care exam ~18 minutes; ~1 in 4 under 12 minutes — https://www.healio.com/news/primary-care/20210121/average-primary-care-exam-lasts-less-than-20-minutes

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.