In Week #259 of the Doctor Penguin newsletter, the following papers caught our attention:
1. Biological Clock. Aging research has primarily focused on adult aging clocks, leaving a gap in understanding biological clocks across the full life cycle.
Wang et al. developed LifeClock, a biological clock that predicts biological age across the entire human lifespan using routine electronic health records (EHRs) from 24.6 million clinical visits across 9.7 million individuals. The system employs EHRFormer, a transformer-based model that converts longitudinal EHR data (184 features including lab tests, vital signs, and metadata) into comprehensive patient representations. EHRFormer is trained via multitask pretraining with five complementary objectives: (1) Masked language modeling for current-visit reconstruction and next-visit prediction, (2) Variational autoencoder-style reconstruction (ELBO), (3) Adversarial cohort discrimination to remove batch effects, (4) Adversarial missing-data discrimination, and (5) Age regression using clinical measurements only (age metadata removed), trained exclusively on healthy individuals to establish normal aging baselines. The research reveals two distinct clocks: a pediatric development clock (under 18 years) driven by aspartate aminotransferase, creatinine, and total protein that predicts growth abnormalities and malnutrition risk, and an adult aging clock (over 18 years) driven by urea, albumin, and red cell distribution width that predicts diabetes, cardiovascular disease, and stroke risk. By leveraging cost-effective routine clinical data rather than expensive multi-omics approaches, LifeClock provides a scalable framework for precision health across the full life cycle.
Read paper | Nature Medicine
2. Biological Clock. AI-driven brain MRI aging clocks have been widely used as clinical biomarkers of neurological aging and cognitive decline. However, this imaging-based aging framework has not been comprehensively applied across multiple organ systems.
The MULTI Consortium applied machine learning to develop seven MRI-based biological aging clocks for multiple organs (brain, heart, liver, adipose tissue, spleen, kidney, and pancreas) using data from over 313,000 individuals. Through genome-wide association studies (GWASs) and post-GWAS analyses, these aging clocks are linked to genetics, plasma proteins, and metabolites, identifying 53 genetic loci and nine druggable genes as potential targets for anti-aging treatments. Notably, the seven clocks predict future disease risk (particularly diabetes) and all-cause mortality, with some counterintuitive findings: higher liver and spleen age appear protective against mortality, possibly reflecting better organ resilience. In an Alzheimer’s disease drug trial, participants with more youthful versus aged brain profiles showed distinct cognitive decline trajectories over 240 weeks, though this heterogeneity couldn’t be fully attributed to the drug itself. The study demonstrates both organ-specific and cross-organ aging interconnections, expanding the multi-organ biological aging framework.
Read Paper | Nature Medicine
3. Digital Intervention. Inflammatory rheumatic diseases cause chronic joint pain, swelling, and stiffness. Beyond physical symptoms, these patients face high rates of depression (15-24%) and anxiety (19-37%). Despite guidelines recommending mental health support, accessing therapy remains difficult due to provider shortages and long wait times, leaving many patients to cope with psychological distress on their own.
Knitza et al. conducted a randomized clinical trial that evaluated a digital psychological intervention for 102 patients with inflammatory rheumatic diseases across Germany, demonstrating significant improvements in both psychological distress and quality of life. Participants receiving the self-guided, web-based cognitive behavioral therapy intervention showed clinically meaningful reductions in psychological distress and improvements in quality of life compared to treatment-as-usual controls, with nearly 60% of intervention participants achieving clinically significant improvement versus only 34% of controls. The intervention also produced substantial benefits in secondary outcomes, including depression, anxiety, perceived stress, self-efficacy, and health literacy, with effect sizes ranging from medium to large. No intervention-related adverse events were reported, and the most common negative effect was the resurfacing of distressing memories—a recognized part of cognitive restructuring in therapy. These findings suggest that scalable digital psychological interventions could effectively address the significant treatment gap for patients with inflammatory rheumatic diseases who experience depression and anxiety, offering an accessible complement to standard rheumatological care that addresses both the physical and psychological dimensions of chronic inflammatory conditions.
Read Paper | JAMA Network Open
4. LLM Phenotyping. Clinicians rely heavily on presenting signs and symptoms to guide initial antibiotic decisions in patients with suspected sepsis. Yet most large observational studies examining antibiotics omit this critical information because extracting signs and symptoms from clinical notes at scale has been impractical. Can large language models (LLMs) address this gap by accurately extracting presenting signs and symptoms from unstructured clinical text?
Pak et al. developed an LLM-based system to extract presenting signs and symptoms from admission notes of 104,248 patients with possible infection across five Massachusetts hospitals. Using LLaMA 3 8B with a straightforward prompt, the system labeled admission notes with a controlled vocabulary of 404 signs and symptoms, achieving accuracy comparable to physician chart reviewers. The 30 most common signs and symptoms clustered into seven clinically recognizable syndromes corresponding to infection sources: skin and soft tissue, cardiopulmonary, gastrointestinal, urinary tract, plus three nonspecific clusters (dizziness, back pain, constitutional). These syndromes revealed distinct risk profiles for antibiotic-resistant pathogens and mortality. For example, patients presenting with skin and soft tissue symptoms were more likely to have methicillin-resistant Staphylococcus aureus (MRSA) infections but less likely to have multidrug-resistant gram-negative bacteria, while urinary tract symptoms showed the inverse pattern. By demonstrating that symptom patterns identify patient subgroups with markedly different pathogen susceptibilities and prognoses, this study suggests opportunities to move beyond uniform broad-spectrum antibiotic protocols toward more personalized empiric treatment strategies that match antibiotic coverage to presenting clinical syndromes.
Read Paper | JAMA Network Open
5. Evaluation. How should we evaluate AI that may soon exceed human expert performance?
This editorial by Gallifant and Bitterman argues that rapidly advancing medical AI systems demand a fundamental shift in evaluation methods beyond traditional question-and-answer benchmarks. The authors highlight that current AI systems can already exceed medical trainee performance on narrow tasks, yet physicians remain skeptical because clinical practice involves far more than knowledge recall. Direct patient contact accounts for only 40% of a physician’s shift, with the rest spent navigating bureaucracy, reviewing disorganized data, coordinating care, speaking with families, and reacting to misaligned system incentives. To address this gap, they propose “Humanity’s Next Medical Exam,” built on three pillars: oral board-style interactive interrogation that challenges models beyond rote knowledge, experiential learning in high-fidelity sandbox environments where AI can learn from consequences through realistic patient scenarios, and real-world continuous learning that enables systems to improve with each patient interaction while potentially exceeding human performance. Crucially, realizing this vision requires not just technical advances but modernized healthcare infrastructure with government-mandated open APIs and data sharing standards to break down silos, without which even the most sophisticated AI tools will remain unvetted for real-world clinical use.
Read Paper | NEJM AI
-- Emma Chen, Pranav Rajpurkar & Eric Topol

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.