RSS Amplifier

Gustavo's Newsletter · Aug 21, 2026

Almost every biomedical paper now uses AI. Almost none of us were taught how

0
Sign in to vote or save

Gustavo Monnerat PhD · Gustavo's Newsletter

Hi,

A preprint went up on arXiv on 11 August with a impressive number: by December 2025, 89% of open-access biomedical papers in PubMed Central showed signs of LLM-assisted writing or editing.

I don’t think that number is a scandal. I think it’s a description of where we are. What I want to write about this week is the gap it opens up, because adoption at that speed with no training attached is how errors get into the permanent record.

What happened

Holzwarth et al built a new way to estimate LLM use from shifts in word frequency. Earlier methods gave a lower bound. This one aims for a direct estimate, which is why the figures jump.

  • Full papers: 52% in 2024, 77% across 2025, 89% in December 2025 alone

  • By section, length-matched: Discussion 68%, Abstract 67%, Introduction 59%, Results 46%, Methods 32%

  • Uncropped Discussion sections hit 78% in December 2025

  • By first-author country in 2025: South Korea 85%, China 82%, UK 28%. Pooled, countries with a majority of native English speakers came in at 37%, everyone else at 72%

  • It’s a preprint, not peer reviewed. English-language, open-access, PubMed Central only. And the method can’t tell a grammar fix apart from a ghostwritten discussion

    My take: the tools are already everywhere, so arguing about whether to use them is a few years out of date. The useful question is what people are doing wrong with them.

This part is my own observation:

1. Fabricated references. Still on of the most common failure, and the easiest to catch. Plausible authors, plausible journal, plausible year, DOI that resolves to nothing or to something else entirely. (Fig from Topaz et al, The Lancet, 2026)

2. Real reference, wrong content. This one worries me more. The citation is genuine, the DOI resolves, the paper exists. But what the sentence claims the paper found isn’t what the paper found. Sample size is off, the endpoint has changed, an association has become a causal claim, or the effect direction is reversed. Reference-checking tools mark it as valid. A reader who trusts the citation trail carries the error forward. This is the failure mode that survives peer review most easily.

3. AI-generated or AI-modified figures. Generative tools now produce blots, histology, micrographs and diagrams that look right at a glance and are anatomically or biologically nonsense. The 2024 Frontiers retraction of a paper with an AI-generated rat figure was obvious enough to become a meme. The dangerous versions aren’t obvious. And “AI enhancement” of a real image, cleaning up background, sharpening a band, filling a gap, sits on the same continuum as manipulation whether the author intended it or not.

4. Models reasoning from the wrong epidemiology. Ask a general-purpose model about disease burden and it will answer from a training distribution that skews high-income, English-language and a year or two stale. Prevalence figures for Latin America, sub-Saharan Africa and South Asia come back wrong or come back as US estimates wearing a different label. Same problem with drug availability, screening coverage and guideline recommendations that vary by country.

5. Fabricated data and the wrong statistical test. Two separate problems that arrive together. Models will happily invent a plausible number to fill a gap in a results paragraph, and they’ll happily endorse a test the data don’t support: parametric tests on skewed distributions, no correction for multiplicity across dozens of comparisons, per-protocol analysis presented as intention-to-treat, a p value quoted without the effect size it belongs to.

6. Knowledge presented as current. Every model has a cutoff. Guidelines get revised, drugs get withdrawn, thresholds change, articles get retracted. A model that writes fluently about a 2022 version of a guideline gives no signal that a 2025 version exists.

1. Don’t paste confidential or unpublished material into consumer models. Manuscripts under review, patient-level data, unpublished protocols, grant drafts, anything under an NDA. Check whether your account trains on your inputs, and assume the default is the worse option until you’ve verified otherwise. Institutional or enterprise deployments with a no-training agreement exist for exactly this. Use them.

2. Match the model to the difficulty of the task. Fixing grammar and summarising a paper you’ve already read are easy. Reasoning about a statistical design, reconciling conflicting trials, or interpreting a result you don’t already understand are not. Small and fast models are fine for the first group and will confidently fail at the second. If the task would take a competent human an hour of thinking, don’t hand it to whatever is cheapest.

3. Give the model the source every time. Upload the PDF, paste the passage, point it at a document set with retrieval. Asking a model what a paper says from memory leads to errors. Grounded in the actual text, the same model gets dramatically more reliable.

4. Verify every output that carries a fact. Open the DOI. Check the abstract, the sample size against the paper. Recompute the percentage. Confirm the guideline version.

5. Write long, specific prompts. Say what the study design is, who the population is, what the endpoint is, what journal and audience you’re writing for, and what you want the model not to do. Most bad output I see traces back to a one-line prompt that gave the model no choice but to guess. Context is the whole game.

6. Disclose. “AI was used” tells an editor nothing. “An LLM was used to edit the English in the Introduction and Discussion; no text was generated de novo; no analysis was performed with AI” is a statement someone can actually assess. Journals are converging on this. Get ahead of it.

7. Don’t outsource the interpretation. Conclusions, Implications, Opnions, and Perspectives should be 100% Human

Written in a personal capacity. Public evidence only. Not medical advice.

No posts

Read the original on gustavonewsletter.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.