Hey friends 👋
Welcome back to the Signal Pro.
I was planning to write this week’s issue on Claude. However, on Tuesday, Substack added a “Scan for AI text” button that required my full attention. So, with that being said, that’s exactly what this piece today is going to be on.
Open any post in Substack, click the three dots in the top-right, tap “Scan for AI text”, and Pangram, the detection tool Substack has partnered with, returns a percentage estimate of how much of the writing was done by AI, AI-assistance, or the human hand alone. It works on posts, notes, replies and comments longer than 100 words, published from July 21 onwards.
Chris Best, Substack’s co-founder, announced the feature in a post called Against Claudefishing, an apt name for dressing up the output of Anthropic’s LLM as human-written.
He illustrated it with an engraving of the Turk, the eighteenth-century chess machine with a human player hidden inside. The problem, as Best captures it in his post, is neither AI use nor the quality of what AI produces, since plenty of AI-assisted work is good and plenty of slop is entirely human-made. The problem lies in the mismatch between what a reader expects and what they get, especially when a human gives attention to a piece of writing with nobody behind it.
So I’ve spent the week working out where I stand, and I think I’ve been able to boil it down into two quite simple questions that lurk beneath the newly added button.
Can a tool actually tell who wrote something?
If it could, would the answer change anything?
Grading human-vs-AI text has been tried at scale for the last three years since the proliferation of ChatGPT in November 2022. In January of 2023, OpenAI, the company behind ChatGPT, built its own detector and ended up retiring it within six months, explaining it was “no longer available due to its low rate of accuracy”. It caught 26% of AI-written text. The company that made ChatGPT could not reliably recognise its own output.
That same summer, Stanford researchers ran 91 TOEFL (Test of English as a Foreign Language) essays, written by people learning English, through seven widely used detectors, alongside essays by American eighth graders. The detectors judged the eighth graders’ work almost perfectly. They flagged more than half of the TOEFL essays as AI-generated, with an average false positive rate of 61.3%, and 97.8% of those essays were flagged by at least one detector. The detectors hunted for predictable word choices. Someone writing in a second language draws from a smaller pool of words, meaning they were therefore at a disadvantage. When the researchers asked ChatGPT to enrich the essays’ vocabulary to sound more like a native speaker, the false positive rate fell to 11.6%. The solution to stop being mistaken for AI was to use AI.
The same thing applies to neurodiverse writers. In 2024, Bloomberg reported the case of Moira Olmsted, a student on the autism spectrum who received a zero on a routine assignment after a detector rated her summary as likely AI-written. Her structured, literal style—the way she has always written—is exactly the style these tools were trained to treat as suspicious.
To be fair to Substack, Pangram is a different class of tool from that first generation. Independent testing by University of Chicago researchers found it was the only detector tested to combine a near-zero false positive rate (≤ 0.5%) with reliable detection of AI text, and Chris Best (Substack CEO) has been clear that the scan is imperfect and cannot measure the care behind a piece of writing.
But this year has already shown what the fallout looks like. In March, Hachette cancelled Mia Ballard’s novel Shy Girl across every market after Pangram scored the book as 78% AI-generated. This figure was disputed by Ballard, who said that a freelance editor introduced the AI text. The Atlantic investigated Pangram in May and found it had flagged a New York Times Modern Love column as more than 60% AI-generated, concluding the tool makes mistakes “perhaps to a greater extent than is currently understood” whilst accumulating the power to end careers.
One error in every 10,000 scans, on a platform where every post, article, note and comment now sits a tap from an AI verdict, creates a steady stream of writers wrongly accused.
I believe AI detection does not work reliably enough to accuse anyone of anything. We must trust what we see with our own eyes before we trust a percentage from an AI tool that confidently guesses. And let’s remember who this impacts the most: people writing in a second language and neurodiverse writers, who have been judged their whole lives for how they speak and write. We must be kind, fair, and exercise our own independent judgement.
There is a deep irony in consulting AI to decide whether a human wrote something, because AI is evil.
In the name of protecting human authorship, we hand over one of the most human capacities we have—reading a piece of writing and deciding for ourselves whether it’s actually any good. Clicking a button and receiving a percentage estimate on the screen eclipses the exact judgement we built to protect.
The score also answers the wrong question. It estimates whether a model produced the words and knows nothing about whether the thinking and linking behind those words involves a human. “The presence of AI does not prove the absence of a human” was one of the most-liked replies under the announcement. And I think it completely nails it.
Let’s run the experiment on ourselves for a second. Think of the last piece of writing that genuinely impacted your life in a meaningful way: a book, a blog, or even a short post on social media. Now imagine a scan telling you, confidently, that AI had a majority hand in it. The words don’t change; the argument still holds. What exactly are you now entitled to think less of?
I’m not asking that rhetorically, either. Earlier this year, The New York Times put a scenario just like this to its own readers in a blind test, human writing against AI writing. Here’s what they found.
In March, Kevin Roose and Stuart A. Thompson built an interactive quiz for The New York Times. It involved five pairs of passages, one from a well-known human writer and one from an AI model, across literary fiction, fantasy, science writing, historical fiction and poetry. Readers picked whichever they preferred with no idea which was which. More than 86,000 people took it, and 54% preferred the AI.
Roose called it “a real moment” and compared it to the Judgment of Paris, the 1976 blind tasting where California wines beat the French wines everyone assumed were untouchable. Roose and Thompson noted that human writing tends to carry clunky phrases, pointing to a rough line of Cormac McCarthy’s “Blood Meridian” with the absence of punctuation: “As well ask men what they think of stone.” Today’s models are now so fluent that awkward syntax has become a hint you’re reading a person. For years the machine gave itself away by sounding wrong, and now the human does.
Noam Brown, a research scientist at OpenAI, replied to Roose’s post of the results, asking why coders are generally fine with AI-generated code whilst writers resist AI-generated writing?
The difference is what each gets judged on. Code gets judged on whether it runs, writing gets judged on how it reads. The output of writing is the input of the writing itself. But the output of code is much further away from the input. It’s abstracted through how the product or service operates, not line-by-line syntax.
And people are very protective of their taste. AI prose is smooth, clear and competent, and it lacks the element of surprise that the human hand often achieves. A model writes by trending towards the most probable next word. So left to itself, the output gravitates to the average of everything it was trained on. Human writing keeps the unexpected word or the crooked phrase that makes you stop, analyse and reread.
Over the 200-word examples, smoothness wins. The quiz used short, isolated passages. But across a full book, or a year of newsletters from your favourite writer, the absence of surprise compounds on the one side, whilst on the other, the averageness starts to bleed through and materialise via accumulated disinterest.
A teenager with ideas and no instrument can now produce a full album in their bedroom. People who can’t write a line of code are shipping working tools today. Writing is simply the latest craft where the gap between having something to say and saying it well has just completely collapsed. For someone with dyslexia, or someone writing in their third language, a model that turns their thinking into clear prose removes a barrier that never had anything to do with the quality of their thinking.
None of this moves me off human writing. But if I learnt tomorrow that a piece I admired had been made with the help of AI, thoughtfully, my judgement of the writing wouldn’t change.
The writing is the evidence. What creates a sinking, disheartening feeling in the pit of my stomach is when a piece has been clearly one-shotted—or at least when phrases, paragraphs, and attributes with AI tells have spread like an unwelcome smell throughout the piece.
A writer who seeds the model with their own writings, views, notes, opinions, ideas, and rambles with all the rough edges, then uses it to sharpen the delivery, ends up with something that belongs to them. But it remains in their judgement to analyse, read, interpret and assign for themselves what to do with that output, much like a sculptor creating a statue.
Substack released a second feature alongside the scanner: a “How I make this” statement. It’s a gentler instrument than the button, a way to “come clean” without a number attached. But why have it in the first place?
A piece of writing should be taken for exactly what it is—a piece of writing. Embarking on an adventure with someone giving you a piece of paper at the start, exclaiming “sorry but here’s kind of how it works and how it’ll go and why it plays out this way” is no longer an adventure.
The opportunity of reading something, and either being completely taken aback or completely disheartened, is what makes reading exciting. You go into reading the piece trusting the person who’s providing the take. If you don’t like it, that trust disappears. If you do, that trust stays, and you return, on a new adventure of a path uncertain and unclear—but that’s what adventures are all about.
We must use our human brains, strive to be curious, question things, and do so with the intent of arriving at the truth.
Substack is great. AI is great. But human judgment is the greatest.
See you tomorrow for the Sunday issue.
— Alex
💡 If you enjoyed this issue, share it with a friend.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.