Welcome back to The Third Hemisphere, where I try to make sense of how AI is reshaping work, thinking, and creativity, often by watching my own assumptions get upended.
If you were forwarded this and want to subscribe, click below. If you want to support a real human writing about AI, upgrade to paid.
Update, July 31 2026: I wrote a companion post inspired by and addressing the lively comment section on this post: “What a Better AI Detector Doesn’t Fix.”
With the news of Pangram coming to Substack, I figured I should weigh in, because I've very publicly changed my mind about the reliability of AI detection after a social media dust up with Pangram’s CEO (not a research method I recommend). Since then, I’ve spent more time reading and writing and speaking to reporters about AI detection then I ever wanted to. So consider this a few lessons from my, ahem, learning experience. Here’s A) what I look out for when assessing people’s arguments about AI detection, B) what feel to me like more interesting questions about AI and authorship. Apologies in advance for the length; in the words of Blaise Pascal that eternally haunt my writing life: “I have made this post longer than usual because I have not had time to make it shorter.”
Lumping all AI detectors together
A common move is to take a company that is less rigorous than Pangram, or one whose business model involves selling people humanizing tools, and treat its failures as evidence about AI detection generally. The performance gaps between tools are enormous. There have been several independent evaluations of Pangram to date, and they all suggest Pangram is very good (a little more nuance on that below), but the one I’m most familiar with is from Brian Jabarian and Alex Imas, economists at University of Chicago. They ran four detectors across nearly 2,000 passages, from Amazon reviews to novel excerpts. In their study, Pangram’s false positive rate was essentially zero at every threshold they tested. The open-source RoBERTa, meanwhile, barely separated human from AI text at all.
So when someone posts a screenshot of an AI detector erroneously determining that their thesis from 1995 or the U.S. Constitution is AI, and concludes that therefore AI detectors are snake oil, the first question is which tool produced the laughable finding. It probably wasn’t one of the good ones.
Citing outdated research
Relatedly, a lot of the research on AI detection is out of date, and the two papers I see posted in replies as if they are slam dunks are actually both from 2023, assessing technology that works very differently than Pangram. I want to be clear these are perfectly good papers, and they served a very useful purpose early on, when universities and schools were eager to address the AI problem with subscriptions to dubious detection services. I think it was crucial to show that this early generation of detection tools was unreliable, and demonstrate how detection resulted in false accusations and generally frayed trust between students and teachers. In doing so, research like this forced a more nuanced conversation in education circles around effort and pedagogy and so forth in the age of AI.
However, AI detection technology has materially advanced, so this early literature no longer applies. If a paper isn’t including a state-of-the-art tool like Pangram, then its findings are largely obsolete.
Using false negatives to suggest that AI detection is worthless
This one is more subtle but important. You’ll see posts like, “I just changed 5 words and now Pangram says it’s human. See, these things are bullshit.” Hundreds of likes and clap emojis ensue. The problem with that reasoning is that a detector can be wrong in two ways: labeling human writing as AI (false positive), or labeling AI writing as human (false negative). It cannot minimize both at once, and Pangram is tweaked to minimize false positives. Most “gotcha” examples are false negatives. I think it’s better to prioritize a low false positive rate because that results in far fewer false accusations (though not zero—which raises serious moral questions about automating detection on the high volume of writing on Substack, with what’s at stake, reputationally). However, minimizing false positives also means false negatives are relatively easy to produce. The implication is that a “human” verdict from an AI detector is a much weaker claim than an “AI” verdict.
This is a bummer, because in a sense Pangram isn’t giving people what they really want. People really want a tool that guarantees their version of “authenticity,” which is often, in their view, human writing without any involvement of AI. But what Pangram does instead is tell you when prose was fully generated by a machine. It does not tell you when someone sort of half used AI, or muddled with AI-generated text until they got the detector to say “fully human.” In other words, “100% human” isn’t a stamp of authenticity, but “100% AI” is a stamp of AI text (at least for essay-length texts; Pangram is far less reliable on paragraphs or shorter posts). These are real limitations, but not the same thing as the tool being worthless.
I want to add, finally, that there is a sizable group of people sowing doubt about AI detection not for the sake of good-faith argument, but for their own benefit. I’m not going to name names because I don’t want to get in online fights, but I’m referring to people who have gained a large audience in recent years posting AI-generated text under the guise of human authorship, but not disclosing it. They do not want a button that says “fully AI-generated” that people believe. This is not to weigh in on their work, or the appropriateness of AI use and disclosure in writing, but simply to say that just as Pangram has an interest in promoting the infallibility of its tool, many posters/influencers/thought leaders have an interest in promoting its fallibility.1
In my view, policing prose targets the least consequential form of AI influence, and essentially lets Pangram set the terms of acceptable use. I’ve written about this extensively, so I’ll just reprint the thrust of my arguments from recent pieces. From my post “Detect and Punish,” back in April, I relayed an exercise I do with science journalists in workshops:
I then hand each reporter one of two AI-generated research reports on collagen supplements. Same underlying studies, same data. Report A opens with positive clinical findings and mentions industry funding as a limitation in paragraph thirteen. Report B opens with the funding bias analysis, loudly labels which results are industry funded. I ask: what story are you primed to write? The answers are different, depending on which report they read. Report A primes a “does collagen work?” story. Report B primes a “why you can’t trust collagen research” story. Both reports are defensible summaries of the same literature. The AI just decided how to order and present the same information.
A reporter who reads Report A before doing any of her own research is more likely to spend the rest of her reporting anchored to the frame that collagen basically works and industry funding is a footnote. A reporter who reads Report B first will spend her reporting anchored to the frame that the research is corrupted. Both will write every word of their articles themselves. Both will pass any AI detector. Meanwhile, a reporter who did her own research and then asked an AI to help clean up her prose would get flagged. I’d argue she was the least influenced of the three, yet the moral intuitions of writers are that this last reporter betrayed the craft the most.
So instead of drawing a bright red line at prose—conveniently policeable by Pangram—here’s what I’d rather see people spend their energy on:
I think it’s fine not to use AI tools and think it is healthy for the journalistic and literary communities to have abstainers; my contention is that some people will use AI, and it would be better to understand and shape that use through evidence and realistic professional norms rather than drive it underground by shaming. Take using AI for research, which many reporters seem to support: I’d like to know how AI-generated research summaries shape story framing compared to, say, searching for original papers via academic databases like PubMed. I’d like to know whether the framing effects are stronger or weaker than what writers are already exposed to via Google’s search algorithm or PR pitches. AI’s influence on media and publishing is likely to be profound and long-lasting, and that is a world I hold dear. But I’m concerned that the efforts spent poring over text for detector-fueled pile-ons—all that time, energy, and fury to assemble a circular firing squad—is aimed at the form of AI influence that, in the long run, may be least important.
It’s OK to change your mind. A lot of people are digging in on AI detection, and again, I’m not saying it’s perfect. Nor am I supporting its use on Substack (or in academia). That’s a separate, normative question. But let’s at least debate the wisdom of AI detection with the empirical facts at hand—which have changed. And if you think I’m just some rando on Substack, take it from the considerably more well-known computer scientist Arvind Narayanan, authoer of AI Snake Oil, who this week posted that after trying a recent version of Pangram he’s “more-or-less ready to change my mind as an AI-detection skeptic.”
No shame in that, guys.
There’s a fourth move that is almost too clickbait-y to refute so I will relegate it to a footnote. The argument goes like this: Using AI to detect AI is somehow, intrinsically, absurd. I get the superficial irony here, sure, but I think some people imagine the process is “Hey Chat, is this AI?” and you get a hallucinated response in return. Not quite. Pangram is a neural classifier trained on loads of text to learn features that allow it to infer whether a passage of text is more likely human- versus AI-generated. Yes, this is a form of AI. But people train these kinds of classifiers successfully to do all sorts of tasks, like differentiating a “1” from a “7” when you scan a check into your banking app.
Don’t get me wrong, the underlying mechanism matters because it might give you insight into likely failure modes (failure due to shortcomings or quirks in the training data, for example). But someone who dismisses Pangram because it is “AI” and their knowledge of “AI” ends at memes about ChatGPT on Instagram isn’t to be taken seriously. In the end all that matters is empirical evidence. Can Pangram differentiate human vs AI text or not in properly designed tests? How does Pangram fail or succeed in the real world? All the rest is a distraction.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.