Welcome back to The Third Hemisphere, where I try to make sense of how AI is reshaping work, thinking, and creativity, often by watching my own assumptions get upended.
If you were forwarded this and want to subscribe, click below. If you want to support a real human writing about AI, upgrade to paid.
First, apologies to readers who are sick of AI detection. Trust me, I am sick of it too, and pine for the halcyon days when there wasn’t AI to detect in the first place. I truly can’t wait to write about something uncontroversial and light-hearted next week, perhaps a hot take on the Fauci hearings.
Nevertheless, here we are. For anyone who missed it, I recently wrote about why “It’s OK to Change Your Mind About AI Detection.” The narrowish point I wished to convey is that Pangram is pretty damn good at detecting AI text under third-party testing conditions, and it was driving me crazy that serious people with large platforms wouldn’t acknowledge this basic empirical point.
But, after reading and engaging with many reactions, including a great discussion in the comment section of the last post, I’m realizing my last post came off to some readers as too promo for Pangram, or even naive, as though I hadn’t thought about the other issues at stake, or implications in play. This, of course, also drives me crazy, so here I am with a companion post. The point here, inspired by reader comments, is to lay out why I think even a quite good Pangram won’t fix a number of other important issues.
I should have been more precise in my last post, so let me borrow a distinction from the biomedical world, where I spend a lot of my time. In medicine, efficacy is how a treatment performs under the controlled conditions of a clinical trial: selected patients, careful protocols, close monitoring. Effectiveness is how it performs in the real world, where patients have three other conditions, or happen to be, um, women—you know, that minor subgroup the FDA actively barred from early-phase trials until 1993. The point is a drug can be efficacious in clinical trials and still fall short in the real world. It will not work has expected in some people, there will be unexpected side effects in others, etc. This is why FDA approval comes with strings attached. Manufacturers have to report adverse events for as long as a drug is on the market, and when a drug is cleared on more provisional evidence, through accelerated approval, the FDA can require confirmatory trials to verify the benefit is actually there.1
Using that vocabulary: Pangram is efficacious. That’s amazing! Under third-party testing conditions, it performs similarly to its maker’s claims. But efficacious is not the same as effective, and this is where I differ, a lot, from Pangram, whose position seems to be that its validation studies prove the tool “works” in the real world. What concerned me most about Substack’s “Against Claudefishing“ announcement was the absence of any reported mechanism for real-world monitoring. Sure, writers can dispute individual verdicts, but a dispute box is case-by-case recourse, and it’s unclear what even happens, if anything, behind closed doors of Pangram and Substack. To my knowledge, Pangram and Substack do not have a plan to estimate ongoing, aggregate error rates from actual deployment.
This is a major oversight. One particular concern of mine is that even though Pangram’s overall false positive rate is low, false positives might cluster in real-world subgroups not identified in the validation studies. Liam Dugan, a University of Pennsylvania graduate student whose PhD dissertation is on AI detection, told me when I spoke to him for a Slate article that “for most people, they might never, ever get a false positive. And for other people, the false positives are sort of disproportionately allocated on them because they just happen to write like AI.” Non-native speakers were the subgroup people rightly worried about early on, and there’s progress on that front. A paper published last month by researchers in Brussels took forty master’s theses submitted to their own faculty before 2019, every one written by a non-native English speaker, and submitted them to Pangram, which flagged none of them. This is progress! I’d still like to see more evidence before we put the non-native English speaker discrimination issue to rest, but even if we do, I doubt non-native English speakers are the only subgroup more likely to trip up detectors.
There may be clusters nobody has thought to look for yet—some writing whose combination of genre, topic, and style just by chance tends to inhabit spaces near the AI/human boundary. We don’t know exactly how Pangram performs on real people writing in 2026, in multiple languages, with all of the weird ways they use AI and have absorbed AI-isms. Pangram works a lot better than AI detectors used to, but it will still fail in unpredictable ways.
The second issue a better detector doesn’t fix: people have different thresholds for how much false-accusation risk they’ll tolerate. This is fine! But it is a normative question, and accuracy data can’t settle it. (I mean, it is partly an accuracy question, in that we lack sufficient real-world data. But the efficacy data exists, and we can start making decisions with it.) For how to think about living with an imperfect but useful measurement, let me borrow from an area I’ve done a lot of previous reporting on: forensic science.
Some forensic science is genuinely junk science. Bite-mark analysis is theoretically unsound—there is no good evidence that human dentition is unique, or that skin records it faithfully—and when examiners have been put to the test, their error rates are far too high. It should not be allowed in court. Fingerprinting, you might be surprised to learn, is pretty good but has a noticeable error rate: in one of the better studies to date, FBI-affiliated examiners made false matches at a rate of about 1 in 1,000. DNA testing is excellent, though it too has a small but real error rate in real-world use, from contamination and sample mix-ups, and adjacent techniques like mixed-sample interpretation are dicier. Does this mean we throw out all of forensic science as junk? No. We accept that error rates exist and deal with them, in one of the highest-stakes setting society has.
Some people want perfection, or close to it: they want Pangram to be like DNA analysis. Others smear it as if it were bite marks. Really, Pangram is more like fingerprint analysis. Pretty good, not perfect. This gray area means people react very differently on a normative grounds. One commenter reacted this way:
One thing that intrigues me is why people are worried about false positives so much. Yes, it's a nuisance to be accused of something you are not. But outside of contexts where an AI accusation carries formal sanctions, like in academia, I feel like it's really a matter of taste (or editorial preference), especially in commercial settings like Substack. For the most part, AI or not is just another dimension publishers and readers are free to judge your writing on. And if they miss out on something great just because they prejudged you, that's on them.
Whereas for another commenter declared:
The facts are it's not 100% so innocent people get attacked and possibly ruined. That's enough.
These are both acceptable positions, but they are normative ones about what is an appropriate level of risk, who bears it, and so forth. My simple plea in the first post (and this isn’t about these commenters specifically, just in general) is that any normative argument about how to deploy AI detection starts from acknowledging the substantial empirical data we have about Pangram, not pretending it doesn’t exist or citing old studies.
There is one big problem with both my clinical and legal analogies: Substack is not a court of law or a hospital. As Ruv Draba, who critically quoted my last post put it:
I think this is an important point. What makes a 1-in-1,000 fingerprint error tolerable in court is the adjudicating apparatus, however imperfect: rules of evidence, cross-examination, a judge deciding what a jury may hear, several levels of appeals.2 What makes releasing a drug into the population reasonable is its use is overseen and monitored by experts in relatively controlled medical settings. On Substack, the process is a single Pangram score, often poorly understood, and then potentially screenshotted and thrown into a feed. This commenter is correct that talking about the tool in the abstract, without the sociotechnological context of its deployment, while not exactly “misleading,” is probably insufficient.
Which brings me to what I feel is the core problem with this partnership: It doesn’t actually solve, nor is it capable of solving, the issue. I don’t doubt the sincerity of the people behind it. Chris Best of Substack framed “Claudefishing” as a mismatch between what a reader expects and what they get, which is a very real problem. Max Spero of Pangram has been consistent that his tool should “never be the ending arbiter” but a starting point for a more thorough investigation. I believe both of them mean it. I also think both of them are wildly naive about how AI detection will function in the actual world we live in. In their vision, readers scan judiciously, click through to reports, understand the strengths and limitations of AI detection, keep up to date on the literature, carefully consider notions of authorship, etc. In the real world, people’s behavior will be driven by an interface that invites none of this nuance. I predict two spectacular kinds of failures, which I’ll illustrate with examples.
The first is on the writer’s side. People say they want less AI writing, but not everyone has considered the full range of AI use and what it can enable. I urge anyone who automatically recoils against AI use in writing to read Emma Klint. She’s a non-native English writer with a self-described neurodivergent brain, and she uses AI for everything she publishes:
I have strong verbal processing and low working memory. AI lets me use that verbal processing to compensate for what my working memory can’t hold. That helps me access thoughts that were already there, that I lost track of while I was still thinking them. Simple as that.
People often warn that using AI will “steal your voice,” but Emma describes using AI as the opposite:
When I started writing with AI, I wasn’t afraid of losing my voice, because I didn’t really have one. I had lots of thoughts and a point of view, sure, but not a writing voice. The process of writing through dialogue, and thinking with AI, is how I found it.
I’ll admit, sometimes in Emma’s writing I hear echoes of Claude, and aesthetically the English-speaking writer snob in me doesn’t like it. But I think it would be shallow of me to take that reaction very seriously. Emma’s posts clearly have a tremendous amount of thought put into them, and judging her valuable chronicling of AI use on the basis of a few linguistic tics strikes me as the wrong metric. As Emma herself put it: “I’m not arguing that Pangram will get the percentage wrong. Even a perfect score would answer the wrong question.” My concern is that the Pangram integration may discourage writers like Emma Klint from writing at all—and I don’t think that’s a net win.
One clever fix proposed in my comments is that Substack sidestep detect-and-punish with detect-and-deprioritize: let readers flip a switch that says “prioritize human writing in my feed,” so that a false positive might limit a writer’s algorithmic reach but spare them a public accusation. Of the design alternatives I’ve heard, this seems reasonable until I think of writers like Emma Klint. A writer like her would struggle to gain traction, without her knowing quite why. Writers who use AI to pump out half-baked content, on the other hand, would just rewrite until they slip past the filter. Users who flipped the “deprioritize” switch would then end up with a feed of human writing blended with undisclosed hybrid writing anyway. This deployment would screen out the Klints and yet reward the evaders.
The second failure is on the reader side. I think of a post from Marc Watkins, a lecturer of writing and rhetoric at the University of Mississippi, who chronicled a week using Pangram’s Chrome extension labeling every post in his social feeds. From his post, “How An AI Detector Made Me Trust People Less“:
I noticed my behavior change. Dramatically. I stopped interacting and reading posts and instead focused on labels. When we allow a company to place labels on our social media interactions, we cede some agency to opaque systems, ultimately giving them a great deal of power over our interactions.
Watkins is a sophisticated reader, aware of what these labels can and cannot claim. And yet, his behavior changed anyway.
Both of these failures reveal that the core problem not just technical but social: People want a determination on provenance, authenticity, and originality, but what they get is a determination on text. So, probably, all of this is part of a longer social process of redefining things like authorship and authenticity, which is really hard to do, and which no scan, however efficacious, will do for us.
So let me state my position as plainly as I can, since I failed to in my last post: Pangram is efficacious and people should stop making sloppy or outdated arguments to the contrary; and also: I think this version of the Pangram/Substack integration will do more harm than good, and I think Max and Chris are a bit naive about the social implications of integrating a powerful weapon in call-out culture. I think it is perfectly coherent to accept the empirical reality that Pangram works fairly well (with major limitations, of course) and normatively be against its deployment (for a variety of reasons).
In announcing the partnership, Best wrote that when readers have to wonder whether what they’re reading is real, it “undermines trust in authorship and threatens the livelihood of writers.” I agree with the problem, but not the solution. A Pangram score cannot differentiate a Claudefisher from a Klint, and a feed of color-coded metrics led Watkins to start trusting people less in a week. The Substack/Pangram partnership was built to restore trust between writers and readers, but the irony is, I suspect, that it will only serve to further fray it.
Again, reality asserts itself: in practice, courts have almost never excluded fingerprint evidence as unreliable, and juries almost never hear an error rate at all
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.