RSS Amplifier

Prompt-Led Product | For PMs Building in the AI Era · Aug 6, 2026

Substack’s AI Detector Cannot Tell a Poem From a Release Note. That Is the Real Problem.

0
Sign in to vote or save

Elena | AI Product Leader · Prompt-Led Product | For PMs Building in the AI Era

I ran Substack’s AI detector on my own posts last week and some of my technical posts flagged as partly AI. The posts that flagged were the ones with numbered steps, database details, and before-and-after metrics. The personal stories passed clean.

That pattern is the whole story. I document AI builds for a living, which means my writing is structured, precise, and repetitive because the genre demands it. A billing bug writeup and a poem do not share a linguistic fingerprint, and any detector trained to find one pattern will misread the other one.

Substack shipped a single generic scanner for every genre on the platform, and that one product decision turned a trust feature into a style tax on its own paying publishers.

Poetry, personal essays, technical documentation, and marketing copy have completely different structural signatures. One model cannot judge all four honestly, and the writers whose income depends on this platform are now editing for a detector instead of for their readers.

  • My technical posts flagged and my stories passed: What my own scan results reveal about who this feature punishes

  • One scanner cannot serve many genres: Why detection without genre context is a product failure by design

  • One paragraph, two verdicts: The experiment that shows what the detector actually measures

  • Publishers are now writing for the scanner: The incentive Substack accidentally shipped to the people who fund it

  • The feature Substack should have built: What a genre-aware trust system looks like

Let’s get into it. 👇

Hey, I’m Elena! 👋

AI Product Manager, builder, and founder of DraftKit.app and the AIAdventChallenge.com. I write Prompt-Led Product, a newsletter for PMs who are tired of just writing requirements. This is for the product leader who wants to actually ship. The kind where you stop managing tickets for a second, open a prompt, and build the logic yourself.

Built by you. Led by your product taste. 🚀

If you’re new here, welcome! Here’s what you might have missed:

Join a community of builders from companies like Reforge, Product School, OpenAI, Lovable, Pendo, and Miro shipping AI products.

Quick context if you missed the news. On July 21, Substack integrated Pangram, an AI detection tool, announced by CEO Chris Best in a post titled Against Claudefishing. Any reader can now scan any post, Note, or comment over 100 words published after that date and get a percentage verdict of human versus AI writing.

The scan uses a three part scale, and that scale is the first thing worth noticing. It separates AI, AI-assisted, and Human, which sounds exactly right for how modern writers actually work. Then it scored my technical draft:

Verdict label: “Fully AI-assisted text” AI: 100% AI-assisted: 0% Human: 0%

That draft documents my own builds with my own numbers, and I say publicly that AI is part of my writing process. I am the textbook case for the AI-assisted middle category. The scanner had a category built for me and gave it zero.

A classifier that cannot place an openly AI-assisted technical writer in its own AI-assisted bucket is failing at the one job its scale promised.

And the reason has nothing to do with honesty. Technical writing is structured, precise, and repetitive because the genre demands it, and every one of those traits reads as machine output to a general purpose model.

The detector did not find AI in my writing. It found structure. And structure is what my readers pay me for.

Strategic advice: Before you trust any classifier verdict about your own work, segment your content by type and scan each group separately. A single aggregate score hides the pattern that actually explains what the model is reacting to. This applies to AI detectors today and to every quality score your own product shows its users.

Share

Here is the product argument nobody in this week’s outrage cycle is making.

What most builders assume: AI detection is one problem, so one good model with a low error rate solves it for an entire platform.

What is actually true: Detection accuracy is genre-dependent. Prose that is deliberately structured, precise, and repetitive overlaps statistically with AI output, so a general-purpose scanner systematically misreads exactly the genres built on those traits.

Think about what lives on Substack:

  • Poetry that breaks grammar on purpose.

  • Personal essays with a voice you could pick out of a lineup.

  • Financial newsletters built from the same weekly template.

  • Technical writeups generated from release notes, changelogs, and build logs.

  • And much greater diversity in writing, which has led Substack built features like recipe embed tool that lets writers insert interactive "how-to make" recipe cards into their posts.

These are different languages wearing the same alphabet. A detector that treats them as one population will always be most accurate on expressive personal prose and least accurate on disciplined technical prose.

This has evidence behind it beyond my scans. A 2025 study on AI detectors found that some detection models flagged text by neurodivergent writers more often than text by neurotypical writers, because rhythm, repetition, and structural consistency read as synthetic to a model trained on typical prose. Different writing patterns, same failure mode: the detector punishes structure it was not trained to expect.

Substack shipped it to every reader anyway, with one model for every genre and a report button as the remedy.

AI detectors measure predictability, not authorship. In plain terms, they check how expected each word choice is given the words around it. Good technical writing is predictable on purpose: consistent terms, repeated structures, no decorative variation, because ambiguity in documentation is a bug. The detector reads that discipline as machine output. The better you are at technical writing, the more synthetic you look.

The draft scan told me the detector’s conclusion. I wanted to know what it actually measures, so I ran the simplest test I could design, directly on Pangram.

I wrote a short note about this exact situation and scanned two versions. Same claims, same receipts in both: DraftKit, a Supabase billing bug debugged at 11pm, 180 legacy blueprints cleaned without a dev team.

Version one, 64 words, edited tight the way I publish everything: 100% AI Generated Version two, 87 words, same content plus hedges, filler, and a PS with an emoji: 100% Human Written Both results carried the same small print: confidence limited, short text

The 23 words I added were pure noise. A “So” to open, an “I’m pretty sure,” an “I think for me,” and a throwaway PS line joking that the note would get flagged. No fact changed. No authorship changed.

The detector rewarded me for making the writing worse. The scanner reads discipline as AI and sloppiness as humanity.

Every edit I normally make before publishing, cutting filler, removing hedges, tightening sentences, moves my score toward machine.

A tool whose verdict flips on filler words is a style detector wearing a trust badge. It cannot tell you who wrote something. It can only tell you whether the text performs humanness the way the model expects, and polished technical prose does not perform.

Now follow the incentive, because this is where the product mistake compounds.

Substack’s revenue is a cut of paid subscriptions. The publishers this feature makes anxious are the same people who generate that revenue. And anxious publishers do predictable things.

They pre-scan every draft. They strip out the numbered steps and tables that make technical posts useful. They add stylistic noise to sound more “human” to a model. They spend editing time on the scanner’s opinion instead of the reader’s experience.

Every one of those behaviors makes the writing worse for the actual audience.

Strategic advice: Any time you ship a score, users will chase the score. That is the most reliable law in product design. If the score measures a proxy instead of the real value, you have just redirected your best users’ effort away from the thing your business depends on.

Ask what behavior your metric buys before you ship it, because you will get that behavior at scale.

The debate is already visible across the platform.

Dina Litovsky called the rollout a civil war, and she noticed something that confirms my experiment from the other direction: her post scanned as fully human even though she deliberately embedded an AI-written sample paragraph inside it as an exhibit.

Style context decided the verdict, in both directions.

When a platform’s paying supply side and its quality enforcement collide, the platform always thinks it is policing bad actors.

What it is actually doing is repricing risk for good actors.

Every legitimate technical writer on Substack now carries a small probability of public misclassification on every post, and they were given no way to buy that risk down except writing differently.

The slop problem is real. I am not defending newsletters generated wholesale by a prompt. But detection without genre context was the cheapest possible answer to an expensive problem, and cheap answers to trust problems create new trust problems.

A serious version of this feature starts from a different question: what does authenticity look like in each genre? For a personal essay, voice consistency over time. For a technical post, whether the builds, numbers, and claims check out.

A fact-check layer would tell readers something a style scanner never can: whether the release note describes a release that actually shipped.

One scanner, one threshold, every genre: that choice was the failure, before the model made its first mistake. Next Thursday I am going full product spec on the alternative: how I would have implemented AI inside Substack, feature by feature, because a detector is the least imaginative thing you can build with this technology on a writing platform.

Check out the Build Series. Part 1 | Part 2 | Part 3 | Part 4 | Part 5

If you write technical content on this platform, do not let a genre-blind classifier restyle your work. Your structure is your value. The moment you blur your writing to look human to a model, the scanner has cost your readers something and gained them nothing.

Two things you can do right now:

1️⃣ Run the two-scan test on your own catalog. Scan your most structured post and your most personal post, and compare the verdicts. If the gap is large, you now have proof that the detector is reading your genre, and you can decide what to change (usually nothing) with data instead of anxiety.

2️⃣ Send this to one technical writer who got a score they did not deserve. The loudest voices this week are essayists debating authenticity. The people with the most at risk are the ones documenting products, code, and research, and almost nobody is arguing their case.

And as always, stop guessing. Start building. 🚀

One question for you: What did your most technical post score versus your most personal one? Drop both numbers in the comments. I will tell you exactly how I would read the gap for your specific publication.

Leave a comment

If you want me to audit how your own product’s scores and classifiers shape user behavior before they backfire, reach out at hello@elenacalvillo.com.

Read the original on promptledproduct.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.