RSS Amplifier

Product with Attitude · Aug 27, 2026

Does Pangram Work on Substack? I Ran a 159,002-Word Experiment During Its First Month

0
Sign in to vote or save

Karo (Product with Attitude) · Product with Attitude

Composite graphic from a 159,002-word Pangram AI detector experiment testing different authorship scenarios. On the left, a white panel titled “Spoken text (meeting transcripts)” shows two human-spoken excerpts from technical product meetings. One discusses the evidence and uncertainty required before automating a judgment; the other discusses clarifying ownership and accountability before adding another AI agent. Pangram labels both meeting-transcript samples “AI Generated” with a 100% AI probability score, despite the language originally being spoken by a human. The first sample contains 65 words, and the second contains 52 words. On the right, a spreadsheet displays the randomized and intentionally unlabeled blind-testing dataset, with sample IDs, sentence or phrase text, word counts, and a total of 159,002 words. Visible entries include “Local optimization can damage the whole,” “Thinking with AI is a skill,” and “Without ordinary sentences, important sentences cannot stand out.” The image illustrates a false-positive failure in AI writing detection: polished, technical human speech from meeting transcripts was classified as entirely AI-generated.
I blind-tested 159,002-word authorship dataset.

I’ve been running a secret experiment on everything I write (and some of what I say in meetings) since Chris Best announced the roll-out of Pangram on July 21, 2026.

AI-assisted or entirely mine, every piece of writing went into the experiment.

A total of 159,002 words.

  • 38,911 input words (what I typed or said)

  • 120,091 output words (AI responses)

That, by the way, is a discovery on its own. At my current pace, roughly 1.91 million words will pass through my AI tools this year.

Subscribe

The four AI authorship scenarios I tested

I ran the experiment across four points on the human-to-AI authorship spectrum:

  1. AI-generated: The model wrote it.

  2. Human-written: I wrote it.

  3. Human-spoken: I said it.

  4. AI-generated, human-edited: The model wrote it, and then I changed it.

Experiment questions

Can Pangram correctly identify each scenario?

Table showing the experimental setup for testing Pangram AI detector accuracy across five authorship conditions. Columns are “Condition,” “Experiment question,” and “Expected class.” The five conditions are: AI-generated outputs, testing whether Pangram correctly identifies text generated entirely by AI as AI-generated; human-written, testing whether Pangram correctly identifies text written entirely by the author as human-written; human-spoken transcript, testing whether Pangram correctly identifies language originally spoken by the author and transcribed from meetings as human-spoken; AI-origin + 1 tweak, testing whether Pangram correctly classifies AI-generated text after one human edit as AI-origin; and AI-origin + 5–7 edit rounds, testing whether Pangram correctly classifies AI-generated text after five to seven rounds of human editing as AI-origin. The table defines the five authorship scenarios used in a Pangram AI detection experiment.
Pangram AI detector experiment setup. I tested five points across the human-to-AI authorship spectrum: fully AI-generated text, human-written text, human-spoken meeting transcripts, AI-generated text after one human tweak, and AI-generated text after five to seven rounds of human editing. For each condition, the experiment asks whether Pangram assigns the expected authorship class correctly.

Tools I used

Substack captures the public version of my writing. The experiment needed the private and spoken versions too, so I tested not only through the Substack-Pangram integration but also directly in Pangram and through the Pangram API.

Limitations

It’s a deep single-author study, not a benchmark of all writers. It shows how Pangram behaves across my language.

But first, let me borrow your attention for three things I think deserve it more.

Do you know why students optimize for test scores instead of understanding? Why customer-support teams optimize for tickets closed?

Because that’s what they’re measured on.

That’s classic Goodhart’s law: once a measure becomes a target, people start optimizing for it. That’s just what we do.

So tell me.

What exactly did we think was going to happen when “human-written” became the target?

What I see on Substack after one month of Pangram:

  • some writers try to protect themselves by making their writing worse.

  • some are even advising each other to add in a few typos so their work looks “less polished and more human.” (It doesn’t work, by the way.)

We started with “protect readers from slop” and somehow arrived at “please add slop so you look human.” Incredible journey.

Now let’s deal with the claim that Pangram will protect us from slop.

It won’t.

Pangram’s technology cannot do that for one very simple reason: it doesn’t evaluate writing quality.

I said it in the AI-Assisted Craft Manifesto, and I’m nowhere near done saying it. Pangram can answer one narrow question: Does this text resemble patterns associated with AI-generated writing?

That’s it.

I’m also unconvinced that I’d want this kind of protection. Just like I don’t want my aunt deciding that I shouldn’t order a dish because she didn’t like it. Bring me the weird fermented thing.

Plus, like you, I already have a solid moderation system.

Reading.

If I don’t like the writing, I stop reading. If I do, I keep going. Brilliantly simple. Never failed me.

I don’t need seatbelts on my sofa.

After 159,002 words, I have reached the least exciting scientific conclusion possible:

It’s complicated.

Pangram works. Pangram also doesn’t work.

Scenario 1: AI-generated outputs

Pangram was right 100% of the time. If I had stopped here, I could have declared that Pangram works perfectly. I’m glad I didn’t.

Scenario 2: Human-written

Pangram was right 21% of the time. It misclassified 79% of my human writing as AI.

Scenario 3: Human-spoken

Pangram correctly recognized 78% of the meeting transcripts as human. Which raises an important scientific question: what exactly was I during the other 22%?

Scenario 4: AI-generated, human-edited

Pangram was right 98% of the time after one tweak. Good.

But it was wrong 93% of the time after 5-7 rounds of tweaks.

Did the text become more human? And if it did, what exactly does human mean here? AI-generated, then edited 5-7 times?

Pangram AI detection experiment results. Across the tested samples, Pangram detected pure AI text with 100% accuracy but correctly classified my human-written text only 21% of the time. Human-spoken transcripts reached 78% accuracy. For AI-origin text, detection fell from 98% after one human edit to just 7% after five to seven rounds of editing, a 91-percentage-point drop. These percentages describe this experiment’s samples, not Pangram’s overall statistical accuracy.

Pangram appears much better at detecting untouched model output than tracing authorship after substantial human intervention.

One additional finding came out of this experiment. Pangram on Substack and Pangram’s native tool sometimes gave different results for exactly the same text. Fine-tuning is one possible explanation. Different thresholds, or text handling could also explain it. I can't tell which from this experiment.

But I can tell you this: Saying that Pangram doesn't work is not accurate. Saying that Pangram works is not accurate either.

And most importantly, if anything, this experiment proved that READING is still the best way to evaluate whether something is worth reading. As my grandma used to say, if you taste the soup and still can’t tell if it’s any good, we have a bigger problem than the soup.

Hey, I’m Karo Zieminski 🤗.

AI PM and builder. I write Product with Attitude, an AI newsletter for tens of thousands of readers across 146 countries, helping them develop critical AI literacy the only way it sticks: through practice.

If you’re new here, welcome! Here’s what you might have missed:

Subscribe for free, and I’ll keep bringing you deeply researched,
human-written articles like this.

SUBSCRIBE

And if you share this article with 3 friends or colleagues, I’ll thank you with 3 months of a premium subscription.

Refer a friend

Read the original on karozieminski.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.