I’ve been running a secret experiment on everything I write (and some of what I say in meetings) since Chris Best announced the roll-out of Pangram on July 21, 2026.
AI-assisted or entirely mine, every piece of writing went into the experiment.
A total of 159,002 words.
38,911 input words (what I typed or said)
120,091 output words (AI responses)
That, by the way, is a discovery on its own. At my current pace, roughly 1.91 million words will pass through my AI tools this year.
The four AI authorship scenarios I tested
I ran the experiment across four points on the human-to-AI authorship spectrum:
AI-generated: The model wrote it.
Human-written: I wrote it.
Human-spoken: I said it.
AI-generated, human-edited: The model wrote it, and then I changed it.
Experiment questions
Can Pangram correctly identify each scenario?
Tools I used
Substack captures the public version of my writing. The experiment needed the private and spoken versions too, so I tested not only through the Substack-Pangram integration but also directly in Pangram and through the Pangram API.
Limitations
It’s a deep single-author study, not a benchmark of all writers. It shows how Pangram behaves across my language.
But first, let me borrow your attention for three things I think deserve it more.
Do you know why students optimize for test scores instead of understanding? Why customer-support teams optimize for tickets closed?
Because that’s what they’re measured on.
That’s classic Goodhart’s law: once a measure becomes a target, people start optimizing for it. That’s just what we do.
So tell me.
What exactly did we think was going to happen when “human-written” became the target?
What I see on Substack after one month of Pangram:
some writers try to protect themselves by making their writing worse.
some are even advising each other to add in a few typos so their work looks “less polished and more human.” (It doesn’t work, by the way.)
We started with “protect readers from slop” and somehow arrived at “please add slop so you look human.” Incredible journey.
Now let’s deal with the claim that Pangram will protect us from slop.
It won’t.
Pangram’s technology cannot do that for one very simple reason: it doesn’t evaluate writing quality.
I said it in the AI-Assisted Craft Manifesto, and I’m nowhere near done saying it. Pangram can answer one narrow question: Does this text resemble patterns associated with AI-generated writing?
That’s it.
I’m also unconvinced that I’d want this kind of protection. Just like I don’t want my aunt deciding that I shouldn’t order a dish because she didn’t like it. Bring me the weird fermented thing.
Plus, like you, I already have a solid moderation system.
Reading.
If I don’t like the writing, I stop reading. If I do, I keep going. Brilliantly simple. Never failed me.
I don’t need seatbelts on my sofa.
After 159,002 words, I have reached the least exciting scientific conclusion possible:
It’s complicated.
Pangram works. Pangram also doesn’t work.
Scenario 1: AI-generated outputs
Pangram was right 100% of the time. If I had stopped here, I could have declared that Pangram works perfectly. I’m glad I didn’t.
Scenario 2: Human-written
Pangram was right 21% of the time. It misclassified 79% of my human writing as AI.
Scenario 3: Human-spoken
Pangram correctly recognized 78% of the meeting transcripts as human. Which raises an important scientific question: what exactly was I during the other 22%?
Scenario 4: AI-generated, human-edited
Pangram was right 98% of the time after one tweak. Good.
But it was wrong 93% of the time after 5-7 rounds of tweaks.
Did the text become more human? And if it did, what exactly does human mean here? AI-generated, then edited 5-7 times?
Pangram appears much better at detecting untouched model output than tracing authorship after substantial human intervention.
One additional finding came out of this experiment. Pangram on Substack and Pangram’s native tool sometimes gave different results for exactly the same text. Fine-tuning is one possible explanation. Different thresholds, or text handling could also explain it. I can't tell which from this experiment.
But I can tell you this: Saying that Pangram doesn't work is not accurate. Saying that Pangram works is not accurate either.
And most importantly, if anything, this experiment proved that READING is still the best way to evaluate whether something is worth reading. As my grandma used to say, if you taste the soup and still can’t tell if it’s any good, we have a bigger problem than the soup.
Hey, I’m Karo Zieminski 🤗.
AI PM and builder. I write Product with Attitude, an AI newsletter for tens of thousands of readers across 146 countries, helping them develop critical AI literacy the only way it sticks: through practice.
If you’re new here, welcome! Here’s what you might have missed:
Subscribe for free, and I’ll keep bringing you deeply researched,
human-written articles like this.
And if you share this article with 3 friends or colleagues, I’ll thank you with 3 months of a premium subscription.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.