RSS Amplifier

Designing Intelligence · Aug 4, 2026

The Wrong Game Comes for Writing

0
Sign in to vote or save

Jason Prunty · Designing Intelligence

Two weeks ago Substack gave every reader a button that estimates how much of any post was written by AI. I did what I suspect a lot of writers did in private. I ran my own work through it.

One hundred percent AI. Fair enough about the words on the page. It says nothing about what it took to get to them, and none of that started with a thin prompt.

So I set out to beat the grader.

One hundred percent AI. Fair enough about the words on the page. It says nothing about what it took to get to them, and none of that started with a thin prompt.

Detectors like Pangram, the one Substack runs under the hood, are pattern readers. They respond to texture: sentence lengths in too narrow a band, paragraphs that all resolve with the same tidy cadence, transitions that glide where a person would lurch. Once you know the grader’s taste you can write for it. It took me a couple of hours. I varied rhythms, roughed up transitions, let paragraphs end without a bow. And I ran the passes through several different AI tools, which is the joke at the center of all this. I was feeding text to models to make it stop looking like it came from a model.

The number went to 64, then 22. The red robot turned green. Mostly Human.

I won.

Winning made me ambitious, which is how most bad ideas start.

What I had was a repeatable process. Diagnose the tells, apply the moves, watch the number fall. A repeatable process wants to be software, and not only for me. Every writer sitting under this new scarlet letter is watching a percentage they did not ask for get attached to work they were proud of. The thing practically designs itself.

I got as far as sketching it. Then I went back and read what the process had already produced.

The piece was not better. The argument had not moved an inch, and the styling got worse so it would read as human-made. I had a reliable method for making prose look like a person had struggled with it, applied to prose a person had already struggled with. It could lower a score. It could not make one idea easier to understand or supply a fact that was missing. Software would have scaled exactly that.

It could lower a score. It could not make one idea easier to understand or supply a fact that was missing.

The ceiling was obvious too. Everything I had learned is learnable by a model, and models had already done most of the work. Every signature a classifier keys on becomes a training target the moment it is known. Detectors tighten, the process adapts, detectors tighten again, and the only thing that grows is the effort everyone spends on texture. I never built it. I spent the next few days on a better question instead.

In May I wrote about Princeton ending 133 years of unproctored exams. Faced with new capability loose in the world, the university tightened the rules of a game it had already decided was the right one. Two months later a publishing platform reached for the same fallback. Not the same instrument. Princeton’s proctor is coercive, and it works, since you really can stop a student using a phone by watching the room. What the two share is the reflex underneath: when the thing you care about gets hard to measure, you measure the conditions of production instead.

The goal is the right one. Slop is real, feeds are filling with competent-sounding weightless prose, and readers came to Substack to escape exactly that. Chris Best says the company is “not against people using AI to assist their work,” and there is a new “How I make this” statement where authors describe their own process. The stated posture is transparency.

But what shipped is a scanner. A reader runs a check on a writer’s post and gets a percentage back. Writers are enrolled by default, and a writer who switches it off gets “AI detection unavailable” printed where the number would have been, so declining the scan is itself a visible act. Whatever the intent, that is the shape of surveillance. A nutrition label you decline to print is simply absent. A scan you refuse announces the refusal.

A nutrition label you decline to print is simply absent. A scan you refuse announces the refusal.

The backlash came fast. In 404 Media, Sam Illingworth called it a witch hunt and worried about non-native English speakers and neurodivergent writers whose ordinary prose trips the meter. Pangram’s CEO puts the false positive rate around one in ten thousand, which sounds small until you multiply it by every post on the platform, every week.

I sit inside that group, though not the way you would expect. I am a neurodivergent writer. Getting from my thinking to the sentences you are reading is the slow part, slow enough that an unassisted piece would take months, which in practice means it would not exist. AI did not make my writing faster. It made it possible. So the meter is not misfiring when it reports that I leaned on a machine. It is correct, and being correct is the least interesting thing about it.

Suppose the detector never misfired at all. Suppose it could tell you with certainty which sentences a model produced. What would you know?

Something about manufacture. Not provenance, though the two get treated as one thing. Provenance is the history of a work: who made it, under what conditions, with what care. Manufacture is the production method, and a classifier can only estimate it in three coarse grades: human, partly human, machine. It cannot tell you whether a person thought the thoughts, or edited, or stands behind the result.

Manufacture also does not tell you value. People wrote oceans of slop before language models existed, and some of the most useful things I read this year were written by people leaning openly on AI. You cannot find better things to read by inspecting how content was made.

Goodhart’s Law, arriving on schedule. When a measure becomes a target, it ceases to be a good measure, in Marilyn Strathern’s compact version. Install a score on a platform full of ambitious people and the score becomes the work. I know. I had a working method and came close to shipping it to everyone else. Goodhart is usually where this conversation ends. I think it is where the useful part starts.

The abandoned app left a better question behind: what feedback about AI-era writing would a writer and a reader both actually want? A writer wants to know whether the work is any good and where it is thin. A reader wants to know whether the next fifteen minutes will be worth spending. Neither is asking about manufacture. What both are circling is originality, and this took me longest to accept because it undercuts the app I had just talked myself out of. If a piece holds a point of view that could not have been assembled from what already exists, the texture is beside the point. A distinctive argument written with heavy AI assistance is worth more than a derivative one typed by hand at midnight, and the detector scores those two backwards every time.

I put a number on this back in June, talking with Christopher Wink for a Technical.ly feature on whether AI will ever make great art. The tools have made it easy to get eighty percent of the way to a competent picture. The remaining twenty, concept, context, intention, narrative, is where the meaning lives, and it has gotten harder precisely because everyone can now reach the eighty.

Writing is in the same position and got there faster. Ask a model for an essay on any subject and you get the consensus of the written record, competently arranged. That is the eighty, the same eighty anyone can assemble from a model and a Google search, which is why it is suddenly worth so little. A capability that becomes ambient stops being the craft, which is what photography did to painting and what I traced at length in The Recurring Revolution. The detector is policing texture at the exact moment texture is becoming the least meaningful place to look.

Some of the twenty lives in voice. Arrangement, timing, what you leave out. Any writer would laugh at me for pretending otherwise. But voice is contested ground now, and this essay has already shown why: two hours of work moved a document from machine-read to human-read without changing a single idea inside it. Style is what these systems absorb fastest. The part of the twenty they cannot reach is the material.

The model's substrate is the written record, and this material is not in the record yet. You were there. It was not.

The not-yet-written is less mystical than it sounds. It is the meeting where you watched a decision go wrong for a reason nobody named out loud. The user you saw fail with an interface in a way no dashboard recorded. The small experiment you ran because you were annoyed. No model can produce this, and not for want of some human spark. The model’s substrate is the written record, and this material is not in the record yet. You were there. It was not.

That is the only real claim this essay has on your attention. Until I beat a detector and walked away from the product that would have followed, the experience did not exist to be written about. Going somewhere and coming back with a report is the oldest job description in writing, and the model has made it the whole job by taking over everything else. AI is very good at the eighty: retrieving, structuring, pressure-testing, tightening. The twenty is a supply problem, and the discipline is unglamorous. Keep field notes. Run small experiments. Stay in rooms where things happen. A writer with a notebook full of firsthand observation and a model to organize it holds a real advantage. A writer with only the model holds what everyone holds.

Goodhart usually gets read as a warning: be careful what you measure. Read that way it sends you hunting for a proxy nobody can game, and that hunt is lost from the start. Motivated people game everything. I took apart a state-of-the-art classifier in an afternoon with tools anyone can rent by the month.

The way out is a test worth cheating on.

Ask of any measure: what does a person have to do to raise their score? If the answer is something other than the thing you wanted, the measure is broken no matter how sophisticated. If the answer is the thing you wanted, the people gaming it are doing your work for you. I found this in a classroom first. Inside the Right Game came down to peer-uptake: score a student on whether the rest of the team built on their contribution, and faking that over a quarter requires most of the behaviors of learning the material.

A detection score fails the question completely. Raising your number means doing what I did in that afternoon, which produced a worse piece with a better score. Gaming is the only available move, because the measure was never attached to anything worth having.

Originality passes, but only built with two halves. Distance from consensus is not the same axis as good, and the cheapest way to sit far from any baseline is to be strange or trivial. So the measure has to hold two questions at once: how far does the argument sit from what already exists, and is the departure any good, scored apart so neither can cover for the other. Clear both and the shortcuts run out. To move that number you have to find something the model does not have and be right about it. A writer optimizing hard for that score is a writer out gathering material, and a platform full of writers gaming it is a platform where more gets said than was said before.

Goodhart is usually where this conversation ends. I think it is where the useful part starts.

So that is what I built instead, at isthisoriginal.com. It takes a piece of writing, extracts the thesis and its claims, and builds the eighty on purpose out of search results and model output. Then it shows where your argument stops being conventional, names the claim or connection that departs most, and puts the nearest precedents beside it. There are numbers underneath, distance from the baselines, prompting effort, and, scored separately, quality. They exist to locate the work rather than grade it.

It is not a verdict, and cannot be. A writer who fabricates the meeting or the experiment sits outside both baselines too, so verification and reputation keep doing the work they have always done. What it can do is show a writer where they stand while the work is still moving.

The Wrong Game argued that education should stop hiring proctors and start designing like dungeon masters, reading play in progress instead of grading a finished artifact. Publishing has less excuse than education here, because publishing already knew. Editors were never manufacture-checkers. A good editor asks where you got that, what you saw, who told you, what you are adding to the pile. The honest version of this argument might end there: editors, reputation, readers who leave. That machinery worked for a century and it does not reach a feed. No editor stands between writer and reader on Substack, which is the entire proposition of the place. Something will fill that gap. The only live question is what it asks.

I raised this worry about students first, and it applies here too. A model will argue with you, hold a structure still while you rearrange it, show you where the reasoning goes thin. It will also let you make more than you otherwise could, and reach versions you would never get to if every one of them had to come out of your own typing. A writer who backs away from all that to keep a clean number has declined a tool that changes how the work happens. And the mark lands hardest on whoever is least willing to game it. I proved in an afternoon that it comes off for anyone who tries.

So I am publishing this one with detection switched on.

It will not score well. I have used AI throughout, the way I always do. I could fix that in an afternoon, the process is written down in a file. I could switch detection off, but the off switch prints its own small notice, and I would rather be scored badly than be seen declining to be scored.

Let it say what it says. It will probably be right about how the words were made. What it cannot tell either of us is whether we got somewhere worth going, you and I, in the time you just spent. That is the only question worth answering, and you have already done the one thing that answers it.

Sources and attributions:

Jason Prunty, The Wrong Game: Why AI Needs Dungeon Masters, Not Proctors, Designing Intelligence (May 2026). Jason Prunty, Inside the Right Game, Designing Intelligence (June 2026). Jason Prunty, The Recurring Revolution, Designing Intelligence (2025).

The 80/20 framing comes from Christopher Wink, The 80/20 problem: Will AI ever create great art?, Technical.ly (June 28, 2026), edited by Danya Henninger, in which I am interviewed.

On Substack’s AI detection feature, launched for posts and notes published after July 21, 2026 and powered by Pangram: Substack support documentation.

On writer reactions, including comments from Sam Illingworth and Substack CEO Chris Best: 404 Media, Substackers Say New AI Detection Tool Is a ‘Witch Hunt’ (July 2026).

Goodhart’s Law: Marilyn Strathern’s compact form, “When a measure becomes a target, it ceases to be a good measure.”

On Princeton’s May 2026 faculty vote ending 133 years of unproctored exams: Inside Higher Ed, The Daily Princetonian.

Read the original on designingintelligence.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.