RSS Amplifier

Writing Secrets · Jun 5, 2026

Why Every New AI Model Looks Identical (Until You Push It Off a Cliff)

0
Sign in to vote or save

Christopher Kokoski · Writing Secrets

I got a comment on my latest Opus 4.8 video this week that stung: “It’s the same as before.”

And honestly? I get why someone would feel that way. New Claude, new version number, another author runs another test. If you’ve watched a few of these, a launch-day demo can start to blur together.

But that comment is pointing at something real — and it’s worth unpacking, because it’s the exact trap that makes it nearly impossible to tell whether a new AI model is actually any better than the last one.

Here’s the not-so-secret secret of most “new model” demos, including plenty of mine: they test at easy mode.

Write me a premise. Outline this. Draft a chapter. At that level, GPT, Gemini, this year’s Claude, last year’s Claude — they all hand you something competent. They all look the same. Because the easy tasks stopped being hard for these models a long time ago.

So when a new version drops and someone runs it through the usual paces, of course, it feels the same. The test isn’t capable of showing a difference. You’re timing two sprinters with a stopwatch that only counts in whole minutes.

If there’s a real jump in capability, easy mode is the last place you’ll see it.

If you want to know whether a model actually improved, you don’t ask it to do something easy slightly better. You ask it to do something it should fail at — and watch how it fails.

That’s the whole reason I didn’t run another chapter demo this time. I asked Opus 4.8 to write a horror musical — story, songs, meter, sustained dread, a separate voice for every character, two genres that openly fight each other — all at once. It’s the hardest creative test I know how to run, and I break down exactly why in the video.

A test like that separates “they all look the same” models in seconds. Easy mode hides the gap. The edge exposes it.

This isn’t really about my video. It’s a tool you can use every single time a new model launches and the whole timeline starts arguing about whether it’s “actually better.”

Don’t judge it on the easy stuff. Judge it on the one thing your writing genuinely needs and that AI usually botches — your hardest 5%. For most novelists, that’s one of three: holding a distinct character voice across a whole book, sustaining tension without it going flat, or matching a specific prose style instead of defaulting to beige.

Run that test. If the new model is genuinely better, that’s where you’ll feel it. If it isn’t, you just saved yourself an upgrade you didn’t need.

So…was the new Claude actually different, or was that commenter right? I’d rather show you than tell you. The full test, warts and all, is here:

Watch it here →

Watch where it cracked and where it genuinely surprised me, and decide for yourself.

— Christopher

No posts

Read the original on writingbeginner.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.