Some weeks, a new AI model lands before I've finished testing the last one.
Opus 5, Fable 5—and that's one company. Every release arrives with a loud crowd insisting it changes everything, and a quieter question none of them answer for you (least of all the people hyping it). Which one is right for what you're doing right now?
A couple of weeks ago, Anthropic posted its own answer. It was a four-card carousel on Instagram, titled "Should I change the model or the effort?"
The idea is, when you don't like what Claude hands back, you have two dials. Change the model when you need more raw capability. Raise the effort when you want it to work harder.
For this issue, I figured it'd be fun to test it with a real prompt—the kind of question you'd bring Claude on a normal working day—while I changed first the effort, and then the model.
What came back wasn't what the guide promised.
For this experiment, I reused a prompt from an issue I wrote last October, where I turned my bookshelf into an expert roundtable. You hand Claude a decision and a few nonfiction authors weigh in, argue with each other—simulated, of course—and talk you through it. (Claude channels each author from their books; the real people aren't involved and haven't endorsed a word of it.)
I made two changes to keep the test fair.
First, I stripped the prompt down to a single question. The original roundtable runs in multiple steps, the kind of prompt that asks you things and works through stages, and I wanted Claude to answer in one clean pass so I could line every run up against the others. The situation itself was made up—a friend weighing a startup offer.
Second, I ran everything in incognito, so Claude couldn't pull anything from my setup, and nothing from the test would carry into a future session. Beyond that, I kept everything identical and changed one thing at a time. First the effort, with the model held still. Then the model, with the effort held still.
Next, I started turning dials.
I started my experiment where most people start, on Opus 4.8 at High, the default you get without touching anything. I kept the model there the whole way; all I moved was how hard I told it to work.
At High, Claude gave me a good answer. It noticed the "decision" was really three tangled together—the job offer, the shift from designing to managing, and the studio dream she keeps deferring—and brought in five authors to work through them, Herminia Ibarra and Cal Newport among them. Balanced, thoughtful, thorough, and useful.
Then, I bumped it to Extra. The answer got a little sharper; it tightened the synthesis, cleaned up the through-line, and started borrowing decision-making tools to structure the advice. But the heart of it hadn't moved: the same read on the situation, mostly the same authors, saying mostly the same things.
Next, Max. This was the strongest of the three, and the gap was noticeable, but small. The panel got bolder. Nassim Taleb turned up with his barbell strategy, and Oliver Burkeman closed on the idea that every choice quietly forecloses the others. It also stopped refereeing. Where High laid the options out evenly, Max took a side and named the lead role as the one to be most skeptical of.
So, the effort dial did do something. By turning it up, I got a deeper version of the same answer, with bolder authors, a sharper opinion, and a cleaner structure. It was still the same advisor Anthropic described, just working harder.
To get a different answer, I'd have to change who was in the room.
For this part of the test, I held the effort at High and swapped the model, using the same prompt. As mentioned above, I started on Opus 4.8, the answer you already saw, then moved up to Opus 5, then across to Fable 5.
To my surprise, Opus 5 changed the cast.
Where 4.8 leaned on identity and the maker-to-manager question, Opus 5 brought in Paul Graham on the maker's versus manager's schedule and Morgan Housel on what the raise is actually for. It gave the sharpest answer of the whole test, and it found the line none of the others did. It called the lead role a paid apprenticeship in the real work of running a studio, the managing and clients and money she'd never touched.
I hadn't thought of it that way, and it was right.
Next, Fable 5. Again, a different room. Oliver Burkeman opened on the fact that every choice closes doors. Annie Duke pushed on how reversible the move really was. And the output, perhaps most usefully of all, ended on the simplest question yet: whether she'd feel relief or regret when she pictured saying no. Fable 5 was warmer, more compressed, and more philosophical than analytical.
So, the model did more than the effort dial did. It swapped the people in the room and changed the register, and Opus 5 found something the others missed. This was the closest the test came to a real difference.
Still, it wasn't much. Underneath the different casts, all of them were handing me the same advice in different outfits. They were good answers, but none of them was the leap the guide implied I'd get.
That's both dials Anthropic sells, turned as far as they go, and neither was the answer. What finally worked was something I'd got wrong in every prompt up to that point.
I'd run the whole test in incognito, with memory off, as the whole point was to have a clean room.
But then I ran Fable one more time with memory switched back on, with the same model and effort I'd just used, on the same prompt.
The only difference was that this time Claude could reach my knowledge base, the notes and highlights I've been saving from books for years.
(I've written about this at length, including why your second brain stays silent in every AI session, how to wire it into Claude, and why I tore mine down and rebuilt it three months later.)
Claude opened by telling me it was "pulling from your knowledge base," and it came back sharper than anything the dials had produced. It brought in Steven Pressfield to ask whether the studio was a real calling or a comfortable one she keeps safe by never testing it; it cited Paul Nutt's research showing that straight yes-or-no decisions fail 52% of the time; it named the kids' ages as the least reversible thing in the picture, the one cost the raise was quietly hiding.
The model hadn't changed, the effort hadn't changed, and the prompt was word for word the same. The only thing I'd added was the option to pull from my knowledge base, and it moved the answer further than either dial had.
Line the three up, and the pattern is hard to miss. Turning up the effort gave me the same answer, deeper. Changing the model gave me a different room and a sharper voice in it. Handing Claude my Second Brain gave me the answer that fit her situation.
Two of those dials belong to Anthropic; they cost money, and they shift with every release. But the third one is mine: the reading I've done, the podcasts I've listened to, the notes I've kept for years, sitting there waiting to be handed over.
And it's the one that works no matter where you're typing.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.