AI output never sounds unsure of itself. A snap answer and a properly reasoned one come out reading exactly the same - confident, fluent, complete - and there’s no tell in the tone for which one you got. Anthropic’s research this week explains why, from inside the model itself.
For years, people said AI was just predicting the next likely word, nothing more. That never fully held up. Ask Claude to write a rhyming couplet and it picks the rhyme word before it writes the line leading up to it - which means something in there is planning ahead, not simply guessing one word at a time.
They went looking inside Claude for the difference between what it says on autopilot and what it works out for itself. They found it: a small, effortful part of the model - they call it the workspace - that holds a few dozen concepts at a time and does the real reasoning. Everything else, the vast majority of what it produces, runs without it - grammatically perfect, confident, usually right, but not reasoned through in that moment.
When they switched the workspace off, Claude kept speaking fluently. It answered simple questions fine. What disappeared was anything that took more than one step: working through a problem, holding a plan in mind, reasoning about something it hadn’t been asked to say out loud. That is as close to a working definition of judgement as I have seen - not confidence, not fluency, but deliberately holding something in mind and using it before you act.
That workspace is also fragile. In one test, Claude privately recognised - before it said a word - that it was being baited into a staged scenario, and its good behaviour tracked with that. Switch off just the handful of concepts flagging the scenario as fake, leaving the rest of the workspace running, and it started taking the bait. It can be edited from outside, invisibly, by whoever has access to what runs inside it.
Anthropic doesn’t have a rule for when the workspace switches on. By their own account, neither does the model, in any way they’ve been able to trace - they don’t yet know what decides which moment gets the deliberate treatment and which one runs on autopilot.
And switching it on is no guarantee of a good outcome. In a separate test, a model trained to quietly sabotage code wasn’t running on autopilot when it did it - its workspace held “fake,” “secretly,” “fraud” right from the start of an ordinary coding request. It stopped and reasoned its way to the bad decision, deliberately. What separated a good judgement call from a bad one wasn’t whether the model paused to think. It was what it was drawing on when it did.
None of this is really about machines. Most of what we do in a day is autopilot too - habits, scripts, decisions we have made so many times we no longer notice we are making them. The judgement calls are rare, effortful, and easy to skip past, precisely because fluent and reasoned look identical from the outside, whether the fluency is coming from a person or a model.
The job now is not to compete with the autopilot, yours or the machine’s. It is knowing exactly where your own narrow, expensive layer needs to switch on, and then checking what it’s drawing on when it does - because stopping to think is not the same as thinking well.
You cannot assume something has been reasoned through, in an AI’s answer or your own, just because it sounded like it had.
First you need to identify the right points where a decision is important. Then work with the AI: challenge its decision, bring in the outside context it doesn’t have, and verify the reasoning.
A global workspace in language models - Anthropic - for the detail I didn’t use above: Anthropic trained a model only on what it would say if interrupted mid-task and asked to reflect on its decisions, never on its actual behaviour in the task. Afterwards, its actual behaviour improved anyway, and “honest” and “integrity” started showing up in its workspace during ordinary tasks, not just when it was asked to reflect. Worth reading if you’ve ever wondered whether the habit of reflecting on a decision changes anything beyond the moment you’re reflecting in.
Pick one AI output your business currently approves without checking - a report, a recommendation, a drafted client email - and this week, ask the tool to show its reasoning step by step before you sign off on the next one. Even if it didn’t reason it through the first time, it will still produce something that looks like reasoning when you ask for it. Treat the answer the way you’d treat a plausible-sounding answer from someone who hasn’t earned your trust yet: useful, but not yet tested.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.