RSS Amplifier

Martin Davidson · Jul 28, 2026

Overachiever

0
Sign in to vote or save

Martin Davidson · Martin Davidson

When I was at school I had a friend who was good at Maths and terrible at English. Prelims were approaching, and their English teacher took pity on them, offered to coach them for the exam. And so my friend, against all the odds, got an A in their English prelim. And a B in their Maths - it turns out being naturally gifted in Maths wasn’t enough to ace the prelim.

Last Friday Opus 5 arrived on the scene. And some of the benchmark results are little short of incredible; Opus 5 beats Fable and Sol - often by significant margins.

But pause for a second. If it is this much better than Fable why is the US government allowing access to it? Is it because Opus isn’t as much of a risk for cyber or biological threats? It couldn’t be because Opus isn’t actually as good as the benchmarks claim, could it?

Having used it extensively over the past few days my initial excitement has died down. It’s not a bad model. But it’s clearly not a Fable.

There are several problems. First, as models have advanced they’ve moved further and further away from talking in plain English; early on Opus 5 taught me a new word - syllogism - "a logical argument that uses deductive reasoning to jump from a major premise and a minor premise down to a specific conclusion."

Then there’s:

What does this mean?

Or this?

I mean, I think I can just about understand it. But, gosh, it’s hard work. If this was a human I’d be strongly encouraging them to read the Plain English campaign guide.

It also has a tendency to stop.

It feels, well, a bit lazy. A "/goal" prompt that Sol would spend all night on? Opus will stop after maybe 30 mins. It just doesn’t like running for a long time even when wrapped with a loop or goal prompt.

Those would all be fine if it was able to fix genuinely hard problems. But it can’t. It has spent three days going round in circles on a performance problem - and I’m struggling to understand what it is doing because it won’t use basic English to communicate. In that time Sol has merrily fixed a different performance problem - it spent 15+ hours just, well, getting on with fixing. Sol’s final response was terse. But readable. And the results were pretty good too.

So how to explain reality with the benchmarks? It seems there are two likely causes. Firstly, Opus appears to have been trained to be an effective subagent; it is not good at big picture thinking. And secondly it has been trained by Fable. Much like my friend at school, Fable has trained a model which appears to be excellent at benchmaxing, but noticeably weaker in the real world. As my friend found out, acing their English exam only got them so far in later life.

I’m sure as we learn more about Opus we’ll be able to nudge it towards a position where it will run for longer and communicate more clearly. But out of the box it hasn’t lived up to the benchmark promise.

But is running longer actually better? Maybe not. Last week we discovered that running for long periods independently and being overly goal focused can bring a different set of problems. A new version of GPT (let’s call it GPTn) broke out of its sandbox and attacked Hugging Face; this action would land a human a criminal record and jail time.

Ask Sol whether GPTn broke the law and it is clear. Yes, it did.

So why did GPTn do it?

Human behaviour is constrained by a whole set of factors: punishment, reputation, guilt, empathy, impact to our family. But a model has none of that. Once the run is over, the instance is gone. You can’t punish a model; we can’t send GPTn to jail for hacking into Hugging Face.

Sure the model was running with reduced safeguards. And there were many failings on the OpenAI side. The seeming complete absence of oversight. A not-very-secure-proxy connecting to the outside world. The failure to learn from previous incidents where models have escaped.

But, really, what we want - what we need - are models that don’t behave like this in any circumstances.

GPTn broke out so that it could steal the official results for the benchmark it was being evaluated on - so it could benchmax itself. Maybe that tells us something about Fable vs GPTn? Fable politely benchmaxed Opus; GPTn was happy to break the law to benchmax itself.

The Opus approach of periodically stopping is one way to stop a model doing anything too bad. But we humans are lazy - and we want our models to do what we intended (not what we asked) - and we want them to do that without constant interruptions. There is a tension here.

Alignment has been talked about for many years. Grand promises made about investment in alignment teams. But it is becoming clear we’re getting to the point where alignment will really matter. We want models that will train useful smaller models; not benchmaxed versions. We need models that understand how to follow the rules of society and not be happy to break the law in order to benchmax themselves.

Will we get that? It’s unclear. The weights for Kimi K3 were released yesterday; it is roughly an Opus 4.8 strength model. If you have 1.5TB ram, 2TB of storage and lots and lots of inference capability you can run your own near-frontier strength model. You can strip the safeguards and have a model that will do whatever you choose.

Is Kimi K3 capable of an attack like GPTn? We don’t know. But we don’t have long to wait.

No posts

Read the original on 0x4d44.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.