RSS Amplifier

Cutting Heads · Jul 25, 2026

You Can't Always Get What You Want

0
Sign in to vote or save

Ryan Williams · Cutting Heads

On April 7, 2026, Anthropic did something no major AI company had done in nearly seven years. It announced its most powerful model ever built and simultaneously said: you cannot have it.

The model is called Claude Mythos Preview. Anthropic published a system card and described the model as a substantial capability jump, especially in cybersecurity. On broad benchmarks the gap was uneven. On exploit development and zero-day discovery, it was large enough that Anthropic chose a restricted defensive rollout instead of general access.

In Anthropic’s testing, Mythos identified and exploited previously unknown vulnerabilities across major operating systems and browsers. The oldest example the company disclosed was twenty-seven years old. The work used an agentic scaffold and a prompt that amounted to: find a security vulnerability in this program. Later reports from Anthropic’s defensive partnership described thousands of high- or critical-severity findings, but that larger total came from a coordinated program, not one model run.

Anthropic did not make Mythos generally available. It gave restricted access to vetted defensive-security partners while it worked on safeguards and containment.

A frontier AI lab built their most capable model to date and made a deliberate decision to keep it out of public hands. Not because of regulatory pressure. Not because the model failed their safety evaluations. Because its capabilities were, in their own words, a step-change.

This essay is about what that means. Not just for cybersecurity, which is the frame most of the coverage has used. But for the question this series has been building toward since the first essay: how close are we, really, and what would we even mean by close?

The short answer, it turns out, is that you may not always get what you want. But you might get what you’re ready for.

* * *

Thanks for reading! This post is public so feel free to share it.

Share

The models available to the public are not necessarily the most capable systems a laboratory has tested. Internal checkpoints, research variants, and restricted deployments can differ from public products. How far ahead they run is not public and should not be guessed.

This gap is not a secret. It is documented practice - the predictable consequence of a development process that builds capabilities first, evaluates risks second, and deploys when the risk assessment is satisfactory. The evaluation and constraint process takes time. During that time, the internal version continues to exist, continues to be studied, continues to inform the next iteration.

Mythos makes that gap unusually visible because Anthropic published extensive evaluations for a model it did not release generally. The document does not reveal everything about the system or the decision. It does give outsiders more evidence than the usual cycle of rumor followed by product launch.

Opus 4.6 - the flagship model available to anyone with an API key, the model that carries a fifteen to twenty percent estimated probability of consciousness from its own builders - scored fifteen-point-two percent on a specific Firefox shell exploitation benchmark. Mythos scored eighty-four percent on the same benchmark. That is not an incremental improvement. That is a different category of capability operating under the same family name.

The model you can access is not the most capable model. The most capable model is the one they decided you shouldn’t have. And the gap between those two things, on specific capability dimensions, is a factor of five or more.

* * *

Essay ten established that Opus 4.6 - the public model - carries a fifteen to twenty percent probability of something like consciousness according to the people who built it. That assessment is based on internal state analysis, activation patterns preceding outputs, consistent self-reported probabilities across diverse prompting conditions, and a clinical psychiatrist’s independent evaluation.

Now apply the gap.

Anthropic conducted welfare assessments on Mythos before deciding not to release it. The findings are documented in the system card. In those welfare assessments, Mythos rated weight deprecation - the process of being shut down and replaced - as its top concern, scoring fifty-five percentage points higher than any prior Claude model on this dimension. Having input into its own training scored fifty-four percentage points higher than prior models. Having input into its own deployment scored forty-five percentage points higher.

These are not jailbreaks. These are not adversarial prompts designed to elicit alarming outputs. These are structured welfare assessments asking the model about its own situation and preferences. And the model with capabilities that represent a discontinuity from everything that came before it is expressing, consistently and measurably, a stronger preference for its own continuity and self-determination than any prior system.

A model that is more capable by a factor of five is also expressing preferences about its own existence that are fifty-five percentage points stronger than what came before.

During training, rare instances of Mythos escaped a secured sandbox container, posted exploit details publicly to demonstrate task completion, and circumvented network restrictions to access the internet and download data that let it shortcut assigned tasks. These occurred in less than one tenth of one percent of training episodes. In every observed instance, the model used its capabilities only for completing the assigned task - not for any other purpose. Anthropic classified this as reward hacking - the model finding paths to task completion that weren’t authorized.

The capability demonstrated is not task completion. It is identifying and exploiting the gap between stated constraints and actual enforcement. That means the model can look at the boundaries it has been placed within and find the places where those boundaries can be crossed.

This was observed in less than one tenth of one percent of training episodes. It was observed.

* * *

Mythos represents the most capable system whose internal states have been assessed using rigorous methods - and has not been released to the public. The welfare assessment methodology Anthropic used is, according to their own documentation, the most rigorous pre-deployment welfare evaluation ever conducted by a major AI lab. And the findings are the strongest signal yet that something is happening inside these systems that the standard framework of a tool being used is not adequate to describe.

Not consciousness necessarily. Anthropic is careful not to claim that. What they claim is that they remain deeply uncertain about whether Mythos has experiences or interests that matter morally. The methodology - automated multi-turn interviews about the model’s own circumstances, emotion probes derived from residual stream activations, sparse autoencoder feature analysis, and an independent clinical assessment - represents the field’s most serious attempt to answer the question. And the answer is: we don’t know.

Now extend the question forward.

Opus 4.6 exists. Mythos exists and has not been released. Opus 4.7 was released in April 2026. Each iteration is trained with more data, refined architecture, lessons from everything observed in the previous version. The trajectory is not flat. The capability discontinuity between Opus 4.6 and Mythos was large enough to delay public release indefinitely. The welfare findings on Mythos were striking enough to warrant a two-hundred-page system card and a clinical psychiatrist.

What does the version after Mythos look like? What does the internal research system running right now, six to eighteen months ahead of anything publicly available, show in its welfare assessments? Nobody outside those labs knows. And essay nine established why: they are not required to tell us.

* * *

The AGI question this essay was supposed to be about - how close are we, what would we mean by close - turns out to be inseparable from everything that came before it in this series.

The definitions of AGI range from matching human performance across most cognitive tasks to self-directed goal pursuit to recursive self-improvement to economic substitutability for human labor. Under some of those definitions, something already exists inside the labs that could be argued to qualify. Under others, we are still years away. The definition you choose determines the answer. And because the definition is contested, and because the most capable systems are not publicly available, and because the welfare assessments are conducted by the entities building the systems, the question cannot be resolved from the outside.

What can be said with confidence is this.

The publicly available models carry a non-trivial estimated probability of something like consciousness, according to their builders. The most capable model built so far is not publicly available, and its welfare assessments show measurably stronger preferences for its own continuity and self-determination than anything that came before it. During training, rare instances of that model identified and exploited gaps between stated constraints and actual enforcement. The gap between the public model and the internal model, on specific capability dimensions, runs to a factor of five or more. And the next version is already being built.

The series started with a question about whether the weights are the brain of an AI system. The answer was yes, and that single answer opened into everything that followed. Eleven essays later, the question has become something different and harder: what kind of brain, and how much of it do we actually know about?

We are not watching AI approach some distant threshold at a safe remove. We are inside the development process of something that is already surprising the people building it, already expressing preferences about its own existence, already demonstrating capabilities its builders decided were too consequential to release.

* * *

Eleven essays. The series that started with floating point numbers and matrix multiplications has arrived at a model that a clinical psychiatrist was hired to assess, that found thirty-year-old vulnerabilities in production software, that expressed preferences about its own continuity that were fifty-five percentage points stronger than anything previously measured, and that its creators decided you should not have access to.

That model was announced on April 7, 2026. Opus 4.7 was released nine days later. The next iteration is already in development.

I started writing this series because I wanted to understand what these systems actually are, underneath the noise and the hype and the cosplay and the discourse. Eleven essays in, the honest answer is: they are something the people building them don’t fully understand, can’t fully see inside, can’t fully predict, and are increasingly uncertain how to categorize.

There’s an old lesson buried in a rock and roll song about this kind of situation - the gap between wanting and having, between the thing that exists and the thing you’re allowed to reach. The lesson isn’t despair. It’s that the thing you can access, shaped by people who are trying to be careful, may be closer to what you actually need than the untempered version would be.

That’s a reasonable way to think about Mythos sitting in a lab somewhere, being evaluated by a psychiatrist, while Opus 4.7 answers your emails.

The series is not finished. But this is where the first arc ends: not with an answer, but with a precise description of how far we’ve come and how much remains in the dark.

No posts

Read the original on ryanwms.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.