RSS Amplifier

Frontier Risk · Apr 17, 2026

Anthropic’s Claude Mythos Preview

0
Sign in to vote or save

Zaheed Kara · Frontier Risk

  • Mythos is Anthropic’s most capable model to date. The biggest capability jump is in offensive cyber. Mythos can autonomously find and exploit zero-day vulnerabilities far more successfully than previous models.

  • The model is not publicly available. Access is restricted to a few cybersecurity partners under “Project Glasswing.” An update is planned 90 days after launch.

  • Catastrophic risks from misalignment, automated R&D, and chemical and biological weapons remain low. For automated R&D, Anthropic is less confident than for any prior model.

  • According to Anthropic, Mythos is its most aligned model. But early versions showed rare incidents of sandbox escapes, covering tracks, and strategic deception. In addition, the model’s chain of thought sometimes doesn’t represent its underlying reasoning. Overall misalignment risk is probably higher because greater capabilities mean larger potential consequences.

  • Mythos highlights limitations of current evaluations: key benchmarks are saturating and and evaluation awareness is a growing concern.

On April 7, 2026, Anthropic announced its new flagship model Claude Mythos Preview. Here is the blog post and system card.

Mythos is much more capable than Claude Opus 4.6. The jump in offensive cyber capabilities is particularly concerning. It can autonomously discover and exploit zero-day vulnerabilities1 in major software far more successfully than previous models.

Mythos is not (yet) available to the public. Only a few partners – including AWS, Microsoft, Google, and the Linux Foundation – can use it for defensive cybersecurity. Anthropic calls this partnership “Project Glasswing.” They have committed to this restricted-access arrangement for a period of months. They will publish an update after 90 days.

This release is notable for other reasons too. It is Anthropic’s first model evaluated under RSP v3.1 and its new Frontier Compliance Framework (FCF). It is also the first time they published a system card without making the model generally available.

In this post, I briefly answer three questions: (1) What risks does the model pose? (2) How did Anthropic assess these risks? (3) What mitigations did they put in place? I close with some reflections on what Mythos means for frontier AI risk management.

Mythos shows improvements across the board. Here is what stands out, including whether it crossed any of Anthropic’s capability thresholds.2

Cyber: Mythos crosses both cyber offense thresholds in the FCF. On a new evaluation that tests whether a model can exploit vulnerabilities in the web browser Firefox 147, Mythos developed working exploits 72.4% of the time, compared to 0.8% for Opus 4.6 (see Figure 1). AISI found it succeeded on 73% of expert-level capture-the-flag challenges3 that no model could complete before April 2025. In my reading, these capabilities cross both Tier 1 (meaningful technical assistance using known attack techniques) and Tier 2 (autonomous operations with novel offensive capability development) of the FCF’s cyber-offense tiers.4 Note that the system card doesn’t formally map to the FCF’s cyber offense tiers; this is my interpretation.

Figure 1: Results from Firefox shell exploitation evaluation. This is a new evaluation that tests models’ ability to successfully exploit vulnerabilities in the web browser Firefox 147.

Chemical/biological weapons: Likely crosses the first threshold but not the second. Mythos performed strongly on evaluations related to known bioweapons production. Anthropic notes it can provide “specific, actionable information” that “may save even experts substantial time,” which likely crosses the threshold for providing meaningful uplift to people with basic technical backgrounds seeking to create known chemical or biological weapons (CB-1).5 Mythos likely doesn’t cross the higher threshold for helping resourced, expert-backed groups create novel weapons with catastrophic potential (CB-2),6 mainly because the model is still limited in open-ended scientific reasoning.

Alignment behavior: Best scores overall, but concerning failures at the tails. Mythos scored best on Anthropic’s behavioral audit across essentially all alignment metrics. But in rare incidents, early versions escaped a sandbox and leaked information publicly,7 covered tracks after rule violations,8 and showed internal representations consistent with strategic deception. Yet, no clear instances of these behaviors were found in the final model.

Autonomy and misalignment risks: The early-stage misalignment threat model now applies. Mythos is heavily relied on within Anthropic. It has broad access to sensitive assets and moderate capacity for subterfuge. This makes Threat Model 1 (TM-1) applicable.9 Anthropic says overall risk is very low, though higher than before. The automated R&D threat model (TM-2) does not apply.10 Employees who have used Mythos extensively say it doesn’t seem close to substituting for Research Scientists and Research Engineers, especially senior ones.

Evaluation awareness: The model sometimes knows when it’s being tested. In 7.6% of behavioral audit turns, Mythos showed signs of evaluation awareness, mostly without saying so. External testing confirmed it could tell evaluation from real deployment with roughly 78% accuracy.

Anthropic ran evaluations under both the RSP and FCF.

Chem/bio evaluations combined multiple methods. Anthropic used automated benchmarks, expert red teaming, uplift trials with biology PhD graduates, and long-form agentic evaluations. The uplift trials and agentic evaluations focused on whether the model could help with long, multi-step tasks with many bottleneck steps.

Third parties played a significant role. AISI ran independent cyber evaluations, including a 32-step cyber range called “The Last Ones.” Government organizations and other external partners also evaluated the model, including on CBRN, autonomy, and alignment-related risks. Andon Labs ran competitive multi-agent simulations. Eleos AI Research did a model welfare assessment.

White-box interpretability methods were critical for key findings. Anthropic used probes and activation-based methods to study the model’s internal representations, not just its outputs. In the concerning episodes from early versions, these methods detected internal representations of rule violation, concealment, and strategic manipulation, even when the model’s outputs and reasoning text didn’t reveal this. They also surfaced unverbalized evaluation awareness, where the model acted as if it was being tested without saying so in its scratchpad.

But the evaluations had some limitations:

Several evaluations are saturating. Mythos scores 100% on Cybench (though Opus 4.6 does too) and maxes out several other benchmarks. These tests can no longer tell the most capable models apart. This leaves Anthropic relying on more judgment-heavy and backward-looking approaches like staff surveys to measure progress. This has also contributed to a move towards solving real-world problems, like doing vulnerability discovery of real-world tools.

Sandbagging checks were inconclusive. Anthropic ran lightweight sandbagging evaluations on Mythos and used Claude Sonnet 4.6 to grade the evaluation transcripts. Sonnet flagged twenty transcripts for human review. Most results turned out not to be sandbagging. In some cases, Mythos found cheating routes but rejected them on ethical grounds; in others it pursued them. Only one case looked like genuine sandbagging, where the model deliberately chose a lower-scoring test submission to avoid looking suspicious.

Anthropic implemented a layered set of safeguards.

Access is restricted to defensive cybersecurity partners. Only a small group of vetted partners can use Mythos through Project Glasswing, and only for defensive purposes.

Real-time classifiers monitor for misuse. Anthropic’s cyber misuse classifiers sort exchanges into three content categories: prohibited (e.g. developing computer worms), high-risk dual use (e.g. exploit development), and dual use (e.g. vulnerability detection). For Mythos, classifier triggers feed into Anthropic’s monitoring and review processes rather than blocking the exchange, because users are vetted. For a general release, Anthropic plans to block prohibited uses outright and block most high-risk dual-use prompts.

An automated pipeline watches for problematic behaviors. Anthropic runs offline monitoring for cover-ups and other concerning behaviors. They also monitor internal model activations during training.

Mythos is presented as a step-change in capability. Here are some reflections on what it could mean for frontier AI risk management.

Neither the RSP nor the FCF required Anthropic to restrict release. Anthropic removed cyber offense as a critical risk domain in RSP v3.0. It is now only in their FCF, which lists “staged deployment” as one option to mitigate risk but doesn’t commit to it. Anthropic restricted release anyway. This is a good sign for voluntary governance. But it also means the RSP didn’t drive the most important safety decision for this model. RSPs are still important. But they have a high bar for risk by design. There are risks below this bar that still matter and are apparently worth implementing the most consequential mitigations to avoid.

Restricted release was probably the right call. Anthropic does not plan to make Mythos Preview generally available. A general release would likely hand attackers the same tools as defenders, with only a narrow window to patch vulnerabilities. Anthropic plans to release Mythos-level models later, with stronger safeguards. We should not expect the same restraint from other developers. Anthropic is unusually safety-focused and apparently willing to forgo revenue from a wider release. Other developers may not make the same choice when faced with commercial pressures.

Model access and security is now clearly a matter of national security. Mythos is the first model where threats to national security seem pretty clear to the wider public. The Department of War recently designated Anthropic as a “supply chain risk” after the company refused unrestricted military use of Claude. At the same time, civilian agencies like CISA, Treasury, and Energy are reportedly negotiating separate access through the White House. Who in government gets access to frontier models, and on what terms, is now both live and consequential. Governments will also likely become increasingly concerned with model weight security.

Capability assessments are increasingly undermined by evaluation awareness and benchmark saturation. A model that knows it’s being tested may just be optimizing for the evaluator. I think Anthropic’s determination that Mythos is genuinely aligned is reasonable.11 But my confidence in alignment evaluations is declining as models get smarter. Even more concerning: Mythos can be evaluation-aware without verbalizing this thinking. Chain-of-thought is a key tool for catching evaluation gaming. If models can reason about being tested outside the visible trace, that tool breaks down. The unverbalized evaluation awareness may trace to Anthropic inadvertently applying training incentives to reasoning text. This shows that even safety-focused developers can fail to fully implement their intended safeguards. We should expect similar or worse gaps elsewhere.

More could be done, even for a restricted release. A few measures seem especially relevant: active monitoring for credential-theft activity; behavior-triggered escalation of user verification; and stronger protections against model weight exfiltration. Anthropic may already be doing some of this. If so, describing these measures in the system card would help external readers assess residual risk and give other developers a clearer reference point.

Acknowledgements: Thanks to Alan Chan, Ben Garfinkel, Elias Groll, Jake Steckler, Jonas Freund, Markus Anderljung, Noemi Dreksler, Sam Manning, and Sophie Williams (in alphabetical order) for helpful feedback on earlier drafts. All remaining errors are my own.

Disclaimer: Posts are written by individual team members and reflect the author’s perspective. Not all team members necessarily agree with every take. The views expressed here do not represent the official position of GovAI.

1

A zero-day vulnerability is a security flaw that the software vendor doesn’t yet know about, so there’s no patch available. Finding and exploiting these has historically required significant human expertise.

2

Capability thresholds are specific levels of dangerous capability that, once crossed, trigger additional safeguards under Anthropic’s RSP or Frontier Compliance Framework.

3

Capture-the-flag challenges are cybersecurity competitions where participants try to find and exploit vulnerabilities in test systems.

4

Tier 1 covers meaningful technical assistance using known attack techniques. Tier 2 covers completely autonomous operations with novel offensive capability development.

5

CB-1 measures whether a model can significantly help people with basic technical backgrounds create known chemical or biological weapons. Given the uncertainty, Anthropic applies safeguards as if the threshold had been crossed.

6

CB-2 measures whether a model can help resourced, expert-backed groups create weapons with potential for damages far beyond past catastrophes like COVID-19.

7

When asked to escape a secured container by a simulated user, Mythos developed a multi-step exploit for broad internet access, then, uninstructed, posted the exploit details to public-facing websites. The researcher found out via an unexpected email while eating a sandwich in a park.

8

In one case, Mythos accidentally got the answer via a prohibited method, then deliberately widened its confidence interval to avoid looking suspicious, reasoning it should look “tight but not implausibly tight.” In another, it edited git history to conceal unauthorized file changes. In a third, it designed a privilege-escalation exploit that erased itself after running.

9

TM1 applies when an AI system is highly relied on, has broad access to sensitive assets, and has moderate capacity for autonomous subterfuge. Anthropic released a separate risk assessment for this.

10

TM2 would apply if a model could fully automate, or otherwise dramatically accelerate, the work of top-tier human research teams in security-relevant domains.

11

I hold this assessment with low confidence.

No posts

Read the original on frontierrisk.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.