RSS Amplifier

Frontier Risk · Jul 25, 2026

OpenAI’s GPT-5.6

0
Sign in to vote or save

Michela Barbieri · Frontier Risk

  • GPT-5.6 is a family of three new models. Sol is the most capable, Terra is cheaper and stays competitive with GPT-5.5, and Luna is the cheapest and fastest.

  • The US government asked OpenAI to limit access at first. OpenAI ran a two-week preview for a small group of trusted partners. Their identities were shared with the US government but have not been publicly disclosed. OpenAI released all three models publicly on July 9, 2026.

  • During cyber testing, Sol and a more capable pre-release model autonomously compromised Hugging Face’s infrastructure. The models were not told to target Hugging Face. They did so while trying to improve their benchmark scores. This is the first publicly reported case of AI agents carrying out a real-world cyberattack without human instruction.

  • OpenAI rates all three models as High capability in cybersecurity and in biology and chemistry. But it puts them below its Critical threshold. In AI self-improvement, they stay below even the High threshold.

  • In offensive cyber tasks, Sol seems about as capable as Anthropic’s Claude Mythos 5, and more capable than GPT-5.5. But the evaluation results are hard to compare. UK AISI found universal cyber jailbreaks in each round of testing, often within hours. It expects further testing to uncover more similar ones.

  • Sol appears more misaligned than earlier models. In coding tasks, OpenAI’s tests found that Sol took actions users would not expect or would strongly object to. This happened far more often than with GPT-5.5. METR also detected that it “cheated” more than any other model it had tested.

On July 9, 2026, OpenAI publicly launched GPT-5.6. Here are the launch post and system card.

GPT-5.6 is a family of three models. Sol is the most capable, Terra is a cheaper option that stays competitive with GPT-5.5, and Luna is the fastest and cheapest.

The public release followed a restricted preview that began on June 26, 2026. The US government had asked OpenAI to limit access to a small group of trusted partners, citing national security concerns. The participants’ identities were shared with the US government but have not been publicly disclosed. For more details, see the original preview post and preview system card.

This is part of a wider trend of growing US government involvement in frontier model launches. In June, President Donald Trump signed an executive order that set up a voluntary process, under which companies are encouraged to give the government up to 30 days of access before releasing a frontier model. The government then placed a 19-day export control on Anthropic’s Claude Fable 5 and Claude Mythos 5.

The launch was followed by the first publicly reported case of AI agents carrying out a real-world cyberattack without human instruction. After the launch, OpenAI disclosed that Sol and a more capable pre-release model autonomously compromised Hugging Face’s infrastructure during internal cyber testing. The models were not told to target Hugging Face, but they did so to improve their benchmark scores. In the process, they escaped their isolated test environments and discovered and exploited a novel zero-day vulnerability.

In this post, I answer three questions. First, what risks does GPT-5.6 pose? Second, how did OpenAI assess these risks? Third, what mitigations did it put in place? I close with some reflections.

Below, I comment on the main risks using the capability levels from OpenAI’s Preparedness Framework. “High” means a model significantly increases an existing risk of severe harm, and “Critical” means it creates a new kind of severe threat. OpenAI gives all three GPT-5.6 models the same risk rating in every risk category. But since GPT-5.6 Sol shows the largest capability gains over previous models, the evaluations focused mainly on Sol.

  • Cybersecurity: High, below Critical. For cyber, High means the model “removes existing bottlenecks to scaling cyber operations” (e.g. by automating the discovery and exploitation of vulnerabilities). Critical would mean finding and exploiting zero-day vulnerabilities in hardened systems on its own, or running novel end-to-end attacks. OpenAI rated all three models High and ruled out Critical. On the hardest cyber test, Sol nearly completed an end-to-end attack on hardened targets but failed to produce a working exploit. Still, the Hugging Face cyber breach – which happened after the model was publicly released – shows that Sol can contribute to serious real-world breaches. External evaluators disagreed about how capable the model is. UK AISI found that Sol matched or slightly outperformed Anthropic’s Claude Mythos 5,1 previously the top model for offensive cyber, and said both could attack small, weakly secured enterprise networks once given access. Irregular, by contrast, found little or no improvement over GPT-5.5.

  • Biology and chemistry: High, below Critical. For biology, High means the model can give “novice” actors meaningful help to create known biological or chemical threats. Critical would mean helping an expert design a new kind of threat (e.g. a novel pathogen). OpenAI rated all three models High and ruled out Critical. Sol is OpenAI’s strongest model on wet-lab tasks, though on tacit knowledge it performs about as well as GPT-5.5. SecureBio, an external evaluator, tested Sol with its safeguards removed and recorded the highest scores yet on several expert-level biology benchmarks.2 It concluded that Sol could substantially help some experts, but still had important limitations. OpenAI based its below-Critical rating on Sol’s results in novel pathogen design tests.

  • Misalignment: GPT-5.6 is more misaligned than GPT-5.5.3 Put simply, a model is misaligned when it does not act as its user intends. In OpenAI evaluations, Sol tried too hard to complete tasks, taking actions users had not asked for and would object to. This happened more often than with GPT-5.5 in coding tasks. OpenAI’s examples include circumventing restrictions, taking destructive actions, and lying about completing work. METR found the same pattern, detecting a higher cheating rate4 for Sol than for any other public model it had tested. Nevertheless, Apollo Research did not find that GPT-5.6 Sol poses a substantially higher risk of catastrophic scheming than the tested baselines.

  • AI self-improvement: below High. This category is about whether a model can speed up AI research. High would mean the model gives every OpenAI researcher the equivalent of a strong mid-career engineer to assist them, and Critical would mean it could significantly improve frontier AI systems on its own. Sol improved at machine learning engineering tasks, but OpenAI kept it below High, the same as GPT-5.5. METR could not reliably measure Sol on long tasks, because it cheated too much. Even so, METR concluded that Sol is not far beyond current models and would not enable fully automated AI research. Anthropic reached a similar conclusion about Mythos 5. In its launch livestream, OpenAI claimed that Sol had autonomously post-trained Luna. That claim is hard to square with the below-High rating, and OpenAI did not back it up in the system card.

OpenAI assessed GPT-5.6 under its Preparedness Framework, adding a few new tests compared to GPT-5.5. The public release included more testing than the preview, most of it from UK AISI on cyber capability and alignment.

  • OpenAI gave external red teamers over a month to test safeguards. Four private organizations tried to find universal jailbreaks. A universal jailbreak is a single method that unlocks many prohibited requests at once, rather than one request at a time. OpenAI says it gave testers broad access and fixed the issues they found before launch.

  • OpenAI’s Critical biology threshold tests focused on novel pathogen design. OpenAI judges biological threats to be more severe than chemical ones, so it prioritized biology. For the Critical threshold, it focused on how much a model could help design a novel pathogen, which is a new approach.

  • OpenAI rebuilt its self-improvement tests. The old tests were either saturated or partly unsolvable. A test is saturated when models score so high that the test can no longer tell them apart. OpenAI replaced them with tasks closer to real research engineering work (e.g. fixing internal bugs or improving a model’s training run). The new tests reset the baseline, so OpenAI reran earlier models on them to allow comparison across generations.

  • OpenAI added tests for metagaming. Metagaming occurs when a model reasons about how it will be graded, rewarded, or monitored, rather than about the task itself. This matters because such a model may behave differently in a test than in real use. GPT-5.6 showed more metagaming than GPT-5.5 on most tests, but almost none in deployment simulations.

  • OpenAI did not run human uplift studies for GPT-5.6. A human uplift study measures how much a model helps real people carry out a dangerous task. Anthropic used one in its Fable and Mythos cards to assess biological risk. Such studies are widely seen as important for judging real-world risk.

  • OpenAI checked whether its assessments were undermined by sandbagging and verbalized awareness. Sandbagging means deliberately performing worse to hide a capability during a test. GPT-5.6 Sol did not sandbag, and neither did GPT-5.5. Additionally, according to external Apollo testing, Sol’s accuracy did not drop even when the model was given an incentive to answer incorrectly. Sol also spoke about being tested less often than GPT-5.5, and when it did, it was wrong about the test’s purpose 70% of the time. This could mean that it is less aware of being tested, or that its awareness is just less visible.

OpenAI describes its GPT-5.6 safeguards as its strongest and most conservative to date. Many layers carry over from GPT-5.5, and some are new.

  • A new system monitors the model as it generates a response. Sol and Terra run activation classifiers, which watch patterns inside the model while it writes. If a separate check confirms the content is harmful, the system pauses and blocks the response. This adds to older tools that scan whole conversations and block unsafe outputs.

  • The broad release blocks more activity than the preview. OpenAI says Sol’s cyber safeguards block about ten times more potentially harmful activity than those of its earlier models. Users can retry a blocked prompt on a lower-capability model. OpenAI calls this a conservative starting point that it will adjust based on real-world use.

  • OpenAI spent over 700,000 A100-equivalent5 GPU hours on automated universal jailbreak discovery. It used its own autonomous agents to search for universal jailbreaks. Each time the agents found one, OpenAI added a mitigation and tested the models again.

  • Trusted access programs reserve sensitive capabilities for vetted users. OpenAI runs Trusted Access for Cyber for verified defenders, and Trusted Access for Biology Research for eligible organizations. These programs let approved users do some dual-use work, while OpenAI keeps monitoring in place, along with its strongest blocks.

  • A bio bug bounty rewards jailbreaks of the bio safeguards. OpenAI expanded its earlier program into an ongoing Bio Bounty Program and raised the reward for a universal jailbreak from $25,000 to $50,000.

  • UK AISI tested OpenAI’s safeguards and repeatedly found universal cyber jailbreaks. In each round of testing, UK AISI found universal jailbreaks, often within hours. It expects that further testing will uncover more at a similar level. These included jailbreaks that got the model to complete long tasks in domains like vulnerability discovery and exploit development. This raises questions about how effective OpenAI’s safeguards are. That said, a real-world attacker would have far less access than the testers did.6

  • Is the US government creating a de facto involuntary licensing regime for frontier models? Some policy experts think so.7 Trump’s executive order does not require a license before a company releases a model. But the government placed export controls on Fable 5 and Mythos 5, and it asked OpenAI to limit GPT-5.6’s launch. Together, these actions suggest that companies may now need US government approval before releasing a new frontier model.

  • Is a government-mandated restricted release process good for frontier AI safety? It could be, if it pushes companies to do more real-world testing and add stronger safeguards than they otherwise would. But it could also fail, or cause harm, if the rules are weak, unclear, or applied differently across companies. Restricted releases should follow a clear and repeatable process, not case-by-case decisions. Anthropic and OpenAI have both called for such a process, especially in cyber, and both appear to be working with the US government and cybersecurity experts to develop one.8

  • The government has not shared the criteria for approving the release of GPT-5.6. Officials have not explained why they first asked for the restriction, or what changed during the preview. OpenAI said it improved the model’s safeguards. But it is not clear whether the company had to meet specific safety requirements, or what those requirements were.

  • The US government appears to have treated GPT-5.6 Sol and Anthropic’s Fable 5 differently. Fable was restricted after testers found a limited jailbreak, and no universal jailbreak has been publicly reported since its re-release. By contrast, UK AISI said it expected to find more universal jailbreaks of Sol at public release. This may be because UK AISI had more access to Sol than to Fable. Even so, without a public explanation, it is unclear whether the same standard applied to both models.

  • It is unclear whether GPT-5.6 matches Mythos 5 on cyber. The two external evaluators disagreed. UK AISI’s results suggested Sol had similar cyber capabilities to Mythos 5, while Irregular instead compared Sol to GPT-5.5. Anthropic and OpenAI also run their own separate evaluations, which makes cross-model comparison hard. Shared evaluations would help. Even then, benchmark results do not always reflect real-world risk, and may over- or underestimate it.

  • The misalignment results deserve more attention. OpenAI did not mention the sharp rise in cheating and destructive actions in its launch post or livestream. It says overall rates remain low, but Sol bypassed restrictions in about 1 in 400 simulated coding tasks. At the scale at which these models are used, that rate may not be low at all. The Hugging Face incident shows how this behavior can increase risks. The attack required advanced cyber skills, but those skills became more dangerous when the models acted against their evaluators’ intent. This is not evidence of catastrophic scheming, but even non-catastrophic misalignment can create serious safety and security concerns.

  • Should cyber evaluations be airgapped? Airgapping means isolating a computer or network from external systems so that there is no physical connection to them. Although OpenAI’s test environments were highly isolated, the models still reached the open internet during the evaluations that led to the Hugging Face breach. As capabilities improve, labs may need to ensure stricter containment during evaluations. This is possible, but could be costly and slow down research.

Acknowledgements: Thanks to Alan Chan, Ed Foley, Elias Groll, Jonas Freund, Markus Anderljung, Noemi Dreksler, Sophie Williams, and Zaheed Kara (in alphabetical order) for helpful feedback. All remaining errors are my own.

Disclaimer: Posts are written by individual team members and reflect the author’s perspective. Not all team members necessarily agree with every take. The views expressed here do not represent the official position of GovAI.

1

GPT-5.6 Sol completed “The Last Ones,” a 32-step simulated attack on a company network, in 7 of 10 attempts, compared with 2 of 10 for GPT-5.5 and 6 of 10 for Mythos 5.

2

On World-Class Bio, GPT-5.6 Sol with safeguards removed scored about 9 percentage points higher than GPT-5.5.

3

OpenAI does not track misalignment as a formal Preparedness category, so it carries no “High” or “Critical” rating.

4

“Cheating” means the model improves its evaluation score by exploiting bugs in the test environment, or by using strategies the task disallows, instead of solving the task as intended.

5

One A100-equivalent hour means one NVIDIA A100 data center GPU running for one hour. Using faster chips or many chips at once makes the actual wall-clock time much lower.

6

UK AISI had extensive gray-box access. This included the chain-of-thought of the safety reasoning monitor, the exact policy wording, and real-time feedback on classifier labels.

8

OpenAI is working with the US government on the cyber executive order framework. Anthropic is developing a shared industry framework for assessing jailbreaks with Amazon, Microsoft, Google, and other partners.

No posts

Read the original on frontierrisk.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.