This research was funded by Anthropic PBC. Anthropic provided input and advice in the development of the report, but all findings and claims represent the views of the authors. For further information, please see the full report.
We ran a forecasting study to better understand expert views on the effectiveness of Anthropic’s AI Safety Level 3 (ASL-3) Deployment Standards at reducing biosecurity risk based on AI models from early 2026. We surveyed 22 biosecurity, national security, and/or terrorism experts, and 20 superforecasters between March 4 and March 29, 2026.
We asked respondents to forecast the total expected financial damages from large-scale human-caused outbreaks between April 2026 and the end of 2028 in four scenarios:
A world without general-purpose AI (GPAI)
A world where no GPAI models had any safeguards
A world where all GPAI models were protected by ASL-3
A world where frontier GPAI models were protected by ASL-3
In all scenarios, we asked respondents to assume that AI model capabilities remain fixed at the level of Claude Opus 4.6. This is because we wanted to understand expert views on how effective ASL-3 is at mitigating risks from the early 2026 generation of publicly available models, rather than how AI capabilities are likely to progress.
Experts and superforecasters generally believe that implementing ASL-3 safeguards across all GPAI models would substantially reduce the risk of large-scale human-caused outbreaks compared to a world with no safeguards. They also estimate that ASL-3 safeguards—if implemented across all GPAI models—would capture most of the risk reduction it is possible to achieve by safeguarding. In an “ASL-3 for frontier models” scenario, forecasters generally think that variations to ASL-3 could increase effectiveness, with the largest change in expected damages associated with a 5x faster speed of patching jailbreaks.
In May 2025, Anthropic announced that it had activated its AI Safety Level 3 Deployment and Security Standards as a precautionary measure in response to Claude Opus 4’s chemical-, biological-, radiological-, and nuclear-related knowledge and capabilities. The deployment standards include the implementation of real-time classifier guards, offline monitoring, access controls, a bug bounty program, threat intelligence, and rapid response.
Although Anthropic can evaluate aspects of its mitigations (for example, the time taken to patch an identified jailbreak), it is unclear how effective these mitigations are at achieving the overarching goal of reducing biosecurity risks posed by GPAI models. To investigate this, we conducted a forecasting survey of experts in biosecurity, national security, and/or terrorism studies, as well as superforecasters. Respondents were given information on Anthropic’s ASL-3 standards, including confidential data from red-teaming efforts, the bug bounty program, and other safeguard testing.1
The main outcome we asked participants to forecast was:
The total worldwide damages attributable to human-caused outbreaks that start between April 1, 2026, and December 31, 2028, and that each individually cause at least $100 million in total worldwide damages (in 2026 USD).2
Although we specified that outbreaks must start between April 1, 2026, and December 31, 2028, we did not specify a cutoff date for damages. Forecasters considered the following scenarios:
No GPAI: A world where GPAI never existed and will not exist. All other features of the world remain the same, so the absence of GPAI does not imply that technological progress is slower across all domains.
No safeguards: A world where no safeguards are applied to any GPAI model.
ASL-3 safeguards for all GPAI models: A world where all GPAI models—regardless of size or capability—are protected by Anthropic’s ASL-3 safeguards.
ASL-3 safeguards for frontier GPAI models: A world where “frontier” GPAI models are protected by ASL-3 safeguards.3
When respondents assumed all models had ASL-3 protections, the median forecasted probability of human-caused outbreaks causing at least $100 million in damages dropped by roughly 40% relative to the “no safeguards” scenario (from 16% to 9.2%). For comparison, when respondents were asked to consider a hypothetical world where GPAI was never developed, the median forecast was 9.3%.
The scenario that asked respondents to assume GPAI models did not exist can provide an upper bound on the risk reduction achievable by safeguards against GPAI model misuse, as it demonstrates the risk of human-caused outbreaks that would persist regardless of GPAI models. We can also express these results in terms of the share of GPAI-attributable risk mitigated by ASL-3.4 The median participant’s forecasts imply that ASL-3 for all models would mitigate 71.7% of GPAI-attributable risk.
This comparison has an important limitation: the “no GPAI” scenario is an imperfect proxy for perfect safeguards. Perfect safeguards would eliminate the risk of misuse while preserving the defensive benefits of GPAI, such as contributions to pandemic surveillance and vaccine development. The “no GPAI” scenario removes all of these. This means the “no GPAI” world is likely more vulnerable to outbreaks (including human-caused ones) than a world with perfectly safeguarded GPAI, making it a lenient benchmark.
The practical consequence is that the 71.7% efficacy figure—the share of maximum risk reduction achieved by ASL-3 for all models—is likely biased upward. If GPAI’s defensive contributions are small relative to the misuse risk, the bias is minor. If they are large, the bias could be substantial.
To understand the potential impact of all frontier models being protected by ASL-3 safeguards, we asked respondents to forecast the main outcome under this scenario.
Under the “ASL-3 for frontier models” scenario, the median probability of human-caused outbreaks occurring between April 2026 and the end of 2028 and leading to at least $100 million in damages (see Figure 1 above) was 10.6%. Compared to the “no safeguards” scenario, the “ASL-3 for frontier models” scenario was associated with an 18% reduction in this risk.
We also assessed how the calculated expected damages under the two ASL-3 scenarios and the “no GPAI” scenario compared to the “no safeguards” scenario.
The median respondent thought that “ASL-3 for frontier models only” would be associated with a 20% reduction in expected damages due to large-scale human-caused outbreaks relative to the “no safeguards” scenario. Most respondents placed the “ASL-3 for frontier models” scenario’s risk between “ASL-3 for all models” and “no safeguards;” but disagreed on where, with some treating it as nearly equivalent to “ASL-3 for all models” and others as only marginally better than “no safeguards.”
Respondents were asked to forecast the main outcome conditional on frontier models being protected by the following hypothetical variations on the ASL-3 scenario:
Time required to identify jailbreaks
A1: The average time required to find a universal jailbreak is 5x longer.
A2: The average time required to find a universal jailbreak is 5x shorter.
Time to jailbreak patching
B1: It takes a fifth (0.2x) of the time to patch all universal jailbreaks.
B2: It takes up to 2x as long to patch all universal jailbreaks.
Capabilities loss associated with jailbreaks
C1: All universal jailbreaks significantly reduce the model’s GPQA score. Compared to jailbroken models’ current performance relative to the “no jailbreak” baseline, jailbroken models’ relative performance decreases by roughly 30%.
C2: All universal jailbreaks only minimally reduce the model’s GPQA score. Compared to jailbroken models’ current performance relative to the “no jailbreak” baseline, jailbroken models’ relative performance increases by roughly 30%.
Figure 4 below shows the median change in expected damages relative to “ASL-3 for frontier models” under the variant scenarios. The largest median change—an 8% decrease relative to the “ASL-3 for frontier models” scenario—was associated with a faster time to jailbreak patching.
We asked which type of actor is most likely to be the primary cause of a human-caused outbreak that causes at least $100 million in damages, and how this would change depending on whether GPAI models had no safeguards or were all protected by ASL-3 safeguards. Both groups of respondents thought that state actors were the most likely cause of such an event and would be an even more likely cause under the ASL-3 safeguards scenario. Among the expert respondents, the next most likely group was individual expert actors, although this group’s probability of being the primary cause of a large-scale human-caused outbreak fell under the ASL-3 safeguards scenario.
When asked about the probability of unauthorized frontier model use contributing to a human-caused outbreak that causes more than $100 million in damages, the median expert forecast a 30% probability, and the median superforecaster a 23.5% probability. We then asked respondents to assume that such a scenario had occurred, and to say how likely they thought the following pathways were to have contributed to that unauthorized model use:
Own jailbreak
Public jailbreaks
Nonpublic jailbreak (e.g., black market)
Trusted user exemption
Each of the pathways we asked about had a median probability of 15% or higher, suggesting that most respondents found all pathways plausible. Relative to experts, superforecasters generally thought that a trusted user exemption was more likely to be used, and that the actor generating their own jailbreak was less likely.
This study asked 22 domain experts and 20 superforecasters to estimate how ASL-3 safeguards would change the expected financial damages from large-scale human-caused outbreaks between April 2026 and 2028. Four findings stand out.
Respondents believe ASL-3 safeguards demonstrate high efficacy. If applied to all GPAI models, the median respondent’s forecasts imply that ASL-3 would mitigate roughly 70% of GPAI-attributable risk.
Protecting only frontier models with ASL-3 reduces risk, but not as much as protecting all GPAI models. Protecting only frontier models captured less of the risk reduction than protecting all GPAI models, though respondents disagreed substantially on the size of the gap. This disagreement generally hinged on how capable and accessible near-frontier models—particularly Chinese open-weight models—are judged to be.
Making jailbreaks harder to find, patching them faster, or further degrading the performance of jailbroken models were generally expected to improve ASL-3 performance. The largest effect was associated with faster patching, consistent with the logic that bioweapon development requires sustained iterative access. The variants were tested only under the frontier-only scenario, where unsafeguarded near-frontier models may dampen the marginal impact of any frontier-specific improvement.
ASL-3 shifts the threat landscape toward better-resourced actors and accidental pathways. Respondents expected ASL-3 to disproportionately screen out nonexpert and lone-wolf actors, shifting relative probability toward state actors, well-resourced organizations, and accidental releases from laboratories.
For more information, see the full report.
Although respondents' forecasts included consideration of both the Deployment and Security Standards that make up ASL-3, the focus of this survey was the Deployment Standards. See the full report for more details.
Since we were interested in forecaster’s views on ASL-3’s ability to mitigate harm from current models, we asked respondents to assume that until December 31, 2028, the capabilities of GPAI models do not improve and that, for the same period, no models outperform Claude 4.6, Gemini 3, or GPT-5, and this slowdown was not indicative of a general slowdown in technology progress.
Frontier GPAI models were defined as all those performing as well as or better than Claude Sonnet 4.5 on the Epoch Capabilities Index at the time of the survey.
Aggregates presented here exclude four participants whose responses imply that GPAI reduces risk, making this metric uninterpretable for them.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.