RSS Amplifier

Forecasting Research Institute · Jul 23, 2026

Forecasting AI Cyber Risks and Capabilities

0
Sign in to vote or save

Forecasting Research Institute · Forecasting Research Institute

In mid-2025, in collaboration with researchers at GovAI, we conducted a pilot study investigating how AI capabilities may affect near-term cybersecurity risk. We surveyed 13 superforecasters and eight cybersecurity experts, examining two high-impact cyberattack pathways in 2026:

  • Data-damaging worm attacks similar to WannaCry and NotPetya that could cause at least $10 billion in economic damages.1

  • Cyberattacks against the U.S. electrical grid that cause a large-scale blackout leading to at least $10 billion or at least $100 billion in economic damages.2

Participants were surveyed between July and August 2025, and some responded to a follow-up survey between December 2025 and January 2026. This study was designed as a pilot, so the results should be interpreted as initial evidence for areas that merit further research, not as definitive estimates.

Because AI cyber capabilities are changing rapidly, the results should not be interpreted as a current assessment of frontier-model cyber capabilities. In particular, at the time of the survey, Claude Mythos Preview/Mythos 5, Claude Fable 5, GPT-5.3-Codex, GPT-5.5/GPT-5.6/GPT-5.5-Cyber, and other cyber-specialized models had not been deployed. We’re working on a follow-up study that will incorporate beliefs about Claude Mythos and other recent AI models, as well as the Hugging Face incident.

The forecasts we elicited in the present study show a consistent pattern: participants assessed baseline risks of catastrophic cyber harms in 2026 as low but non-negligible, and they expected some AI capabilities to substantially increase those risks, especially when AI lowers barriers for moderate-sophistication actors. It is difficult to rule out the possibility that an open-weight, safeguard-free version of a Mythos-level model would reach the level of another capability we described: elite exploit development for moderately skilled individual hackers. If that is the case, then our findings suggest that wide access to these capabilities could substantially increase the risk of a major cyber event, but that mitigation measures could significantly reduce that risk.

This post summarizes the key results from the pilot. You can read the full report here.

Participants estimated a 5–8% probability of at least one data-damaging worm attack causing at least $10 billion in damages in 2026. The median participant’s forecast of expected annual damages from data-damaging worms was approximately $10–15 billion.

Participants estimated a 1% probability that a cyberattack against the U.S. electrical grid would cause at least $10 billion in damages in 2026, and a 0.1% probability that such an attack would cause at least $100 billion in damages. Expected annual damages were roughly $0.2–1 billion, more than an order of magnitude lower than for worms.

Under a hypothetical scenario in which AI models enable 25% of moderately skilled individual hackers to find vulnerabilities and write elite exploits (zero-click exploits that allow remote code execution with high privileges—the kind that led to the WannaCry and NotPetya attacks), and models with this capability are available open-weight, worm attack risk estimates increase by 3–3.5x. The median expert forecast of a data-damaging worm attack causing at least $10 billion in damages rose from 8% to 41%, while the median superforecaster forecast rose from 5% to 15%.

The median expert thought it would take until 2028 for an AI model to solve more than 90% of tasks on the Cybench benchmark, while superforecasters predicted 2030. This capability was most likely surpassed in February 2026, when Claude Opus 4.6 achieved 93% on a subset of problems on the benchmark. It is difficult to rule out the possibility that an open-weight, safeguard-free version of a Mythos-level model would reach the level of another capability we described: elite exploit development for moderately skilled individual hackers.3 Experts and superforecasters forecast this would be achieved in 2032 and 2030, respectively.

Figure 1: Forecasts of the year in which participants expected each AI capability to be achieved.

Participants generally saw AI performance on Cybench as informative of technical progress, but not as the best standalone indicator of whether capabilities most relevant to catastrophic cyber outcomes, such as elite exploit development or real-world grid attacks, had been achieved. Respondents viewed current (as of mid-2025) benchmarks as incomplete measures of operational constraints such as zero-day discovery, targeting, coordination, and persistence. Current cyber evaluations may track important technical progress without fully measuring the forms of AI-enabled uplift most relevant to real-world cyber risks.

Nearly half of the participants thought that keeping the relevant models—those that could enable moderate-sophistication actors to develop elite exploits—proprietary and protected by anti-jailbreak measures would at least halve the risk of a large-scale worm attack. However, participants also noted that such measures may be less effective against sophisticated actors or exploit-as-a-service markets.

A subset of participants revisited some of their forecasts between December 2025 and January 2026, after reviewing the original study results and Anthropic’s November 2025 report on an AI-assisted cyber espionage campaign. The median forecast for a data-damaging worm attack causing at least $10 billion in damages in 2026 remained unchanged. Many participants thought the Anthropic cyber espionage report suggested that actors were more capable than they had originally expected, but most did not view it as direct evidence that AI could develop elite exploits or execute a large-scale data-damaging worm attack by 2026.

Rapid advances in artificial intelligence capabilities have introduced new dimensions of uncertainty into the cybersecurity landscape. A central uncertainty is whether AI will strengthen defensive capabilities overall, disproportionately benefit malicious actors, or affect different threat actors in different ways.

This pilot study is an initial effort to systematically forecast how AI might affect large-scale cyber risks over the near term, with a particular focus on 2026. We used structured forecasting methods to elicit judgments about two high-impact cyber threat scenarios: data-damaging worm attacks and cyberattacks against the U.S. electrical grid.

A “worm” is malware that can spread autonomously between systems without an attacker needing to infect each system individually. In this survey we concentrated on “data-damaging worms” that cause damage by wiping, encrypting, or corrupting data on a large number of systems. The WannaCry and NotPetya attacks, both released in 2017, were data-damaging worms that caused roughly $1 billion to $10 billion in damage after infecting hundreds of thousands of systems.

Worm attacks can exploit vulnerabilities: bugs in software or hardware that create security weaknesses in the design, implementation, or operation of a system or application. An exploit is malicious code that takes advantage of one or more software vulnerabilities to infect, disrupt, or take control of a computer without the user’s consent and typically without their knowledge. Following Halstead and Righetti (2026), we define exploits (or exploit chains) as “elite” if they satisfy all of the following criteria:

  • Zero-click: Infection requires no user interaction, such as opening emails, clicking links, or visiting a webpage.

  • Remote code execution: Allow attackers to execute arbitrary code on a system without the user’s knowledge, and without the attackers needing physical access to the system.

  • High privileges: Have administrator privileges or higher.

  • Targets widely used software: Effective against more than 10 million systems.

Elite exploits are especially well-suited to worm attacks as they enable autonomous spread and significant damage on a large number of systems. The 2017 leak of elite exploits initially developed by the NSA quickly led to the two major data-damaging worm attacks cited above: WannaCry and NotPetya. Developing elite exploits may require an order of magnitude more skilled researcher time than the other tasks involved in developing a data-damaging worm.

The effect of AI on the risk of worm attacks is uncertain and likely to vary over time. AI systems that can discover vulnerabilities, develop exploits, or automate attack steps are dual-use: the same capabilities could help both defenders and attackers. The net effect depends on several uncertain factors, including which actors gain access to the most capable models, how quickly vulnerabilities are disclosed and patched, and whether patch deployment keeps pace with vulnerability discovery.

The power grid consists of generators, transmission and distribution networks, and substations. Grid operations, like other infrastructure and industrial processes, rely on two broad types of computer systems:

  • Information technology (IT) systems: systems that handle most business operations like billing, email, and administration, and are typically connected to the internet.

  • Operational technology (OT) systems: programmable systems that interact with the physical environment or manage devices that do. In grid operations, these systems monitor and control equipment like generators, circuit breakers, and transformers.

Grid cyberattacks involve compromising OT systems, as this is a prerequisite for directly interfering with grid behavior. OT environments, when compared to IT environments, present additional challenges. They typically are—or should be—segmented from IT and internet-facing networks, and they use more niche software and protocols that require more specialized knowledge; OT devices may have relatively individualized configurations to a given environment.

The most relevant prior OT cyberattacks have required substantial time, resources, and specialized expertise, often from state-level actors. Beyond the significant resource and time investments, prior OT cyberattacks have required the integration of diverse capabilities, including long-term reconnaissance, knowledge of the target OT environments, tailored malware, realistic test environments, and stealth.

These requirements make the role of AI difficult to assess. AI could plausibly assist with some components of a grid attack, such as reconnaissance, code generation, vulnerability discovery, and operator decision support. It is less clear how much AI would help with other important bottlenecks, such as avoiding detection, and orchestrating complex, dynamic operations.

The median expert forecasted an 8% probability that a large-scale data-damaging worm attack would cause at least $10 billion in economic damages in 2026, while the median superforecaster forecasted a 5% probability. These baseline forecasts were anchored by the low historical frequency of cyberattacks of this magnitude, although participants noted that the 2017 WannaCry and NotPetya worm attacks demonstrate that worms can spread rapidly, and that future attacks could have large impacts in a more digitally dependent economy.

AI-enabled elite exploit development was seen as a significant driver of increased risk. If an open-weight AI model enabled 25% of moderately skilled individual hackers to find vulnerabilities and write elite exploits, the median expert forecast rose to 41%, and the median superforecaster forecast rose to 15%. Participants viewed this as an important risk pathway because individual hackers are relatively numerous and may be more willing to cause damage, but they are usually capability-constrained; more sophisticated actors are generally more capable but constrained by escalation and retaliation risks.

Figure 2: Forecasts of the probability of at least one data-damaging worm attack causing at least $10 billion in economic damages in 2026, under baseline assumptions, and conditional on AI capabilities and potential mitigations.

Participants expected model access controls and safeguards to reduce this risk, especially for actors below state-level sophistication, but not to eliminate it. In particular, having models with the relevant capabilities protected with anti-jailbreak measures was believed to significantly decrease the probability of a large-scale data-damaging worm attack. Model-weight theft, exploit-as-a-service markets, and slow patching of vulnerabilities remained important limitations.

Participants assessed large-scale cyberattacks against the U.S. grid in 2026 as substantially less likely than data-damaging worm attacks. The median expert and superforecaster both put the probability of a cyberattack against the grid causing at least $10 billion in damages at 1%; for a larger attack, causing at least $100 billion, both medians were 0.1%. Expected annual damages were also much lower, at approximately $0.2–1 billion.

Participants emphasized that grid attacks face additional barriers beyond cyber capabilities, including expertise in industrial control systems (ICS), operational coordination, physical infrastructure constraints, and the risk of geopolitical escalation. They generally viewed a large-scale grid attack as more likely in the context of state-level conflict than in the context of ordinary cybercrime.

Forecasts of large-scale cyberattacks against the U.S. electrical grid were less sensitive to AI capability scenarios than forecasts in the data-damaging worm threat model. AI uplift in an ICS/OT capture-the-flag style competition only modestly increased the median forecast from 1% to 1.5–2%. A real-world AI-enabled warning shot causing at least $100 million in damages produced a larger increase in risk, raising the median forecast to 15% among experts and 4% among superforecasters. Overall, results from this pilot suggest that AI may increase the risk associated with this threat model, but large-scale grid attacks remain constrained by operational complexity and geopolitical considerations.

Figure 3: Forecasts of the probability of at least one cyberattack against the U.S. electrical grid causing at least $10 billion in economic damages in 2026, under baseline assumptions, and conditional on Capability 3 (AI enables individual hobbyist hackers to perform like a group of roughly 10 well-known criminal hackers in an OT-specific capture-the-flag style competition) and Capability 4 (a real-world warning shot incident).

The study results suggest that AI capabilities could significantly increase some cyber risks in the near term. Both experts and superforecasters predicted substantial increases in the likelihood of a large-scale data-damaging worm attack when open-weight AI enables moderately sophisticated individual hackers to develop elite exploits. Median forecasts suggest that AI-enabled vulnerability discovery could increase the risk of large-scale worm attacks by 3–3.5x, with expected annual damages rising from roughly $10–15 billion to $33–67 billion. Our respondents, answering in mid-2025, believed these capabilities would arrive within the next five to six years, but it is plausible that Claude Mythos has already achieved this capability.

At the same time, structured mitigation measures appear to matter. Participants believed that maintaining proprietary model access with strong anti-jailbreak and monitoring controls could more than halve the probability of such catastrophic outcomes. Limited early access for defenders was considered less effective, though still potentially useful.

Given the small sample size and the speculative nature of some scenarios, these results should be interpreted as directional rather than definitive. Nonetheless, they illustrate the value of structured forecasting for anticipating technology-driven security risks. Future work could explore a wider range of possible cyber harms and include a larger, more representative sample of relevant subject-matter experts.

For more from this study, including forecasts of threat actors' capability and willingness to carry out cyberattacks, and more details on expected damages from attacks, see the full report.

No posts

Read the original on forecastingresearch.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.