Some AI companies and researchers support having the option to collectively slow (or pause1) frontier AI R&D.2 One potential worry is that increasingly rapid AI progress could outpace our ability to keep AIs safe and under human control. Slowing down could buy time to address those safety problems.
Assessing whether a slowdown would be good—and preparing for it if so—requires more clarity on what it would entail.
This piece outlines 3 components of a slowdown: aims, what is slowed, and structure. To keep this piece contained, I will not address how such slowdowns can be coordinated or enforced.3 Overall:
Aims: During a slowdown, various actors (e.g. AI companies, governments) could aim to do safety R&D and implement safeguards, take policy actions (e.g. improve societal resilience), and/or simply use the time to deliberate on next steps.
What is slowed: Slowdowns could restrict R&D itself (e.g. directly prohibiting training runs), its inputs (e.g. frontier compute production), and/or its outputs (e.g. ≤ x% improvement in compute efficiency, caps on the frequency of model releases).
Structure: Slowdowns can specify start and stop conditions (e.g. lift a slowdown upon widespread implementation of DNA synthesis screening), stages (e.g. a slowdown gradually lifts over time rather than all at once), pacing (e.g. new model releases only once every 3 months), and scope (e.g. industry-wide vs. company-specific, global vs. US-only).
I summarize some key considerations for whether and how to slow down:
Whether to slow down:
Slowing capabilities research could also slow safety research.
Many societal resilience activities do not need further AI progress and thus could be effectively carried out during a slowdown, such as implementing mandatory DNA synthesis screening and improving AI expertise within government.
Counterintuitively, slowdowns could hinder societal resilience in defense-dominant domains, where further AI progress strengthens defenders more than attackers (cyber is a potential example).
Slowdowns could also reduce or forgo the benefits of AI progress, including economic gains that may make a society better able to absorb risks.
How to slow down:
Pacing (e.g. new model releases only every 3 months) may be especially useful for enabling policy actions because it can align with legislative cycles and provide a deadline for such actions.
Without accompanying restrictions on R&D inputs (e.g. compute) or a sufficiently gradual phase-out, some slowdowns could lead to large capability jumps when activities recommence.
Broadly scoped slowdowns (e.g. industry-wide, global) better address race dynamics but are harder to coordinate. Narrower slowdowns could be more feasible and address risks in a targeted way, but might allow competitors and/or irresponsible laggards to catch up and make future safety coordination harder.
Lastly, I run through some illustrative examples of slowdown designs motivated by different aims:
A slowdown for pacing R&D to policy cycles.
A narrow slowdown of specific capabilities focused on achieving societal resilience to those capabilities.
A slowdown that tries to maintain a lead between leaders and competitors.
During a slowdown, we could:
Do more safety work and implement safeguards, such as by:
Running thorough risk assessments or audits
Doing alignment research
Building secure (e.g. SL5-level) data centers
Take policy or other actions to address AI risks, such as by:
Building up government AI expertise, such as in AI safety institutes
Passing legislation
Improving societal resilience, such as by stockpiling PPE or mandating DNA synthesis screening
Use the time to deliberate, such as by:
Thinking through model release decisions more carefully
Watching how risks develop with models that are already deployed
Building consensus across stakeholders
Gathering public opinion on risks
A slowdown could target R&D itself, its inputs (e.g. compute), or its outputs (e.g. model releases, a level of capability improvement).
An R&D slowdown typically refers to restricting or delaying capabilities R&D.4 Examples include:
Restrictions on the compute given to pre-training teams, which could be redirected to safety teams.
Prohibitions on building training environments that advance potentially concerning capabilities, such as long-horizon agency.
Prohibitions on pre-training runs above a certain size, to allow more time to research and implement safeguards.
For an R&D slowdown to be effective, R&D inputs may also need to be slowed. Otherwise, large capability jumps could occur once R&D returns to its former pace. One mechanism is a compute overhang: while training runs are paused or slowed, the quantity and quality of compute could keep growing. The first resumed run can use this entire stock of improved compute at once, training models that are far more capable, and potentially riskier, than before.
Slowing R&D also involves some trade-offs. First, slowing capabilities R&D could also slow safety R&D because:
Safety methods need to be tested on capable models, if such models are eventually to be released.
Historically, some methods have advanced both capabilities and safety to a large degree (e.g. RLHF). More generally, work that makes a model more steerable will benefit both capabilities and safety, since a more steerable model is also a more useful one.
Safety and capabilities R&D often rely on the same infrastructure, such as for running experiments or building evaluations.
Second, R&D could be difficult to target. There is genuine disagreement about what counts as safety research, and many projects improve both safety and capabilities. If restrictions hit particular teams, developers could relabel projects, move researchers between teams, or recast capabilities work as safety-motivated.
Inputs include compute, data, and talent. Some examples of slowing inputs are:
Restrictions on the rate of growth of frontier compute production, so as to slow large training runs.
Data licensing requirements, to limit the data that can be used to improve capabilities.
Headcount limits on research staff at frontier companies, to cap the amount of R&D activity.
Compared to R&D itself, R&D inputs could be simpler to target. For example, whether a chip counts as “frontier” comes down to measurable physical characteristics, such as processing power and memory bandwidth. In contrast, the distinction between safety and capabilities R&D is more conceptual and fuzzy.
At the same time, targeting R&D inputs could be overly broad. If we are most worried about specific capabilities (e.g. bio or persuasion), we could target the R&D activities aimed at advancing those capabilities (e.g. the construction of particular training environments) rather than restricting inputs across the board. Slowing frontier compute production, by contrast, would affect all frontier AI development and reduce AI benefits like economic growth.
Slowing R&D outputs could mean:
Capping specific improvements to models (e.g. no more than x% gain on a persuasion benchmark), to target a particular risk.
Restricting (both external and internal) model deployment (e.g. no deployment at all, or deployment only to smaller groups like safety teams).5
In practice, capping improvements likely has to work backwards through limits on R&D and its inputs. Improvements are difficult to constrain directly because we rarely know what a new algorithm will yield until we test it. Such caps could also give developers an incentive to sandbag or conceal capabilities.
Restricting model deployment could be useful for a number of reasons:
Slowing external deployment can help to address misuse. Companies could have more time to test safeguards internally before deployment. Delayed deployments could also leave more time to improve societal resilience (e.g. improving cyber defenses).
Slowing internal deployment could provide more time to understand and mitigate risks from AIs sabotaging R&D or company operations, such as by poisoning the training of successor systems or tampering with safety evaluations.
However, it may not be clear what should count as a new deployment. Developers should be able to update their models with safety patches without contravening deployment restrictions. At the same time, some patches might improve both safety and capabilities. Developers could game the restrictions to make incremental advances that increase capabilities and risks.
Furthermore, if R&D and its inputs are not slowed in parallel, external deployment slowdowns could lead to large capability jumps once new models are deployed.
Slowing deployment also involves some trade-offs:
External deployment helps build awareness of AI capabilities and motivates policy action. Demos for policymakers could capture some of this benefit, but external deployment also keeps the public informed and gives them a say in what level of risk is tolerable.
Internal access to the most capable models can accelerate safety research, and wider use can surface problems faster simply because more people are using the model.
Slowdowns can specify start and stop conditions, stages, pacing, and scope.
Start and stop conditions could involve any event or observation, including:
Thresholds on the pace of AI progress
Capability thresholds (e.g. x% on a benchmark)
Whether particular safeguards have been implemented
Whether particular policy actions have been taken
The size of the lead between leading and lagging companies
For example, training runs could be paused once a model crosses a capability threshold, resuming only after sufficient safeguards are implemented. As another example, a slowdown might have no stop condition at all, such as a perpetual cap on frontier compute production.
Slowdowns can have different stages, each triggered by different start and stop conditions. For example, a narrower slowdown (e.g. only specific R&D activities) can turn into a broader one (e.g. R&D activities and inputs) if the former proves insufficient or its aim goes unmet. To avoid large capability jumps, slowdowns could also be lifted in stages rather than all at once. For instance, a slowdown could permit training runs only up to a certain size and then raise that ceiling by small steps over time.
A slowdown can permit certain activities to occur only at regular intervals. For example, model releases can be permitted only every 3 months. Pacing can align with legislative cycles and provide strict deadlines, which could be especially effective for enabling policy action. At each interval, decision-makers can watch how deployed models behave and decide whether to continue, tighten, or relax the slowdown.
Pacing likely needs to be paired with limitations on R&D and/or its inputs. Without such limitations, the gains between intervals could be uneven or large (e.g. because of compute overhangs).
The scope of a slowdown can vary widely, from a single company to the entire industry, and from one jurisdiction to many. Different scopes introduce trade-offs:
A broader scope is harder to coordinate but directly addresses race dynamics. Industry-wide and global slowdowns ease the pressure to race ahead by locking in the current competitive landscape, and could also stop less careful rivals from overtaking today’s leaders. But they also foreclose competition, including on safety innovation.
Narrower slowdowns can be more precise and easier to coordinate, but do not address race dynamics. If only one company triggers a capability threshold, slowing just that company targets the risk directly. However, other companies could catch up to or overtake their rival. Similarly, a US-only slowdown would cover most frontier progress today and be easier to enact than a global agreement. At the same time, it would let geopolitical competitors catch up and, by widening the field of players, may make future safety coordination harder.
To illustrate how different aims favour different slowdown designs, I’ll walk through several high-level examples. These examples are meant to be illustrative: they are not necessarily net good or feasible.
Aim: Give policymakers time to deliberate on and pass good legislation, while avoiding capability overhangs.
Summary of design: Pace model releases and capability improvements, and slow R&D activities and inputs.
Activities slowed:
Pace model releases (e.g. every 4 months), to allow time to deliberate on legislation while still providing a deadline for action.
Limit capability improvements for each model release (e.g. ≤ x% every 4 months), by slowing R&D activities and compute production.
Structure:
Scope: Industry-wide.
Start condition: E.g. a capability threshold.
Stop condition: A sunset date, or a decision to lift the slowdown based on other considerations (e.g. the benefits of AI progress).
Aim: Buy time to improve resilience in one domain (e.g. biosecurity), where doing so does not need further AI progress.
Summary of design: A ladder of restrictions that starts narrow and broadens only if needed.
Activities slowed:
Stage 1 (narrower): Industry-wide slowdown of R&D that specifically advances AI bio capabilities, covering relevant training environments, evals, data generation, etc.
Stage 2 (broader): If Stage 1 proves insufficient (e.g. because general capability improvements spill over to bio capabilities), escalate to a broader restriction on AI R&D activities and inputs.
Structure:
Scope: Industry-wide.
Overall start condition: E.g. a bio uplift eval crosses a defined threshold.
Stage 1 → Stage 2: The same uplift eval rises by ≥ x% despite the restrictions in Stage 1.
Stop condition: Sufficient resilience to biological risks is obtained. E.g. DNA synthesis screening is mandated and adopted across providers covering ≥ x% of global synthesis capacity, and medical countermeasure stockpiles reach a level sufficient to protect essential workers for Y days.
Aim: Obtain more time to develop safeguards, while ensuring that US frontier labs stay ahead of Chinese competitors.
Summary of design: Slow down only if you are sufficiently ahead.
Activities slowed:
Slow R&D activities and frontier compute production.
Structure:
Scope: US industry-wide.
Start condition: The estimated lead over competitors is ≥ 12 months and a specific safety objective is pending (e.g. building SL5 data centers).
Stop condition: The objective is met or the lead has fallen to ≤ 6 months, whichever comes first.
Slowdowns have costs. One cost is delaying or reducing the benefits of AI progress, such as productivity gains. A poorer society may be less able to absorb AI risks (e.g. from the models already deployed). Slowdowns could also hurt societal resilience in defense-dominant domains, where further AI progress would strengthen defenders more than attackers. Cyber is a potential example: if AI systems eventually let us “fix all of the bugs”, accelerating progress and giving defenders access to frontier systems first might be better than slowing down.
There also remain many unanswered questions about slowdowns and their implementation.
Scope:
Should slowdowns also cover non-frontier AI R&D, e.g. in academia?
Implementation:
Who should adjudicate when start and stop conditions have been met?
How do we measure start and stop conditions? E.g. what does it mean for safety research to have progressed sufficiently?
What new technologies do we need to implement and enforce slowdowns?
Costs of a slowdown and alternatives:
How significant are the costs of a slowdown? E.g. how significant are the economic costs of slowing AI progress?
Which types of slowdowns might be most feasible and desirable?
What are the best alternatives to specific types of slowdowns?
Many thanks to GovAI staff and Tom Reed for feedback that helped to improve this piece.
In this piece, I consider a pause to be a sufficiently extreme slowdown.
In some sense, slowdowns already exist. For example, company safety frameworks stipulate safety mitigations that must be in place before particular capability thresholds are reached, which slows development. More broadly, we impose slowdowns of various kinds in other domains. The FDA’s clinical trial process slows the development of drugs so that we can better understand their safety and efficacy.
Although less common, an R&D slowdown could also cover other kinds of AI progress (e.g. efficiency improvements).
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.