RSS Amplifier

Frontier Risk · May 6, 2026

Comparing Anthropic’s RSP and compliance framework

0
Sign in to vote or save

Sophie Williams, Jonas Freund · Frontier Risk

  • Anthropic’s Responsible Scaling Policy (RSP) was originally a form of voluntary self-regulation.

  • Several laws have since made safety frameworks mandatory, including California’s SB-53, New York’s RAISE Act, and the EU AI Act / GPAI Code of Practice.

  • Anthropic responded by publishing a separate framework: the Frontier Compliance Framework (FCF). The FCF is designed to meet obligations under those laws, while the RSP remains voluntary.

  • Having a separate compliance framework may allow Anthropic to make more ambitious commitments in the RSP. It might also help keep the RSP easy to understand and update.

  • But it might also turn the FCF into a box-ticking exercise, and could create confusion about which framework actually governs day-to-day decisions, especially where the frameworks overlap or contradict each other.

Frontier AI companies are now required by law to maintain a safety framework that describes how they assess and mitigate catastrophic risks from their models.

California’s SB-53 and New York’s RAISE Act require companies to write, publish, and comply with such frameworks. The EU GPAI Code of Practice – which concretizes the EU AI Act – goes further by specifying minimum standards and requiring certain measures to be state-of-the-art.

Most frontier AI companies had already published safety frameworks before these laws existed. They now face a choice: adapt their existing framework to satisfy these laws or create a separate compliance framework. Anthropic is the first company to take the latter approach.

In December 2025, Anthropic introduced its Frontier Compliance Framework (FCF), designed to meet its regulatory obligations. It then published an update in March 2026.

The FCF sits alongside Anthropic’s Responsible Scaling Policy (RSP), which was first published in 2023 and most recently updated in April 2026.

Anthropic describes the RSP as its voluntary safety framework. It goes beyond what the law requires. For example, it includes Risk Reports, a Frontier Safety Roadmap, and industry-wide recommendations. We provide an overview of the last major update here.

In this post, we answer three questions: (1) What are the main similarities between the RSP and FCF? (2) What are the main differences? and (3) Why did Anthropic create a separate compliance framework? We close with some reflections on what the split means for Anthropic and other frontier AI companies.

The two frameworks share a common foundation.

  • Both focus on catastrophic risks from frontier models. They both describe how Anthropic intends to assess and mitigate potentially catastrophic risks from its most capable models.

  • Both define risk thresholds. Thresholds play a central role in both documents. The RSP refers to them as “capability or usage thresholds” (and uses AI Safety Levels to label some of the current mitigations in place),1 whereas the FCF refers to them as “risk tiers.” The descriptions are almost identical across the two documents (e.g. for CBRN, high-stakes sabotage, and automated R&D).

  • Both link mitigations to capabilities. The RSP sets out the mitigations needed to keep risk at an acceptable level at each threshold. The FCF states that mitigations will be proportionate to the risk level, but does not specify what those mitigations would be.

  • Both commit to public reporting. The FCF commits to Model Reports, which contain information about the capabilities and mitigations of a specific model. The RSP commits to Risk Reports, which assess the overall level of catastrophic risk posed by Anthropic’s models.

  • Both involve external input. The RSP commits to external review of Risk Reports in some cases. The FCF commits to drawing on commissioned research and domain experts when assessing risk.

Although the two frameworks overlap, they differ in important ways.

  • The FCF is legally binding; the RSP might not be. The FCF is intended to comply with several regulatory obligations. If Anthropic violates them, regulators in California, New York, and the EU may take enforcement actions, such as issuing fines. It remains to be seen how regulators will treat the RSP.

  • They define risk differently. The RSP focuses on extreme risks, such as “existential threats or fundamental destabilization of global systems.” The FCF uses lower statutory thresholds: more than 50 fatalities, or more than $1 billion in damages. As a result, the two frameworks may trigger action under different conditions.

  • The risk categories do not fully match. The RSP covers chemical and biological weapons, high-stakes sabotage, and automated R&D. The FCF covers all of these but also adds cyber offense and harmful manipulation.

  • The RSP distinguishes unilateral commitments from industry-wide recommendations. The RSP distinguishes between what Anthropic will do regardless of its competitors (e.g. maintaining ASL-3 protections), and what it considers necessary across the industry (e.g. RAND Security Level 4). Many of the RSP’s more ambitious commitments fall in the latter category. Anthropic will only take certain actions if competitors have strong safety measures or Anthropic is in the lead. The FCF does not have this structure; all of its commitments are framed as Anthropic’s own.

  • They take different positions on pausing. Under the FCF, Anthropic will only proceed with development and deployment if risk is deemed to be at an acceptable level. The RSP no longer implies that Anthropic will pause if it can’t keep risks below acceptable levels. Holden Karnofsky, who led the RSP v3.0 update, acknowledges that this is the biggest change compared to previous versions.2

  • The RSP makes additional transparency commitments. In addition to Risk Reports described above, the RSP also commits to publish a Frontier Safety Roadmap setting out non-binding safety goals. By contrast, the FCF’s Model Reports are calibrated to regulatory requirements.

Anthropic’s decision to split the FCF from the RSP was likely driven by the following reasons.

  • To reduce legal exposure from ambitious commitments. The RSP includes ambitious commitments and industry recommendations, such as RAND Security Level 4 protections and “eyes on everything” monitoring of internal AI development. A separate voluntary framework lets Anthropic make such commitments with less enforcement risk. But regulators might still consider the RSP when assessing compliance, not just the FCF.

  • To leave less severe risks out of the RSP. Anthropic seems to think that some risk areas are less important from a safety perspective. Most notably, the RSP doesn’t cover risks from harmful manipulation and cyber. Though Anthropic may have changed its views on cyber after the Mythos release.

  • To keep the RSP easy to read. Anthropic was likely concerned that adding legal language to the RSP would have made it longer and harder to read. The split allowed Anthropic to keep the RSP clean. However, we think Anthropic could have achieved the same goal by moving legal content to footnotes and appendices. It’s also unclear how much legal language the regulations actually require.

  • To update the RSP faster. The RSP has been updated six times since September 2023 and it is intended to be a “living document.” Keeping the RSP separate from the FCF might make it easier to update in the future, since its voluntary commitments might not need to be reviewed as closely by its legal team. That said, the RSP may still require legal sign-off given Anthropic says it “may serve some regulatory requirements.”

The RSP-FCF split is a significant departure from current industry practice. Here are some reflections on what the implications might be.

  • The RSP seems to be the main framework. Anthropic seems to treat the FCF mostly as a compliance exercise, and the RSP as its “real” framework for managing catastrophic risk. We expect regulators in California, New York, and the EU to push back against this.

  • Reading either document alone gives you an incomplete picture. The RSP fills some gaps left by the FCF (e.g. on whistleblowing and anti-retaliation). There are also hidden dependencies between the two documents (e.g. the FCF says Anthropic will tailor mitigations to risk levels, and the RSP provides that mapping). As a result, regulators may end up treating both documents as part of Anthropic’s compliance framework, which would partly defeat the purpose of the split.

  • The two frameworks should stay consistent. If Anthropic keeps two separate frameworks, they should use the same concepts and definitions, and ensure that they do not contradict each other. Anthropic should also coordinate updates across teams to keep the frameworks aligned over time.

  • Other companies will face the same choice. SB-53, the RAISE Act, and the EU AI Act also apply to OpenAI, Google DeepMind, and others. Whether they follow Anthropic’s approach or keep a single framework will shape the future of safety frameworks.

  • Overall, we think the split was probably not necessary. Anthropic could have addressed most concerns by moving legal language and “less important” commitments to footnotes or appendices, and using conditional language to reduce legal exposure.

Acknowledgements: Thanks to Aidan Homewood, Elias Groll, Markus Anderljung, and Zaheed Kara (in alphabetical order) for helpful feedback on earlier drafts. All remaining errors are our own.

Disclaimer: Posts are written by individual team members and reflect the author’s perspective. Not all team members necessarily agree with every take. The views expressed here do not represent the official position of GovAI.

1

The ASL concept plays a smaller role in the latest RSP than in previous versions, where it was used to define required safeguards for future capability levels. Read about the latest major RSP update here.

2

In the v3.1 update, Anthropic clarified that it would “strongly consider pausing development and/or deployment” regardless of its competitors.

No posts

Read the original on frontierrisk.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.