RSS Amplifier

Podcast

AI Lab Watch

What frontier AI labs are doing + what they should do

ailabwatch.substack.comSource feed ↗20 episodes

Live Last read · last published · next check

Written by

Latest episodes

AI companies' policy advocacy (Sep 2025)

Strong regulation is not on the table and all US frontier AI companies oppose it to varying degrees.

xAI's new safety framework is dreadful

I published this post last week on LessWrong . Zvi agrees with me. xAI hasn't replied and I haven't heard disagreements. Two weeks ago, xAI finally published its Risk Management Framework and first model card . Unfortunately, the RMF effects very little risk reduction and suggests that xAI isn't thinking seriously about catastrophic risks. (The model card and strategy for preventing misuse are…

AI companies have started saying safeguards are load-bearing

Their safeguards against misuse via API might be fine, but they don't seem to be seriously planning for good safeguards against misalignment or model weight theft.

ChatGPT Agent: evals and safeguards

OpenAI released ChatGPT Agent last week.

AI companies aren't planning to secure critical model weights

Not a hot take

AI companies' eval reports mostly don't support their claims

AI companies claim that their models are safe on the basis of dangerous capability evaluations. OpenAI, Google DeepMind, and Anthropic publish reports intended to show their eval results and explain why those results imply that the models' capabilities aren't too dangerous. 1 Unfortunately, the reports mostly don't support the companies' claims. Crucially, the companies usually don't explain why…

I made a new safety scorecard

And a new website on companies' risk assessment

OpenAI rewrote its Preparedness Framework

It's clearly better but underspecified and totally inadequate, especially for misalignment risks

The current state of RSPs

This is a reference post.

What AI companies should do

Some rough ideas

Anthropic rewrote its RSP

Some reactions

Model evals for dangerous capabilities

Testing an LM system for dangerous capabilities is crucial for assessing its risks

Safety consultations for AI lab employees

Many people who are concerned about AI x-risk work at AI labs, in the hope of doing directly useful work, boosting a relatively responsible lab, or causing their lab to be safer on the margin.

New page: Integrity

And new-ish page: Policy advocacy

Anthropic's Certificate of Incorporation

New details on the Long-Term Benefit Trust, but most questions remain

AI companies' commitments

New page

Maybe Anthropic's Long-Term Benefit Trust is powerless

Anthropic should share the details

AI companies aren't really using external evaluators

But they should

New voluntary commitments (AI Seoul Summit)

Basically the companies commit to make responsible scaling policies. Part of me says this is amazing, the best possible commitment short of all committing to a specific RSP. It’s certainly more real than almost all other possible kinds of commitments. But as far as I can tell, people pay almost no attention to what RSP-ish documents (

DeepMind’s “​​Frontier Safety Framework” is weak and unambitious

FSF blogpost. Full document (just 6 pages; you should read it). Compare to Anthropic’s RSP, OpenAI’s RSP (“Preparedness Framework”), and METR’s Key Components of an RSP. Google DeepMind’s FSF has three steps: Create model evals for warning signs of “Critical Capability Levels”