
AI Text Watermarking Is Free And Good
Scott Aaronson, while working at OpenAI, largely solved AI text watermarking together with Hendrik Kirchner.
A world made of gears. Doing both speed premium short term updates and long term world model building. Currently focused on weekly AI updates. Explorations include AI, policy, rationality, medicine and fertility, education and games.
Live Last read · last published · next check

Scott Aaronson, while working at OpenAI, largely solved AI text watermarking together with Hendrik Kirchner.

This was a week of quiet aftermath, an opportunity to process recent events and start to figure out the path forward.

OpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision.

I am grateful that Anthropic is producing periodic Risk Reports.

Some podcasts are self-recommending enough that I look to break them down if I have the chance.

The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters.

As AI has escalated increasingly quickly, more and more of my posts have ended up focusing on AI.

Pre Post Mortem

In the wake of the letter calling on us to prepare to potentially Pace the Frontier, there has been much discussion of when pacing the frontier would be prudent, and whether it makes sense to prepare to do so.

Today I am taking the time to write the shorter, simpler version of What Happened.

How does the situation keep turning out to be worse than we know?

What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.