RSS Amplifier

Podcast

Crossing the Rubicon

Posts about AI alignment. Victory or death.

crossingtherubicon.substack.comSource feed ↗10 episodes

Live Last read · last published · next check

Latest episodes

Predictions Over What?

Accuracy is only half the battle

Catching Misalignment vs. Fixing It

What happens if evaluations, interpretability, and monitoring succeed?

Emergent Misalignment, Inoculation Prompting, and Unsolved Problems

The AI Alignment ecosystem working as intended?

Solutions to Corrigibility, Deceptive Alignment, Reward Tampering, Monitorability, and Reward Hacking

I am not a crackpot

Probably Safe AI

A Case for Marginal Progress on Superalignment Theory

Myopic Goals without Myopic Capabilities

How we can get AI to set up a good future without caring about it

Intelligence Explosion Macroeconomics

Economists believe in an intelligence explosion, even if they don’t realize it yet

Safe Predictive Agents with Joint Scoring Rules

It's time to talk about what I've been working on

Myopia Is All You Need (For Alignment)

That's not strictly true, but myopia is still incredibly useful

Simplifying Corrigibility – Subagent Corrigibility Is Not Anti-Natural

Ensuring corrigibility in an AI and ensuring corrigibility in the agents it creates are two completely distinct problems.