Video and transcript of talk on writing AI constitutions
From a talk at Yale Law School in March 2026.
Joseph Carlsmith's Personal Website
From a talk at Yale Law School in March 2026.
My take on slowing down AI.
AIs will face philosophical questions humans can t answer for them.
From a talk I gave at Constellation in December 2025.
AIs with alien motivations can still follow instructions safely on the inputs that matter.
On a career move, and on AI-safety-focused people working at AI companies.
On blocking paths to power, and on making deals.
From a talk at UT Austin in September 2025.
A four-part picture.
From a public talk on long-term equilibria post-AGI, given at Mox in SF in July 2025.
An overview of my take on AI welfare as of May 2025, from a talk I gave at Anthropic.
On seeing and not seeing souls.
From a talk at Anthropic in April 2025.
It s really important; we have a real shot; there are a lot of ways we can fail.
We should try extremely hard to use AI labor to help address the alignment problem.
On the structure of the path to safe superintelligence, and some possible milestones along the way.
Examining the conditions for rogue AI behavior.
Introduction to an essay series about paths to safe, useful superintelligence.
Also: to avoid it? Handle it? Solve it forever? Solve it completely?
When the line pulls at your hand.
What can we learn from recent empirical demonstrations of scheming in frontier models?
An attempt to distill down the whole Otherness and control series into a single talk.
Extra content includes: AI collusion; the nature of intelligence; concrete takeover scenarios; flawed training signals; tribalism and mistake theory; more on what good outcomes look like.
Extra content includes: regretting alignment; predictable updating; dealing with potentially game-changing uncertainties; intersections between meditation and AI alignment; moral patienthood without consciousness; p(God).
Garden, campfire, healing water.
Examining a certain kind of meaning-laden receptivity to the world.
An intro to my work on scheming/ deceptive alignment.
Examining a philosophical vibe that I think contrasts in interesting ways with deep atheism.
What does it take to avoid tyranny towards the future?
Let s be the sort of species that aliens wouldn’t fear the way we fear paperclippers.
Who isn t a paperclipper?
Examining Robin Hanson s critique of the AI risk discourse.
On the connection between deep atheism and seeking control.
On a certain kind of fundamental mistrust towards Nature.
AIs as fellow creatures. And on getting eaten.
Introduction and summary for a series of essays about how agents with different values should relate to one another, and about the ethics of seeking and sharing power.
My report examining the probability of a behavior often called deceptive alignment.
Superforecasters weigh in on the argument for AI risk given in my report on the topic.
It was, she said, a great discovery, albeit my real life.
How worried about AI risk will we be when we can see advanced machine intelligence up close? We should worry accordingly now.
Building a second advanced species is playing with fire.
My dissertation.
On looking out of your own eyes.
Who needs a system if you re free?
Nearby is the country they call life.
Can the epistemology of consciousness save moral realism and redeem experience machines? No.
If you find a button that gives you a hundred dollars if a certain controversial meta-ethical view is true, but you and your family get burned alive if that view is false, should you press the button? No.
Report for Open Philanthropy examining what I see as the core argument for concern about existential risk from misaligned artificial intelligence.
Video and transcript of a presentation I gave on existential risk from power-seeking AI, summarizing my report on the topic.
Final essay in a four-part series on expected utility maximization (EUM). Examination of some theorems that aim to justify the subjective probability aspect of expected utility maximization (EUM), namely: Dutch Book theorems; Cox’s Theorem; and the Complete Class Theorem.