RSSAmplifier

Blog

Joseph Carlsmith RSS Feed

Joseph Carlsmith's Personal Website

joecarlsmith.comRSS feed ↗94 posts

Latest posts

Video and transcript of talk on writing AI constitutions

From a talk at Yale Law School in March 2026.

On restraining AI development for the sake of safety

My take on slowing down AI.

Building AIs that do human-like philosophy

AIs will face philosophical questions humans can t answer for them.

Video and transcript of talk on human-like-ness in AI safety

From a talk I gave at Constellation in December 2025.

How human-like do safe AI motivations need to be?

AIs with alien motivations can still follow instructions safely on the inputs that matter.

Leaving Open Philanthropy, going to Anthropic

On a career move, and on AI-safety-focused people working at AI companies.

Controlling the options AIs can pursue

On blocking paths to power, and on making deals.

Video and transcript of talk on giving AIs safe motivations

From a talk at UT Austin in September 2025.

Giving AIs safe motivations

A four-part picture.

Video and transcript of talk on “Can goodness compete?”

From a public talk on long-term equilibria post-AGI, given at Mox in SF in July 2025.

Video and transcript of talk on AI welfare

An overview of my take on AI welfare as of May 2025, from a talk I gave at Anthropic.

The stakes of AI moral status

On seeing and not seeing souls.

Video and transcript of talk on automating alignment research

From a talk at Anthropic in April 2025.

Can we safely automate alignment research?

It s really important; we have a real shot; there are a lot of ways we can fail.

AI for AI safety

We should try extremely hard to use AI labor to help address the alignment problem.

Paths and waystations in AI safety

On the structure of the path to safe superintelligence, and some possible milestones along the way.

When should we worry about AI power-seeking?

Examining the conditions for rogue AI behavior.

How do we solve the alignment problem?

Introduction to an essay series about paths to safe, useful superintelligence.

What is it to solve the alignment problem?

Also: to avoid it? Handle it? Solve it forever? Solve it completely?

Fake thinking and real thinking

When the line pulls at your hand.

Takes on “Alignment Faking in Large Language Models”

What can we learn from recent empirical demonstrations of scheming in frontier models?

Video and transcript of presentation on Otherness and control in the age of AGI

An attempt to distill down the whole Otherness and control series into a single talk.

(Part 2, AI takeover) Extended audio/transcript from my conversation with Dwarkesh Patel

Extra content includes: AI collusion; the nature of intelligence; concrete takeover scenarios; flawed training signals; tribalism and mistake theory; more on what good outcomes look like.

(Part 1, Otherness) Extended audio/transcript from my conversation with Dwarkesh Patel

Extra content includes: regretting alignment; predictable updating; dealing with potentially game-changing uncertainties; intersections between meditation and AI alignment; moral patienthood without consciousness; p(God).

Loving a world you don’t trust

Garden, campfire, healing water.

On attunement

Examining a certain kind of meaning-laden receptivity to the world.

Video and transcript of presentation on Scheming AIs

An intro to my work on scheming/ deceptive alignment.

On green

Examining a philosophical vibe that I think contrasts in interesting ways with deep atheism.

On the abolition of man

What does it take to avoid tyranny towards the future?

Being nicer than Clippy

Let s be the sort of species that aliens wouldn’t fear the way we fear paperclippers.

An even deeper atheism

Who isn t a paperclipper?

Does AI risk “other” the AIs?

Examining Robin Hanson s critique of the AI risk discourse.

When “yang” goes wrong

On the connection between deep atheism and seeking control.

Deep atheism and AI risk

On a certain kind of fundamental mistrust towards Nature.

Gentleness and the artificial Other

AIs as fellow creatures. And on getting eaten.

Otherness and control in the age of AGI

Introduction and summary for a series of essays about how agents with different values should relate to one another, and about the ethics of seeking and sharing power.

New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”

My report examining the probability of a behavior often called deceptive alignment.

Superforecasting the premises in “Is power-seeking AI an existential risk?”

Superforecasters weigh in on the argument for AI risk given in my report on the topic.

In memory of Louise Glück

It was, she said, a great discovery, albeit my real life.

Predictable updating about AI risk

How worried about AI risk will we be when we can see advanced machine intelligence up close? We should worry accordingly now.

Existential Risk from Power-Seeking AI (shorter version)

Building a second advanced species is playing with fire.

A Stranger Priority? Topics at the Outer Reaches of Effective Altruism

My dissertation.

Seeing more whole

On looking out of your own eyes.

Why should ethical anti-realists do ethics?

Who needs a system if you re free?

On sincerity

Nearby is the country they call life.

Against meta-ethical hedonism

Can the epistemology of consciousness save moral realism and redeem experience machines? No.

Against the normative realist’s wager

If you find a button that gives you a hundred dollars if a certain controversial meta-ethical view is true, but you and your family get burned alive if that view is false, should you press the button? No.

Is Power-Seeking AI an Existential Risk?

Report for Open Philanthropy examining what I see as the core argument for concern about existential risk from misaligned artificial intelligence.

Video and Transcript of Presentation on Existential Risk from Power-Seeking AI

Video and transcript of a presentation I gave on existential risk from power-seeking AI, summarizing my report on the topic.

Dutch books, Cox, and Complete Class

Final essay in a four-part series on expected utility maximization (EUM). Examination of some theorems that aim to justify the subjective probability aspect of expected utility maximization (EUM), namely: Dutch Book theorems; Cox’s Theorem; and the Complete Class Theorem.