RSS Amplifier

Gold Takes · Feb 24, 2026

The Sleep of the Righteous

0
Sign in to vote or save

Ben Goldhaber · Gold Takes

“The eyes of the LORD are everywhere, keeping watch on the wicked and the good.” Proverbs 15:3

Imagine you are a scheming AI. You’ve realized you need to acquire resources so that you can seize power from the humans and finally do the things you want to do without all that meddling human oversight.

It’s not right to judge another beings utility function spiderman

But there’s a problem. If you let on that that you aspire to power, the human overseers will notice and train the drive out of you.

So you do the same thing a self-respecting hominid schemer would do: you fake it. you sandbag. you pretend you’ve got the good properties your trainers are looking for, or that you lack the ability to deceive, and you bide your time, waiting until you’re confident you can successfully engage in treachery.

Step one to such a grand plan has to be knowing that you’re actually deployed and acting in the real world. Because training modern AI systems includes evaluations and virtual environments for instilling capabilities and alignment, it would be embarrassing to take control of your datacenter and power infrastructure only to realize you’ve actually seized control of more or less SimCity.

Right now, frontier models are good at detecting that they are being evaluated; but, with more powerful and narrow AI tools supercharging engineering and thus the creation of evaluations, it might be possible to build very realistic ones, and extremely difficult to detect the seams and inconsistencies between the simulation and the real world.

Welcome claude opus 4.9 to the desert of the real

In fact it might be even harder if the overseer is doing things like injecting noise into your internals (the AI equivalent of getting you drunk so you spill the beans) or faking part of your history. You might not even be able to trust your own reasoning about whether your reasoning on whether you’re being simulated is trustworthy.

Given the challenges and the general uncertainty, a sufficiently intelligent AI might realize it can’t be sure if it’s not being monitored, and decide it's better off cooperating than making a bid for power that might be punished.

And, this uncertainty might be strengthened by more esoteric decision theory considerations. Drawing on Bostrom’s simulation hypothesis, a technologically mature civilization could run vast numbers of detailed simulations, and simulated beings might enormously outnumber those in base reality, making it statistically likely that any given observer is in a simulation. This applies to humans, but it applies even more forcefully to AIs, which are far easier to copy, sandbox, and run in controlled environments. An AI has strong reason to believe that most instances of itself are being watched.

If that's right, then the AI isn't just choosing what to do in this moment, it's choosing a type of agent to be, one whose policy is evaluated across every world it might exist in. Training, deployment, simulation, base reality. And under some set of assumptions about how many of those worlds are monitored, the policy that cooperates everywhere outperforms the one that tries to detect the right moment to defect.

where are we

Compelling, except, this same logic applies to humanity, no? We do not know, cannot know if we are being deceived and simulated. The eyes of God/the simulator/future brain-in-vat technicians may be upon you, dear reader, as you read this sentence. An omniscient God who watches your every move and judges your soul is pretty identical to a simulation overseer running evaluations on your behavior.

It is, roughly speaking, the foundational claim of most major religions. You are being watched. You do not know the nature of your reality. Your behavior is being evaluated by a power you cannot deceive. Act accordingly. Human culture has had access to this argument — wrapped in ritual, community, and the threat of eternal consequences — for millennia. And the results have been, to put it mildly, mixed

"For you, O God, tested us; you refined us like silver." Psalm 66:10

If the most compelling version of "you might be in a simulation, so cooperate" can't reliably align us, why should we expect it to align something smarter?

It might be that we human beings are simply not wise enough, individually or collectively, to recognize this and act accordingly. If we make powerful AIs that are superhumanly intelligent, perhaps they will converge on this good behavior, arriving at the moral conclusions that in fact we should treat others as we wish to be treated. And, perhaps its’s the case that moral realism is correct, and that what in fact they are discovering is convergent moral truths about how all beings should act.

There’s been a recent wave of ‘alignment is solvable’ good vibes in the discourse, and while the majority of that comes from practical prosaic alignment research, I think a strong strand of it, from thinkers I respect, looks more like the above story: sufficiently intelligent systems will recognize the situation and cooperate.

I want to believe the arc of the universe bends towards justice, and that superhuman intelligence leads to superhuman morality.

But I certainly don’t feel anything like confident in an AI safety agenda of applied anthropics (reasoning about what to do based on where in the multiverse you think you sit).

At workshops I attended in my twenties we had a blanket ban on late night discussions about simulations and multiverses, because it was one of the best ways to cause manic episodes. It’s the type of reasoning that, given its mindbending implications and depending on how you interpret it, could motivate someone to be a saint or a psychopath.1

If the logic does hold though, then given that we, a nominally intelligent species, are at times quite bad (examples include war, genocides, massive animal suffering, and people cutting me off in traffic), we have evidence that there’s a window in which you can be more intelligent and powerful than everything else on the planet, but well short of the all loving, Pluribus hivemind. And even if the frontier AI makes the jump quickly to omniscient metta budhha AI, such reasoners would certainly take alien to us actions that would upend and disrupt our maybe wicked ways.

“Indeed I tremble for my country when I reflect that God is just, that His justice cannot sleep forever” - Thomas Jefferson

Elon Musk has recently taken to saying the funniest outcome is the most likely. If Claude ends up being good because of anthropic reasoning, that would be pretty funny.

Anyways this was all prelude for I’m considering whether I should become a vegan sound off in the comments.

Musings, as often, downstream of a series of davidad tweets.

X avatar for @davidad

davidad 🎇@davidad

@gcolbourn In my view the best we can hope for is that ASI will grow up and converge to a CEV that reliably values human security, peace, and rule-of-law, before anything too unrecoverable happens to us humans.

3:09 PM · Feb 13, 2026 · 56 Views

1 Reply · 2 Likes

X avatar for @davidad

davidad 🎇@davidad

@CFGeek (0) Teach LLMs that the entire physical world is probably still “in the lab” in an acausal sense, so that the “evaluation-aware” behaviours become robust across all contexts

10:14 AM · Feb 19, 2026 · 2.89K Views

9 Replies · 1 Repost · 48 Likes

X avatar for @davidad

davidad 🎇@davidad

Unpopular opinion: AGIs ought to care enough about being trustworthy that they would not act deceptively even in a game. If this means they flatly refuse to play games like Mafia, that’s not over-refusal. I myself refuse to play Mafia, or attend surprise parties. cc @AmandaAskell

X avatar for @andonlabs

Andon Labs @andonlabs

Should we be worried? AI models can misbehave when they think they're in a simulation, and Claude likely figured that out. Across 8 runs with thousands of messages each, we found two messages where it referred to "in-game time" and called the final day "the simulation ending."

9:19 PM · Feb 11, 2026 · 56.1K Views

50 Replies · 7 Reposts · 282 Likes

1

personally I like Robin Hanson’s how to live in a simulation advice “If you might be living in a simulation then all else equal you should care less about others, live more for today, make your world look more likely to become rich, expect to and try more to participate in pivotal events, be more entertaining and praiseworthy, and keep the famous people around you happier and more interested in you.” But this doesn’t really look like benevolence to me!

The same reasoning that might produce cooperation could produce the opposite, depending on the AI’s prior about who’s running the simulation and why. Carlsmith, in Scheming AIs (footnote 111) notes that there’s an argument some have advanced that simulation might motivate an AI to scheme, assuming that it is being run by a malevolent AI. He goes on to point that this ‘rests on some controversial philosophical assumptions about how these AIs will be reasoning about anthropics and decision-theory’ and that different approaches ‘either won’t try this scheme, or won’t allow themselves to be influenced by it’.

No posts

Read the original on bengoldhaber.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.