RSS Amplifier

On Brains, Minds, And Their Possible Uses · May 17, 2026

The suspense budget

0
Sign in to vote or save

Jan Hendrik Kirchner · On Brains, Minds, And Their Possible Uses

An update of approximately zero bits. Stylised to preserve people’s privacy.

In case anybody was wondering what I am reading in my free time, it’s a (slightly cheesy) spy novel about an AI researcher who is tasked with aligning a superintelligence at a top-secret government facility. I went into it with slight trepidation (see above), but it was an overall fun read. Max Harms gets the texture of what it’s like to work on these projects pretty right, and the crumbs of his corrigibility theory sprinkled throughout are enjoyable.

But something still fell flat1. In this post, I’ll use math to show that this is not just, like, my opinion. (Also, mild? spoiler warning2.)

You cannot expect to change your mind in a predictable direction. I mean, you can, but you shouldn’t! If you know what you are going to believe in the future, you really should be believing it today already. This is Conservation of Expected Evidence (CoEE) and it’s one of the principles of rationality that I personally find quite hard to internalize in everyday life3. (In research and finance, CoEE comes somewhat more naturally.)

Yet there’s one non-technical domain where CoEE comes naturally: fiction. If you read enough thrillers, you stop being surprised by the third-act betrayal. If you read enough romance, you stop wondering whether they end up together. Genre-awareness lets you skip past the emotional arc of a slow realization and adopt the resulting belief immediately4.

Let’s do some math to this.

One way to model reading a novel is to think in terms of belief updating. We start with some prior expectations from knowing the author or the genre, and then receive updates about some central question with each chapter. For simplicity, let’s consider w.l.o.g. a coin that is biased toward heads or toward tails (you do not know which), and someone is showing you some subset of flips. What can you learn about the bias?

Building a schematic like this used to take hours and now I just have an image model do it for me! (Ignore how stuff isn’t centered properly.)

A trustworthy & polite author just shows you the flips. They built a world, things happened in it, and they report them in roughly the order a reasonable witness would notice them. Your belief converges smoothly with big updates early when you are uncertain, small ones late when you have things mostly figured out. This is what e.g. AI 2027 and gwern/clippy do.

The other extreme is the murder mystery à la Christie, Knives Out, Death Note. The author is not trustworthy or polite; she shows you a curated sequence designed to maximize the number of times your suspicion swings from one suspect to the next. The first time you read Murder on the Orient Express (and watch Knives Out) (and get to L's reveal in Death Note) your mind is blown roughly twelve times before lunch. That is certainly fun, but this arbitrage cannot last forever5.

Once you have seen sufficient examples from a given genre (your tenth Christie), the author's signal is no longer direct evidence about the coin; it's evidence about the author. The evidence now becomes conditioned on what you know about the genre, and in the case of a Christie your belief barely moves when the butler turns out to be involved in chapter 126. This is CoEE in action.

Three example trajectories of the reader’s belief evolving over time, each starting with a prior of 0.6 (top), and the total squared change in belief for the example trajectory (color) and averaged over a thousand trajectories (solid black) (bottom).

Now here is a curious observation. Simulate a bunch of belief trajectories across many author policies, and one quantity stays preserved for a reader that respects CoEE: the total sum of expected squared changes in belief, which is roughly how much you expect your beliefs to change while reading the story7. That quantity intuitively corresponds to how much you are ‘on the edge of your seat’; call it “suspense”8.

The kicker is that the total suspense is not only conserved, it is actually a straightforward quantity that only depends on the reader’s prior probability.

(With CoEE, beliefs μ₀ … μₙ form a martingale, so the increments are uncorrelated and the expected squared moves telescope:

\(\sum_t \mathbb E\left[(\mu_{t+1} - \mu_t)^2\right] = \mathbb{E} \left[\mu_n^2\right] - \mu_0^2,\)

where n is the final index. If the story fully resolves the binary question9, μₙ ∈ {0,1}, so E[μₙ²] = E[μₙ] = μ₀, and the whole sum collapses to μ₀(1 − μ₀) which is just the Bernoulli variance of the prior.)

For a genre-savvy reader who has coherent beliefs about the author’s strategy, no clue, no reveal, no act break can add suspense above what the reader arrived with.

And it gets worse.

We've established that you are no longer susceptible to whodunit shenanigans if you're genre-savvy and coherent. But if the author does their work well10, at least we can still expect some nontrivial amount of suspense as we read the book. Unless, of course, it’s a book about superintelligence.

On its face, superintelligence is the perfect suspense engine since the whole premise of the thing is that it does what you, a merely human, did not see coming. However, unpredictability of moves is not the same as uncertainty about outcomes. Even if the author is very smart indeed and fakes a decent superintelligence, ‘I cannot predict what it will do next’ does not imply ‘I do not know how this ends.’ The superintelligence will win, that’s what it means to be superintelligent. (Yudkowsky calls this Vingean uncertainty)11.

A well-done whodunit (that doesn’t pick the butler more often than implied by a uniform distribution) still has substantial suspense to offer even for a reader who is genre-savvy (left). Superintelligence fiction, however, has kind of painted itself into a corner (right).

For a genre-savvy reader of superintelligence rationalist fiction, there's almost no suspense left to distribute. The reader’s prior about the outcome is strongly concentrated. The total available suspense is determined by the variance of this prior, μ₀(1 − μ₀), which is very small. Nothing that happens in the book matters in terms of suspense.

The last section was almost entirely bad news, and yet there is superintelligence fiction that works. Friendship is Optimal is a ton of fun. What I did in the hedonium shockwave, by Emma, age six and a half hits you right in the feels. How is that?

Move the question. The total suspense budget is computed relative to some question, and nothing forces the story to spend its belief mass on a question that is as well-trodden and pre-constrained as ‘will the AI take over?’ The Chorus opens Romeo and Juliet by telling you, in the prologue, that the lovers die so we can focus on the how and why. The hedonium shockwave story’s inevitability is conceded at the outset, and the live question becomes who the narrator is and how her world metabolizes a doom it cannot avert.

Break the reader’s model of you. The total suspense budget is only invariant if the reader has a calibrated model of the author’s policy. If the author subverts those expectations, then there is extra suspense to be had. That is, at least until the audience prices the subversion in, at which point you must subvert the subversion12. planecrash does this well by subverting the mechanism of superintelligence. The gods are superintelligent and qualitatively above mortals but their attention is splintered across the whole of reality and they never bring the full brunt of their intelligence to bear.

Expand the set of outcomes. The total suspense budget also assumes that the set of possible outcomes is known to the reader. Instead of resolving the central question in the expected way, the author might also say mu and reject the framing. This one is easy to mess up (if the reveal of a third alternative is done clumsily then it can feel like betrayal), but one successful instance of mu is Asimov’s The Last Question.

So did Harms write an unexciting book? No, I don’t think so. If you’re a reader who has spent years marinating in frontier-AI development and rationalist fiction then you can see the dramatic beats of the story coming from a mile away. But to most everyone else this criticism doesn’t apply. What remains is a pretty enjoyable story about one of the most important topics of our times that doesn’t dumb down the technical bits. You should recommend it to an uncle or someone.

1

Obviously I enjoyed the book a ton, given that I’m writing a whole blog post about it on a Saturday!

2

Although by the end of this post I’ll have argued there wasn’t much to spoil.

3

This consistency checking across time is not how the mind wants to work. “Delaying a decision”, “self-deception”, and “intentionally not looking” are pretty natural mechanisms for protecting oneself from painful realizations. (And these short-term mechanisms can then turn maladaptive over the long term, because they have no well-defined end-condition.)
But also, taking “the scenic route” to a realization kind of is a big part of what life is about. Jumping immediately to the conclusion seems wrong when emotions are involved.

4

I believe this is why a pre-mortem is so powerful: it is a forcing function for obtaining CoEE coherence. So is reading HPMOR for the seventh time.

5

Abram Demski showed that any computable Bayesian reasoner who updates on logical evidence without modeling its source can be driven arbitrarily up and down, forever, by an adversary who only ever says true things. He called it “all mathematicians are trollable”, and I referenced it in my 2021 XMas special Belief-conditional things. Demski’s trolled mathematician is not failing to model the source out of naivety; they are failing because a computable prior on logic cannot be fully coherent (Gödel et al.), so there are always unrecognized consequences for the troll to selectively reveal.

6

Demski called the construction “an untrollable mathematician.” The untrollable mathematician’s belief stays inside a bounded band no matter what the troll picks, because the troll’s selection is itself evidence and the reasoner is using it.

7

While a CoEE coherent agent can’t expect their beliefs to change in a particular direction, they can totally have expectations that their beliefs will change a lot or a little.

8

This whole argument is due to Ely, Frankel, and Kamenica who take a bog-standard result about martingales and give this wonderful interpretation in terms of suspense and surprise.
You may know Kamenica from Bayesian persuasion (with Gentzkow, 2011), which is the sibling result: a sender commits to an information policy, a rational receiver Bayes-updates on it, and the sender’s only constraint is that the receiver’s expected posterior equals their prior. That constraint is the martingale property. Suspense-and-surprise is what happens when you run Bayesian persuasion forward through time and ask what the receiver experiences along the way.

9

Without it the conserved quantity is E[μ_T²] − μ₀², and an ambiguous ending has a smaller budget.

10

The EFK paper actually gives us a recipe for how to write suspense-optimally. They assume that utility is concave in suspense (somewhat plausible), then the optimum is to distribute suspense evenly across the entire book (same argument as here!). This does explain why The Da Vinci Code would have been so wildly popular. It also reveals that this is not a model of good literature per se (see again The Da Vinci Code).

11

The budget identity only considers your prior over the terminal question (does it get out, does alignment hold) never the entropy of the path that gets you there! “I can’t predict its next move” is path-entropy; suspense, in the sense the identity conserves, is variance in the outcome belief. Vingean uncertainty is exactly the regime where the first is maximal and the second is near zero, which is why the intuition that unpredictability buys suspense fails precisely here.

No posts

Read the original on universalprior.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.