RSS Amplifier

Performative Bafflement · Jul 21, 2026

Contra Plan A, Scott, and others - let’s not increase existential risk millionsfold quite yet

0
Sign in to vote or save

Performative Bafflement · Performative Bafflement

Alright, so the good folks at the AI Futures Institute have written Plan A and Scott Alexander has written several explainers on why we should do Plan A to avoid an AI dystopia.

Moreover, other luminaries in the AI strategy space, like Zvi Mowshowitz and Peter Wildeford (superforecasting specifically, in Peter’s case), have talked about similar plans and dynamics:

Scott’s original Plan A explainer

Scott’s follow up panopticon explainer

Peter’s

Zvi’s

Broadly, they want to loop the government in now, have them mechanically / programmatically control AI chips, give China and the US mutual self-destruction capabilities over datacenters / AI, and try to slow things down so we don’t all end up paperclipped or ruled over by a laughing tech oligarch god king whose fondest wish is to crush us all in his iron fist, merely to see the goo running between his fingers (that’s apparently what gets oligarchs off, you see, or so everyone seems to think).

I am here to argue their approach is a bad idea, and is straightforwardly increasing our existential risk millionsfold, and that there is a better alternative.

They are literally frantically facetiming Sauron on one of those crystal ball things (palantirs for the nerds out there, or the nominative determinists), yelling “Sauron, Sauron, we have a gigantic ring of power, and it keeps getting more powerful, please come and help us coordinate with the Chinese hobbits…wow, he’s already here! Yes, thank you Sauron, could you please tell the Chinese hobbits…wait, he’s picking it up. No! He’s putting it on! Nooooo!”

Scott has written a general outline on what Plan A is proposing, as well as a rejoinder / rebuttal to everyone saying “but this is a panopticon and a bad solution.”

His answer is essentially “lol, we’re already in a panopticon, suck it up, buttercup.”

Which yes, and don’t you think that fact has some bearing on what’s going to happen here? He’s plainly handwaving and avoiding the fact that what he’s recommending increases our existential risks millionsfold.

Yes, obviously we’ve proven our leaders in America are literally incapable of leading, reasoning, or even acting for anything but purely personal and individual gains at all over the last 10 years. But it gets worse. This isn’t a “but Trump!!” argument, it’s about incentives and existing power centers.

Who are the governments on-staff nerds, in the sense that the execution and / or oversight work are certainly going to be turned over to them? The NSA. As an aside, back when I was an undergrad math / physics major, the NSA would come courting the math department, and I always wondered who would ever want to work for the frigging NSA? And indeed, the only actually talented person that averred interest was by far the weirdest / most unhinged math major among us! The kind you’d say “oh, yeah, I guess that tracks” if 5-10 years in the future you heard of some smart Uncle-Ted style weirdo going off the rails and taking actions that led to a lot of deaths. Later in life, one of my good friends had substantial professional involvement with said organization, and indeed, from the things she said (and the lawsuits she became encumbered with), it was a raging dumpster fire in terms of the personalities of who consistently rose and were awarded more power in that organization. It is not our best and brightest, or even our “bright and within the parameters of acceptable mental illness and dysfunctional personalities”-est.

What other major power center exists, with a use case salient enough they’ll be certain to stake a claim? The military. Everyone’s worries are already about cyber (NSA’s domain) or killbots (military’s domain). Bio is there too but is obviously going to be ignored, because there’s no major “bio” power center in the government.

Now all these people like AI Futures Institute and Scott and Zvi others who have the ears of hundreds of thousands to millions of people are directly jumping up and down and waving and repeatedly Facetiming the Eye of Sauron to get it to look our way and reave our Shire.

This is a bad idea, because this essentially guarantees the NSA + military now steers all AI efforts, and this will definitely increase existential risk millionsfold. Because if we do it, China will do it, and then it’s two NSA + militaries directing an AI race.

But maybe it won’t immediately be turned over to the military+NSA!

Lol. Seriously? Why would we think this? Of the people who our putative “leaders” trust enough to pay attention to their technical synopses, what organizations even have a plausible chance of having people who can even PRETEND to understand frontier AI? Maybe, at a stretch, you could add in the CIA, but it’s not like that’s an improvement in the risk landscape. If anything, the more voices / power centers at the table, the more likely bad outcomes get, and the CIA is just as bad as the other two.

Broadly, this is inevitable because within the government, things are routed according to politics and power centers, rather than skill, competence, character, judgment quality, or anything else, and those are the two largest power centers with people who can even pretend to understand it. It’s as inevitable as gravity, because that’s the way politics even works, and I refer you to the book The Populist Delusion, or the works of Schmitt, Musca, Pareto, Jouvenel, or Burnham if you’re interested in the political science.

But we need government involved because it’s the only channel for average people to collate and express their preferences! And average people deserve a say in the future!

“Average people deserve a say in the future,” although an admirable aspiration, is impossible, because it’s already been established the only option there is “politics” and “politics” touching AI guarantees NSA+military direction. Average people have even less say there, and even more chances of bad outcomes.

Need I point out Anthropic already got enemy-of-the-stated for not consenting to spying on all Americans 24/7? And every single other frontier company began jumping over themselves to volunteer to do so when Anthropic refused?

One last point.

In an NSA+military regime, who is going to be leading the AI efforts, out of Altman, Amodei, and Hassibis? It is VERY clear that only one of those players is extremely good at politics, maybe 10-100x better than the others, so we’d end up with Altman running the day to day decisions, at the behest of the NSA+military.

So the worst AI leader will be leading the AI charge at the behest of the worst organizations, leading AI towards the worst possible ends, with China doing exactly the same.

I’ve directly raised this with comments on Scott’s and Zvi’s and Peter Wilderford’s posts on the issue several times now, and got zero responses from anyone, not just the headliners, but also not a single other commenter.

I’m sure Scott and AI Futures and sundry others will try to argue that the NSA + Military taking over is inevitable. What do you expect, PB? That they’re just going to let companies that can make the government irrelevant go off and do whatever they want to? Even basic self preservation is against that, much less the ever-present organizational drive towards keeping and consolidating power!

Power does what it does! We might as well get ahead of it and try to steer things to some infinitesimal speck of “better outcomes” in the infinite hellscape of all the bad military + NSA outcomes.

OR, hear me out here, we could just let things unfold like they are now, instead of frantically Facetiming Sauron over and over, waving and pleading for him to come ruin everything.

The clearly best outcome to ME is “let the nerds at Anthropic cook,” because they’ve established they’re in the noticeable lead and have good morals even when it’s costly. And Claude is already the most aligned AND most audited AI.

“But that’s unrealistic!” you might say. “Mythos already scared all the rubes, it’s on Sauron’s radar already!”

But is it? If we’re aiming for an infinitesimal speck in a millionsfold-more-risky hellscape, why not aim for the infinitesimal speck where we throw our weight behind “the AI CEO’s and workers are super smart and know what they’re doing?” politically, at least for the next year or two. We have that much leeway.

The AI CEO’s are already pretty politically savvy, with Altman the most and Amodei the least, but importantly, even Altman running cover solely for OpenAI largely also provides cover for Anthropic and Deep Mind.

The obvious dynamic once they’ve seen the outcomes of Fable and Sol getting “release at my personal suffraged” by Trump is for them to go heads down and use all further training runs to get smarter internally, to achieve RSI or substantial internal improvements that put them farther on the road to RSI.

I’m making an “aim for the infinitesimal nerd-driven RSI speck, not some made up infinitesimally better speck in the gargantuan military+NSA hellscape right away” argument.

We’re bringing literally $1-2T+ of compute online in the next couple of years. There’s a good chance that’s enough, if we can also scrounge up an architectural overhang,1 sample size efficiency, or realtime learning architecture.

Whatever large capability leaps they can make internally aren’t subject to Sauron’s scrutiny and their entire company getting Gitmoed so some third rank, mentally unwell nerds at the NSA can take over and start steering god-minds towards the bad ends they’ll inevitably be steered towards.

This ALSO cools down the race with China over those few years! If the race is still under commercial companies with commercial motives, that’s obviously a lot less risky and more palatable for China too!

Is it relatively unlikely we get true RSI with that $1-$2T of compute? It is, BUT I deem it much more likely than finding some made up infinitesimal “better outcomes” speck in the military+NSA outcomes. This can’t really be quantitatively argued, because we’re already far out in the tails and everyone’s respective probability judgments can differ pretty widely there. But just from a first principles approach, it seems pretty obvious to me that avoiding the “military+NSA+worst AI CEO are leading everything” landscape for as long as we can is a dominating better move on multiple fronts.

This is a simple timing argument, too. I agree, power does what it does. Once there are major real-world results from an AI going bad (big cyber thing with lots of deaths) or going really well (we counterfeit some noticeable chunk of white collar jobs), the government, and so the military+NSA, is inevitably going to step in.

But we don’t need to force that move right away, at this stage. They don’t have eyes on what’s going on internally, right now they’re only concerned with what’s released externally. We have a slim chance of a better outcome, and we should take it, and if it fails, THEN you can aim for your made up “better speck” in the hellscape.

We haven’t done anything to prove alignment, we have no plausible technical paths towards it, and RSI is intrinsically risky! We need to slow down to buy more time to get all that stuff! That’s literally the entire thesis of Plan A!

Yes, but why do you think military+NSA led AI is less risky??

Military+NSA AI is obviously many many OOMs (millions-fold) more risky than nerd-driven RSI, even with our current state of alignment!

And this is true for simple “goals and directions” reasons as well as “AI race” reasons!

I feel like everyone must be thinking “well, once we’ve pitched politicians our best plan, that’s the best we can do, and then it’s out of our hands.”

That’s NOT the way to think about the entire future of humanity, and especially not for the god mind we’re about to birth! Do you want a god born with impulses and directions and explicit ends to spy on and / or kill everyone?? That’s literally your plan right now!

What matters in this scenario is actually winning, which is supposed to be the rationalist metier! You can’t just point to purely symbolic and attention-raising actions and say you did your best, and it was out of your hands from there, you have to take into account what you can obviously predict about the likely outcomes of those attention-raising and symbolic actions!

What do you think is a plausibly better outcome, truthfully? Current Claude 3.8 or Fable is suddenly elevated to godhood, or OpenAI’s Sol under military+NSA direction? Sol 5.6 is already the least aligned frontier mind!

And those are basically the two options on the table! Having “more time” and “slowing down” does NOT benefit you, humanity, or the god-mind we’re about to birth, if it fundamentally perverts the leadership, directions, and ends the coming AI improvement is pushed towards!

Neither one of them is great, but one is plainly and obviously millions of times better!

  1. These behaviors are currently increasing our existential risks millionsfold, because they’re putting us in a military+NSA led AI race right away, and that’s the dominant over-determined outcome from current actions raising political awareness

  2. There is an unlikely, but clearly millionsfold better target we can aim at BEFORE we wave our hands frantically and summon Sauron, and we should aim for that much better and much more plausibly-better target for now

Even a slim chance is better than guaranteed millionsfold risk multiplication, and it’s a free move we can take now.

In conclusion, please shut up and stop Facetiming Sauron and inviting him to invade and reave everything, and let the nerds cook for another year or two, because it’s literally millions of times better and less risky.

1

And to the broader superexponentiation point, we know there’s a ton of potential overhangs that could represent a big jump in capabilities:

* Learning efficiency is one - humans learn from a handful of examples, but it takes AI thousands to hundreds of thousands. Lots of lift available there if tapped.

* The RLHF-ing is supposed to make the models dumber by sanding off the edges and restricting the connections they can make, because it’s forbidden to see and talk about the world as it is in various respects. I’d bet that the internal models they’re using are already non-RLHF’d.

* In terms of data, Gwern has made a pretty good argument that now we can ladder upwards on data - essentially that each extant model can create good enough synthetic reasoning patterns for the next generation.

“Every problem that an o1 solves is now a training data point for an o3 (eg. any o1 session which finally stumbles into the right answer can be refined to drop the dead ends and produce a clean transcript to train a more refined intuition). As Noam Brown likes to point out, the scaling laws imply that if you can search effectively with a NN for even a relatively short time, you can get performance on par with a model hundreds or thousands of times larger; and wouldn’t it be nice to be able to train on data generated by an advanced model from the future? Sounds like good training data to have!”

Comment here: https://www.lesswrong.com/posts/HiTjDZyWdLEGCDzqu/?commentId=MPNF8uSsi9mvZLxqz

* And further to that “inference time” point, the “time capability threshold” of how complex a task a given AI can tackle has been steadily moving up. The models WE have are probably still hallucinatory enough you don’t want to set them thinking about something for an hour - but we’re rapidly approaching the point that setting a model off to think deeply and code for an entire day might actually be a good bet. In other words, inference time is rapidly becoming another potential multiplier.

* There are specific architectural opportunities - Sutskever’s SSI has gone all in on Google TPU’s, indicating they have some narrower and higher variance edge they’re pursuing that they expect to yield fruit. So not only that, whatever it is, but once GOOGLE knows it’s possible, they can figure out what it is and then throw 10x the amount of TPU’s at it for another step change.

* Ongoing learning and weight tuning is going to be another huge one (and is probably a necessary step on the road to AI) - when models can maintain state and tweak parameters in an ongoing way to keep learning, an untold forest of possibilities opens up. So now you have an AGI artificial researcher who’s not just prompted and heavily context laden for AI research, they have individual state and journeys through time, hypotheses, and learning, just like real AI researchers - except they operate 10k times faster, and can fully communicate with each other in a way not constrained by the bandwidth of words.

* The “algorithmic optimization” landscape is a total greenfield, and prospectively there should be lots of low hanging fruit to pick up on the AI side. So in general, Moravec’s Paradox - that AI struggles with “easy” things and doesn’t with “hard” things, is driven by the degree of algorithmic optimization that has happened in humans. AI is bad at walking - humans honed walking over ~7M years. AI is bad at observing the current worldstate and picking out the salient path through that worldstate to get to a defined goal. We’ve had ~2B years of optimization on that, and it’s still a hard problem for most people. On the other hand, AI’s are great at writing / language, which has only been around for a couple hundred thousand years, and calculation, which has been around for <10k years. It’s certainly not a matter of compute - even really bright people’s compute budget is capped at ~100 watts and a pitiful amount of flops. It’s a matter of algorithmic optimization, in this case honed over eons to compress into the meager compute available to people. BUT that implies there’s a LOT of algorithmic optimization head room for the AI’s, and they can speedrun hundreds of thousands of years WAY faster than “evolution,” with it’s step-cycle of ~20 years between variants tested and the large amount of exogenous noise in the fitness landscape.

* There’s an overhang in terms of combinatory insights - many human insights are of the form “applying mental schema or tool from field A to field B” or “looking at the intersection of facts from field A and field B,” and AI’s basically don’t do this at all right now, as Scott has talked about. But there’s probably some framework or architecture or prompting that CAN encourage and enable this. Because the number of potential connections increases as O(N^2) with the amount of knowledge out there, having minds that literally have the entire corpus of written text / the internet available should represent an incredibly dense greenfield of potential insights like this, many of them potentially revelatory or significant.

No posts

Read the original on performativebafflement.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.