RSS Amplifier

Arachne · Aug 2, 2026

Scattered Theses on RSI, AI Math, and Doomerism

0
Sign in to vote or save

Nathan Witkin · Arachne

Prometheus, 1930, José Clemente Orozco

While I tend to reserve my Substack for more polished, long-form articles, I’ve written a pair of lengthy X posts in the past couple of days that I think it’s worth publicizing more widely in the hopes of spurring discussion (I also previously posted the first as a Substack note). I’ve pasted lightly edited versions of both below (you can read the originals here and here); I will warn that they assume a bit more background knowledge about recent debates in AI than is typical for my writing on the topic.

I don’t plan to make a habit of reposting other writing on Substack, but thought these two posts covered interesting-enough ground to be worth it. You should continue to expect long-form, Substack-exclusive articles from Arachne.

This is a fantastic paper that I think all who are interested in the concept of recursive self-improvement, or RSI should read.

Some reactions I had, which I should emphasize are not inferences narrowly licensed by the paper itself, but mix in my own interpolation.

1. Cutting edge research is insanely hard, and a lot of folks conjecturing willy-nilly about its “automation” are underestimating just how hard it is. To produce a paper ready for NeurIPS or a top journal requires numerous iteration cycles over the ideation, design, execution, writing, and review process. It is not analogous to most other forms of white-collar labor.

2. Handing these processes over to an AI model means reckoning with dozens of emergent failure modes it is impossible to predict in advance. Human experts have an impressive knack for using high-level judgment and tacit knowledge to refine their work over time, whereas AI models are often just as likely to introduce new errors via iteration as to solve them.

3. It is difficult to see how these limitations could be overcome in the near or even medium-term. There are no verifiable reward signals for “good” research, and it’s hard to see how we could create them. The issue is that “good” is not just field and sub-field specific; it’s also constantly changing as new research comes out!

What about RLHF? Can we get top human researchers to provide reward signals? There are at least two major reasons to think the answer is no.

3.1. Researchers disagree on what counts as good, and given the rarity of significant breakthroughs, most of them are wrong! Averaging together the perspective of a bunch of top researchers does not get you the proper reward signal.

3.2. If researchers are spending time judging and rewarding model outputs, itself a very time-consuming, difficult task, then they’re not keeping up with the field, meaning their judgment will get worse. There is just not enough time in a day for top researches to do RLHF on the side while remaining top researchers.

What about pre-training? Why not think think models will become so smart that they develop their own superhuman research taste? At least two reasons.

3.3. Mode collapse. Research taste is dependent on originality. So long as model outputs tend toward the middle of their training distributions, AI-driven iteration cycles will push output away from the tails where juicy new directions are.

Notice that I’m not saying that models are incapable of creativity. When it comes to one-shot outputs and high-level ideation, models do show some level of creativity. But that is very different than continually bringing a novel perspective to bear on the entire research workflow, such that over many iterations, that novelty is maintained under rigorous constraints.

3.4. Tacit knowledge. There is just no data out there making explicit what the inputs are to good research taste. We know almost nothing about what goes on in researchers’ heads when they make fine-grained calls on e.g. how to phrase a specific concept, or how to tune a statistical model, or how to clean up a data set.

Some of the reasons for that are discussed above: the inputs to these decisions are field-specific, context-specific, and maybe most importantly in a constant state of change. Pre-training cycles are not even close to fast enough, and arguably never will be, to hard-load the relevant inputs to good research taste into model weights.

This is why I’ve been so harsh on the concept of RSI in the past: almost every appeal to RSI I encounter betrays a lack of foresight about the challenges with getting models to do “good” research, not to mention basic domain knowledge about what “good” research means.

One thing I’ll add, which was absent in the original post, is that a lot of these questions come down to a broader one that is helpfully flagged in the paper, and that those interested in RSI would do well to think about: “how much of frontier AI development depends on open-ended research rather than hill-climbing on well-specified objectives” (p. 8, Table 3). In other words, if we can collect a sufficient number of quantifiable metrics correspondent to progress in machine learning, then it becomes much more straightforward to post-train models to perform machine learning research. Many of the considerations I flag above, such as those to do with tacit knowledge, become less daunting in that case, since models may be able to develop the relevant tacit knowledge as an emergent consequence of post-training.

I hasten to caveat, though, that an extraordinary amount of hill-climbing on ML-related metrics has already been done, as the table below (also from the paper) makes clear, and this has not translated into a human-like capacity for open-ended research.

The release of this paper by OpenAI, sharing ten new results achieved by an internal model on open problems in math and theoretical computer science, has predictably led to a cancerous profusion of ice-cold takes about our impending doom and/or intellectual obsolescence on X. These are some of my reactions:

1) AI capabilities continue to be jagged even within pure math, and it naively underestimates the scope and internal variegation of the field to infer from the recent round of Astra results that “math is solved,” or will be in the near future. There will be plenty of work for human mathematicians well into the foreseeable future.

2) As AI continues to make blistering progress in math and beyond, human verification will become a huge bottleneck. Already, there are very few mathematicians qualified to verify OpenAI’s new results. As progress continues, that number will approach zero. The implications of this dynamic are underappreciated: on the current trajectory, we will be increasingly ill-positioned to distinguish genuine results from highly convincing slop, and as a result will have a hard time leveraging AI’s research output to important practical ends. At a minimum, we will often need to wait for human understanding to catch up. I see this as a very strong reason to doubt we’re heading for any kind of AI-powered scientific take-off. In many if not most domains, we will only be able to proceed at the pace of human understanding.

I’ll add that, as of now, the idea of autonomous AI agents verifying and executing on each other’s work without a human-in-the-loop is pure fantasy, and anyone who disagrees needs to give the recent literature on this issue a much closer read (at minimum, ask your favorite LLM for a summary).

3) Even coming from doomers self-defined, I see quite a lot of revelry thinly disguised as doom-saying. Many of you can barely contain your excitement at the prospect that AI might soon outperform humans at everything.

Besides the fact that there’s little evidence this eventuality is worth taking seriously given extreme jaggedness, the role of human-to-human contact in much of public life, the continual need for human verification, underwhelming performance on strong benchmarks by frontier models, widespread and growing political disquiet surrounding AI, and so on, it’s frustrating how little agency this perspective presumes on the part of our species.

It’s hard not to see this attitude as stemming from blinders the doomer set has about the wider social world outside of Silicon Valley and its satellites. When you’ve spent most of your life working with computers, entities that until very recently couldn’t talk back, you’re going to have a hard time understanding the difference between partial and general equilibria. In fact, for the software engineer, it’s plausible that these often collapse into one another: you have extraordinary autonomy to intervene in a system that is well-defined and relatively transparent, and there is very little risk of that system turning around and intervening in you. The implicit assumption, then, is that AI will relate to us as we related to its more acquiescent ancestors.

At the risk of stating the obvious, this greatly misunderstands human social systems, which do not take large parametric changes lying down. Go read about the litany of world-shaking crises we’ve passed through in just the pass century, in periods where we were far, far less capable and organized than we are now, and tell me with a straight face that AI is going to destroy the world.

Frankly, it’s a wonder we still reserve as much social grace as we evidently do for misanthropes making flailing evidence-free pronouncements about how screwed we all are. While I don’t agree with the common take that frontier labs and others in the industry play up AI risk for the purpose of driving investment, I think there’s a more plausible version of this thesis driven by implicit rather than explicit motivations.

Clearly, a lot of folks now have a mix of financial, personal, and career incentives to construe AI as something more explosive than it is. In another reality, the truth would be explosive enough: AI is transformative general purpose technology rivaling the internet or the steam engine, one set to reshape numerous facets of social and economic life.

But the nature of public discourse now is that absent the imputation of apocalyptic potential, one’s pet issue risks losing mindshare (think climate change). The result is self-styled rationalists adopting the discursive stylings of every other single-issue slop merchant, but with a maddening veneer of false rigor (see AI 2040, one of the worst recent offenders). The reason I tend to be more outspoken about doomerism and related ideas like RSI is because these ideas do real damage to public understanding.

I fully agree we need to do some serious thinking in advance about the prospective impacts of AI, but poorly evidenced sensationalism makes for bad policymaking, politics, philanthropy and much else. The better view is to say, yes, AI is a very big deal, but we have the capacity to meet whatever challenges it poses and then some. The message that we're probably screwed is not only wrong, but makes its own thesis more likely to be true.

No posts

Read the original on arachnemag.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.