RSS Amplifier

Unpublishable Papers · Mar 4, 2026

Automation will set science free

0
Sign in to vote or save

Eli Stark-Elster · Unpublishable Papers

“If I had an hour to solve a problem and my life depended on the solution, I would spend the first 55 minutes determining the proper question to ask… for once I know the proper question, I could solve the problem in less than five minutes.” — Albert Einstein

Last week, Claude — the large language model developed by Anthropic — made me go slightly insane. I had seen on social media that some scientists were now deploying it to conduct literature reviews, build computational models, run data analyses, and in some cases package everything into a publishable scientific paper. They suggested that large language models were now research assistants possessed of infinite patience, speed, knowledge, and potential. Scientists could automate most of their day-to-day labor.1

Big if true, I figured. That would mean the end of science as we know it.2

So I started messing around with Claude. I told it to write a detailed literature review on shared principles in perception and memory. I told it to find out if precipitation variance affects the technological development of human societies. I told it to build a Python script that could extract acoustic features from MP3 music files. I wrote up a rough sketch for a cognitive psychology experiment, then told it to code the experiment into an HTML prototype.

And it worked. Every single request was granted. Ultimately, I fed Claude a proposal for a minor research project I had been working on and asked the machine to conduct it from start to finish: experimental design, analysis, literature review, figure generation, paper writing. After a couple of hours, I had an entire scientific paper. For obvious reasons, I don’t intend to submit it, but I suspect that it would be published.

So, science is over. Right? I should drop out of my PhD program and apply to McKinsey.

Well, maybe. Or maybe science has just begun.

For the past century or so, scientific success has been measured in papers. No matter what kind of research you’re doing, the final product is a published article that people can read, assess, and build from in the future. Prestige, jobs, and grants all largely ride on your publication record. Scientific institutions are very conservative — so this is unlikely to change, despite the grand predictions made by some commentators.

So when we think about how artificial intelligence could transform science, we should think about how artificial intelligence could transform the production of scientific papers. Which parts of the game have changed?

We can roughly identify two kinds of labor that go into a paper: empirical labor and theoretical labor. Empirical labor means creating a methodology to answer a question, following the methods, then analyzing the results. Per the above epigraph, it is the work that Einstein wanted to spend five minutes doing. Theoretical labor means contextualizing and identifying a question, exploring the import of the answer, and synthesizing everything into the broader literature. It is the work that Einstein wanted to spend 55 minutes doing.

To some extent, artificial intelligence eases the burden of theoretical labor. For instance, it is very good at reviewing prior work on some subject of interest. Rather than spending weeks clicking through Google Scholar, you can now ask Claude to give you a report on everything pertinent to your research question, and it will do a decent job.

But for now, at least, artificial intelligence is still pretty mediocre at two key theoretical tasks. First, it is bad at picking important research questions. By “important,” I mean a question for which the answer fundamentally contributes to the expansion of human knowledge. Second, it is bad at using existing research to build or contribute to scientific theories — that is, the frameworks we derive from our scientific labors to help us understand the world.

Arguably, these are the two most important tasks in all of scientific practice. We’re trying to build pyramids of knowledge, not scattered piles of rocks. If we’re not spending our scientific time trying to advance theories, then we should throw away our lab coats (or notebooks, in my case) and find something less trivial to do.

For reasons that I don’t understand, though, artificial intelligence is not very good at laying out the blueprint for a pyramid, or even identifying where the next brick should be placed. Maybe it’s hard to figure out what kind of knowledge is meaningful to humans if you don’t have a human body. Regardless, theoretical labor remains our province alone.

Empirical labor, on the other hand, is kaput. Artificial intelligence can do almost all of it instantly, and it will only improve with more computational power.

In a paper, the outcomes of empirical labor fall into two sections: Methods and Results. A standard Results section takes quantitative data and reports statistical analyses from it. For example, an author might calculate differences between the performance of two groups in an experiment and declare that the gap is statistically significant.

A large language model can handle that easily. Anthropic co-founder Jack Clark says that the company will likely rely on Claude to produce 99% of its code by the end of this year. If artificial intelligence is handy enough to write the actual code used by artificial intelligence companies to improve their multi-billion dollar product, then presumably, it’s good enough to write a data analysis pipeline in Python.

The Methods are more varied and complicated. Obviously, Claude can summarizes the techniques that were applied to answer the inquiry driving the paper. But the interesting question is not whether a large language model can, say, summarize your methods from an outline — it can — but whether it can actually carry out those methods on its own or even independently come up with the protocol.

Increasingly, though, this also looks like a possibility. Alberto Acerbi recently published a preprint of an AI-written paper that generated three computational models to formalize one of his theories. Paul Novosad has similarly noted that artificial intelligence can easily replicate a vast swath of the research conducted in economics. Literature on so-called “self-driving labs” also suggests that even basic physics and chemistry experiments can be identified and carried out by artificial intelligence.

As for the various exceptions one might want to raise — you still need research assistants to dissect the mice! — always keep in mind that we’re at the very beginning of the agentic growth curve. What we’re seeing right now is the worst that these agents will ever be at science. Presumably, if you told certain ENIAC3 technicians that high-powered computers would someday fit in your pocket, they would have laughed in your face.

In effect, artificial intelligence has eradicated (or will soon eradicate) virtually all the time and effort that goes into empirical labor. Data collection for some fields, of course, remains firewalled from this turbo booster. Anthropologists (lucky me!) will not see Claude traveling to the Kalahari and shacking up with the Ju|‘hoansi. But for many disciplines, methods are now cheap or, in some cases, free. Once you have your data — which again, large language models can increasingly fetch for you — you can now find out the answer to your scientific question within days or hours, when once the pursuit might have taken years.

Robot dog trains on White Sands dunes for future Mars exploration
Hello, fellow humans! Please tell me about your foraging habits.

But artificial intelligence has still not changed much about the time and effort that goes into theoretical labor. Theory is still expensive. No matter how you prompt it, the current iteration of Claude is not going to discover a new paradigm for psychology or suggest the next great experiment.

But theoretical labor was always the important part. Artificial intelligence doesn’t cut down from Einstein’s 55 minutes of theory; instead, it drops his five minutes of methods down to almost zero.

Okay, cool. 55 minutes is still close to an hour. Are we really going to get this worked up over an ~8% boost in efficiency?

My PhD advisor likes to ask a scary question. If you pitch him a research idea, or talk about an ongoing project, there is a good chance he’ll nod slowly, lean back dangerously far in his chair, and say: “Right, word, okay, for sure. What are the stakes for that?”

The science behind why you fall back in your chair | Daily Mail Online
Near-optimal position for asking someone to justify their research agenda

By “stakes,” he means the proof that there is real theoretical labor behind your research — that answering the research question would reshape our scientific theories, or expand them, or generate a new one entirely, or at least resolve some obviously interesting unknown in the world. Insofar as scientists should spend 55 minutes out of their gun-to-head hour identifying a question that has real stakes, they should have no trouble replying.

Yet many capable, hard-working, and very intelligent scientists regularly seem stumped by his question. And I think that if he went around and asked every working scientist in the world to explain the stakes of their research, the majority would not know how to reply. That’s because the scientific world is dominated by one kind of researcher: the artisan.

An artisan is a master of their craft. They know all the best techniques, and they love to use them. Sometimes, they even invent new ones. They possess the soul of a blacksmith, or carpenter, or cobbler. They relish the act of production for its own sake. Like all great artisans, the actual outcome is rarely the point. Cobblers don’t think about all the shoes they make, or how they’ve advanced the theory of shoe-making. They just make the goddamn shoes.

Science is full of artisans. I can’t imagine a statistic that would verify my claim — perhaps the boutique literature arguing that most research is totally useless — so I’ll ask you to take my word for it.

At present, most working researchers possess the knack for methodology and analysis; they are great at empirical labor. But their theoretical labor tends to fall by the wayside, because it’s hard to be great at both. Hence why no one can tell my advisor why their research matters.

But if theoretical labor is the important part, why don’t we instead have a scientific world maniacally focused on the stakes above all else?

As always, the answer is the incentive: publish or perish. If you want to succeed in science, you must publish papers. It’s better if they’re good or even groundbreaking. But something is better than nothing. And right now, empirical labor — not theoretical labor — is the main barrier to publication. Data analysis is slow. Iterating experiments is hard. If you can’t do those things right, you can’t do them at all.

You can, however, pick an unimportant research question for which to carry out those tasks. That bit is easy, albeit unadvisable. So those skilled in empirical labor, despite their lack of theoretical skill, can still regularly publish and thus receive jobs, grants, and prestige. And because artisans are usually the ones judging other artisans, poor theoretical labor can easily go unnoticed.

Their counterparts, however — those skilled in theoretical labor, but not empirical labor — just won’t cut it. It isn’t enough to pick a good question. If you want a paper, you need to pour the tears and screams into working out the solution. Those unwilling or unable to do so will not usually find their way to the faculty lounge.

The incentive to publish widely, paired with the difficulties of empirical labor, makes artisanhood the most viable strategy for success. The success of artisans then elevates more artisans into prestigious positions, because artisans (like anyone else) are unlikely to disdain their own path to greatness. Hence we find ourselves in a scientific world full of people who opt to spend 5 minutes on the question and 55 minutes on the answer.

With all that preface, I can finally explain why artificial intelligence spells the end of science as we know it: it lets anyone work like an artisan. With Opus or Codex on your side, it isn’t hard to publish decent papers. I could easily pump out sixty papers in a year if my theoretical standards were low. That isn’t a flex. You could do it too — in fact, a large language model isn’t even half-bad at identifying a workable question.

I imagine some researchers will get in on that action and publish at a ridiculous rate. But that won’t last long, because the benefits of artisanmaxxing are about to vanish. If everyone’s an artisan, no one is. Empirical labor no longer distinguishes you from anyone else. Large language models did not just shave 5 minutes from the hour of work. They dropped almost the entire shebang.

The artisan is dead; long live the artisan.

So what comes next?

Empirical labor is now easy. But theoretical labor is still hard. And despite our recent history of artisanal dominance, grand visions of science have always centered on the theory. We worship the people who dared to see further; but in practice, it is hard to make your way onto the shoulders of the giants if your back hurts from hunching over R Studio.

What ate your thoughts on Ferra & Torr? : r/MortalKombat
Isaac Newton standing on Plato’s shoulders

No longer. Theoretical labor is the only game in town. That means scientific success will accrue to a new group of researchers: the seers.

A seer understands that the purpose of science is the pursuit of knowledge, not self-actualization4. So they pour their labors into picking out questions that will produce useful knowledge and building theoretical frameworks that refine and guide their efforts. They always know the stakes. Seers refuse to work on anything that lacks a clear theoretical purpose.

To be clear, the most influential scientists have always been seers, at least partially. The problem is that any successful seer also needed to pair their theoretical labor with intense empirical efforts; they needed the soul of a seer in the body of an artisan.

This brings to mind a famous quote from Stephen Jay Gould: “I am, somehow, less interested in the weight and convolutions of Einstein’s brain than in the near certainty that people of equal talent have lived and died in cotton fields and sweatshops.” I am similarly interested in the near certainty that people with lots of great research questions have lived and died as consultants, because they just couldn’t bear to spend two years writing up an experiment in Javascript.

With the advent of artificial intelligence, though, the only distinction that remains between scientists is that of seerdom: who can ask the right questions, and who can figure out how to put everything together? Finally, the seer soul is free to roam.

By moving into an era in which empirical labor is near-instantaneous, we also move to an era in which theoretical labor is the main difference between scientists in the race for prestige. Suddenly, the stakes of scientific research cannot be ignored or brushed aside. Stakes are all that remain of the scientific enterprise; though of course, they were all that the enterprise ever should have been.

But institutions do not necessarily evolve at the same pace as technologies. As better writers have pointed out — special shout-out to Adam Mastroianni, whose work made me want to start a Substack in the first place — the current scientific publishing system is stupid. Peer review massively slows down the publication process, while simultaneously failing to make papers better and adding further unpaid labor to the lives of scientists. To compensate for the months or years it takes to see papers in print, scientists get to pay exorbitant publishing fees to make sure people can actually, you know, read their work. Just as musicians pay their label to let them release an album. Right? Is that not how it works?

Most importantly, the existence of the Internet (check it out, super useful) raises the question of why we need scientific journals at all. It seems that we could speed up the scientific process, avoid publishing fees, abolish the burden of peer review, and improve science in a million other ways by simply publishing papers as blog posts or uploading them onto free host sites.

Oh, right. We tried that. Most researchers don’t even get their papers from journals anymore; they download preprints uploaded by the authors or pirate them from SciHub. And yet, here I am, still trying to optimize my paper for Nature.

Our ongoing failure to fix scientific publishing suggests that the mere emergence of useful technologies does not mean that institutions will adjust to support them. The scientific ecosystem is built for the 20th century. Even the Internet couldn’t make it budge.

What then of artificial intelligence? The number of scientific papers seeking publication is about to skyrocket. Peer reviewers are delegating their work to large language models. Researchers may soon automate most or all of their labor, which raises the question of what ‘authorship’ should even mean in the age of Claude. These changes are seismic — I can’t imagine reforms to the current system that would allow it to withstand the quake.

So I won’t propose any. Instead, I’ll make a bold prediction. The automation of science will eventually prove too much for scientific journals to handle; they cannot, nor should they, arbitrate between ten million papers each year. The system will collapse. We should start thinking about what to build from the rubble.

In sum, then: the automation of science, and specifically the automation of empirical labor, will set science free, both from the demands of artisanhood and the absurdities of scientific publishing.

But serious questions remain. What happens if artificial intelligence becomes great at theory, too? Really, what I’ve described here is not true automation, but instead the distribution of empirical resources to a much broader subset of scientific thinkers. For the near future, this seems like the most plausible outcome. But it seems possible that we might someday get true automation — that artificial intelligence may take over the entirety of Einstein’s hour.

Whether that possibility should scare us or not depends on what we think science is for. The goal of literature, for example, is arguably to facilitate the experience of another human mind. Automating literature with large language models would therefore defeat the point of the entire enterprise.

By comparison, the goal of firefighting is to put out fires. If large language models were better at putting out fires than people, we would hopefully endorse the automation of firefighting. It doesn’t matter who puts out the fires; what matters is that the fires get put out as quickly, efficiently, and safely as possible.

Some people probably think science is more like literature than firefighting. If that’s your perspective — and I don’t disdain you for it — then the prospect of automated science is horrifying.

But much as I empathize with this romantic view, I disagree. I think the point of science is to expand our understanding of the world. If God descended from heaven and offered us answers to every burning scientific question, no labor required, I would hope that we’d take him up on it.

It doesn’t matter who does the expanding. Science was never about the scientists.5

1

Thanks to Lindsey Cannon, Patricio Cruz y Celis Peniche, Madison McCartin, Benji Elster-Stark, Tracy Pan, Manvir Singh, Louise Toutee, and Ethan Wellerstein for helpful conversations and feedback on some of these ideas.

2

The revolution I’m describing here is maybe marginally more relevant for people in the social sciences, at least right now, mainly because we work so much more often with big datasets and the like.

3

For the uninitiated, ENIAC was the world’s first programmable computer. It weighted 30 tons and had one byte of RAM.

5

Regarding this (possibly inevitable) shift in scientific practice, it’s worth reading this recent piece from Asimov Press on the ‘legibility problem.”

No posts

Read the original on unpublishablepapers.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.