RSS Amplifier

Fetch Decode Execute · Aug 7, 2026

Can agentic AI really do research?

0
Sign in to vote or save

Michael Lones · Fetch Decode Execute

I’ve been finding it hard to get my head around the capabilities of agentic AI systems in scientific research. The low signal-to-noise ratio caused by endless hype has made it challenging to work out whether they’re about to replace human researchers or are just bleh.

Recent findings by the group1 of my REFORMS-collaborators at Princeton suggest it may be more the latter than the former. They asked a bunch of models hosted on various agentic scaffolds to shadow two research groups who had submitted papers to NeurIPS by telling the system the original research questions and asking it to solve them independently. The results were then reviewed by the authors of the original papers, and they weren’t exactly glowing.

Basically the agents did a lot of the things you’d expect from an inexperienced human researcher: they failed to adequately check existing literature, they didn’t properly assess approaches before moving on, they wrote impenetrable papers, they drew strong conclusions from weak results, and they ignored feedback. All things which are achingly familiar from years of trying to prod students in the right direction.

Something that agentic systems do have on their side is speed and availability, and the paper also hints at this as an explanation for why AI-generated work has been seeping into the scientific literature. If an agentic model can write hundreds of papers, and “authors” are allowed to submit as many as they like to leading conferences, then some are bound to get through the noisy process of conference review.

This also explains why these conferences are starting to place limits on the number of papers an individual author can submit. The volume of papers submitted to the top AI conferences has grown enormously in recent years, and this is placing an unmanageable burden on conference organisers and reviewers. Agents might be able to write a paper in no time, but it still takes hours for a human to read and review it.

Assuming, of course, that reviewing is being done by humans. Reviews I’ve been receiving back from my own papers have many of the hallmarks of AI-generated text, and there’s a feeling amongst the community that the use of AI in writing reviews has gotten out of hand. This is partly a result of conference rules: people who submit are expected to review, and they don’t necessarily want to2, so they take shortcuts. And, more generally, young researchers are keen to get reviewing on their CV.

But things aren’t quite that simple. Last year, another of the big AI conferences, AAAI provided AI-generated reviews to authors and asked for feedback. According to the resulting write-up, authors actually preferred these AI-generated reviews to human reviews. Though perhaps this says something about the standards of human review, which has long been a subject of criticism, rather than the capabilities of agentic AI.

Similar findings came out of a more recent analysis of reviews for Nature journals, where the AI-generated reviews were also deemed superior, on average, to human reviews. However, the authors of this study also found that AI reviewers are more homogenous than humans and perform poorly on niche topics. This suggests that whilst they might perform well on mainstream topics, they might fail miserably on the kind of innovative niche research that actually drives science forwards.

And this is where a lot of my worries about the use of agentic AI in research lie. Much of research is quite iterative, and it’s no great surprise that AI is helpful with this when all human endeavours in a particular domain of research have been trained into its parameters. But currently there’s not much evidence to suggest that AI can help with truly original research. By replacing human researchers with AI, we risk losing this, especially given that the reward structures and finances of many research centres promote short-term productivity over long-term scientific endeavour.

I find it far too easy to imagine a future where research is being “done” by academics whose main skills lie in using AI systems, with those who have genuine research skills finding themselves on the scrapheap due to poor productivity. No original research will get done, and nobody will be left who can tell whether the products of AI-generated research are meaningful, with those remaining having cognitively offloaded whatever research skills they had to their AI some time ago.

Which is not exactly an upbeat conclusion. So let’s imagine that future generations of AI will overcome the deficits of the current generation, and will do groundbreaking research whose proceeds will enable us meagre humans to live a life of careless luxury. Hurrah to AI!

1

Led by Peter Kirgis and Sayash Kapoor, but with a host of other contributors.

2

Which is a bugbear of mine. A lot of academics say they don’t have time to review papers, but they seem to have time to submit plenty of them. Who do they expect to review them?

No posts

Read the original on fetchdecodeexecute.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.