RSS Amplifier

The Curious Wavefunction · May 6, 2026

The next frontier for agentic AI - actual science

0
Sign in to vote or save

Ash Jogalekar · The Curious Wavefunction

A year ago, agentic AI in science felt like magic. You could describe a problem in plain English and watch a system spin up dozens of tools, query databases, run models, and return something that looked like the output of a small research team. For anyone who has spent years stitching together brittle pipelines in drug discovery, biology or chemistry, the experience was almost surreal. What once took weeks could now be compressed into hours. It still feels like magic to me.

But that phase doesn’t last long. With any powerful technology, there’s a predictable shift: what begins as amazement quickly becomes a baseline. And once it becomes a baseline, the question changes. It is no longer “What can it do?” but “Can we trust it?” and, even more importantly, “Can it actually do science?” As agentic AI is increasingly applied to scientific problems, what becomes clear is that as it stands, it’s a testament to clever engineering. Whether it can produce clever science remains to be seen.

That shift between process and product reveals two problems that sit at the heart of agentic AI, problems that are easy to miss precisely because everything else is working so well. The first is that science is fundamentally non-linear and long-horizon, while agentic workflows are still largely linear and short-lived. The second is that science depends on a kind of epistemic discipline, an integrity of knowledge over time, that these systems do not yet possess.

Start with the first. Most agentic workflows today are, at their core, elegant pipelines. They take a task, decompose it into steps, execute those steps in sequence, and produce an answer. The sophistication lies in how many tools they can call and how seamlessly and blindingly fast they can move between them. But structurally, they are still doing step one, then step two, then step three.

Plainly speaking, that is not how scientists work.

A real scientific process, especially in a complex discipline like drug discovery, looks much messier. You begin with a hypothesis, but you don’t march forward in a straight line. You test part of it, skip ahead, circle back, revise assumptions, combine ideas that weren’t initially connected, and often abandon entire lines of thinking when the data refuses to cooperate. The process is recursive and adaptive. It is less like executing a plan and more like feeling your way through an uncertain landscape. It’s the opposite of deterministic, with only a loosely defined high-level structure imposed on it.

Crucially, it unfolds over time. A medicinal chemist might design a handful of molecules, wait weeks for them to be synthesized and tested, receive partial and noisy data, and then update their mental model and priors before starting again. Each cycle changes not just the next step, but the interpretation of everything that came before. First fix activity, then fix solubility, then fix permeability and, oops, need to fix activity and solubility again. It’s the classic multi-parameter whack-a-mole, and the “state” of the problem is constantly evolving.

This is where current agentic systems hit a wall. They are extraordinarily good at compressing time, but real science is not just about speed. It is about maintaining coherence across long stretches of time. To participate meaningfully in a scientific process, an agent would need to remember not just what it did, but why it did it, what assumptions were embedded in those decisions, and how new data should reshape those assumptions. It would need to carry a living, evolving model of the problem forward over weeks or months without drifting into inconsistency. It should be able to reconcile contradictory pieces of data from January 8, 2026 and September 20, 2026, and perhaps reconcile those results with data from the literature and from internal databases from ten years ago. It’s an extraordinarily difficult task for a human, which is why we need agentic systems to address it urgently.

Operationally this is a very different challenge from executing a complex workflow. It is closer to sustaining a continuous line of reasoning in the face of uncertainty and delay. Rather than trying to compress six months into days, it necessarily involves working like a sentinel for six months; ingesting, monitoring, analyzing, correcting. And we do not yet know how to do that well.

Even if we solved that problem, however, a second and arguably more dangerous problem emerges. As agentic systems scale, they generate enormous volumes of outputs - molecules, hypotheses, analyses, predictions. The bottleneck quickly shifts from generation to verification. It becomes harder and harder to answer a simple question: is any of this actually correct?

This is where epistemic integrity comes in. Science is not just about producing answers; it is about producing answers that can be trusted, traced, and, when necessary, falsified. A human scientist may be slow and biased, but the process is anchored in certain disciplines. Data has provenance. Experiments can be revisited. The right statistical tests are run. Contradictions are eventually surfaced, if not always immediately resolved.

An agentic system operating at scale can easily lose that discipline. It may pull data from multiple sources without preserving where each piece came from. It may mix simulated results with experimental observations in ways that are not clearly labeled. It may carry forward early-stage assumptions that were statistically weak, allowing them to shape downstream conclusions. And because the outputs often look polished and coherent, these failures can be hard to detect.

Over short workflows, this is manageable. A human can inspect the results, spot inconsistencies, and course-correct. But over long horizons, when a system is effectively running an extended research program, the risk compounds. Small epistemic errors accumulate and blow up over time. Weak signals get reinforced. Entire lines of inquiry can drift away from reality while still appearing internally consistent.

The core issue is that the system lacks a durable sense of a consistent world model: what it knows, how it knows it, and how confident it should be. It does not reliably distinguish between what was measured and what was inferred, between what is well-supported and what is speculative. It does not always maintain a clean audit trail that would allow a human to reconstruct its reasoning months later. And it is not naturally inclined to revisit and revise its own priors when confronted with new, contradictory data.

In other words, it lacks the very qualities that make science self-correcting.

These two problems - long-term, non-linear experimentation and epistemic integrity - are deeply intertwined. A system that can iterate over months but does not preserve the integrity of its knowledge will simply drift, becoming more elaborate and more wrong over time. A system that maintains strict epistemic discipline but cannot engage in extended, adaptive experimentation will remain brittle and shallow, unable to explore complex problems in the first place.

What we ultimately need is something much harder: an agent that can evolve its understanding over time while remaining anchored to reality, an agent that can explore without losing track of what it has learned, and revise without erasing the past. That is not just an engineering challenge; it is a conceptual one. It requires rethinking what it means for a machine to “do science.” Part of it also involves changing our own mindset, prizing accuracy rather than speed in an age where time compression and workflow acceleration seem to mean everything. In other words, we should be ok waiting for an answer because that answer promises to be so good.

If we fail to solve this problem, the likely outcome is not that agentic AI disappears. It will spread, because the productivity gains are too large to ignore. But we will find ourselves in a world where the volume of scientific output has exploded while our ability to trust it has not kept pace. The result is a kind of high-throughput ambiguity of more answers and less certainty that would be worse than the certainty of wrong answers.

That would be a strange inversion of the scientific enterprise, which has always traded speed for reliability. Agentic AI gives us the speed back. The question now is whether we can recover the reliability.

Note: Thanks to Mark Murcko for helpful discussions.

No posts

Read the original on medchemash.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.