RSS Amplifier

A Causal Affair · May 20, 2026

AI and the Research O-Ring

0
Sign in to vote or save

Paul Goldsmith-Pinkham · A Causal Affair

There is a Chinese curse which says “May he live in interesting times.” Like it or not, we live in interesting times. They are times of danger and uncertainty; but they are also the most creative of any time in the history of mankind. And everyone here will ultimately be judged - will ultimately judge himself – on the effort he has contributed to building a new world society and the extent to which his ideals and goals have shaped that effort.

-Robert F. Kennedy (1966, University of Capetown)

We live in interesting times. Economists’ research toolkit is changing rapidly, and we are part of a large group of skilled laborers who are continually told that their jobs will soon not exist.

This post is a longer version of a response I gave to a moderated discussion at the NBER Applications of AI in Healthcare meeting. My comments reflect two realities: first, the impact of agentic AI is hard to predict and fraught with uncertainty. Second, while AI has sped up parts of the research pipeline, this acceleration has amplified the human frictions in the rest.

The evidence for continued and dramatic AI improvement is convincing. The figure above replicates the graph made by the non-profit METR (Model Evaluation & Threat Research), which benchmarks the time that it takes different AI models to do tasks, and the current iteration documents exponential growth in ability to do more complex tasks (“the 50%-time horizon is the duration at which an agent is predicted to succeed half the time.”) Frontier models continue to break through barriers rapidly, especially in code and math.

There are two serious sources of uncertainty. The first is the direct level of uncertainty about these measures, and how to extrapolate further. Consider the error bounds on these models have as we expand the horizon. There is tremendous growth, and equivalent uncertainty (any economist tasked with forecasting exponential growth will sympathize!). Small variability can imply very different long-run paths.

The second source of uncertainty is from how to map these concepts into real-life outcomes. We’ve seen limited evidence of significant labor market disruptions — although adjustments to AI-threatened sectors could quickly once the broader economy experiences a serious disruption. It is just hard to know how the analog world will react to this very digital product.

My second comment is on the research pipeline. I think that research follows an O-ring production process in the style of Kremer (1993). This model assumes that production comes from multiple inputs of quality q that work together to ensure the production of output. Hence, if two high quality inputs with q = 0.99 work together, they are expected to succeed with probability 0.98. If instead there is one high quality input with q1 = 0.99 and q2 = 0.5, then the success rate falls to 0.495. As the chain grows, one single weak input can ruin the pipeline. To quote Kremer:

[The] O-ring production function differs from the standard…formulation of labor skill, in that it does not allow quantity to be substituted for quality within a single production chain.

You can see the issue. If AI is extraordinary at some things but not others, there will still be certain tasks where humans’ inputs remain important. Moreover, beyond Kremer’s model, humans face real frictions doing those inputs well if they played no role in the rest. Imagine that writing, presentation, and idea generation remain human work. Those jobs get harder if you’ve spent no time on any of the other parts of the paper — understanding what your agent did is itself costly time.

Using the O-ring model to predict AI’s influence on research production requires assumptions about how many inputs there are to research, and how each input is supplemented, influenced or replaced by AI. One insight we can learn here comes from returning to the METR benchmarks above. The traditional benchmark focuses 50% reliability – likely a relatively poor input to the O-ring production mode. The figure below plots the 80% reliability measures as well. Here, we can see that even 80% reliability is growing quickly, but with levels significantly below 50% reliability. The question is whether reliable tasks will grow at the same rate.

Part of the adjustment is learning to use the tools. But my guess is that some things, like coding, should improve across the board for everyone eventually. That suggests that every paper’s quality will improve quite a bit. Does that inherently mean that people will write more papers?

This question of paper farms caused by AI (provocatively laid out by Scott C. and others) is hard for me to answer credulously. I think reputational norms prevent any research from credibly posting hundreds of papers to their website and submitting them to journals. Even in other fields, where researchers have far more papers than us, there is still an expectation that researchers have a coherent understanding of their work (even if this norm is sometimes violated).

The broader discussion of what counts as research predates LLMs. Agentic AI has lowered the cost of producing the surface artifact so fast that the conversation is now unavoidable, and working through it will be healthy.

Before closing, let me quickly give an example in health economics of how these tools can rapidly improve our ability to do research. In our paper studying radiologists, Alex Zentefis and I needed to read signing radiologists’ scan reports and identify their diagnoses (and differentiate chronic vs. acute pulmonary embolisms). Five years ago, this would have been extraordinarily difficult. Today, it’s trivial, using the local LLMs that the hospital has installed on its research servers.

So: beware nihilism. Instead, consider RFK’s advice above: “[these] are times of danger and uncertainty; but they are also the most creative of any time in the history of mankind.” The interesting question is which problems we choose to deploy these tools against.

No posts

Read the original on paulgp.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.