Science is full of forecasts disguised as ordinary research language. A grant proposal predicts that a mechanism will work. A laboratory chooses one experiment because its members expect a useful result. A paper identifies a promising direction and quietly assigns some probability to its success. Research institutions distribute money and attention according to judgments about what will become possible, important, or true. These expectations usually remain embedded in prose. Their probabilities are rarely stated. Their deadlines remain flexible. Once the result arrives, memory reorganizes the original prediction around what eventually happened.
Artificial intelligence has made this old feature of science easier to examine. Models can search literature, propose hypotheses, design experiments, and generate plausible research programs. Their fluency raises a sharper question: given only the evidence available at a particular moment, can they foresee which scientific developments will actually occur?
Four benchmarks released in 2026 have begun to test that question directly. They place models behind historical knowledge cutoffs and ask them to forecast experiments, future papers, technical directions, and the timing of scientific advances.
Current systems can produce convincing accounts of possible scientific futures, with little indication of which trajectory will eventually materialize.
Most scientific evaluations give a model a problem whose answer already exists somewhere in the research record. The task may be difficult, yet the surrounding informational environment has absorbed the result through papers, citations, terminology, reviews, and subsequent discoveries.
A forecasting benchmark reconstructs an earlier state of uncertainty.
The model receives information published before a fixed date. Later evidence stays outside the task. Its prediction is compared with events that occurred after the cutoff. The system has to commit while the answer remains unavailable.
CUSP, or Cutoff-conditioned Unseen Scientific Progress, applies this method across 4,760 events in AI, biology, chemistry, physics, and medicine. Its tasks cover feasibility assessment, mechanistic reasoning, solution design, and the timing of future progress.
Plausible direction selection emerged as one of the models’ stronger capabilities. Their estimates of eventual realization and timing remained unreliable. Performance also varied substantially across disciplines, with developments in AI proving easier to time than advances in biology, chemistry, and physics. The models displayed recurring overconfidence and strong response biases.
CUSP exposes a basic difficulty in scientific prediction. A field can support many coherent futures at once. Each may follow from existing literature, fit current theories, and resemble the trajectory of previous work. Reality eventually selects a much narrower path.
PreScience studies scientific forecasting through the production of research itself. Its dataset contains 98,000 recent AI papers, connected to a broader graph of 502,000 publications with author histories and citation relationships.
The benchmark divides scientific forecasting into four interdependent tasks: predicting future collaborators, selecting relevant prior work, generating future contributions, and estimating impact. It asks whether a system trained on the scientific record up to a fixed date can produce a credible account of what comes next.
In contribution generation, GPT-5 received an average similarity score of 5.6 on a ten-point scale. When the researchers combined the forecasting components into a simulated twelve-month scientific corpus, the resulting papers showed lower diversity and novelty than the research humans produced during the same period.
Scientific fields develop through continuities, accidents, rediscoveries, institutional choices, unusual collaborations, and ideas imported from elsewhere. A model may extrapolate the most visible trajectory while missing the contribution that later reorganizes the field.
Prediction becomes especially difficult near moments of divergence. Several directions look viable, resources are distributed unevenly, and the decisive experiment may not yet have been conceived. The future of a discipline depends partly on choices that the existing literature cannot fully contain.
A forecast of science therefore requires some account of novelty itself. The forecaster has to estimate which combinations are likely to emerge, which neglected ideas may return, and where current assumptions will break under experimental pressure.
SciPredict brings the problem closer to empirical work. The benchmark contains 405 tasks drawn from recent studies across 33 specialized areas of physics, biology, and chemistry. Each task describes an experimental system and asks the model to predict the observed outcome.
Frontier models achieved accuracies between 14 and 26 percent. Human experts reached approximately 20 percent overall. The most revealing difference appeared when participants assessed the predictability of the experiment itself.
Human accuracy rose from approximately 5 percent on experiments judged difficult to predict to around 80 percent on those judged highly predictable. Model accuracy remained close to 20 percent across different confidence levels and across the models’ own assessments of whether physical experimentation was necessary.
Model-reported confidence showed little relation to prediction accuracy.
Research decisions depend on estimates whose reliability can be understood. A laboratory can work with uncertainty when that uncertainty is visible. Poorly calibrated confidence can redirect attention, funding, and experimental effort toward weak directions.
SciPredict also found that expert-curated background knowledge improved model performance. Background material generated by the models themselves frequently reduced accuracy through irrelevant details and misleading assumptions. The results locate part of the difficulty in the selection and weighting of evidence.
ForeSci examines forward-looking judgment within AI research. It contains 500 tasks across four fast-moving domains and four decision families. Systems must identify bottlenecks, choose technical directions, rank research plans, and position projects within a wider field.
Each task is paired with an offline knowledge base aligned to its historical cutoff. Papers published after that date remain hidden during generation and are used later for validation.
Explicit evidence organization improved traceability and factual support. The researchers also identified a recurring pattern they call evidence-decision decoupling: an agent may retrieve relevant evidence and still forecast the wrong research object.
Real sources, coherent analysis, and accurate local citations can culminate in a forecast that misreads the direction of the field.
Scientific foresight depends on how evidence is weighted, which uncertainties are recognized, and which course of action is selected. Literature retrieval supplies material for that judgment. The central decision remains unresolved.
The same problem appears throughout human research. A well-cited proposal may rest on a weak estimate of feasibility. A fashionable direction may accumulate evidence while its central assumptions remain untested. A precise description of the present offers limited protection against a mistaken future.
Scientific institutions already make forecasts constantly.
Peer review contains judgments about whether a method will work. Grant panels estimate future value. Research roadmaps assign timelines to technological development. Laboratories choose experiments according to expected information gain. Journal articles end with predictions about generalization, replication, and future applications.
These judgments shape careers, funding, and the distribution of scientific attention. Their original form often disappears.
A broad statement may later be remembered as a precise prediction. A failed expectation can be reframed as an exploratory suggestion. A successful one can acquire a confidence its author never expressed.
A clearer record would preserve the expectation at the moment it was made.
Each forecast could include a specific claim, a probability, a deadline, a resolution method, and the evidence available at the time. Updates could remain visible as new information arrived. Once the outcome became known, the record would show how the prediction changed and how closely its confidence corresponded to reality.
Such an archive would reveal forms of scientific judgment that publication alone cannot capture. Some researchers may excel at identifying feasible experiments. Others may understand timing. A model may perform well in one discipline and remain unreliable in another. Groups may detect broad shifts early, while specialists retain an advantage on narrow technical questions.
The history of changing probabilities would also show how scientific belief moves before consensus settles.
Episteme is developing an environment for recording scientific claims while their outcomes remain unresolved.
Each hypothesis is structured around defined parameters, a time horizon, and resolution terms. Participants take positions as new evidence emerges, while the platform tracks changes in the aggregate probability of the claim. This evolving record can preserve how expectations changed before the outcome became known.
Over time, such an archive could make different forms of scientific judgment comparable. Researchers, communities, and AI systems could encounter the same question, deadline, and resolution criteria, leaving their estimates open to evaluation after the result.
The record would support a different history of science, composed of expectations, revisions, hesitation, misplaced confidence, and early recognition.
Which forecasters understand the limits of their knowledge?
Which domains produce reliable collective judgment?
What kinds of evidence lead to useful revisions?
Where do experts and models diverge?
How early can a genuine change in scientific direction be detected?
These questions concern the formation of scientific expectations, their response to evidence, and the conditions under which uncertainty becomes collective knowledge.
AI systems are becoming more involved in hypothesis generation, research planning, and experimental design. As their outputs enter decisions about funding, laboratory time, and scientific priorities, prospective records will offer one way to evaluate their judgment through outcomes unfolding in real time.
Scientific archives preserve results after uncertainty has closed. A forecasting record can preserve the earlier landscape of expectation: the available evidence, the probabilities implied by collective judgment, and the revisions made while several futures remained possible.
Episteme begins with the attempt to build that record.
The four benchmarks discussed in this article were released as preprints in 2026 and may be revised.
Sean Wu et al., “Forecasting Scientific Progress with Artificial Intelligence”, arXiv, May 21, 2026.
Anirudh Ajith et al., “PreScience: A Benchmark for Forecasting Scientific Contributions”, arXiv, February 24, 2026.
Udari Madhushani Sehwag et al., “SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?”, arXiv, April 12, 2026.
Qiuyu Tian et al., “ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment”, arXiv, May 30, 2026.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.