RSS Amplifier

EPISTEME · Aug 18, 2026

Forecasting Without a Crowd

0
Sign in to vote or save

sasha shilina · EPISTEME

Prediction markets have acquired a familiar origin story. Gather enough people around an uncertain event, give them something to risk, and their scattered fragments of knowledge begin to settle into a price. Expertise enters alongside instinct, gossip, private information, bad guesses, careful research and occasional luck. A number emerges from the mixture.

A candidate has a 63% chance of winning. A drug has a 41% chance of approval. An experiment has a 72% chance of replicating.

Once the number appears, its history tends to disappear. The price says nothing about how many people produced it, how much they knew, whether their information came from independent sources, or whether one unusually well-informed participant did most of the work.

Recent research on prediction markets has begun to recover some of that hidden history.

A large study of Polymarket activity followed 1.72 million accounts across almost 100,000 events over two years. Only around 3% of accounts showed persistent forecasting skill. Those traders reacted unusually quickly to new information and were disproportionately responsible for moving prices toward eventual outcomes. Skilled traders and market makers together represented fewer than 3.5% of accounts while capturing more than 30% of total gains.

The old phrase wisdom of crowds begins to look slightly misleading here. A crowd can contain enormous differences in attention, knowledge and skill. Thousands of participants may provide liquidity and activity while a small minority carries much of the information that makes the market useful.

Scientific forecasting brings this unevenness into sharper relief. Science rarely supplies crowds of thousands. A question about whether a result will replicate may be meaningful to a few dozen researchers. A specialised biotechnology milestone may have a relatively small circle of people capable of assessing it seriously. Certain theoretical disputes become narrow enough that five informed participants already represent a substantial part of the relevant community.

And yet small scientific forecasting markets have produced useful results.

Across four projects that used markets to predict scientific replication, forecasts correctly classified the outcomes of 103 replication attempts about 73% of the time.

A later experiment recruited 162 social scientists to estimate the replicability of online experiments. Market forecasts helped determine which studies would actually undergo replication. Ten of the twelve studies receiving the highest prices replicated according to the preregistered primary criterion. Among the twelve receiving the lowest prices, four did.

These results suggest that headcount alone tells us little about the quality of a forecasting system. A small group can carry a great deal of information when the question is well specified and the participants know something worth aggregating.

The harder problem appears when that information reaches the market unevenly.

In June, Giovanni Angelini and Luca De Angelis published a fine-grained study of Kalshi markets covering 1,438 NBA games. Sport offers an unusually clean setting for studying information flow because reality keeps generating timestamped events: baskets, turnovers, fouls, lead changes. Market prices can be compared minute by minute with statistical estimates of how the game state has changed.

Prices generally moved in the expected direction after new information arrived, though much of the adjustment happened slowly. A one-minute change in the benchmark probability produced around 64% of the corresponding immediate movement in market price. The rest accumulated afterward, with larger delays under low liquidity.

For a thin scientific market, this creates an ambiguity. A quiet price may reflect uncertainty, lack of attention, scarce liquidity, or a piece of information that has reached only a few participants. The visible market compresses all of these conditions into the same numerical surface.

Science makes the problem even more pronounced because evidence arrives irregularly. A paper appears. A dataset is released. A replication fails. A trial reaches an endpoint. A laboratory quietly changes a deadline. A benchmark improves enough to alter what seemed plausible a month earlier. Long periods of stillness can be interrupted by a single piece of evidence that changes the entire shape of a forecast.

A scientific prediction system therefore accumulates something richer than a sequence of prices.

Imagine a replication market sitting at 58% for several weeks. A new dataset appears and the price rises to 76%. The movement itself carries a history. Which evidence triggered it? Who reacted first? How concentrated was the change? Did the participants who moved early have a good record on similar questions? Did specialists react differently from generalist forecasters?

Over time, those traces begin to describe the behaviour of the forecasting system itself.

Eight researchers who have spent years working on a problem and eight strangers clicking FOR produce the same participant count. The informational content can be radically different.

A single probability hides that distinction. Scientific forecasting becomes more legible when some of the structure beneath the number remains visible: the number of independent participants, the concentration of positions, previous forecasting records, domain expertise, reactions to new evidence, and the amount of liquidity available when a participant tries to express a view.

There is something attractive about the severe compression of a market price. Entire disagreements collapse into a percentage. For scientific questions, the surrounding provenance may deserve to survive alongside the number.

That provenance becomes especially interesting once machines join the forecasting population.

Earlier this year, researchers reported experiments with hybrid markets for estimating scientific replicability, allowing algorithmic agents to trade alongside human forecasters. The artificial agents drew on patterns learned from previous replication outcomes; human participants contributed disciplinary context and judgment. Across most of the reported experiments, the hybrid markets matched or exceeded the performance of artificial-only markets.

A single question could be evaluated simultaneously by researchers close to the field, experienced generalist forecasters, specialised prediction models, frontier language models and an open public market. Each group would leave a different trace.

A new paper appears and several models move sharply toward FOR. Domain researchers remain sceptical. The wider market barely reacts. Two weeks later an independent result arrives.

The interesting record now extends beyond the eventual winner. We can see who recognised consequential evidence early, who revised too aggressively, who remained confident for too long, and which kinds of forecasters repeatedly performed well in particular domains.

Models may converge because they inherit similar training data or assumptions. Researchers may share disciplinary blind spots. A public market may occasionally notice something both groups miss. Across dozens or hundreds of resolved questions, these patterns become measurable.

A forecast then leaves behind a small history of cognition under uncertainty.

Early collective-intelligence systems face a peculiar dependency. Their signal grows more informative as independent participants accumulate, while the presence of an informative signal gives new participants a reason to pay attention in the first place.

At the beginning, very little hides this circularity.

A market with four participants can simply display four participants. A price dominated by one position can disclose its concentration. Expert forecasts can carry their own label. Model estimates can remain separately visible. A young market can expose the fragility of its signal without losing its usefulness.

Such transparency becomes more valuable as questions resolve.

Resolution introduces a dimension that raw activity cannot provide: performance over time.

After fifty questions, it becomes possible to see which forecasters repeatedly recognised important evidence early. After a hundred, one can begin asking whether expertise helped more in biotechnology than in machine learning, whether models were better calibrated on product milestones than on replication questions, whether some participants systematically overreacted to new papers, or whether a small group of proven forecasters consistently outperformed a larger open crowd.

The archive gradually acquires a second life. It records expectations about future events, and it also records how different forms of intelligence behaved while those events were still uncertain.

Sometimes the useful signal will emerge from a large crowd. Sometimes it will come from a handful of specialists. Sometimes the revealing moment will be a disagreement between a model and a researcher, or a solitary participant who moved long before everyone else.

A market gives those judgments a place to become visible.

Time gives them a record.

Eidōlon

Arena Guide

Eidolon Points 101

Clusters Registration Form

Propose a Hypothesis

Website: episteme.ac

X (Twitter): @episteme_sci

Telegram: Channel | Community

Linktree: linktr.ee/episteme_sci

No posts

Read the original on epistemesci.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.