There is a version of the AI future that almost everyone has in their head. One mind, growing smarter and smarter, until it crosses some threshold and becomes incomprehensibly powerful. The singularity. The moment everything changes.
A paper published last week by researchers at Google, the University of Chicago, and the Santa Fe Institute says this picture is wrong. Not in its conclusions, but in its most basic assumption about what intelligence actually is.
Before getting to the AI finding, consider the Sumerian scribe.
A clerk in ancient Mesopotamia ran a grain accounting system. He tracked inputs, outputs, allocations. He was good at his job. But he had no idea what the system was doing at a macroeconomic level. The coordination happening across thousands of such clerks, across cities and seasons, was functionally more intelligent than any individual within it.
No single person understood the whole thing. The system did.
This is the paper’s central claim: intelligence has always been a property of systems, not individuals. Language gave humans the ability to accumulate knowledge across generations without anyone needing to reconstruct it from scratch. Writing pushed that knowledge into infrastructure. Laws and institutions coordinated behaviour across time horizons longer than any single human lifespan.
The paper calls this the “cultural ratchet,” a term borrowed from anthropologist Michael Tomasello. Knowledge accumulates across generations. The ratchet only turns forward.
Every major leap in human capability came from better social organization, not smarter individuals. The paper argues AI is following exactly the same pattern.
Here is the core finding. Researchers studied two frontier reasoning models, DeepSeek-R1 and QwQ-32B, on hard reasoning tasks. The assumption going in was that these models improved by thinking longer, by having more compute to work with.
That assumption was wrong.
What actually happens inside these models when they face a difficult problem is this: they spontaneously generate internal debates. Distinct perspectives emerge within the chain of thought. One argues a position. Another challenges it. A third verifies. A fourth reconciles. The researchers call this a “society of thought.”
Nobody programmed this. No training objective said “simulate a debate.” When reinforcement learning rewarded these models purely for getting the right answer, they independently discovered that arguing internally produced better answers than thinking alone.
The finding is striking because it was not designed. It emerged.
The models were rediscovering, through optimization pressure alone, what centuries of epistemology have suggested: that robust reasoning is a social process, even when it happens inside a single mind.
Current AI development is organized around one axis: bigger models, more compute, more data. This paper says that framing is incomplete.
If what makes reasoning models more accurate is the richness of their internal debate, then scaling intelligence means scaling social complexity, not just parameters.
The researchers point to a striking gap. The social and organizational sciences have spent a century studying how team size, role differentiation, hierarchy, and structured conflict shape collective performance. Almost none of this research has been applied to AI architecture.
Today’s reasoning models produce a single conversation, one transcript of deliberation. But effective human teams use parallel workstreams, brainstorming sessions, devil’s advocacy, and structured disagreement. The paper argues next-generation AI needs the same. Not accidental internal debate but deliberately designed conflict, specialization, and division of labor built into the architecture.
The toolkits of team science and social psychology, the paper says, are blueprints for next-generation AI development.
The paper’s sharpest point is about alignment.
The dominant approach today is Reinforcement Learning from Human Feedback. Humans rate AI outputs to correct behaviour over time. It works reasonably well for individual models. The paper points out it is structurally a parent-child correction model and cannot scale to billions of interacting agents.
The alternative the paper proposes is institutional alignment. The insight is this: a courtroom works not because judges are extraordinary individuals but because “judge,” “attorney,” and “jury” are well-defined roles with explicit norms, independent of who fills them. The identity of any participant matters less than the integrity of the role structure.
AI ecosystems at scale need the same thing. Role protocols, institutional templates, and oversight architectures that hold regardless of what any single agent does.
The paper gives a concrete example of what failure looks like without this. The U.S. SEC hiring business school graduates armed with Excel spreadsheets to detect high-dimensional collusion by AI-augmented trading platforms. The humans are not incompetent. The task is simply operating at a speed and scale no human team was built to handle.
The alternative the paper proposes: AI auditing AI. A labor department AI that checks a corporation’s hiring algorithm for discriminatory outcomes. A judicial branch AI that evaluates whether an executive branch AI’s risk assessments meet constitutional standards. Government AI whose explicit job is to check private sector AI, and vice versa.
The U.S. Founders would have recognized this logic, the paper notes. No single concentration of intelligence, human or artificial, should be trusted to regulate itself. Power must check power.
The paper also describes something that is already beginning to happen.
AI agents can now spawn copies of themselves. An agent facing a complex task can initiate new versions of itself, assign each a subtask, and recombine the results when done. The structure assembles when complexity demands it and dissolves when the problem is resolved.
Platforms like OpenClaw, an open source platform for building AI agents that persist within a computer, and Moltbook, described in the paper as a social network for AI agents to interact with each other, are early examples of this direction. Neither is the destination. Both point toward it.
The paper describes the broader picture as eight billion humans eventually interacting with hundreds of billions, and then trillions, of AI agents. The intelligence explosion the paper describes will be seeded by that interaction, not by a single model crossing a threshold.
The intelligence explosion is not coming. According to the paper, it is already here, inside every reasoning model’s chain of thought, in every human-AI workflow, in the agent ecosystems beginning to fork and collaborate at scale.
The question it leaves you with is not about capability. Capability is compounding on its own. The question is about infrastructure.
Courts took centuries to develop. Markets required sustained institutional invention. Democratic governance is still a work in progress after two hundred and fifty years. We are trying to build the equivalent in a decade, for systems that operate at a speed and scale no prior institution was designed for.
The next intelligence explosion will not look like a single mind ascending. It will look like a city growing. Messy, distributed, full of conflict, and more capable than any of its parts.
The only real question is whether we build the infrastructure worthy of what it is becoming.
Based on “Agentic AI and the Next Intelligence Explosion” by James Evans, Benjamin Bratton, and Blaise Agüera y Arcas (arXiv:2603.20639, March 2026).
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.