The debate has recently found two prominent champions. Fei-Fei Li, with World Labs, is building what she calls “spatial intelligence.” Yann LeCun, with AMI, is explicitly positioning his work as an alternative to the language model paradigm that has produced ChatGPT, Claude, and Gemini. Both are collecting billions of dollars before shipping a single product. Both are making the same foundational claim: that language models are not enough, and that what we need instead are world models — systems capable not merely of describing reality but of structuring it causally, of predicting consequences, of acting rather than narrating.
The claim is correct. But it is correct in a way that goes much deeper than the debate currently acknowledges, and the depth of the correctness is precisely what makes the proposed solution uncertain.
Hume noticed it in the eighteenth century. No finite sequence of observations logically entails a universal law. From a thousand white swans you cannot derive “all swans are white” — the inference is a leap, not a deduction. Quine radicalized this three centuries later: any set of empirical data is compatible with infinitely many theoretical systems, provided you are willing to adjust your peripheral beliefs accordingly. The data underdetermines the theory. Always. Without exception.
This is not a limitation of language models. It is the structural condition of any knowing system — biological or artificial, connectionist or symbolic. A world model trained on physical data will construct a causal representation of reality. But which causal representation, among the infinite number compatible with the training data? The problem does not disappear with better architecture. It migrates to a deeper level.
Judea Pearl formalized this precisely. His do-calculus — the algebra of intervention — allows us to compute the effect of acting on a variable, rather than merely observing it. But the computation requires knowing the causal graph: which variables cause which, in what directions. And that graph, in general, is not identifiable from observational data alone. You need experiments. You need interventions. You need, in a word, agency — a body that touches the world and watches what happens.
The implication is uncomfortable: a world model trained passively on video, text, or even sensorimotor data may construct a causal representation that is internally consistent and empirically adequate while remaining structurally false. The Ptolemaic system was internally consistent and empirically adequate. For fifteen centuries.
Alison Gopnik, at Berkeley, has spent her career studying how children learn. Her conclusion is that children are not passive statisticians. They are experimenters. They manipulate objects, test hypotheses, update their models based on the results of their own interventions. Learning is active and interventionist at its core — not because activity is pedagogically useful, but because causality is only legible through action. You cannot infer a cause by watching. You infer it by pushing.
Josh Tenenbaum, at MIT, has mapped what he calls “intuitive physics” and “intuitive psychology” — the implicit representations of objects, agents, forces, and intentions that human infants deploy well before they acquire language. These representations appear to be learned through embodied interaction with the world in the first months of life. The crucial point: language does not generate these representations. It presupposes them. It expresses and refines a prior cognitive architecture that was built through touch, gravity, resistance, and the persistent identity of objects across occlusion.
If Tenenbaum is right — and the evidence increasingly suggests he is — then language models are missing something more fundamental than causal structure. They are missing the pre-linguistic layer on which language itself is built. A world model trained on text or video may reproduce the same gap, one level deeper.
The fashionable version of this argument invokes Kahneman: language models are “System 1,” fast and associative; world models would enable “System 2,” deliberate and causal. The map is partially wrong. System 2 in humans is not computationally pure — it is permeated by System 1 biases, emotional inflections, heuristic shortcuts. A “pure System 2” artificial reasoner would be something that does not exist in nature and whose properties we genuinely do not know. The analogy illuminates and misleads simultaneously.
In Being and Time, he distinguishes between two modes of relating to objects. Vorhanden — present-at-hand: the mode of theoretical contemplation, in which an object is isolated, represented, analyzed. Zuhanden — ready-to-hand: the mode of practical engagement, in which a tool is not represented but used, its properties absorbed into the transparency of fluent action. The hammer, when it works, disappears. You do not think “hammer” — you think “nail.” The tool vanishes into the task.
Understanding, for Heidegger, is primarily Zuhanden. It is not the possession of a correct representation. It is the capacity for fluent, situated action — knowing how to go on, in the particular place, with the particular resistances and affordances that the environment provides.
A language model is constitutively Vorhanden. It represents the world without inhabiting it. It describes hammers without the weight of them, describes gravity without having fallen. A world model, if it remains a static representation queried episodically, replicates the same structure at a different level of abstraction. The transition from Vorhanden to Zuhanden requires not a better model but a different relationship to the world — embodiment, feedback loops, the irreversibility of action.
Merleau-Ponty goes further. The body is not a vehicle for a mind that could in principle exist without it. The body constitutes perception. The schema corporel — the implicit representation of one’s own body as a field of motor possibilities — is the condition of possibility of spatial understanding. “Spatial intelligence,” the phrase Li uses for World Labs, is precisely this: not a geometric model of space, but a lived, motor-inflected relationship to it. The question is whether that can be learned without a body that has actually moved through space with stakes attached to the movement.
In the Philosophical Investigations, he does not answer the question “what is understanding?” He dissolves it. To understand is to know how to go on — to use an expression correctly in new contexts, to participate in a practice. “Following a rule” is not a private mental event. It is a form of social life. Correctness is constituted by participation in a community of practice, not by correspondence between an internal representation and an external state of affairs.
This is scomodo — uncomfortable — for both sides of the current debate.
For those who defend language models: a system that continues texts correctly has not demonstrated understanding in the Wittgensteinian sense, because that sense requires participation in embodied social practices, not mastery of statistical patterns. The Turing test fails as a criterion precisely because it confines the test to the medium in which LLMs are strongest.
For those who propose world models: a system with an accurate causal representation of physical reality has also not demonstrated understanding, because understanding requires normative integration into social practices. A robot with a perfect world model that never interacts with humans in normatively structured contexts — contexts where correctness is socially calibrated, where mistakes have social consequences — does not understand in the relevant sense.
The real threshold is not language model versus world model. It is isolated system versus system integrated into sociotechnical ecosystems. This is the distinction the current debate almost entirely ignores.
LeCun is Chief AI Scientist at Meta. Meta has lost the frontier LLM race to OpenAI and Anthropic, despite LLaMA’s open-source success. The narrative “LLMs are fundamentally limited” serves a precise competitive function: it delegitimizes the advantage that OpenAI and Anthropic have built in the current paradigm, and positions Meta as the protagonist of the next one. The narrative is not false — but it is not neutral either. Its production is economically rational independently of its truth.
Li built ImageNet — the data infrastructure that enabled the deep learning revolution in computer vision. “Spatial intelligence” is a bet on the convergence of vision, 3D representation, and robotics: precisely the domain where her intellectual assets and network position are maximally valuable. Again, the bet may be right. But the incentive structure that produces it must be held alongside the claim.
A billion dollars collected before shipping a product does not validate a scientific hypothesis. In a venture capital market still processing the returns of early OpenAI and Anthropic investors, it signals excess liquidity seeking differentiation. The LLM market is already oligopolistic. Capital is rational to purchase options on alternative paradigms, even at high premiums, even with low individual probabilities of success. The billion is an option, not a verdict.
Chickering showed in 2002 that finding the optimal causal DAG from observational data is NP-hard in general. Peters, Mooij, Schölkopf and others have established that causal identifiability from observational data requires strong functional assumptions — additive noise models, linear relationships — that are rarely realistic. The sample complexity of causal structure learning grows exponentially with the number of variables in the general case.
This is not an argument against world models. It is an argument about what world models will actually be able to do: they will learn causally structured representations in restricted domains, under strong assumptions, with experimental data or very dense sensorimotor feedback. They will not learn general causal structure of the world from passive observation. The ambition must be calibrated to the mathematics.
Current LLMs, meanwhile, generalize out-of-distribution more robustly than critics acknowledge — probably because natural language already contains implicit causal structure. Human narratives describe causes and effects. A system trained on enough of them acquires a compressed, imprecise, but operationally non-trivial causal model. The line between language model and world model is less clean than the paradigm war suggests.
So: what is actually happening?
Two serious scientists, positioned by their institutional and competitive contexts to propose a paradigm shift, are raising real philosophical and technical problems — problems that the AI field has been avoiding because confronting them honestly would complicate the narrative of continuous progress through scaling. The problems are genuine. The solutions proposed are uncertain. The capital flowing toward those solutions is responding as much to narrative as to evidence.
What we are witnessing is not the announcement of the next paradigm. It is something more interesting: the moment when a field is forced to confront the question it has been deferring. What does it mean for a system to understand what it is doing? And the honest answer, which neither the language model defenders nor the world model proponents fully acknowledge, is: we do not yet have the conceptual vocabulary to answer that question rigorously. We are still arguing, in 2026, about distinctions that Heidegger drew in 1927, that Wittgenstein drew in 1953, that Merleau-Ponty drew in 1945 — and that the engineering tradition absorbed only partially, selectively, and often incorrectly.
The billion dollars is not funding a product. It is funding a question that philosophy asked a century ago and left, correctly, open.
The world is not a text. But neither is it a model. It is the thing against which both texts and models break — and from which, if we are lucky, we learn something about what breaking means.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.