This paper examines Judea Pearl’s causal revolution in science via structural causal models and evaluates its force in the context of the reasoning capacity of large language models (LLMs). It places Pearl’s thesis in dialogue with five further bodies of thought: Richard Sutton’s Bitter Lesson, which argues that general learning from data always outperforms expert-crafted structure; Bas van Fraassen’s constructive empiricism, which contests the metaphysical realism Pearl’s framework presupposes; the embodied cognition tradition of Andy Clark, David Kirsh, and Michael Spranger, which argues that genuine physical causal understanding requires sensorimotor coupling and social interaction; the agent-based simulation programmes of Brian Skyrms and J. Doyne Farmer — together with their philosophical foundations in Michael Bratman’s planning theory of intention, Daniel Dennett’s intentional stance, the BDI architecture of Rao and Georgeff, and Joshua Epstein’s generative social science — which provide the functional analogue of embodied intervention for the social and economic domain; and Niklas Luhmann’s systems-theoretic account of causality as a contingent, highly selective and system-specific complexity-reducing attribution schema reproduced by social systems. The paper draws on recent empirical benchmarking of reasoning LLMs — including the Corr2Cause and CausalBench studies, the executable counterfactuals research of Vashishtha et al., the NeurIPS 2024 work of Liu et al., and the constructive findings of Kıcıman et al. — to argue that Pearl’s challenge remains valid in a qualified form. A pure LLM trained only on text is unlikely to be sufficient for robust artificial general intelligence, but neither are LLMs irrelevant to it: they are powerful discourse engines that may serve as components in hybrid architectures. A key contribution is the argument that the domain boundary between physical and social causality maps onto distinct methodological requirements: robotics and embodied agents for the former; agent-based simulation populated by socially constituted goal-oriented agents for the latter — the latter requiring the philosophical apparatus of intention, belief, and plan that Bratman, Dennett, and BDI architectures supply. The path forward is the principled integration of LLM capabilities with structured causal models, embodied robotic systems, and agent-based social simulation environments, drawing on each tradition for the domain to which it is appropriate.
The paper culminates in a systems-theoretical sociological critique of the very notion of a non-domain-restricted, unitary AGI: drawing on Luhmann’s sociology of functional differentiation as practically relevant first philosophy, it argues that the very conception of artificial general intelligence as a single all-purpose peak intellectual competence is a false expectation and misguided goal. Advanced modernity is constituted by operationally closed, code-distinct functional systems whose incommensurable discourse-practices cannot be synthesised into a unified rationality. The proper conception of advanced artificial intelligence that emerges is therefore not unitary AGI but a constellation of specialised artificial competences, each developed in deep co-evolution with the functional system it serves.
Keywords: causal inference, structural causal models, do-calculus, large language models, hybrid AI architectures, embodied cognition, robotics, agent-based simulation, BDI architecture, intentional stance, planning theory of intention, generative social science, evolutionary game theory, complexity economics, constructive empiricism, systems theory, functionally differentiated society, second order observations, self-reproducing structures, polycontexturality
In 2018, Judea Pearl and Dana Mackenzie published ‘The Book of Why’, a work that translated Pearl’s decades of technical research into a pointed provocation: that artificial intelligence, despite its spectacular advances, remained fundamentally limited by its confinement to the first rung of what Pearl calls the Ladder of Causation — the level of statistical association. Pearl’s claim was not merely technical, but philosophical, as a matter of logic: that no amount of data, and no learning algorithm operating solely on observational distributions, could ever substitute for the formal apparatus of causal modelling. His slogan — ‘no causes in, no causes out’ — captures the argument in its starkest form.
This paper takes Pearl’s thesis as its point of departure and subjects it to a multi-fronted philosophical dialogue. The first interlocutor is Richard Sutton, whose 2019 essay ‘The Bitter Lesson’ represents perhaps the most influential counter-position in contemporary AI research: the inductive historical claim that general learning methods, scaled with compute, consistently outperform approaches that encode domain knowledge. The second is Bas van Fraassen, whose constructive empiricism offers a sustained philosophical challenge to the metaphysical realism about causal structure that Pearl’s framework presupposes. The third — introduced here as a productive middle ground between Pearl’s formalism and Sutton’s empiricism — is the embodied cognition tradition, represented by Andy Clark, David Kirsh, and Michael Spranger, which argues that genuine causal and spatial understanding requires physical situatedness and social interaction. The fourth is the agent-based simulation tradition, represented by Brian Skyrms and J. Doyne Farmer and grounded philosophically in Michael Bratman’s planning theory of intention, Daniel Dennett’s intentional stance, the BDI computational architecture, and Joshua Epstein’s generative social science, which provides the analogue of embodied causal discovery for the social and economic domains where physical robotics is inapplicable. The fifth is Niklas Luhmann, whose systems-theoretic account of causality as a contingent, system-specific complexity-reducing attribution schema reframes the entire debate in terms of social epistemology, whereby sociology – in my view rightly – claims the mantle of a practically relevant prima philosophia.
The paper proceeds in nine sections. Section 2 reconstructs Pearl’s framework. Section 3 examines its empirical status in light of current reasoning LLMs, drawing on the most recent benchmark evidence on both sides of the debate. Section 4 analyses the Pearl-Sutton confrontation. Section 5 brings van Fraassen’s constructive empiricism to bear. Section 6 introduces the embodied cognition tradition and its implications for physical causal AI. Section 7 introduces the agent-based simulation programmes of Skyrms and Farmer as the functional analogue for social causality. Section 8 presents Luhmann’s framework and its challenge to all prior positions. Section 9 summarizes where the argument stands up to here. Section 10 concludes by arguing that an all-purpose AGI is a misguided goal in an irredeemably polycontextural social world.
Pearl’s starting point is a logical observation about the expressive limits of probability theory. Standard probability calculus can represent the conditional probability P(Y|X) — the probability of Y given that we observe X — but it has no notation for the distinct claim that X causes Y, nor for the quantity P(Y|do(X)) — the probability of Y were we to intervene and set X to some value. This is not merely a notational gap. It reflects a genuine conceptual distinction: observing that patients who take a drug recover more often than those who do not is different from knowing what would happen if we administered the drug.
The Inadequacy of Statistical Language
Pearl’s demarcation line is stark: ‘causal relations cannot be expressed in the language of probability, and hence any mathematical approach to causal analysis must acquire new notation — probability calculus is insufficient’ (Pearl, 2009, p. 40). The implication is that any AI system confined to learning statistical associations from data — however vast the data, however sophisticated the algorithm — operates below the threshold at which genuine causal reasoning becomes possible. Pearl organises causal reasoning into a three-rung hierarchy. Each rung corresponds to a qualitatively different cognitive activity and a distinct class of query that cannot be answered from data at a lower rung without additional assumptions. Here is Pearl’s ‘Ladder of Causation’:
Rung 1 — Association (Seeing): The domain of standard statistics and machine learning. What does observing X tell me about Y? Contemporary deep learning operates primarily at this level.
Rung 2 — Intervention (Doing): What happens to Y if I actively change X? This requires the do-operator. Randomised controlled trials are the canonical experimental instrument; Pearl’s do-calculus provides the formal tools for accessing this rung from observational data combined with a causal model.
Rung 3 — Counterfactual (Imagining): What would have happened to Y if X had been different, given that X actually took value x? This requires the full structural causal model. It is the rung most characteristic of sophisticated human reasoning about causes, attribution, regret, and responsibility.
A fundamental result is that higher-rung queries cannot in general be answered using only lower-rung data. This is not a practical limitation but a mathematical one — and it grounds Pearl’s claim that AI systems confined to statistical learning face a hard ceiling. Pearl’s formal apparatus centres on structural causal models (SCMs), which combine a directed acyclic graph (DAG) encoding causal structure with structural equations specifying how each variable is determined by its parents and noise terms. The do-calculus — a complete set of three inference rules — formalises the distinction between P(Y|X) and P(Y|do(X)), operationalising the act of intervention by deleting all arrows into X and fixing its value from outside the system (Pearl, 1995; 2000). Together, SCMs and the do-calculus constitute the two languages of the causal revolution: graphical models for representing causal assumptions transparently, and the do-calculus for deriving causal and counterfactual conclusions from those assumptions combined with data.
Richard Sutton’s 2019 essay surveys decades of AI research to argue a recurring pattern: methods incorporating human domain knowledge produce short-term gains but are consistently overtaken by general methods — search and learning — that scale with compute. The lesson is ‘bitter’ because AI researchers repeatedly forget it. The practical implication is a methodological empiricism: do not encode structure, let the system learn its own representations from data.
The Depth of the Disagreement
The Pearl-Sutton confrontation is not a disagreement about empirical facts alone — it is a disagreement about the epistemic type of the constraint Pearl identifies. Sutton’s lesson is an inductive generalisation: it has always been true that scaling wins. Pearl’s argument is deductive: there is a structural mathematical reason why associational data cannot identify causal quantities without causal assumptions. An inductive generalisation, however well-supported, is in principle defeasible. A mathematical impossibility result is not. Pearl is not predicting that scaling will probably fail; he is arguing that certain queries are definitionally unanswerable from certain data, regardless of the learner’s power.
The Empirical Falsification Zone
The current evidence does not support a Suttonian resolution. The 25–40% accuracy drop from interventional to counterfactual tasks in state-of-the-art reasoning models suggests the Rung 2/3 boundary is not eroding under scaling in the way Sutton’s framework predicts. Whether this gap will close with the next generation of models or prove persistent is genuinely open as of 2026 — constituting an empirical falsification zone where the Pearl-Sutton debate will be partially adjudicated over the coming years.
Sutton’s Hidden Architecture Assumption
There is a further point that Sutton’s framework tends to obscure. The claim that ‘general learning methods’ are architecture-neutral is itself questionable. Every learning system operates with a specific architecture, loss function, inductive bias, and training distribution. These design choices constitute the system’s own observational schema. Sutton’s ‘general methods’ are not observations of the world without a standpoint; they are highly specific second-order observations that happen to be very powerful. The bitter lesson may be better read as a lesson about which observational schemas scale well, not as a lesson that observation requires no schema — a point that anticipates what we shall learn below from Niklas Luhmann’s analysis.
Why LLMs Complicate Pearl’s Original Challenge
LLMs are trained statistically, but their training material contains huge amounts of human causal knowledge: scientific explanations, legal reasoning, technical manuals, everyday narratives, historical accounts, economic analysis, design rationales, and philosophical argument. So an LLM can acquire linguistically encoded causal schemas. The simple dismissal that ‘LLMs only do correlation’ is therefore too crude. Pearl himself acknowledged in a 2023 interview that LLMs trained on text introduced a complication to his original characterisation, conceding that ‘the ladder restrictions do not hold as strictly anymore, because the data is text, and text may contain information on levels two and three’ (Pearl, cited in Pearl and Mackenzie, 2023). A widely cited 2023 paper by Kıcıman, Ness, Sharma, and Tan found that GPT-3.5 and GPT-4 performed surprisingly well on several causal-reasoning tasks, including causal discovery, counterfactual reasoning from text, and event causality. The authors argued that LLMs can help generate causal graphs and extract background assumptions from unstructured text. But they also stressed that LLMs have unpredictable failure modes and that the most promising direction is to combine LLMs with principled causal methods rather than replace causal inference (Kıcıman et al., 2023). So the debate has shifted. The question is no longer whether LLMs can ever utter correct causal claims — they clearly can. The harder question is whether they possess a robust, manipulable causal world-model, or are merely borrowing causal patterns from text without reliable causal competence.
Evidence Supporting Pearl’s Sceptical Side
Several benchmark studies support Pearl’s worry. Jin et al.’s Corr2Cause benchmark asked whether LLMs can infer causation from correlational statements through reasoning rather than commonsense memory. Seventeen LLMs performed close to random; fine-tuning improved in-distribution performance but failed to generalise robustly under perturbations (Jin et al., 2024). CausalBench similarly found that LLMs show promise on simple causal structures but lag behind traditional causal-learning algorithms on larger networks, especially those with more than fifty nodes. It also found that LLMs struggle with collider structures and tend to rely on semantic associations with familiar entities rather than reasoning from numerical distributions or causal structure directly. A NeurIPS 2024 paper, ‘Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?’, is frequently cited by sceptics. Its central finding is that apparent causal competence often collapses under controlled perturbations, suggesting that some LLM ‘causal reasoning’ is pattern-matching or memorised correlation rather than stable structural reasoning (Liu et al., 2024). Recent work on executable counterfactuals (Vashishtha et al., 2025) operationalises Pearl’s three-step counterfactual process — abduction, intervention, and prediction — in code and mathematical problems. The results reveal a substantial accuracy drop of 25–40% when moving from interventional (Rung 2) to counterfactual (Rung 3) tasks for state-of-the-art models including o4-mini and Claude Sonnet. The abduction step is the specific point of failure — precisely the step that has no analogue in forward statistical reasoning.
Clinical reasoning studies of o1, o3, Gemini, Claude, and DeepSeek identify ‘the absence of a robust world model — an internal causal representation that enforces physical and logical constraints, maintains latent state, and supports counterfactual reasoning’ as the underlying cause of systematic failures in medical scenarios designed to disrupt familiar pattern-matching (Griot et al., 2025). Together, this body of evidence supports a qualified Pearlian conclusion: LLMs can imitate, retrieve, and recombine causal reasoning, but they are not yet reliably causal reasoners in the strong structural sense.
Evidence Against a Hard Barrier
The opposing view is not that Pearl is wrong about causality, but that LLMs may become components within causal systems. Several arguments support this position.
First, human causal knowledge is itself often learned from language, testimony, diagrams, examples, and institutionalised discourse. Pearl himself has acknowledged that humans learn much causal information from books, although he insists humans also bring innate causal templates and embodied interaction with the world. The asymmetry between LLM and human causal acquisition may be one of degree and supplementation rather than of fundamental kind.
Second, LLMs can act as causal assistants: proposing variables, identifying possible confounders, drafting causal graphs, explaining assumptions, checking counterfactual scenarios, and translating expert knowledge into formal causal models. This is not AGI by itself, but it is genuinely valuable. The most promising current research direction is hybridisation: LLMs combined with causal discovery algorithms, simulators, experimental design tools, embodied agents, reinforcement learning, tool use, and explicit structural causal models.
Third, reasoning models in 2025–2026 have improved substantially at multi-step problem solving, code execution, tool use, and self-checking. This does not dissolve Pearl’s challenge, but it weakens the older claim that neural systems are merely passive curve-fitters. They can increasingly search, test, revise, simulate, and call external tools. The open issue is whether this becomes a genuine causal architecture or merely a more elaborate association machine wrapped around tools.
Reading the Current Balance
Pearl’s challenge remains valid, but the phrase ‘hard barrier’ needs refinement. A pure LLM, trained only to predict or transform text, is unlikely to be sufficient for robust AGI if AGI means autonomous scientific, legal, engineering, political, or everyday practical reasoning under novel interventions. Such a system needs stable representations of actions, mechanisms, constraints, affordances, temporal processes, and counterfactual alternatives — and the empirical evidence consistently shows that current LLMs lack precisely these stable representations. But this does not imply that neural models or LLMs are irrelevant to AGI. The more plausible view is that LLMs are powerful discourse engines and knowledge interfaces, but AGI will require them to be embedded in architectures with causal modelling, memory, active experimentation, world interaction, simulation, and explicit decision procedures. Pearl is probably right against AGI by correlation-crunching alone, but not necessarily against AGI architectures that use LLMs as one component within a broader causal-agentic system. This refined position has significant implications for the comparisons developed in the remainder of this paper. The architectures most likely to produce genuine causal AI are precisely those that embed LLM components within frameworks providing the elements LLMs lack: structured causal models in the Pearlian tradition; embodied robotic systems for physical causality, in the tradition of Clark, Kirsh, and Spranger; and agent-based simulation environments for social and economic causality, in the tradition of Skyrms and Farmer. The frontier is not a contest between paradigms but their principled integration.
Bas van Fraassen’s The Scientific Image (1980) argues that the aim of science is empirical adequacy rather than truth. A theory is empirically adequate if it correctly describes what is observable; whether its theoretical posits correspond to reality is a question we need not answer. Van Fraassen’s anti-realism is epistemological: he does not deny that unobservable entities might exist, only that we have adequate grounds for believing in them on the basis of theoretical success.
Van Fraassen’s Challenge to Pearl
Pearl’s framework makes strong metaphysical commitments. SCMs are intended as representations of real causal structure. DAGs are supposed to encode genuine asymmetric causal relationships, not merely useful predictive summaries. The do-operator presupposes facts of the matter about what would happen under interventions that never occur. Counterfactuals assert truths about worlds that did not come to pass. Van Fraassen would subject each commitment to critical scrutiny. A constructive empiricist can accept a DAG as a tool organising inference about observables without committing to the claim that its arrows represent real causal powers. On counterfactuals — paradigmatically unobservable — van Fraassen has no clean way to ground such claims in experience. His earlier work on explanation (van Fraassen, 1980, ch. 5) argues that what counts as the cause of an event depends on the contrast class pragmatically determined by the questioner, not fixed by the world. Pearl’s do-calculus, by contrast, aims to give objective, context-independent answers once the causal graph is fixed.
Partial Convergence and the Instrumentalist Reading
SCMs can be understood instrumentally: not as literal maps of nature’s causal joints, but as representational tools enabling a certain class of inferences. This reading is more congenial to van Fraassen than Pearl himself would endorse, but the practical framework of do-calculus and identification theory is largely independent of the metaphysical question. The deeper convergence concerns the inadequacy of pure statistical methods: van Fraassen’s pragmatic account of explanation, while different from Pearl’s causal realism, shares the diagnosis that explanation cannot be read off from distributions alone.
The embodied cognition tradition represents a third position in the debate between Pearl’s formalism and Sutton’s scaling optimism — one that has received insufficient attention in discussions of AI causality. Andy Clark, in works from Being There (1997) to Supersizing the Mind (2008) and Surfing Uncertainty (2016), has consistently argued that cognition is not a process occurring inside a skull but is constitutively extended across brain, body, and environment. Bodily action is not merely the output of computation — it is, in Clark’s formulation, among the means by which computational and representational operations are implemented. The mind extends into the world.
The implication for causal cognition is direct and under-explored. If causal understanding in humans is grounded in sensorimotor interaction with the world — in pushing, pulling, lifting, dropping, releasing, and observing consequences — then an AI system trained exclusively on text or even video is missing the medium in which causal knowledge is originally constituted. It has access to the verbal and symbolic residue of causal experience without the experience itself. Pearl’s ladder, on this reading, cannot be climbed from inside a language model because the ladder is built from embodied intervention, not from description of intervention.
Kirsh and the Epistemic Role of Action
David Kirsh’s work at UCSD on situated and distributed cognition develops this point with particular precision. In his influential paper on epistemic versus pragmatic actions (Kirsh and Maglio, 1995) and in his subsequent work on embodied cognition and interaction design (Kirsh, 2013), Kirsh distinguishes between actions performed to achieve a physical goal and actions performed to improve the cognitive conditions for solving a problem. Agents — human and potentially robotic — restructure their environments not just to accomplish tasks but to create better epistemic affordances: to make causal structure visible, tractable, and manipulable. This has a direct bearing on Pearl’s hierarchy. Rung 2 — the do-operator, the intervention — is not merely a logical operation performed on a causal graph. It is, in its original cognitive form, a physical act: doing something to the world and observing what follows. Kirsh’s framework suggests that the capacity to perform genuine interventional reasoning is inseparable from the capacity to perform genuine interventions. A system that can only predict what would happen under intervention, without ever having intervened, may be performing sophisticated Rung 1 inference about Rung 2 content — not genuine Rung 2 reasoning.
Spranger and the Evolution of Grounded Spatial Language
Michael Spranger’s The Evolution of Grounded Spatial Language (2016), published by Language Science Press as part of the Computational Models of Language Evolution series, provides a particularly compelling empirical demonstration of these principles. Spranger reports a series of robotic experiments in which agents equipped with sensorimotor capacities and placed in shared social environments evolve spatial language — systems of reference for describing the location of objects relative to landmarks, frames of reference, and projective directions — through communicative interaction. What is striking about Spranger’s results for the present debate is that the spatial conceptualisation strategies the robots evolve are not pre-programmed, nor are they derived from a statistical distribution over language data. They emerge from the interplay of environmental conditions, embodied perception, communicative pressure, and social feedback. The agents must, in Pearl’s terms, reach Rung 2: they intervene in their environment, observe consequences, communicate about those consequences, and revise their conceptual schemes accordingly. The causal and spatial categories they develop are genuinely grounded — they are constituted by the sensorimotor history of the agents in their environment, not by pattern-matching over symbols. Spranger’s work also demonstrates that compositional semantics — the capacity to combine primitive spatial concepts into complex descriptions — emerges naturally from embodied interaction, without being explicitly engineered. This has implications for the grounding problem in AI: the reason language models may struggle to acquire genuine causal concepts is not merely that they lack a formal causal graph, but that their representations are not grounded in the sensorimotor cycles that give causal concepts their content in the first place.
Causal Robotics: Embodied Causal Discovery
A growing body of research in causal robotics operationalises these insights computationally. The 2023 IROS workshop on Causality for Robotics articulated the challenge explicitly: ‘correlation-based methods often lack robust generalisation, are brittle to distribution shifts, and can learn incorrect, spurious relationships in observed data. In contrast, a hallmark of human intelligence — reasoning about cause and effect — can provide clues for the next generation of robot intelligence’ (IROS Causality for Robotics, 2023). Concrete implementations have appeared across several domains. Brawer et al. (2020) demonstrated that humanoid robots can build and learn structural causal models from a mix of observation and self-supervised physical trials for tool affordance — precisely the embodied version of Pearl’s Rung 2. Castri et al. (2022; 2023) applied causal discovery algorithms to model and predict human spatial interactions in social robotics contexts, showing that causal structure can be recovered from the time-series of an embodied agent operating in a social environment. The ADAM framework (2024) integrates causal methods and embodied agents in open-world environments, deploying both LLM-based and intervention-based causal discovery within a system that can navigate, perceive, and act. The Causal-HRI workshop at the ACM/IEEE International Conference on Human-Robot Interaction (HRI 2024) brought together researchers working specifically on causal learning for human-robot interaction, recognising that the social task environment — where a robot must understand human intentions, predict the consequences of joint actions, and explain its own behaviour — is precisely the kind of domain where the limitations of purely statistical AI become most apparent and most consequential.
Embodied Cognition as a Third Way
The embodied cognition tradition suggests a path between Pearl’s formalism and Sutton’s scaling optimism that neither fully anticipates. Pearl is right that statistical learning cannot substitute for causal structure — but he tends to locate the solution in explicit formal models imposed from outside the learning system. Sutton is right that hand-crafted representations are limiting — but he underestimates the degree to which the right kind of general learning requires the right kind of embodied situation. The embodied tradition suggests that causal understanding can be learned — contra Pearl’s formalism — but only by systems that are physically situated, action-capable, and socially embedded — contra Sutton’s architecture-agnosticism. This also illuminates the specific failure mode of current LLMs on Rung 3 tasks. These systems have access to vast quantities of human causal discourse — descriptions of interventions, explanations of mechanisms, counterfactual reasoning — but they lack the sensorimotor grounding that gives such discourse its original content. They are, in a sense, reading the map without ever having visited the territory. The abduction step — inferring latent states from observations in order to reason about what would have happened under different conditions — requires precisely the kind of model of the world that is constituted through embodied interaction, not through symbolic compression.
The embodied cognition tradition makes a compelling case that genuine causal understanding of the physical world requires sensorimotor grounding — the capacity to push, pull, release, and observe in a physical environment. Spranger’s robots evolve spatial language through bodily interaction with objects and space; Brawer’s humanoid robot builds causal models through physical trial and error with tools. This insight carries real force for the domain of physical world causality: understanding that a falling weight causes deformation, that heating causes expansion, that friction causes deceleration, is plausibly constituted through the sensorimotor history of an agent in a mechanical environment.
However, a significant portion of the causal questions that matter most to science, policy, and human welfare belong not to the physical domain but to the social domain: sociology, economics, organisational theory and political science. In these domains, the relevant causal structures are not constituted by sensorimotor coupling with objects in space. They are constituted by the interaction of agents with beliefs, strategies, expectations, and institutional rules — entities that have no direct physical correlate accessible to a manipulating arm or a rolling chassis. No amount of robotic embodiment will teach an agent that raising interest rates causes inflation to fall, that inequality causes social fragmentation, or that network effects cause market concentration. These causal relationships emerge from the dynamics of interacting agents operating within social systems, not from the dynamics of forces acting on physical bodies. This domain asymmetry is crucial for the present debate. It does not undermine the embodied cognition argument for physical causal understanding; it limits its scope. And it raises a question that neither Pearl’s formal SCMs nor Sutton’s scaling optimism nor Clark and Kirsh’s embodied frameworks directly address: what is the functional analogue of embodied intervention, for AI systems that need to discover and reason about causal structure in the social domain? The answer proposed here is agent-based interaction simulation.
Brian Skyrms: Agent-Based Simulation as Causal Discovery in Social Systems
Brian Skyrms’s programme of evolutionary game-theoretic simulation — developed across Evolution of the Social Contract (1996), The Stag Hunt and the Evolution of Social Structure (2004), Signals: Evolution, Learning, and Information (2010), and Social Dynamics (2014) — represents the most systematic philosophical and computational programme for discovering causal structure in social systems through agent-based interaction. The core method is to populate a simulated social environment with agents endowed with simple decision rules and strategic repertoires, allow them to interact repeatedly under various conditions, and observe which social structures, norms, conventions, and communication systems emerge from the dynamics of those interactions. The causal claims Skyrms derives from these simulations are Rung 2 and Rung 3 claims in Pearl’s sense. They are not merely observational correlations between social variables. They are claims about what happens when specific parameters are changed — when the network structure of interactions is altered, when the population size varies, when signals are introduced or removed. In Pearl’s terms, the simulated agent populations are performing do(X) operations: the researcher intervenes on the model environment and observes the resulting distribution of outcomes. The simulations thus generate interventional and counterfactual knowledge about social causal structure that no amount of observational data about real social systems could supply, precisely because real social systems do not permit the controlled interventions that the simulations make tractable.
Skyrms’s programme is also directly relevant to the grounding problem for social causal concepts. Consider the concept of a social norm: the claim that a particular norm causes cooperative behaviour, or that its absence causes defection, is not a physical causal claim. It is a claim about the functional role of a shared expectation in an interactive environment. Skyrms’s simulations demonstrate how such norms can emerge spontaneously from agent interactions without being pre-programmed — how the causal efficacy of norms is constituted through the dynamics of interaction, not through any physical mechanism. This is precisely analogous to the way Spranger’s robots’ spatial concepts are constituted through bodily interaction with physical environments: in both cases, the causal concept is grounded in a history of structured interaction rather than being imposed from outside as a formal stipulation. The specific causal mechanisms Skyrms illuminates — the Stag Hunt coordination problem as the fundamental structure of social contract formation, the role of network topology in enabling or impeding cooperation, the spontaneous emergence of meaningful signaling systems from initial states of noise — are paradigmatic examples of Rung 2 and Rung 3 social causal knowledge. They cannot be obtained by training a language model on descriptions of human cooperation. They require running the interactions.
J. Doyne Farmer: Agent-Based Simulation for Economic Causality
J. Doyne Farmer’s programme of complexity economics at the Institute for New Economic Thinking at Oxford provides the most developed application of agent-based simulation to the causal structure of economic systems. In the foundational manifesto Farmer and Foley (2009), ‘The Economy Needs Agent-Based Modelling,’ published in Nature, Farmer and collaborator Duncan Foley argue that the standard tools of macroeconomics — dynamic stochastic general equilibrium models, representative agent frameworks, equilibrium analysis — are constitutively incapable of capturing the causal mechanisms responsible for the phenomena economists most need to understand: financial crises, endogenous business cycles, systemic risk, and the dynamics of technological change. The core problem Farmer identifies is precisely the one Pearl names: standard economic models are built on statistical associations and equilibrium conditions that do not represent causal mechanisms. A DSGE model that correctly predicts the correlation between interest rates and inflation under normal conditions may entirely fail to predict what happens when a financial crisis disrupts the causal mechanisms underlying that correlation — because the model has captured the correlation without representing the mechanism. An agent-based model, by contrast, builds in the causal mechanisms from the bottom up: heterogeneous agents with bounded rationality, making decisions under uncertainty, interacting through networks of supply chains, credit relationships, and market exchanges. The macroeconomic regularities are emergent properties of these micro-level causal interactions, not statistical regularities imposed from above.
Farmer’s Macrocosm project — a large-scale simulation of the global economy representing individual firms and their assets as distinct agents — operationalises this programme at unprecedented scale. As of 2026, the simulation encompasses tens of thousands of companies and their capital assets, with agents making decisions based on simple rules including imitation and trial-and-error learning, in explicit contrast to the fully-informed optimal agents of computable general equilibrium models (Farmer, 2024; Wikipedia, 2026). The causal ambition is clear: Farmer describes the goal as doing ‘for economic planning what Google Maps did for traffic planning’ — providing a model in which causal interventions can be simulated, their consequences traced through the network of agent interactions, and counterfactual scenarios evaluated. This is a direct operationalisation of Pearl’s Rung 2 and Rung 3 for the economic domain.
Farmer and Axtell’s comprehensive survey ‘Agent-Based Modeling in Economics and Finance: Past, Present, and Future’ (2022) documents the maturation of this programme across financial markets, housing markets, macroeconomic stability, and climate economics. Crucially, Farmer explicitly contrasts agent-based modelling with machine learning in terms of causal adequacy, noting that while machine learning ‘builds on the world as it is and it’s unclear if it can handle an alternative situation where it hasn’t seen data,’ an agent-based model ‘has the causal mechanisms built in’ and can therefore address interventional and counterfactual questions that purely statistical models cannot (Farmer, interview, 2024). This is an independent, practitioner-level restatement of Pearl’s argument, arrived at through the practice of economic modelling rather than through philosophy of causation.
The Intentional Foundations of Social Causation: Bratman, Dennett, and BDI Architectures
To understand why agent-based simulation captures something essential about social causality that statistical learning cannot, it is necessary to address the philosophical question of what kind of causation operates in the social domain. The most direct philosophical inspiration for agent-based social simulation comes from Michael Bratman’s planning theory of intention, developed in Intention, Plans, and Practical Reason (1987). Bratman argued that intentions are not reducible to beliefs and desires — they constitute a distinct mental state with three characteristic properties: they persist over time, settling what an agent will do and freeing cognitive resources for other tasks; they constrain future deliberation, such that rational agents do not constantly re-evaluate settled plans; and they coordinate with other plans, both within an agent’s own plan hierarchy and across agents. This planning theory directly inspired the BDI (Belief-Desire-Intention) computational architecture, formalised by Rao and Georgeff in the early 1990s. In BDI agents, beliefs represent the agent’s model of the world; desires represent motivational states (goals); and intentions represent committed plans currently being executed. The architecture is explicitly designed to capture the distinctive feature of practical rationality that Bratman identified: rational agents do not reason from scratch at every moment — they commit, persist, and revise only when necessary. This is not a peripheral feature of human social cognition; it is constitutive of what it means to be a planning agent capable of stable cooperation, reliable promising, and meaningful coordination.
Daniel Dennett’s intentional stance, articulated in The Intentional Stance (1987), provides a complementary framing that is essential for understanding the methodological status of BDI-based simulations. Dennett argues that attributing beliefs, desires, and intentions to a system is a predictive strategy adopted when it provides the most useful level of description for anticipating behaviour. The intentional stance is instrumentally justified — it does not require commitment to the agent ‘really having’ beliefs in some deep metaphysical sense, only that the stance generates good predictions. It is level-relative — the same system can be described at the physical, design, or intentional level depending on what is explanatorily useful. And it is generative — it licenses attributing whatever beliefs and desires a rational agent in a given situation ought to have. This dual philosophical foundation matters enormously for agent-based social simulation. When researchers model agents with BDI architectures, they are operationalising the intentional stance: they are building systems in which the intentional level of description is causally efficacious within the simulation. The question of whether such simulations capture genuine causal structure then becomes a precise empirical-philosophical question: does the simulation’s intentional-level causation track real-world intentional-level causation? This is the social-domain analogue of the question whether Spranger’s robots’ spatial categories track real spatial relationships in the physical world.
The Strategy Layer: From Intentions to Coordinated Action
Between intentions and observable action lies a crucial intermediate level — strategy — where game theory, norm theory, and institutional analysis enter. Several frameworks operate at this layer and inform the design of sophisticated agent-based models. Schelling’s analysis of focal points and tacit coordination demonstrates that agents with compatible intentions can coordinate without explicit communication by reasoning about what a rational counterpart would expect. This is a paradigmatically intentional-stance phenomenon: the explanation of successful coordination depends essentially on attributing to each agent reasoning about the other’s reasoning, which has no purely physical or statistical reduction. Skyrms’s signalling-game work builds directly on this insight, showing how meaningful communication can evolve from initial states of noise through the dynamics of agents reasoning about each other’s expected behaviour.
Elinor Ostrom and Sue Crawford’s Institutional Analysis and Development (IAD) framework treats institutions as shared rules that structure the action situations agents face. Institutions, on this account, operate precisely at the BDI level: they shape which beliefs are relevant, which desires are legitimate, and which intentions are actionable. Cristina Bicchieri’s The Grammar of Society (2006) provides the philosophical grounding for the analysis of social norms as conditional preferences supported by empirical and normative expectations — themselves intentional-stance objects requiring the recursive attribution of beliefs about beliefs. Sophisticated social agency further requires Theory of Mind layering — recursive intentionality of the form ‘I believe that you intend that I believe...’ This recursion is computationally expensive but essential for realistic modelling of deception, trust, signalling, and coordination. The capacity to reason about another agent’s reasoning about one’s own reasoning is what makes possible the entire architecture of social life: contracts, promises, threats, reputations, and the conditional cooperation that sustains complex institutions. None of these are reducible to statistical regularities in observable behaviour; all of them require the intentional-level description that BDI architectures operationalise.
Generative Social Science: Epstein, Sugarscape, and Cognitive Architectures
The most influential demonstration that macro-level social phenomena can be ‘grown’ from intentional micro-level interactions is Joshua Epstein and Robert Axtell’s Growing Artificial Societies (1996), which introduced the Sugarscape model. Sugarscape demonstrated that complex social patterns — trade networks, wealth inequality, cultural differentiation, demographic dynamics, and even rudimentary forms of inheritance — could emerge from populations of agents endowed with simple beliefs (about resource distribution), desires (to gather sugar, to avoid conflict), and intention-like persistence (movement strategies, cultural tags). None of these macro-patterns were explicitly programmed; they emerged from the dynamics of intentional-level interaction. Epstein subsequently extended this programme explicitly toward BDI-style agents in Agent-Based Computational Models and Generative Social Science (2006), articulating the methodological claim that what cannot be ‘grown’ from agent interactions has not been explained — a position he termed the generativist motto. This is a direct social-domain parallel to Pearl’s claim that what cannot be derived from causal assumptions and the do-calculus has not been causally established. Both Epstein and Pearl insist that genuine causal understanding requires more than statistical capture: it requires the demonstration that the proposed mechanism actually produces the observed phenomenon when implemented.
Schelling’s segregation model — predating but later canonised within agent-based modelling — provides a particularly clear illustration of why intentional-stance attribution is essential at the micro level. Mild individual preferences for neighbourhood composition produce strong macro-level segregation, and the explanation of this outcome depends essentially on attributing to each agent a preference (an intentional state) about its neighbours. No purely statistical description of the macro outcome captures the causal mechanism; only the intentional-stance description of agent preferences does. David Hales’s work on tag-based cooperation extends this further, showing how arbitrary identity markers can become the basis of stable cooperation through agents reasoning about expected behaviour conditional on observable tags.
Cognitive architectures such as SOAR (Newell) and ACT-R (Anderson) push the integration further, embedding BDI-like structures within full cognitive frameworks that include working memory, procedural memory, and learning. These architectures have been applied to model military command decision-making, air traffic control errors, and learning in complex environments — domains where the intentional-level causal structure of the situation cannot be bypassed by statistical methods, because the relevant phenomena (commitment, revision, coordination under stress) are constitutively intentional.
Why Intentional-Level Causation Matters for AI World Models
The philosophical and empirical case converges on a key point: many of the most important causal mechanisms in social systems operate at the intentional level. Physical-level descriptions miss them because the relevant entities — beliefs, intentions, expectations, norms — have no straightforward physical correlates. Statistical-level descriptions miss them because the relevant causal pathways pass through the rational decision-making of agents who interpret their situations rather than merely respond to stimuli. A genuine causal world model for social phenomena must represent agent belief states, including false beliefs, uncertainty, and belief updating; hierarchical goal structures, with immediate desires embedded in longer-term plans in Bratman’s sense; commitment and persistence, capturing the fact that rational agents do not re-optimise from scratch; social and normative expectations, including what agents believe others expect of them; and institutional scaffolding — the shared rules that transform the effective action space.
Current LLM-based approaches gesture toward this through in-context theory of mind, but they lack the causal persistence and commitment structure that Bratman identifies as essential to planning agency. An LLM may produce text that describes an agent forming an intention, but it does not itself instantiate the intentional structure — its representations of agent intentions are linguistic descriptions rather than causally efficacious commitments that constrain its own future processing. Agent-based simulation with explicit BDI architectures offers a more principled path: one in which intentional-level causation is constitutive of the model rather than merely emergent from statistical patterns over text descriptions of intentional behaviour.
This connects directly to the diagnoses of LLM failure surveyed in Section 3. The 25–40% accuracy drop on counterfactual tasks, concentrated at the abduction step, is precisely the kind of failure one would predict for systems lacking explicit intentional-level causal structure. To answer ‘what would have happened if the agent had believed X instead of Y?’ requires that the agent’s beliefs be represented as causally efficacious states whose modification has determinate consequences for subsequent reasoning. This is exactly what BDI architectures provide and what unaugmented LLMs do not.
The deepest methodological challenge for this programme is validation: how to demonstrate that the intentional-level causal structure within a simulation genuinely mirrors that in the real social world, rather than merely producing surface-level outcome similarities. This is a problem requiring both empirical social science — comparing simulation outputs against natural and quasi-experimental data — and careful philosophical analysis of what it means for a model to capture rather than merely predict social causation. It is the social-domain analogue of the validation problem for embodied robotics: confirming that the simulated agent’s grasp of physical causality genuinely transfers to action in the unstructured world. In both cases, the answer is not provided by the simulation alone but by the disciplined dialogue between simulation, empirical observation, and theoretical analysis that constitutes mature scientific practice in the relevant domain.
Agent-Based Simulation as Social Embodiment
With these philosophical foundations in place, the parallel between physical embodiment and agent-based social simulation can now be stated more precisely. In both cases, the claim is that genuine causal understanding of a domain requires the system to engage with that domain through structured interaction rather than through passive observation of descriptions of interactions. For physical causality, the relevant medium is sensorimotor coupling with a physical environment — robots that push, pull, and observe. For social causality, the relevant medium is agent-based simulation of intentionally-constituted agents in a structured social environment — populations that form beliefs, commit to plans, coordinate intentions, signal, bargain, and revise.
In both cases, the interactive process generates a history of structured action-outcome pairs from which causal structure can be extracted. In both cases, the causal concepts that emerge — spatial relations for physical agents, social norms and intentional dynamics for simulated social agents — are grounded in the interaction history rather than stipulated from outside. And in both cases, the resulting causal knowledge supports interventional and counterfactual reasoning that purely statistical analysis of observational data cannot provide. The crucial addition the present section makes to the embodied cognition argument is that social embodiment requires intentional-level architecture: it is not enough to have agents that interact; the agents must interact as planning, believing, intending entities, because that is what social causation is constituted by.
This framing also clarifies a relationship between Skyrms’s social philosophy and Farmer’s economics that is rarely made explicit. Both are engaged in the same epistemic project: the discovery of causal structure in complex social systems through the controlled manipulation of simulated environments populated by intentional agents. Skyrms focuses on the evolutionary origins of norms, cooperation, and communication; Farmer focuses on the emergent dynamics of financial and macroeconomic systems. Both use agent-based simulation as the medium through which social causal structure becomes legible. And both arrive at the same methodological conclusion that Pearl reaches from the philosophy of causation: passive observation of social correlations, however richly data-supported, cannot substitute for the active exploration of causal mechanisms through intervention upon agents whose intentional-level dynamics generate the macro-scale phenomena.
Implications for AI Systems in the Social Domain
The implication for AI systems designed to reason about social and economic causality is direct. Such systems will not acquire genuine causal understanding of social phenomena by training on text descriptions of social events, policy interventions, and economic outcomes — however vast the corpus. They will acquire borrowed causal knowledge: the accumulated causal models of human social scientists encoded in language. Pearl’s concession about text-trained LLMs applies here with full force, but with an important addition derived from the intentional-stance analysis: the failure is not merely that LLMs absorb causal claims without causal architecture, but specifically that they absorb intentional-level descriptions without instantiating intentional-level structure. The causal mechanisms of social life pass through agents who plan, commit, and revise; an AI system without these structures cannot reliably reason about phenomena whose causation is constituted by them.
The path toward genuine AI causal understanding in the social domain runs through agent-based simulation augmented by explicit BDI or comparable intentional architectures: through systems that interact within simulated social environments not merely as statistical learners but as planning agents whose own intentional structure mirrors that of the agents they aim to understand. This is not the same as the physical embodiment of Clark, Kirsh, and Spranger — the relevant environment is social and institutional, not physical and spatial, and the relevant grounding is intentional rather than sensorimotor — but it is precisely analogous in its epistemic structure. The frontier research programme is the integration of the formal apparatus of Pearl’s causal hierarchy with the generative power of large-scale agent-based simulation populated by intentionally-constituted agents, creating systems that can not only describe social causal structure but actively explore, intervene upon, and reason counterfactually about it from within the intentional standpoint that constitutes the structure itself.
From a Luhmannian perspective, there is an additional reflexive dimension to this programme that deserves acknowledgement. Agent-based simulations of social systems are themselves second-order observations — models of how social systems operate, constructed using the distinctions of particular scientific communities. The causal knowledge they generate is system-relative in Luhmann’s sense: it is the causal structure of the social world as seen through the observational schema of complexity economics or evolutionary game theory or BDI cognitive architecture. This does not make it arbitrary or merely subjective; it makes it revisable through the evolution of the scientific system’s internal standards. The productive tension between the formal rigour of Pearl’s causal framework, the intentional-stance philosophy of Bratman and Dennett, and the generative empiricism of Skyrms’s, Farmer’s, and Epstein’s simulations is precisely the kind of internal scientific debate that mature scientific practice requires.
Causality as a Selective Schema
Niklas Luhmann’s earliest and most directly relevant treatment of causality appears in his 1962 article ‘Funktion und Kausalität,’ published in the Kölner Zeitschrift für Soziologie und Sozialpsychologie, 14, pp. 617–644, and reprinted in Soziologische Aufklärung 1 (Springer Verlag, 1970). This foundational essay establishes the key move that Luhmann will sustain throughout his mature systems theory: causality is not a feature of the world to be discovered or a tool to be used, but a functional schema that systems employ to reduce complexity. The world, in Luhmann’s framework, is characterised by overwhelming causal complexity: virtually everything is connected to virtually everything else. No observing system can process this totality. Causal attribution is one of the primary mechanisms through which systems select from this complexity, singling out a small number of antecedents and effects as relevant. This move is sustained and developed across Luhmann’s major works, including Social Systems (1984/1995), The Science of Society (1990/1992), and Risk: A Sociological Theory (1991/1993). Causal attribution, for Luhmann, is not pragmatic in van Fraassen’s sense — not simply a matter of different questioners finding different causes relevant to their interests. It is a constitutive operation of the observing system: it is how the system constructs its environment as an environment.
Second-Order Observation
Central to Luhmann’s account is the concept of second-order observation: observing not the world directly but observing how other systems observe the world. Both first-order and second-order observations are always communicative operations of semi-stable, self-reproducing (autopoietic), evolving social systems with their system-specific communication structures which include specific, highly selective discriminations. Every observation employs a distinction — a schema that divides the world into a marked and an unmarked side — and this distinction cannot itself be observed from within the observation that uses it. The blind spot of every observation is the distinction it relies upon. Applied to Pearl’s framework: a causal graph is an observation made using specific distinctions — cause versus background condition, relevant versus irrelevant variable, direct versus indirect effect. Pearl presents these distinctions as recoverable from the causal structure of the world. Luhmann would say they are produced by the observing system, shaped by its history, disciplinary conventions, and research programmes. This does not make them arbitrary — it makes them system-relative, which is different.
Luhmann’s Challenge to Each Interlocutor:
Against Pearl: Every SCM is a second-order observation — a model of causal structure as seen from the standpoint of a particular scientific community using particular distinctions. Different communities, asking different questions with different schemas, would produce different graphs. There is no view from nowhere that would validate one graph as uniquely correct.
Against Sutton: Every learning architecture embeds distinctions — what counts as a feature, a prediction, an error. These constitute the system’s horizon of observability. Sutton confuses a very powerful second-order observation for a first-order contact with the world as it is.
Against van Fraassen: The distinction between observable and unobservable is itself produced by the observing system’s internal operations, not found in the world independently. This dissolves van Fraassen’s distinction between empirically adequate and literally true theories more radically than van Fraassen himself intends.
Against the embodied cognition tradition: While Luhmann would find the embodied cognition programme more sympathetic than the others — since it foregrounds the system-environment coupling that produces cognitive operations — he would note that even embodied observation is bounded by the distinctions the organic and psychic system brings to its environmental interactions. Spranger’s robots evolve spatial language through embodied interaction, but the conceptual schemas they develop are still selected from the space of possibilities opened by their architecture. There is no schema-free embodied observation any more than there is schema-free statistical learning.
Against the agent-based simulation tradition: Skyrms’s and Farmer’s simulations are themselves second-order observations — models of social and economic systems constructed using the distinctions of evolutionary game theory and complexity economics. The causal knowledge they generate is system-relative: it is the causal structure of the social world as seen through a particular observational schema. This does not undermine the value of such simulations — they remain among the most productive available methods for accessing social causality — but it situates them within the same reflexive condition that applies to all observation.
The Acyclicity Constraint: Pearl’s Ontological Mismatch with Self-Reproducing Social Structures
From Luhmann’s perspective we can launch another challenge to Pearl’s conception and formalisation of causality which must be remedied with respect to the social domain. A feature of Pearl’s structural causal models that has received insufficient critical attention is its most elementary formal requirement: the directed acyclic graph. It is a constitutive constraint on the ontology that SCMs can represent. A directed acyclic graph forbids cycles: no variable may be its own ancestor. This means that every causal pathway terminates; there are no self-sustaining loops, no structures that reproduce themselves through their own operations. The do-calculus, the identification theory, and the three-step counterfactual procedure — abduction, intervention, prediction — all presuppose this acyclicity. Without it, the mathematical apparatus that underwrites Pearl’s causal revolution breaks down: interventions cannot be cleanly isolated, parent sets become ill-defined, and the truncated factorisation that gives the do-operator its content ceases to hold.The question this section poses is: what kind of causal world does the acyclicity requirement presuppose, and what does it exclude?
An Ontology of Events, Not Structures
Pearl’s SCMs are built for a world of discrete variables with determinate values — drug doses, policy switches, temperatures, test outcomes. Each variable takes a value that is determined by its parents and an independent noise term. The resulting ontology is punctuated and event-centred: causes are settings of variables; effects are the values that result. This maps naturally onto the experimental ideal of the randomised controlled trial, where the experimenter sets one variable and measures another. It maps less naturally onto the entities that populate the social and institutional world. Consider a market, a legal system, a social norm, or a scientific research programme. These are not events that occur and terminate; they are structures that persist through time by continuously reproducing the conditions of their own operation. A market does not cause prices in the way a drug causes recovery. A market is the ongoing process through which prices are generated, and the generation of prices is itself one of the operations through which the market reproduces itself as a market. The causal loop is not incidental; it is constitutive.
The standard Pearlian response is temporal unrolling: represent X at time t as causing Y at time t+1, which causes X at time t+2, transforming the apparent cycle into an acyclic sequence indexed by time steps. Dynamic Bayesian networks formalise precisely this manoeuvre. And for many purposes it works: if one is interested in the effect of a central bank’s interest-rate decision at t on inflation at t+1, the unrolled representation is adequate. But the manoeuvre conceals a deeper problem. It describes what happens inside the loop while remaining silent on why the loop exists at all. The temporal unrolling tells us the sequence of states; it does not tell us why that sequence is self-sustaining rather than dissipative, why the structure persists rather than decays. The persistence of the structure — its autopoietic self-reproduction, in Luhmann’s terms — is precisely the phenomenon that requires explanation, and it is precisely the phenomenon that the acyclicity requirement renders formally invisible.
Three Categories of Causal Entity That Resist Acyclicity
The mismatch between DAG ontology and social reality can be made precise by identifying three categories of causal entity that are poorly served by the acyclicity constraint.
First, self-reproducing institutions. Luhmann’s autopoietic systems — law, economy, science, politics, education, art — are defined by the property that they produce and reproduce the elements of which they consist through a network of those very elements. The legal system produces legal communications (judgements, statutes, contracts) through legal communications; the scientific system produces truth-claims through truth-claims. The circularity is not a defect to be linearised; it is the mode of existence of the system. An SCM that represents the legal system must either break this circle — misrepresenting the phenomenon — or violate acyclicity.
Second, structural attractors. In complex systems theory, an attractor is a state or set of states toward which a dynamical system tends to evolve and to which it returns after perturbation. Poverty traps, institutional equilibria, and path-dependent technological standards are attractors in this sense: they actively maintain themselves against perturbation through feedback mechanisms that are constitutively circular. Modelling such attractors as acyclic causal chains captures the local dynamics of any given perturbation-and-return but misses the global feature that defines the attractor: its self-maintenance. The causal power of an attractor consists precisely in its capacity to close the loop, to ensure that departures from the pattern are corrected by forces that the pattern itself generates. This circularity is the explanation, and it cannot be represented in a DAG without remainder.
Third, constitutive relations. In many social phenomena, A does not merely cause B; B is partly constitutive of what A is. A market does not merely cause the price mechanism to operate; the price mechanism is partly constitutive of what it means to be a market. Bratman’s planning intentions do not merely cause coordinated action; the possibility of coordinated action is partly constitutive of what it means to have a planning intention. In these cases, the causal arrow presupposes the existence of the very entity it purports to explain, creating a conceptual circularity that is not temporal but ontological. No amount of temporal unrolling can linearise a constitutive relation, because the relation is not between events at different times but between an entity and its own conditions of possibility.
Consequences for Causal AI in the Social Domain
The implications for the broader argument of this paper are substantial. The acyclicity constraint is not merely a technical limitation that future extensions of Pearl’s formalism might overcome. It reveals that SCMs are ontologically mismatched with the entities that Luhmann identifies as the primary constituents of modern social reality: operationally closed, self-reproducing functional systems. Pearl’s formalism is a theory of causes among events; modern social reality is constituted by structures that persist through self-referential operations. These are different kinds of entity, and they require different formal apparatus.
This does not render Pearl’s framework useless for social science. Within bounded contexts — where one can reasonably treat institutional structure as exogenous background and ask what happens when a specific variable is manipulated — SCMs remain powerful tools, and the do-calculus provides rigorous identification conditions. The point is rather that SCMs operate within a social-structural context that they cannot themselves represent. They are tools for asking causal questions about events against a background of structures, not tools for asking causal questions about the structures themselves.
This reinforces the case developed elsewhere in this paper for agent-based simulation as the appropriate methodology for social causal discovery. The agent-based models of Skyrms, Farmer, and Epstein are not bound by the acyclicity constraint. They can and do model feedback loops, self-reproducing structures, and emergent institutional patterns. A Skyrms model of norm evolution represents precisely the circular causal process by which a norm’s existence generates the behavioural patterns that sustain the norm. A Farmer model of financial markets represents the self-reinforcing dynamics through which market structures create the trading patterns that reproduce those structures. These models operate with a richer ontology than Pearl’s SCMs — one that admits structures, loops, and constitutive relations alongside events and interventions. The cost is that they sacrifice the clean identification guarantees of the do-calculus; the gain is that they can represent the kinds of causal entity that actually constitute the social world.
Bongers et al. (2021) have proposed a formal extension of Pearl’s framework to cyclic SCMs, developing equilibrium-based semantics that can handle simultaneous feedback. Their work demonstrates that interventional and counterfactual reasoning can in principle be defined for cyclic causal structures, though at the cost of requiring additional assumptions about equilibrium selection and structural stability. This is a productive direction, but it also confirms the depth of the problem: the extension requires fundamentally different mathematical machinery from that of the standard acyclic framework, precisely because self-referential causal structures are a qualitatively different kind of entity from the event-chains that DAGs were designed to represent. For the social domain, where Luhmann’s autopoietic systems and Farmer’s self-reinforcing market dynamics are the norm rather than the exception, the cyclic extension is not a peripheral refinement but a necessary precondition for representational adequacy.
Iwasaki and Simon (1994), working within Herbert Simon’s earlier tradition of causal ordering, addressed the problem of simultaneous equations in economic models — systems in which equilibrium conditions create mutual dependence among variables — by deriving causal orderings from the mathematical structure of the equations themselves. Their approach demonstrates that causal direction can sometimes be recovered even in simultaneous systems, but it also makes visible the gap between the simultaneity that characterises institutional self-reproduction and the sequential event-ontology that standard SCMs presuppose. The social world is full of simultaneous mutual constitution: supply constitutes demand as demand constitutes supply; legal precedent constitutes legal authority as legal authority constitutes the bindingness of precedent. These are not merely simultaneous equations awaiting causal ordering; they are structural circularities whose resolution into linear causal paths requires precisely the kind of observational standpoint — a system-specific, distinction-dependent schema of attribution — that Luhmann’s framework identifies as the condition of all causal observation.
The acyclicity constraint thus deepens the Luhmannian critique of Pearl developed above. There, the argument concerned the observer-relativity of causal schemas: every SCM is a second-order observation made using system-specific distinctions. Here, the argument concerns the ontological scope of SCMs: their formal structure excludes the self-reproducing entities that constitute the primary fabric of modern social reality. Together, these two objections — one epistemological, one ontological — establish that Pearl’s causal revolution, for all its power in the domain of event-causation, operates with a social ontology that is too thin for the demands of genuine social-scientific causal reasoning. The path forward requires not the abandonment of Pearl’s insights but their integration with formal and computational tools — agent-based simulation, cyclic causal models, systems-theoretic analysis — that can accommodate the richer ontology that social causality demands.
The Social Stabilisation of Causal Attribution in AI
Luhmann’s most distinctive and practically important contribution concerns the social processes through which causal attributions become authoritative within functional subsystems. Pearl’s framework gives the logic of causal inference, assuming a causal model. Van Fraassen urges epistemological humility about that model. Sutton says derive it from data. The embodied tradition says ground it in physical interaction. The agent-based simulation tradition says generate it through structured social interaction. None of these positions, however, asks: who decides which causal schema is operative in a given context, through what social process, with what consequences for the system’s self-reproduction?
The question of whether an AI system ‘genuinely’ performs causal reasoning — whether it operates at Rung 2 or Rung 3 of Pearl’s ladder, or merely appears to — will not be settled by a technical demonstration alone. It will be shaped by the parallel evolution of multiple functional subsystems: the scientific community’s benchmarking practices, the AI industry’s performance claims and competitive deployment, and the institutional standards that emerge from the interaction between research programmes and the domains where they are tested. Each of these subsystems will stabilise its own criteria for what counts as causal understanding in AI, and these criteria need not converge — though productive friction between them is the very mechanism by which scientific progress occurs in functionally differentiated modernity.
The Luhmann-Habermas Debate and AI Governance
The debate between Jürgen Habermas and Niklas Luhmann — most directly staged in their 1971 exchange Theorie der Gesellschaft oder Sozialtechnologie — is one of the most consequential confrontations in twentieth-century social theory, and it bears directly on how AI governance is conceptualised. Habermas contends that social action requires consensual decision-making grounded in communicative rationality. Social processes can in principle be evaluated and redeemed through discourse oriented toward mutual understanding, provided the conditions of ideal speech are approximated — including equality of participation, absence of coercion, and the force of the better argument. Luhmann, by contrast, contends that social activities are too complex for consensual bartering; they require the impersonal systemic regulation that emerges from functional differentiation. For Habermas, the political system — when functioning properly under conditions of deliberative democracy — can legitimately adjudicate between competing claims in science, economy, and technology, because communicative rationality provides a framework that transcends any particular functional system. For Luhmann, this is a category error: the political system operates according to its own code and cannot serve as the meta-system adjudicating the truth claims of science, the efficiency claims of the economy, or the innovation claims of technology without distorting the functional differentiation on which complex society depends.
In the context of AI and causal intelligence, the Habermas-Luhmann debate becomes urgently practical. The question ‘does this AI system genuinely reason causally?’ is simultaneously a scientific question (answered through benchmarking, formal analysis, and experimental study), an economic question (answered through market performance and cost-effectiveness), a legal question (answered through evidentiary standards and liability attribution), and a political question (answered through regulatory frameworks and public policy). Each of these functional systems processes the question according to its own code, and there is no guarantee — and, on Luhmann’s account, no realistic expectation — that their answers will converge. The penetration of moral or political values into scientific and technological domains that should be governed by the codes of science (true/false), economy (profit/loss), and engineering (functional/dysfunctional) constitutes a form of system colonisation that distorts the functional logic of all involved. When the political system expands into the epistemic domain of science and technology, it does not improve on the outcomes those systems would have produced through their own evolution; it substitutes its own operative logic — risk-aversion calibrated to electoral cycles, legibility calibrated to parliamentary scrutiny, accountability calibrated to litigation — for the logic of scientific inquiry and technological innovation. The result is not safer or more genuinely intelligent AI; it is AI optimised for political legibility rather than epistemic adequacy. This does not imply that political discourse - taking the results of science, technology and markets as premises - has no role in AI governance.
The debate reconstructed and extended in this paper is not resolvable within any single disciplinary framework. It requires philosophy of science (van Fraassen), the theory of computation and learning (Sutton), causal logic (Pearl), embodied cognitive science and robotics (Clark, Kirsh, Spranger), agent-based social simulation (Bratman, Epstein, Skyrms, Farmer), and social systems theory (Luhmann) to be adequately mapped.
Pearl’s core claim — that associational data cannot yield causal conclusions without causal assumptions — remains logically compelling and is substantially supported by recent empirical evidence. The structured pattern of failure at Rung 3 counterfactual tasks in state-of-the-art reasoning models, concentrated at the abduction step, is consistent with Pearl’s theoretical predictions. Sutton’s Bitter Lesson poses a genuine empirical challenge but is an inductive argument confronting a deductive one: the history of AI cannot prove that a mathematical boundary can be crossed.
Van Fraassen usefully strips the debate of metaphysical excess without eliminating its practical force. The embodied cognition tradition of Clark, Kirsh, and Spranger establishes that physical causal understanding is constituted through sensorimotor interaction with the world. But this insight has a crucial domain boundary: the social sciences require a different medium. Agent-based simulation — as developed via the BDI architectures of Rao and Georgeff, by Skyrms for social norms and cooperative dynamics, and by Farmer for economic causality, and grounded philosophically in Bratman’s planning theory of intention, Dennett’s intentional stance, and Epstein’s generative social science — is the social sciences’ functional analogue of embodied robotics. Both programmes share the same epistemic structure: causal knowledge is generated through structured interaction with a generative environment rather than through passive observation of descriptions of interactions. The crucial addition the present paper makes is that social embodiment requires intentional-level architecture: it is not enough to have agents that interact; the agents must interact as planning, believing, intending entities, because that is what social causation is constituted by. This is precisely the distinction Pearl’s Ladder of Causation formalises, and its operationalisation differs appropriately across physical and social domains.
Luhmann’s contribution is the most unsettling for all parties. He does not ask which framework is correct — he asks through what social processes correctness is determined and stabilised differently within society’s differentiated functional subsystems. His answer suggests that the resolution of the technical debate will be, in significant part, a social and institutional achievement, mediated through the internal evolution of the scientific system, the competitive dynamics of technology markets, and the institutionalised distinctions through which different functional systems make causal claims legible to themselves.
The most defensible reading of the present situation is therefore neither pure Pearlian pessimism nor unbounded scaling optimism. A pure LLM, trained only to predict text, is unlikely to be sufficient for robust artificial general intelligence understood as autonomous reasoning under novel interventions. But LLMs are not for that reason irrelevant: they are powerful discourse engines and knowledge interfaces that can serve as components within hybrid architectures that supply what they lack. The most promising path forward is the principled integration of LLM capabilities with structured causal models in the Pearlian tradition, with embodied robotic systems for physical causality in the tradition of Clark, Kirsh, and Spranger, and with agent-based simulation environments for social and economic causality in the tradition of Skyrms and Farmer. Pearl’s challenge stands against AGI by correlation-crunching alone — but not against AGI architectures that embed LLMs within broader causal-agentic systems, drawing on each tradition for the domain to which it is appropriate.
The argument advanced thus far still concedes too much to the framing within which it operates. The very notion of artificial general intelligence as a presumptively unitary, all-purpose, peak intellectual competence — a single system that approximates or exceeds human capability across every domain of cognition simultaneously — is naive and must be rejected. The reasons for this rejection are not technical but sociological in the deepest sense: they follow from a Luhmannian sociology that, taken seriously, functions as a practically relevant first philosophy for any inquiry into the nature of intelligence in modern conditions.
Luhmann’s mature theory posits that advanced modern society is functionally differentiated — organised not as a hierarchy with a unified apex, nor as a totality with a common rationality, but as a constellation of operationally closed, autopoietic social systems each recognisable by, and operating according to, its own binary code: the legal system via legal/illegal, the economic system via profit/loss, the scientific system via true/false (probable/improbable), the technology-engineering system via functional/dysfunctional, the medical system via healthy/diseased, the political system via government/opposition, the education system on pass/fail, the art system on novel/conventional etc. Each system reproduces itself through communications structured by its own code; each generates its own time, its own memory, its own forms of attention; and crucially, none can substitute for the others, nor can any single observation-position synthesise them into a unified whole. This is the condition Luhmann calls polycontexturality: the irreducible multiplicity of incommensurable observational frames that constitutes modern society as such.
Polycontexturality is not a regrettable fragmentation that better integration could overcome. It is the structural condition that makes complex modern society possible. The tremendous productive capacity of advanced societies — their wealth, their knowledge, their technological achievements, their differentiated forms of life — depends on the parallel operation of distinct functional systems that need not, and indeed cannot, be reduced to a single rationality. The legal system’s processing of legality cannot be replaced by the economic system’s processing of profitability; the scientific system’s processing of truth cannot be replaced by the political system’s processing of legitimacy. Each system co-evolves with its own system of distinctions, information processing, discourse-practices, canonical texts, training pathways, modes of dispute resolution, its own highly selective relations to its thereby constituted specific environment. The result is a society of incommensurable competences, each indispensable, none translatable without remainder into another.
This has direct and devastating implications for the dominant conception of AGI. The standard framing imagines a single artificial system that can do everything a competent human can do — reason scientifically, argue legally, design architecturally, treat medically, govern politically, create artistically, teach pedagogically — and do so at or above the highest human level across all domains simultaneously. But this is precisely the conception of intelligence that polycontexturality forbids. There is no unitary cognitive position from which all functional domains can simultaneously be mastered, because each domain requires its own exclusive discourse-practice, fluency with its own code, sensitivity to its own historically evolved modes of distinction-making. Each of these competencies implies simultaneous blind spots. These competences are not reducible to a common substrate of ‘general intelligence’ that could be scaled up to encompass them all. They are constituted by the differentiated systems within which they operate.
This is also why the apparent successes of current LLMs across multiple domains are systematically misleading as evidence for AGI. LLMs trained on text from all functional systems can produce plausible-sounding outputs in legal, medical, scientific, economic, artistic, and political registers — but plausible-sounding outputs are not the same as competent participation in the relevant systems. The legal system does not accept an LLM’s output as a legal judgement; the medical system does not accept it as a clinical decision; the scientific system does not accept it as a contribution to peer-reviewed knowledge. Each system has its own criteria for what counts as a competent communication within it, and these criteria are not satisfied by surface fluency in the system’s vocabulary. The borrowed causal knowledge diagnosis advanced earlier in this paper is itself a special case of this more general phenomenon: LLMs absorb the linguistic residue of competent participation across functional systems without thereby becoming competent participants in any of them.
Luhmann’s analysis of modernity is that is brought to bear here on the question of AGI is not pointing to a logical, cognitive or metaphysical impossibility. It is an empirically grounded theoretical offering from within the science of sociology pointing to the practically highly relevant incommensurability of different discursive-practical function systems understood as systems of communication. These domains are institutionalized via distinct, separate institutions. More-over, AGI projects must take into account that function systems are self-referentially enclosed, each with its specific selectivity via systems of distinctions, evaluative codes, specific media and criteria of success related to the different specific functional contributions these functions systems make to our highly differentiated societies. This is where the crux of incommensurability lies. It is deeper than merely different scientific paradigms within a scientific discipline, and deeper than in the case of different scientific disciplines. Here there can be paradigm integration, and interdisciplinary research. But the very different task orientations of different function systems implies a deeper incommensurability, e.g. the incommensurability between a scholarly legal discourse engaging jurists trying to appraise the status of a certain causal attribution as a matter of law and the appraisal of such an attribution as a matter of e.g. the science of psychology. The point is here that the universe of scholarly legal discourse which includes all common law case law with an unbounded combinatorial cross-referencing potential is so vast that co-evolution of vast LLMs with this universe of discourse is the only practical way forward. The same applies to the vast universe of science, of engineering, of cultural criticism, of politics. To assimilate these domains by referring to them as “cognition” or “intelligence” is misleading. The functional system differentiation, driven to point of incommensurability, is a mechanism of complexity reduction which allows society to evolve further where it otherwise would come to a halt at an otherwise insurmountable complexity barrier. Are we really expecting a single mega LLM to deliver a unitary AGI where advanced world society itself at the aggregate level/dynamic has moved forward via radical differentiation where communication (bounded within function systems) is substituted by co-evolution between function systems?
The proper conception of advanced artificial intelligence that emerges from this analysis is therefore not unitary AGI but a constellation of specialised artificial competences, each developed in deep co-evolution with the functional system it serves. Causal reasoning AI for the sciences will develop through integration with the formal apparatus of Pearl’s causal hierarchy, embodied robotic experimentation, and the institutional practices of scientific peer review. Legal AI will develop through immersion in the codes, precedents, and adversarial procedures of legal systems. Economic modelling AI will develop through agent-based simulation calibrated against empirical economic time series and refined through the disputes of professional economists. Architectural and design AI will develop through engagement with the discourse, historical archive, and material practice of design culture. Each of these will require its own distinctive architecture — its own combination of LLM components, structured causal models, embodied or simulated agents, and domain-specific training environments — appropriate to the functional system within which it operates. There is no master architecture that subsumes them all.
This is not a deflationary conclusion. Far from diminishing the prospects for transformative artificial intelligence, the polycontextural framework clarifies what such intelligence will actually look like and how it will actually be developed. It will be developed, as advanced human intelligence has always been developed, through the parallel evolution of differentiated competences within functional systems that co-evolve with their respective AI augmentations. The frontier is not the construction of a single peak intelligence but the elaboration of an artificial counterpart to the polycontextural society itself — a constellation of artificial systems each calibrated to a functional domain, each communicating in that domain’s code, each subject to that domain’s standards of competence, and each contributing to the parallel reproduction of differentiated modern society.
Luhmann’s sociology, on this reading, functions as a practically relevant first philosophy precisely because it correctly identifies the basic structural condition that all theorising about intelligence in modern conditions must respect. Pearl’s causal revolution, the embodied cognition tradition, agent-based simulation, and even the empirical findings about LLM limitations all gain their proper significance when placed within this polycontextural frame. They are not competing routes to a single peak; they are complementary contributions to the differentiated artificial-cognitive infrastructure that advanced functionally differentiated society requires. The naive conception of unitary AGI — peak general competence across all domains in a single system — is implausible, presupposing a unification of social rationality that has not existed in advanced societies for at least two centuries and whose absence is the condition rather than the obstacle of their tremendous productive achievements.
Note: This paper was researched and drafted with the aid of Claude Sonnet 4.6, ChatGTP 5.4 and Claude Opus 4.6. & 4.7
Axtell, R. and Farmer, J. D. (2022). Agent-Based Modeling in Economics and Finance: Past, Present, and Future. INET Oxford Working Paper No. 2022-10. Institute for New Economic Thinking, Oxford.
Bicchieri, C. (2006). The Grammar of Society: The Nature and Dynamics of Social Norms. Cambridge University Press, Cambridge.
Bongers, S., Forré, P., Peters, J., and Mooij, J. M. (2021). Foundations of Structural Causal Models with Cycles and Latent Variables. Annals of Statistics, 49(5), 2885–2915.
Bratman, M. E. (1987). Intention, Plans, and Practical Reason. Harvard University Press, Cambridge, MA.
Brawer, J., Qin, M., and Scassellati, B. (2020). Causal Reasoning about Robots: How Humanoid Robots Build Structural Causal Models from Observation and Self-Supervised Trials. Proceedings of the 2020 ACM/IEEE International Conference on Human-Robot Interaction.
Castri, L., et al. (2022). Causal Discovery to Predict Human Spatial Interactions in Social Robotics Contexts. Proceedings of IEEE RO-MAN.
Castri, L., et al. (2023). F-PCMCI: Advancing Causal Discovery through Feature Selection and PCMCI. arXiv preprint arXiv:2310.09615.
Clark, A. (1997). Being There: Putting Brain, Body and World Together Again. MIT Press, Cambridge.
Clark, A. (2008). Supersizing the Mind: Embodiment, Action, and Cognitive Extension. Oxford University Press, Oxford.
Clark, A. (2016). Surfing Uncertainty: Prediction, Action, and the Embodied Mind. Oxford University Press, Oxford.
Clark, A. and Chalmers, D. (1998). The Extended Mind. Analysis, 58(1), 7–19.
Crawford, S. E. S. and Ostrom, E. (1995). A Grammar of Institutions. American Political Science Review, 89(3), 582–600.
Dennett, D. C. (1987). The Intentional Stance. MIT Press, Cambridge, MA.
Epstein, J. M. (2006). Generative Social Science: Studies in Agent-Based Computational Modeling. Princeton University Press, Princeton.
Epstein, J. M. and Axtell, R. (1996). Growing Artificial Societies: Social Science from the Bottom Up. Brookings Institution Press / MIT Press, Washington and Cambridge, MA.
Farmer, J. D. (2024). Making Sense of Chaos: A Better Economics for a Better World. Allen Lane, London.
Farmer, J. D. and Foley, D. (2009). The economy needs agent-based modelling. Nature, 460(7256), 685–686. DOI: 10.1038/460685a.
Griot, M., Hemptinne, C., Vanderdonckt, J., and Yuksel, D. (2025). Large language models lack essential metacognition for reliable medical reasoning. Nature Communications, 16(1), 642.
Habermas, J. (1984/1987). The Theory of Communicative Action. 2 vols. Beacon Press, Boston.
Habermas, J. and Luhmann, N. (1971). Theorie der Gesellschaft oder Sozialtechnologie: Was leistet die Systemforschung? Suhrkamp, Frankfurt.
IROS Causality for Robotics Workshop (2023). Workshop proceedings summary. IEEE/RSJ International Conference on Intelligent Robots and Systems, Detroit, October 2023. Available at: https://sites.google.com/view/iros23-causal-robots
Iwasaki, Y. and Simon, H. A. (1994). Causality and Model Abstraction. Artificial Intelligence, 67(1), 143–194.
Jin, Z., et al. (2024). Can Large Language Models Infer Causation from Correlation? (Corr2Cause Benchmark) and CausalBench: A Comprehensive Benchmark for Causal Learning Capability of Large Language Models. arXiv preprint arXiv:2306.05836; arXiv:2404.06349.
Kıcıman, E., Ness, R., Sharma, A., and Tan, C. (2023). Causal Reasoning and Large Language Models: Opening a New Frontier for Causality. arXiv preprint arXiv:2305.00050.
Kirsh, D. and Maglio, P. (1995). On Distinguishing Epistemic from Pragmatic Actions. Cognitive Science, 18(4), 513–549.
Kirsh, D. (2013). Embodied Cognition and the Magical Future of Interaction Design. ACM Transactions on Computer-Human Interaction, 20(1), Article 3.
Liu, H., et al. (2024). Unveiling Causal Reasoning in Large Language Models: Reality or Mirage? Advances in Neural Information Processing Systems (NeurIPS) 38.
Luhmann, N. (1962). Funktion und Kausalität. Kölner Zeitschrift für Soziologie und Sozialpsychologie, 14, 617–644. Reprinted in: Luhmann, N. (1970). Soziologische Aufklärung 1. Springer Verlag, Opladen.
Luhmann, N. (1984/1995). Social Systems. Stanford University Press, Stanford.
Luhmann, N. (1990/1992). The Science of Society. Stanford University Press, Stanford.
Luhmann, N. (1991/1993). Risk: A Sociological Theory. Aldine de Gruyter, New York.
Luhmann, N. (1997/2012). Theory of Society (originally Die Gesellschaft der Gesellschaft). 2 vols. Stanford University Press, Stanford.
OpenCausalAI Lab (2024). ADAM: An Embodied Causal Agent in Open-World Environments. arXiv preprint arXiv:2410.22194.
Pearl, J. (1995). Causal diagrams for empirical research. Biometrika, 82(4), 669–688.
Pearl, J. (2000/2009). Causality: Models, Reasoning, and Inference. 2nd edition. Cambridge University Press, Cambridge.
Pearl, J. (2009). Causal inference in statistics: An overview. Statistics Surveys, 3, 96–146.
Pearl, J. and Mackenzie, D. (2018). The Book of Why: The New Science of Cause and Effect. Basic Books, New York.
Pearl, J. and Mackenzie, D. (2023). Judea Pearl, AI, and Causality: What Role Do Statisticians Play? Amstat News / ASA Magazine, September 2023.
Rao, A. S. and Georgeff, M. P. (1991). Modeling Rational Agents within a BDI-Architecture. In J. Allen, R. Fikes, and E. Sandewall (eds.), Proceedings of the 2nd International Conference on Principles of Knowledge Representation and Reasoning (KR’91), Morgan Kaufmann, San Mateo, pp. 473–484.
Schelling, T. C. (1978). Micromotives and Macrobehavior. W. W. Norton, New York.
Schumacher, P. (2025). Discourse Capitalism. [Publisher details forthcoming].
Skyrms, B. (1996). Evolution of the Social Contract. Cambridge University Press, Cambridge.
Skyrms, B. (2004). The Stag Hunt and the Evolution of Social Structure. Cambridge University Press, Cambridge.
Skyrms, B. (2010). Signals: Evolution, Learning, and Information. Oxford University Press, Oxford.
Skyrms, B. (2014). Social Dynamics. Oxford University Press, Oxford.
Spranger, M. (2016). The Evolution of Grounded Spatial Language. Language Science Press, Berlin. DOI: 10.26530/OAPEN_611695.
Sutton, R. (2019). The Bitter Lesson. Incomplete Ideas (personal blog), 13 March 2019. Available at: http://www.incompleteideas.net/IncIdeas/BitterLesson.html
van Fraassen, B. C. (1980). The Scientific Image. Oxford University Press, Oxford.
van Fraassen, B. C. (2002). The Empirical Stance. Yale University Press, New Haven.
Vashishtha, N., et al. (2025). Executable Counterfactuals: Improving LLMs’ Causal Reasoning through Code. arXiv preprint arXiv:2510.01539.
Willig, M., et al. (2023). Causal Reasoning in Large Language Models: A Comprehensive Survey. arXiv preprint arXiv:2306.05836.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.