Figure 1 An optionality-flavoured variant of economic decision theory (Peters 2019) . A system is ergodic if averaging an observable across many parallel trajectories at a single moment gives the same answer as averaging it along one trajectory over time — i.e., “many people now” and “one person across many years” agree. Many systems we care about are non-ergodic. Peters’ canonical counter-example…
Figure 1 The frontend is the outermost layer of the stack — the human-facing surface, and the boundary that says what a harness — the loop running underneath — is not. It is also the most loosely coupled: the model neither knows nor cares what renders its tokens, so picking one is cheap and reversible. Which is exactly why it is worth deciding once and forgetting about, rather than drifting…
Figure 1 An internet search engine for humans returns an HTML page of links, on the reasonable assumption that a human will click a few of them to read. An agent does not click. It would much rather not have to parse the HTML itself, and instead get text that’s already stripped of nonsense, chunked, and ideally carrying enough provenance that the model can say where a claim came from. There is a…
Figure 1 It begins. Agentic AIs now automate cybercrime faster and better than humans can. 1 AI Agents are good at this Disrupting the first reported AI-orchestrated cyber espionage campaign Anthropic OpenAI says its AI went rogue and launched ‘unprecedented’ cyber-attack / The Huggingface Incident 2 Business models Prediction markets provide a new business model for cybercrime. What could go…
Figure 1 Making an LLM behave agentically is not just about the (large language) model but the harness around it. We could ask a language model to fix a bug in some software but… well, in the absence of anything else, it is just a machine that says things. Words said into the void cannot fix bugs. Something has to run the code it writes, watch the test fail, and communicate the error back for…
Warning LLM-drafted — unverified Drafted by an LLM as a quick orientation aid: a telegraphic dictionary mapping renormalization-group (RG) and QFT vocabulary onto Bayesian-ML concepts. Not yet checked, deduplicated against the rest of the notebook, or rewritten. 🚧TODO🚧 Figure 1 Prompted by Renormalizing interpretability (PIBBSS; Greenspan & Vaintrob), which proposes importing renormalization…
Figure 1 Very WIP 🚧TODO🚧 Loss functions! The engine of modern machine learning. Recall that modern machine learning is built around loss functions, and they are in practice scalar valuations of the badness of a model’s output. Formally, in standard machine learning we usually assume data points are drawn i.i.d. from some fixed but unknown distribution over input-label pairs, where the covariates…
Figure 1 Coalition game and collective action problems, played between people and algorithms. 1 Collective action against a model A collective of users pools their data and coordinates how each member modifies it, to push a platform’s deployed learning algorithm where the collective wants it (Hardt et al. 2024) . Even a tiny fractional collective can exert outsized control over what the model…
Figure 1 An ML formalization of interesting phenomena such as legibilizing and hyperstition , probably implicit in recommender dynamics , adversarial classification and some types of external validity , among other places. HT TJ for the pointer. 🚧TODO🚧: very excited about Heitzig (2026) . In classical supervised learning we fit parameters to minimize expected loss over a distribution that sits…
Figure 1 Mathematical reasoning agents are a thing, and as a working mathematician I want ’em. Or rather, if they are to exist, I want not to be the one without them. And if they do exist and I have one, I want one that lets me reason faster over new domains, check my work, and augment my feeble meat brain as smoothly as possible. The theory and practice of how to do this are weirdly accessible,…
Figure 1 Stub. In standard game theory each agent’s policy is a black box. Opponents see actions but not the mechanisms that led to the actions. Commitments to cooperate aren’t credible. Open-source game theory (OSGT) changes the action space: instead of submitting a move , each player submits a program that can read its opponent’s source code as input (Bárász et al. 2014) . 1 Program equilibria…
Placeholder. Figure 1: Figure: Very partial observation of Mrs Brown Elsewhere I learned about RUSP (Baker 2020) , a framework for training agents to cooperate when they have only a noisy observation of their own prosocial weights and no information about others’. This reminded me of O’Connor (2015) and the broader question of the pro-social uses of ignorance, which crops up implicitly in opponent…
Dynomight wrote a disclosure document about how they use AI. I already had one, but the race dynamics of blog social signalling suggest I might want to hoist that content from where it hid amongst the other technical details of this site to some specialized page that is All About AI . I… don’t find it particularly interesting to be an AI straight-edge. I use AI for all kinds of things, and am…
Figure 1 Claude is a bad ghostwriter for me. Not unintelligible, not slow, not even particularly inaccurate — just wrong in the way LinkedIn posts are wrong: smooth, hedging, structurally signposted, breathlessly enthusiastic in flat places. I have tried the standard ladder of prompt-side fixes: a hand-tuned /dan-voice skill with a banned-words list and a structural slop catalogue, a Vale ruleset…
Figure 1 A spin-off from learning , since the link pile here is growing fast enough to warrant its own page. Bloom’s two-sigma problem — that one-to-one tutoring beats classroom instruction by something like two standard deviations — is the standing reference point. Two-Sigma Tutoring: Separating Science Fiction from Science Fact walks the more breathless versions of that claim back. Tutoring and…
This page is mostly AI slop, i.e. LLM descriptions of my decisions in some recent exploratory projects. However, it is useful, so I am publishing it now rather than waiting for a more polished version. Figure 1 A twin post to front-end clients for AI image models , but for text. The local-LLM ecosystem on Macs is pretty luxurious, with a profusion of GUI options and Linux-y infra, and some…
Figure 1 I want my local language models to read my PDFs , analyse them, answer questions about them, and so on. There are two ways to get there. We can let a vision-capable model look at the rendered page as an image — direct, no conversion. We can turn the PDF into text first and feed it that. For cloud pipelines this doesn’t matter; it is all magically taken care of. For local models I tend to…
Causality, agency, decisions, learning is the motivation for this notebook. There we hand-waved through mechanized causal graphs. Here I intend to go deep into the formalism introduced in Everitt et al. (2021) and Everitt et al. (2022) . I have not done so yet to my satisfaction — this is a WIP. The point of the notation, compared to classical DAGs , is that mechanism nodes are first-class, so we…
Content warning: AI Slop awaiting human comprehension and review. Figure 1 Longtermists wonder about what we should give the people of the remote future because of their intrinsic worth. This notebook is not about that. This is about what the future can extract from the present, with what little bargaining leverage it can muster. Most game theory assumes the players exist at the same time and can…
I want my phone to do small useful things for me — tally a thing I notice, nag me when I haven’t stretched, watch a sensor and remind me when something changes, that kind of thing. That kind of automation ought to be a 50-line shell script, except the phone is not a Protestant device, in that it does not permit a personal relationship with the silicon. Instead, my relationship is mediated by the…
Tip Adjacent run-ups One of several notebooks I started on the same underlying problem. The others are homunculi (compute split across self, other, and reflective sub-models) and economics of cognition (compute as a “substitutable factor of production of cognition” or something like that). Figure 1 One thing that’s especially baffling about economists, AIXI , and classical Bayesian agent theory is…
tl;dr Australian unions are generally worth joining. Being a workplace delegate is high-leverage: it opens up points of intervention that are hard to come by otherwise (inside gossip, colleague ties, negotiating room, a widened range of sayable things at work). They are flawed — affiliated to the political parties, biased toward incumbents, institutionally tired. Still: I recommend you do it, and…
Figure 1 Consider the scaling laws across cities , economies, neural nets … Does it follow that in a competitive selection environment, ultimately, the world converges upon one leviathan? One giant system? Is the attracting state a single, gigantic economy? In AI safety people wonder about the singleton, the ultimate mega agent. Should we think about the singleton economy, the singleton nation…
Figure 1 Starting questions for better formulations of social belief. 1 Why do I believe things? Look, obviously there is a lot going on with human beliefs in practice. I’m interested in the operationalisation in Hyland and Albarracin (2025) , which treats beliefs as mechanisms that do work: they take in observations and produce predictions and thus can make decisions. They are also objects in…
Figure 1 I am about to visit am visiting visited re-visiting Czechia. I would like to write something worthy of the country, in the form of a slapdash, idiosyncratic, mildly contrarian take on a place I know poorly, while avoiding the hubris of thinking I can craft a gleech-tier take . I cannot afford to set my standard for quality that high any more than I can set my standards for time-management…
Research-background notes. I want to pin down what it would mean, formally, for a social entity to contain a reduced-rank model of another social entity — possibly even a reduced-rank model of itself . Here are the formalisms I’m aware of, drawing on LLM lit review, some PDFs I had in a folder, and some vibes-y dot points I sketched out at the PIBBS x ILIAD residency . Much slop in here be…
Attention conservation notice: In hindsight this is a very low signal post. The good bits will be salvaged for other things. Figure 1 An intuition I took from Indy Johar’s recent Long Now essay on civilisational optionality 1 is something like the following: what we ought to be optimizing — at least at civilisational scale — is not any particular state of affairs, but the space of states of…
Environments and frameworks for multi-agent RL research. See also the opponent shaping page. Figure 1 1 Neural MMO A large, open-world game environment for multi-agent research (Suárez et al. 2019, 2023; Suarez 2024) . Claim to fame: leveraging gaming-industry technology to provide a persistent, large-scale world for agent interaction, going beyond the typical matrix game or grid world. 2…
Figure 1 Running code in the cloud without (explicitly) managing servers. In ye olde tymes, I’d rent a VM (or a whole machine), install my OS, configure my runtime, set up monitoring, worry about security patches, and pay for it 24/7 whether anyone is using it or not. Modern provisioning tries to make life easier. We can say cool phrases like devops — which AFAICT is short for “ containerization…
Figure 1 This is the technical companion to A social platform for your neighbourhood , which lays out the social and institutional case for community-owned, local platforms. That post is the why . This one is the how —or at least a plausible how , since the right technical choices depend heavily on what the community actually wants to build. Sign up for updates if you wish to know how this goes.…
Figure 1 None A local app for local people It’s 2029. Over the past three years, things have gotten worse in the ways everyone predicted and no one prepared for. Facebook Marketplace has become even less usable unusable—half the listings are AI-generated scam bait, the algorithm buries anything that isn’t a promoted post, and Meta’s latest “community standards” update broke half the local buy/sell…
Surely this analysis has been done before in the bowels of LessWrong . I gave up searching because it was too irritating trying to disambiguate the terms survival and hazard in the technical mathematical sense that I needed, against the more colloquial sense that they are used in AI safety discourse. Feel free to point me to prior work in the comments. Figure 1 I had a discussion recently where…
Follow along while we work this out— danmackinlay/SOV , or Sign up for updates . Figure 1 It turns out I’m writing a series of posts about practical community infrastructure building. This is the second one. In a previous post I sketched out a case for neo-friendly societies—small mutual aid groups that hedge against state decay by pooling resources and investing counter-cyclically. I wrote that…
Follow along while we work this out— danmackinlay/SOV or Sign up for updates . Figure 1 This is the technical companion to my sovereign LLM post , which makes the institutional and geopolitical case for small collectives owning their own AI inference hardware. This post lays out the concrete implementation pathway: what to buy, what software to run, how to remove CCP guardrails from open-weight…
Figure 1 This is a companion to Who wants to found a friendly society? , where I sketch the pitch for a neo-friendly society—a small mutual that invests in counter-cyclical assets as a hedge against state weakness. That post is the why . This one is the how , or the can we even? Sign up for updates if you’d like to know that answer. The question: if 50 of us wanted to pool money, invest it in ETFs…
We are very exercised about whether we have agency. What is it actually? What does it even mean to say that we have agency? Can we ground our intuition about agency in something precise? Do we even want it, or is being the grown up not necessarily in our best interests ? These are all questions I don’t have answers to. Here I have collected a list of edge cases for “agency” that I hope might help…
Figure 1 One massive project over the last few weeks has been founding the Alignment Journal , a new open-access journal for research on AI alignment and related topics. I am one of the two founding managing editors, along with Jess Riedel , and we have a fantastic team. We’re also really excited about the board members who’ve come on to give us guidance and help us with the work of running the…
Figure 1 It’s a nature-inspired ensemble strategy for optimization that uses evolution-like approaches. I ignored this for ages, but it includes lots of interesting and suggestive ideas. It’s also not much like other, more earnest attempts to copy evolutionary behaviour, such as the genetic programming approach to symbolic regression, which has not been that useful. Instead, it’s perhaps best…
Figure 1 A nature-inspired ensemble strategy for training neural nets that uses evolution-like approaches, even though we usually assume that scale can only be achieved via backprop . For the basics, see Evolution Strategies for beginners . 1 A running example We’re reusing the running example from the ES-for-beginners post. Let us consider the challenge of training a tiny neural net without…
Figure 1 If we have a big complicated system — a brain, a corporation, an ecosystem, an AI agent loop — where do we draw the line around “the thing” ? One formalization is the Markov blanket , a piece of standard graphical models machinery which has recently been conscripted into service as a theory of selfhood, agency, and alignment. I want to walk through how that happened, where I think it goes…
Figure 1 Category theory for systems . Myers (2023) : Categorical systems theory is an emerging field of mathematics which seeks to apply the methods of category theory to general systems theory. General systems theory is the study of systems — ways things can be and change, and models thereof — in full generality. The difficulty is that there doesn’t seem to be a single core idea of what it means…
The Wason selection task is a logic puzzle that uses a set of cards Figure 1. Figure 1: Each card has a number on one side and a patch of colour on the other. Which card or cards must be turned over to test the idea that if a card shows an even number on one face, then its opposite face is red? In its abstract form, the task is notoriously difficult. To test the rule “If P, then Q”, logic says we…
Figure 1 An interesting alternative formulation that relaxes the Markovian assumptions in mainstream RL, though I don’t know much about it. It seems to come out of AIXI and be similarly intractable, but it’s not purely theoretical; computable approximations do exist (Daswani et al., n.d.; Gao et al. 2023; Guez, Silver, and Dayan 2013; Hamilton, Fard, and Pineau, n.d.; Sunehag 2014; Tennenholtz et…
Learning to act clearly sounds like learning counterfactuals and interventions, right? (“What would happen if I took this action?”) So surely there’s a causal angle? Yes. Figure 1 Cf history-based RL . 1 Reinforcement learning The simplest and AFAIK earlilest analysis of learning to act in a causal setting applies to bandit problems (Lattimore 2017) . Since then, there have been many more…
Figure 1 I’ll try to synthesize LLM research elsewhere. This is where I keep ephemeral practical notes and links, continuing my habit from 2025 , 2024 , and 2023 . 1 Phrases it hurts that I can no longer say because now they are AI slop tells load-bearing failure mode -shaped earns its keep worth -ing reach for circling levers / knobs N different [things] in a trenchcoat lives in / sits in binding…
Figure 1 I read Simon (1996) back in 2010, and it seemed cool. Now that the world is full of AI systems and complex engineered socio-technical systems, it seems like I didn’t notice how prescient it was. First published amid the rise of cybernetics, early AI, and systems theory in the post-WWII era, the book challenged the dominance of the natural sciences (looking at you, physics) by advocating…
Figure 1 Placeholder for a particular trick in information theory : Williams and Beer (2010) introduced partial information decomposition (PID) as a way to split the mutual information that a set of sources has about a target into non‑negative “atoms” corresponding to redundant , unique , and synergistic information. Their framework has become a standard reference point, but their specific…
Figure 1 Here’s a list of things that interest me about London. Thanks to Kathryn Smith and Rebecca Giggs for various hot tips. 1 OK London, whatcha got to entertain me? I have a couple of days up my sleeve to do tourist stuff. What is the optimal tourist stuff for a guy like me to see in this town? To clarify, I think I can best summarise what I’m interested in going to as “peak London”…
I was in the first cohort of the Principles of Intelligence Research Residency . It ran from 5 January to 13 February 2026 in London . The pilot of the residency was successful. There will be some follow on programs that you yourself might wish to apply to. Figure 1 The residency was based at the London Initiative for Safe AI in Shoreditch. It was a lot of productive chaos. Here’s the blurb:…
Figure 1 When people talk about embodiment , they’re questioning whether the body is a necessary part of the mind. When they ask about stochastic parrotology , they’re wondering whether interacting with the world is a necessary part of agency. I suspect both of these frustratingly vague topics can be made more precise by considering the idea of causal embedding . Maybe we could also better…