It has been encouraging to see recent opposition to some popular but misguided characterisations of AI. Most notably, the idea that large language models (LLMs) are just a sophisticated form of “next-word prediction.” That characterization has served its purpose of demystifying the new technology to an otherwise deeply mystified public around the time that the ChatGPT moment arrived in late 2022. But the comparison has become ever less accurate and far less fruitful.
The most entertaining criticism of this idea, and one that aligns with my broader argument in this article, comes from the excellent Machine Pareidolia Substack by Jinx, in a fabuolous demolition entitled “The Cathedral and the Stones: What ‘Next Word Prediction’ Actually Means.” The opening sentences encapsulate the key point:
There is a criticism of large language models that sounds sophisticated. You’ll hear it from computer scientists who should know better, from philosophers who definitely should, and from that particular species of online commenter who confuses dismissiveness with rigor. It goes like this:
“It’s just next word prediction.”
Technically true. Also: ballet is just muscle contractions. A symphony is just air pressure variations. The Mona Lisa is just pigment on wood. The word “just” is doing enormous load-bearing work in that sentence, and it is NOT up to the task.
Somewhere, “next-word prediction” coalesced with the catchy critique that AI models are mere “stochastic parrots.” The latter comes from an influential paper from 2021, before the ChatGPT event horizon. In "On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?” Emily Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell present now-familiar objections to the accelerating, enlargening, as I think Dr Seuss would say, “biggering” approach taken by AI companies.
Contrary to how it may seem when we observe its output, an LM is a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning: a stochastic parrot.
For all their concerns about the risks of large language models (LLMs), including environmental costs, biases, and models’ inability to truly understand meaning, Bender and company’s metaphor belies a deep chauvinism toward the 410 species of the order Psittaciformes. I am honour-bound as a behavioural ecologist to challenge this. And I will. But not today.
Since 2021, the deployment–I am tempted to say “parroting”–of that paper’s avian metaphor has become a sign of deep insecurity, an attempt to minimise or, more precisely, preemptively to dehumanise a technology. Why dehumanize? Because AI, and language models in particular, present escalating challenges to our notions of what it means to be human.
It helps to understand how things work in order to understand their consequences. A grasp of how LLMs learn and generate their outputs, and indeed a grasp of the ancient and venerable technology of next word prediction, can help clarify the uses, benefits and costs of LLMs. But these outcomes are not the mere sum of the mechanistic parts. They also entail complex interactions that carry a whiff of the ineffable.
In the ecological and evolutionary study of animal behaviour, including the behaviour of human animals, there exists a rich respect for understanding mechanisms. Studying genes, hormones, neurons, neurotransmitters, receptors and other constituent parts can provide a rich and detailed understanding of how behaviour works. But complementary traditions concern themselves with why organisms behave the way they do, rather than in any of the inifinite other possible ways.
Why do cicadas spend years underground as larvae only to emerge simultaneously as adults to make their ungodly racket together?
Why do male fairy wrens forego breeding in their first year and remain in their natal nest to raise younger siblings?
Why are humans, perhaps alone among mammals, prone to preeclampsia in the latter stages of pregnancy?
The “Why?” in these questions translates to “What is the adaptive benefit of such an adaptation?” That is to say, what features of the system would have favoured any underlying genetic information and the configuration of the mechanisms involved.
You don’t need to know the mechanisms first, or to build from a reductionist to a more holistic understanding. Science can operate in both directions, often simultaneously. To make progress it is sometimes necessary to remain agnostic about the underlying mechanisms. At least it is for a time.
Evolutionary game theorists occasionally adopt “the phenotypic gambit”, an approach in which they consider the evolutionary dynamics around a trait of interest without getting into the details of the genes involved. The same can be true of any mechanistic step, not just the genes. It is a strategy not universally embraced because, of course, sometimes the mechanistic details do make all the difference.
Consider pre-eclampsia, a dangerous condition affecting around 1 in 20 pregnancies, in which a mother’s blood pressure rises late in the third trimester, leading to edema, protein in the mother’s urine, and blood vessel damage. Worse, the damaged vessels can lead to yet higher blood pressures in a feedback loop that risks damage to the mother’s brain and the life of both herself and baby.
Attempts to understand and treat this serious pregnancy complication understandably focused on the last-trimester storm of blood pressure, kidney strain and the ensuing symptoms. Some reports also found that placentas from pre-eclamptic pregnancies were smaller than those from other pregnancies. So were the newborn bodies, other than the heads which were normal size for their gestational age. The symptomotology was so complex and so puzzling that preeclampsia came to be known as “the disease of theories.”
The key to understanding why pregnancies, a literal pathway to reproductive fitness, are so prone to the dangers of preeclampsia comes when we consider the evolutionary interests of everyone involved: mother, father, and fetus. As well as the fact that the placenta is decisively playing on “team fetus”. The answer: a surreptitious battle of physiological wills in the first trimester defines how the placenta will grow, and if the mother gets an advantage and prevents the placenta from winning this battle, the risks of preeclampsia escalate.
Here is a serious example where decades of mechanistic hammering could not find the single causative nail. As deeply non-intuitive understandings of how family relations simmer with conflict emerged from evolutionary biology, notably the work of David Haig, they finally provided a coherent understanding of why preeclampsia happens. The mechanisms weren’t the message, but rather they were intermediaries from a conflict that plays out eight months before the mother’s blood pressure rises and her feet begin to swell.
Consciousness is just another mechanism. An impressive one, to be sure, but a mechanism that exists to fulfil an adaptive function nonetheless.
The tendency to mistake the mechanism, or, worse, a simplified sound-bite account of the mechanism like “stochastic parrot”, for the entire picture seems to be everywhere in AI discourse right now. Sometimes it is worth zooming outward and looking simply at what the machine does and why, while stepping away from the details of how. That is especially true because current-day AIs are such complex multidimensional beasts, and because the mechanisms are always changing. They change even when you hold the architecture constant. That’s learning for you.
Jinx made a related point, and their articulation is pithier:
What’s happening when someone collapses the entire mathematical machinery of a transformer into “next word prediction” is not simplification. It’s a category error. They’re describing the training objective and calling it the capability, which is roughly equivalent to saying evolution is “just reproductive fitness” and therefore everything it produced, consciousness included, is reducible to sex.
As an evolutionary biologist, I am used to making the argument that evolution is ultimately reducible to reproductive fitness1. But, as phenomena like preeclampsia and consciousness illustrate, even if you embrace this simplification of natural selection being much like the “training objective” in AI, the resulting capability can be an incredibly complex and interesting blend of function and mechanism.
Reproduction is the gateway that gives evolution its shape, much like a training objective gives an AI its properties. But organisms evolve all manner of both functions and supporting mechanisms as a result of the downstream reproductive advantages that they may, in a rather loose and probabilistic sense, tend to confer. Even strategically-applied prudishness can be an effective mechanism in the right context, with the function of constraining one’s competitors from unabashedly throwing a leg over and gaining an evolutionary leg-up.
What about consciousness? Is it a mechanism or a function? Consciousness presents one of the biggest hangups in the AI universe. It will likely remain so because people, even those in Silicon Valley, are very fond of their own consciousness and also their conscious accounts of their own experience. Many would like to believe it is a function, a desirable end-point to which AI creators should play midwife-creator.
I disagree. Consciousness is a mechanism, at least for humans. An impressive one, to be sure, but a mechanism that exists because consciousness, and the intermediate trait-steps by which it evolved, historically enhanced the reproductive fitness of those who had better versions of it than their competitors had.
Machine consciousness may yet come to exist. Hell, it may already exist. But even if it does, will it be the same thing as human consciousness? And is there something special about human consciousness that sets it apart from other functions or capabilities?
I can’t see why. But the nature of machine consciousness, and of any other interesting capability, cannot be demonstrated in the mundane details of model architecture or training objectives alone. It will (also) require an answer to the ultimate question:
What does it do?
Yes, at the level of the gene. Not worth getting into kin selection etc here. And reproduction does not equal sex. Creationists of both the religious and cultural variety note the primacy of sex and reproduction in evolutionary accounts, pointing out that not everything is about sex. Hell, not even reproduction is purely about sex, but I am happy to let the prudes off that particular hook for now.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.