An article that deeply resonated with me when I first read it a few years ago was “The Empty Brain” by Richard Epstein, for the way it insightfully broke down ways in which our everyday language equates the brain with a computer and ways in which that metaphor can be quite inaccurate. To set the context, Epstein starts by listing a variety of historical metaphors for the brain: hydraulics, automata, electronics, factories, etc. He links them together by noting they were each the most advanced technology available at the time the metaphor was used.
Some of these metaphors seem a bit ridiculous to us now, but they didn’t seem that way to those who used them, because the language was familiar and most people didn’t stop to consider how it wasn’t accurate. Like how I rarely stop to consider that time is not all that much like money: I can’t put seconds in a bank or work a job to earn months or use years to buy a car or loan a friend a few hours and expect to get them back later. And this sort of metaphorical fuzziness has a real effect on how I live my life: by treating time as a precious resource to be spent I tend to cram as much meaningful stuff into my time as I can — and end up ignoring my personal needs like sleeping or just plain silence.
But of course, no metaphor is perfect. They all highlight some important aspects at the expense of others. It is only when it comes to paradigm-defining metaphors (“Metaphors We Live By” as Lakoff and Johnson put it) that we run into trouble. Some of these historical metaphors for the brain held back medical science for a long time, because it gave the doctors who used it wrong ideas about how the body works. For example: suppose you think the body works by fluid pressure (hydraulics), as many people did for many years. You might think certain problems in the body could be corrected just by increasing or decreasing the pressure in certain fluids — for example, by blood-letting.
So, because talk about the brain as if it were computer is so familiar and pervasive, it is worth examining how it came to be, the ways in which we use it, how it might be accurate or inaccurate, and where it might lead us to think about our lives in helpful and unhelpful ways.
For all of these reasons, “The Empty Brain” was a compelling read for me, but one thing that seems to be missing from the essay is what kind of metaphor we might use in place of a computer. One possibility I’ve recently been considering is artificial neural networks, in part due to my familiarity with them. While at first it might seem odd to turn the artificial neural network metaphor around and use it in the opposite direction of what we’re used to, it will seem more natural as familiarity grows.
We all have a pretty solid grasp of the modern usage of the word computer but I think it is worth considering the origins of the word. When the term was invented, it applied to a profession — calculating sums and products and derivatives for business purposes. Many of the parts of the computer: the processor, the memory, the speakers, and so on were actually named from things that people did, as a convenient shortcut to understanding how that part of the computer was supposed to work.
Of course it is impossible to apply words in new contexts without changing our understanding of the word. By associating the word memory with the computer version, we are more likely to think of all types of memory (including human) as static, fixed-size, encoded, discrete, etc. Think about how characters in Star Trek describe memory as consisting of memory engrams which blur the concepts of memory from people and computers. However much the concepts might seem to resemble each other, they’re actually quite different. To explore the differences, Epstein has a neat little exercise for his students: he has them draw a dollar bill on the board. Of course none can reproduce it very accurately, even though they’ve all seen a dollar bill many times.
I think the fundamental mistake with terms like “memory engrams” or other ways that we conflate the brain with computers is that we conceive of the mind as software running on our brain’s hardware, a dualistic perspective encouraged by Christianity and other religions that think of people as body + soul. This perspective is thoroughly examined by Anthony J. Bell in his piece “Levels and Loops: the Future of Artificial Intelligence and Neuroscience.” He shares a wealth of evidence that there is no fundamental computational unit in the brain, and he states clearly his conclusion that:
A computer is an intrinsically dualistic entity, with its physical set-up designed not to interfere with its logical set-up, which executes the computation. In empirical investigation, we find that the brain is not a dualistic entity. Computer and program may be two, but mind and brain are one. The brain is thus not a machine, meaning it is not a finite model (or computer) instantiated physically in such a way that the physical instantiation does not interfere with the execution of the model (or program).
Matthew Cobb takes this idea — that mind and body are identical — and explains some of the implications in his article “Why Your Brain is not a Computer”:
The materialist working hypothesis is that brains and minds, in humans and maggots and everything else, are identical. Neurons and the processes they support – including consciousness – are the same thing. In a computer, software and hardware are separate; however, our brains and our minds consist of what can best be described as wetware, in which what is happening and where it is happening are completely intertwined.
Imagining that we can repurpose our nervous system to run different programmes, or upload our mind to a server, might sound scientific, but lurking behind this idea is a non-materialist view going back to Descartes and beyond. It implies that our minds are somehow floating about in our brains, and could be transferred into a different head or replaced by another mind.
I love the use of the term wetware here — adopted from science fiction, the genre of my heart — and the way it implies that the meat is inseparable from its computation. People with disabilities have a special awareness of how they are inseparable from their bodies, and that bodies matter! Claiming that one can be 'cured' of hearing loss without altering their identity requires a certain disregard for the importance of the body.
One caveat to the use of wetware is that the particular bits of meat relevant to the mind are hard to pin down. We often say the brain is identical to the mind, but drawing a boundary between the brain and the rest of the neural system is hard. There are pervasive impacts from the gut on thinking, is the gut a part of the mind? And what about counting on your fingers, are they a part of your mind too? And writing notes to yourself or creating calendar events, aren’t those types of artificial memory? And forming opinions based on what you read on the internet, aren’t you outsourcing what could be calmly considered rational arguments to a group mind?
So it might be hard to define what physical matter counts as the mind exactly, but perhaps we could agree at least on what the mind does. Nowadays people talk about information processing as the key sense in which people are like computers. Under this formulation, the stuff that counts as information is expanded beyond its traditional meaning of facts or numerical data, to basically anything that can be experienced. From external senses (sight, smell, touch, proprioception, etc.) to internal senses (feelings, gut instincts, sense of time, etc.) we have way more inputs than we can actually process. And we’re dynamically changing the type of information we collect as we process it, like turning on a light, or counting aloud. If all of this counts as information processing, then pretty much everything is an information processor: from nematode worms to fungi to single-celled organisms. If so many things are information processors, how is it that humans are especially computer-like?
I think there is a more interesting sense in which computers are like brains: both manipulate and operate on units of information known as symbols. There are entire academic disciplines devoted to symbols (e.g. semiotics, symbolic logic, etc.), so I won’t pretend I can do them justice here. But for the purposes of this work, you can think of symbols as discrete entities that represent other entities. As a metaphor, think of high-school algebra, e.g. “x stands for 17.” Traditionally, the “x” and the “17” are entirely unrelated to each other, the symbol could have been anything at all, say “y” or “” or “the_cost_of_my_dinner.”
What computers do is almost 100% symbolic. They take sequences of 1’s and 0’s and convert them into different sequences of 1’s and 0’s. This is not an accident, it is the preferred behavior for computers: that there be a marked separation between the content and the mechanism. Once again, a sort of dualism is baked into the cake.
While a lot of people think symbols work in approximately the same way in humans, a growing minority is not so sure. During my research on the topic of symbols, I found a fascinating paper called “Reconciling symbolic and dynamic aspects of language: Toward a dynamic psycholinguistics” by psychologists Joanna Raczaszek-Leonardia and J.A. Scott Kelsob. The article is rather technical but I wanted to quote part of it here because it fits perfectly with our discussion:
In some traditional frameworks, communication is often described as an exchange of symbols between participants. Symbols have some meaning ‘‘encoded’’ by a speaker and ‘‘picked up’’ by a listener. This ‘‘container’’ metaphor of a symbol (as Lakoff & Johnson (1980) labeled it) seems to be prevalent in many theories as well as in common knowledge about symbols. Some problems with such a view stem from assuming that (a) every time a symbol is used the meaning conveyed is the same (except in the cases of polysemy), and (b) a symbol is an independent and potent entity, whose ‘‘content’’ suffices to evoke a particular meaning in a listener.
If, however, […] in line with more pragmatically oriented frameworks, a symbol is considered to be just an element of a situation of communication, an element which has a function of affecting a listener in a certain way, it becomes clear that symbols almost always underdetermine what is really being conveyed. The rest is supplied by context: linguistic, situational, as well as variables characteristic of both the speaker and the listener. […] Linguists often emphasize the productive side of language—that it is possible to generate an infinite number of sentences. But it is no less amazing that the very same sentence uttered in different situations may have a completely different meaning.
So if computers operate on the bucket metaphor of language, having literal encodings and so on, then I wonder if we’re impeding our understanding of language when we assume language works in people like it works in computers. Nevertheless, it still seems to me that humans have achieved an unusual degree of facility with symbols, and this is something that at least on the surface they share with computers.
So if computers aren’t particularly good metaphors for the human brain, except in a rather narrow sense as symbol manipulators, then is there a better metaphor available or is that the best we’ve got? Well if Epstein is right, perhaps we should look to the most advanced technology of our age: artificial intelligence. There are already communities that have started to use this metaphor, and I imagine it will only become more prevalent, which is why it is worth examining in closer detail. Some examples I’ve heard personally:
Well, it depends what you’re optimizing for.
I’ll update my belief in light of this new information.
That seems to contradict my prior probability estimate.
How many parameters are there in the human brain?
Perhaps I just need more training data to learn this.
It seems to me that whether we like it or not, artificial neural networks as a metaphor are likely to become a part of our language. Like any language changes, there are things that it describes aptly and blind spots that it might encourage. But first, perhaps it’s worth describing how the language of neural networks came to be in the first place. Of course brains and neurons were first used as metaphors for their artificial counterparts, not the other way around.
For those unfamiliar with the concepts of artificial neural networks: the basic idea is to train a set of artificial “neurons” to produce the desired outputs when fed certain inputs, by learning correlations and patterns in the different types of inputs. Each of these neurons is wired to some subset of inputs and/or other neurons, and associates a “weight” to each given input wire. The weights are said to be like the inhibitory or excitatory nature of various types of neurons. When the incoming wires produce a sum above a given threshold, the artificial neuron “fires” and produces an output.
To do the training, examples are presented one after another, and the answer the network produces is compared against the true answer. If the network gets the answer even slightly wrong (90% chance the image is a cat, not 100%), then the neurons that most contributed to the wrong answer are algorithmically computed (with calculus), and they update their weights slightly towards producing the correct answer for that sample. This algorithm is called “back propagation” for the way it works from the outputs towards the inputs.
In practice, this is implemented with a series of matrices that are multiplied together, which do not much resemble neurons. Neurons are not updated back to front in the brain, nor do they read or produce floating point values. They are not arranged in neat layers of matrices but can have intricate and complicated connections. Real neurons are heavily modulated by their environments, with neurotransmitter availability, blood flow, and other bodily processes changing their behavior, while their artificial counterparts are designed to produce the same outputs when given the same inputs.
Given these differences between artificial and biological neurons, I’d like to note that it was not inevitable that brains should come to be the primary source of language for talking about modern artificial neural networks. In fact in “Introduction to Neural Nets (Without the Brain Metaphor),” Mark Reidl describes a good alternative that I think makes it clearer how neural nets actually work: he says artificial neural nets are like an electrical wire network used to connect sensors to control mechanisms. The wires have a resistance which allows more or less current to pass through, and the control mechanisms behave differently based on incoming current.
Another alternative is “differentiable programming,” popularized by famous AI researcher Yann LeCun in a Facebook post, which emphasizes the calculus and computer-science aspects. Basically the idea is that there are many computer operations that we might like to program with examples, rather than hard codes, so we replace these operations with differentiable ones and provide examples to how the program should behave.
All these examples are to show that there are compelling alternative metaphors, and neurons are not necessarily the one I would have chosen for what is now called artificial neural networks. This is not just because the metaphor is inaccurate in some ways, but also because as a metaphor it is evocative, leading us to ascribe intelligence where none is likely to exist. But it is what we’ve got, so let’s look at it deeper.
First, let me relate some things I think might be well explained by this metaphor. Here’s a part of an essay titled “Human Language Understanding & Reasoning” by the computational linguist Chris Manning, which I think does a nice job of showing how a neural network metaphor of language can help us understand meaning and by extension the mind:
I suggest that meaning arises from understanding the network of connections between a linguistic form and other things, whether they be objects in the world or other linguistic forms. If we possess a dense network of connections, then we have a good sense of the meaning of the linguistic form. For example, if I have held an Indian shehnai, then I have a reasonable idea of the meaning of the word, but I would have a richer meaning if I had also heard one being played. Going in the other direction, if I have never seen, felt, or heard a shehnai, but someone tells me that it's like a traditional Indian oboe, then the word has some meaning for me: it has connections to India, to wind instruments that use reeds, and to playing music. If someone added that it has holes sort of like a recorder, but it has multiple reeds and a flared end more like an oboe, then I have more network connections to objects and attributes. Conversely, I might not have that information but just a couple of contexts in which the word has been used, such as: From a week before, shehnai players sat in bamboo machans at the entrance to the house, playing their pipes. Bikash Babu disliked the shehnai's wail, but was determined to fulfil every conventional expectation the groom's family might have. Then, in some ways, I understand the meaning of the word shehnai rather less, but I still know that it is a pipe-like musical instrument, and my meaning is not a subset of the meaning of the person who has simply held a shehnai, for I know some additional cultural connections of the word that they lack.
Of course “network of connections” is a bit more abstract than “artificial neural network,” but it is not hard to see how the metaphor could be easily extended: we have weights for the connections between various concepts and these get updated when we hear and use terms in new contexts. If you assume you can quantify the weights, then perhaps they could form a sort of vector that defines a concept, which could be more or less similar to other concept vectors.
This sort of description of language is designed to explain why words are so hard to define. An example I’ve discussed with my friends before is that of the food chili. We were able to come up with several recipes that have entirely non-intersecting ingredients lists but which all would commonly be thought of as chili: chili con carne, white chicken chili, chili verde, and black bean chili. If the recipe doesn’t define a food dish, what does? Well perhaps “word vectors” or “networks of association” could explain the concept of chili, where a recipe falls short. The foods we call chili exhibit certain features that the word chili conjures up more or less strongly, for example: spicy, soupy, meaty, tomato-y, etc. The idea is that if a food matches enough of these features then it will match somehow the representation of the word in our heads.
Overall, I think the artificial neural network metaphor does some things quite well. It takes into account the fact that associations can change over time and give rise to new meanings. Also, it adds a level of context-dependence to meaning — that some associations might be stronger in certain contexts, and weaker in other contexts. Finally, it helps us understand that different people might have different associations with a word or concept.
Yet for all that it does well, I don’t think a neural network metaphor of meaning entirely succeeds at breaking out of the bucket metaphor of meaning. It still describes words as context-dependent vectors or networks of association, which presumably contain the meaning of the word. With artificial neural networks, the only context that is considered is the words that occur near the relevant word. Several of my grad school lab mates studied the limitations of this approach — one example is that opposites tend to appear in the same contexts, and are sometimes seen as similar words by neural network systems.
The amazing performance of ChatGPT and other recent Large Language Models (LLMs) show that a surprisingly large portion of the contextual cues in written language can be picked up from the text itself. The writing style tells you a lot about the author. The topic tells you something about the historical context of the piece. When communicating by text, authors often make the relevant context explicit, because a written piece is an unusually static form of language in the history of language.
Even in written language, though, there are dynamic elements. Authors will change their style as they learn and grow throughout their lives (e.g. early Wittgenstein vs. late Wittgenstein). Assumptions are made about what concepts a reader will be familiar with, and a reader encountering an unfamiliar word or concept can stop reading to look it up (e.g. Wikipedia tab explosions). Reading is an inherently temporal experience — one normally starts at the beginning and proceeds through the text, though it can easily vary with restarts or breaks.
Many of these dynamic parts are not true of neural networks, or at least not LLMs. First, most neural networks are trained and then frozen for the rest of their shelf lives; they do not develop a unique style or change their perspective. There exists some research on lifelong learning but it remains a niche field in machine learning. Next, unless explicitly trained or prompted to do so, an LLM will not look up unfamiliar terms, and certainly wouldn’t be so distractible as to look up terms unrelated to the current query. And finally, modern LLMs do not read the text from start to finish, but rather pay attention to parts of the text as they are relevant (or not) to whatever word it is currently generating in response.
To sum up, however good LLMs are (or seem to be) at generating coherent text, they do not read or write or learn like people do. They are generalization machines that are built to predict word distributions, and they do this very well. I have heard them described as “calculators for writing” to borrow a metaphor from math, and I think this is a nice way of conceptualizing them.
Are there other metaphors for the mind that break out of the assumption that meaning works like a bucket? Going back to meaning as it pertains to our minds, I think Ludwig Wittgenstein was on the right track. He repeatedly suggests to look at a word’s use to determine the meaning:
For a large class of cases of the employment of the word ‘meaning’—though not for all—this word can be explained in this way: the meaning of a word is its use in the language.
Of course this appeals to me, as a Pragmatic Buddhist, along with his zen-like sayings, such as: “Don’t think, but look!” by which he means we should avoid excessive generalization and appreciate the uniqueness of individual cases.
So what does this mean for the mind? The mind defies any boundaries you might try to draw around it, definitely exceeding the brain that some people have equated it with. But not because it is some abstract ethereal spirit or soul that transcends space and time, but rather because of the very real processes playing out on fingers, on paper, on computers, on neural networks, etc. These things reflect the properties of our minds to some degree because our minds are inhabiting them.
Sometimes artists will say they “put themselves into” a work of art. Perhaps an artist making a collage or a musician improvising would be a better metaphor for the mind, the way they playfully use the various tools at their disposal to explore the possibilities of space and time. Like these creative folks, minds follow certain rules, but when the rules no longer serve, minds can break the rules in new and creative ways.
To be clear, the purpose of this essay isn’t to advocate that we should or shouldn’t use artificial neural networks as a metaphor for the human brain, but only to observe what seems to me a trend in talking about human cognition and to see where such a metaphor already pops up. If there is a point to this essay, then it is perhaps to motivate us to notice our language a little more and to hold on to our metaphors a little less.
Bell, Anthony J. “Levels and Loops: the Future of Artificial Intelligence and Neuroscience.” Philosophical Transactions of the Royal Society B, 1999.
Cobb, Matthew. “Why Your Brain is not a Computer.” The Guardian, 2020.
Epstein, Richard. “The Empty Brain.” Aeon, 2016.
Lakoff, George and Johnson, Mark. “Metaphors we Live By.” University of Chicago Press, 1980.
Manning, Chris. “Human Language Understanding & Reasoning.” Daedalus, 2022.
Raczaszek-Leonardia, Joanna and Kelsob, J.A. Scott. “Reconciling Symbolic and Dynamic Aspects of Language: Toward a Dynamic Psycholingistics.” New Ideas in Psychology, 2008.
Riedl, Mark. “Introduction to Neural Nets (Without the Brain Metaphor).” Medium, 2017.
Wittgenstein, Ludwig. “Philosophical Investigations.” Oxford: Blackwell, 1953.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.