RSS Amplifier

Cutting Heads · Jun 13, 2026

The Orchestra and the Instrument

0
Sign in to vote or save

Ryan Williams · Cutting Heads

The comparison gets made constantly, and it’s almost always wrong in the same direction.

Someone points to a model’s parameters. Someone else points to the brain’s synapses. The counts get placed beside each other as if they measure the same thing, and the remaining gap becomes a roadmap: make the model larger until it becomes a brain.

But a parameter is not a synapse, and neither count tells you what the surrounding system is built to do.

This is comparing an orchestra to a single instrument and concluding that the instrument needs to get bigger.

The brain is not one system doing one thing at a scale that current AI hasn’t reached yet. It’s dozens of specialized systems - each with its own architecture, its own connectivity patterns, its own effective dimensionality for its specific job - all running simultaneously and integrating their outputs through mechanisms we still don’t fully understand. The visual cortex processes spatial relationships in ways that have almost nothing in common with how the hippocampus consolidates memory, which works differently again from how the cerebellum handles timing and coordination, which is different again from how the prefrontal cortex manages planning and abstract reasoning.

Comparing the whole brain to a language model is comparing an orchestra to one very sophisticated instrument. The instrument may be extraordinary. It is still one instrument.

* * *

Thanks for reading! This post is public so feel free to share it.

Share

The parameter-count comparison is misleading because the units are not equivalent and the systems are organized differently.

The brain distributes its resources across specialized, deeply interconnected systems. Vision, movement, memory, planning, emotion, and autonomic regulation run continuously and influence one another. A language model concentrates computation on a narrower problem. Raw counts conceal that architectural difference.

Language is not confined to a neat set of dedicated regions. The familiar labels - Broca’s area, Wernicke’s area, the angular gyrus - identify important parts of a much broader network involving perception, memory, motor planning, and context.

That makes a clean numerical comparison impossible. The useful comparison is functional: a language model can outperform people on some language tasks while lacking the biological systems that give human language continuity, stakes, sensation, and a world to refer to.

The question isn’t whether AI has matched the brain. It’s which parts of the brain the comparison is actually being made against — and most people are making it against the wrong parts.

If you think current AI is just a smaller version of the brain that needs to get bigger, the path forward looks like scaling - more parameters, more compute, more data. And scaling has produced remarkable results. But it hasn’t closed the gap on the things that make AI feel fundamentally different from human intelligence, because the gap isn’t primarily about size.

* * *

You are never only processing what’s directly in front of you.

Right now, reading this, you are simultaneously aware - at some level below conscious attention - of the temperature of the room. The ambient sound. The quality of the light. Whether your body is comfortable or slightly stiff. The emotional residue of whatever happened earlier today. The vague background sense of what time it is and how long you’ve been sitting. Whether the environment feels safe or subtly off. Whether you’re hungry. Whether something smells wrong.

None of this is what you would describe as thinking. It’s not what you would point to if someone asked what you were doing. But it is running. Constantly. Below the threshold of conscious attention, your nervous system is maintaining an enormous ambient feed of sensory and proprioceptive information that grounds every conscious thought you have in a continuous experience of physical reality.

This background processing isn’t separate from cognition. It shapes it. The same information lands differently depending on your physical state. The same argument reads differently when you’re rested versus exhausted, comfortable versus in pain, safe versus threatened. The ambient feed isn’t neutral context. It’s an active ingredient in every thought.

A language model has none of this. Not a reduced version of it. Not a simplified approximation of it. None of it.

When you send a prompt, the model exists for the duration of that inference pass and nothing else. It has no body. No continuous experience. No ambient world. No proprioception. No sense of time passing between conversations. No emotional residue from yesterday. No background hum of physical reality grounding its processing in anything beyond the text in front of it.

It’s not that the model is like a brain in a vat. It’s more like a brain that flickers into existence at the moment of the prompt, processes the text, produces an output, and ceases to exist until the next prompt arrives. The continuity of experience that humans take so completely for granted that they never think to mention it - the fact that you are always somewhere, always in a body, always embedded in a physical and social environment that is always providing information - is entirely absent.

This gap is more significant for the AGI question than the parameter count gap. Not because embodiment is philosophically necessary for intelligence - that’s a debate I’m not qualified to settle. But because the ambient sensory grounding that humans have shapes cognition in ways that pure language processing structurally cannot replicate.

Consider what it means to truly understand the concept of cold. Not the definition. Not the associations. Not the way the word is used in context. The understanding that comes from having been cold - from the specific unpleasantness of it, the way it affects attention and mood and decision-making, the involuntary physical responses, the particular quality of relief when warmth arrives. A model trained on every piece of text ever written about cold has processed vastly more information about cold than any human ever has. But it has never been cold. And there is something in the having-been that the reading-about cannot fully replace.

This is not a sentimental argument about the specialness of human experience. It’s a technical argument about the information content of embodied existence versus textual description of embodied existence.

The text about cold is a compressed, lossy representation of the experience of cold, produced by humans trying to communicate something that language can only partially capture. The model learns from the representation. The human learns from the thing itself. These are not equivalent inputs, no matter how much representation you accumulate.

* * *

The scaling story gets complicated here.

If the gap between current AI and human-level general intelligence were primarily a parameter count gap, the path forward would be relatively clear even if expensive - train bigger models on more data with more compute. The field has done this repeatedly and it works. Capability scales with parameters and data in ways that are predictable enough to plan around.

But if a significant part of the gap is the ambient sensory grounding problem - the continuous embodied experience of a physical world that shapes cognition in ways that text-based training cannot replicate - then scaling the current architecture doesn’t close it. You can train a trillion parameter model on every text ever written and it still flickers into existence at the prompt and ceases at the output. The embodiment gap doesn’t narrow with scale. It’s architectural.

The research directions that take this seriously are the ones working on embodied AI - systems that operate in physical environments, receive continuous sensory input, develop something like proprioception and spatial awareness through interaction with the world rather than through text description of the world. Robotics research. Continuous learning systems. Multimodal architectures that go beyond text and image to include audio, haptic feedback, spatial reasoning grounded in physical navigation.

These are harder problems than scaling a transformer. They require different architectures, different training paradigms, different evaluation frameworks. And they’re less mature than the language modeling work that’s produced the current generation of impressive systems.

But they’re pointed at the right gap.

The brain comparison that actually matters for AGI isn’t the parameter count comparison. It’s the architecture comparison. The human brain is a massively parallel collection of specialized systems, continuously fed by an enormous ambient sensory stream, integrated through mechanisms that produce something we call consciousness and that we don’t yet know how to engineer. Current AI is one very sophisticated language-processing instrument, existing only in the moment of inference, with no body and no world.

Getting from here to there isn’t a matter of building a bigger instrument. It’s a matter of building the rest of the orchestra — and then figuring out how to make it play together.

* * *

I’ve been sitting with this essay longer than the others before writing it, because the peripheral awareness argument feels both obviously true and surprisingly underarticulated in the public conversation about AI.

Everyone who works with these systems long enough notices the flicker. The way the model has no memory of yesterday. The way context resets. The way it has no stake in the conversation beyond the conversation itself - no tiredness, no hunger, no ambient mood shaping the response, no body that the answer has to go back to living in. It’s always mentioned as a limitation of current systems, something to be engineered around with memory tools and persistent context.

It’s more fundamental than a missing feature. It’s a description of what the system is. And what the system is - a sophisticated language processor that exists only in the inference pass - is different from what the brain is in ways that parameter counts don’t capture and scaling doesn’t address.

That’s not a criticism of current AI systems. They are extraordinary at what they are. The ability to process language at this level of fluency and apparent understanding, from a training process that is still only partially understood, using an architecture that was invented to solve a different problem - this is one of the more remarkable things that has happened in the history of technology.

Remarkable at what they are is not the same as approaching what the brain is. The gap is real. It’s specific. And it’s not primarily a size gap.

They are extraordinary instruments. The orchestra is still missing most of its sections. And nobody has yet written the score that would tell all the sections how to play together.

That’s where the next piece goes.

No posts

Read the original on ryanwms.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.