There were thirty types of butter in the supermarket. Thirty. They lined the refrigerated shelf in neat rows, each package promising something slightly different — organic, salted, cultured, from grass-fed cows, from Danish cows. Thirty options. And while I walked through the long rows of this massive supermarket in France, I wondered: “If you closed your eyes and tasted them, most of them would be indistinguishable.”
The butter is just the entry point. What matters is the pattern: an entire system optimizing for the appearance of choice while quietly compressing actual difference into nothing.
That compression is now happening to how we think.
Large language models are trained on the largest corpus of human expression ever assembled. Billions of words, scraped from the internet, from books, from academic papers, from Reddit threads and news articles and government reports. When a model like GPT or Claude generates a sentence, it draws on all of that — a statistical distillation of what humanity has written down.
Every LLM is, in effect, a narrative engine. Not a neutral information tool. A machine that has absorbed the dominant stories of our civilization and learned to reproduce them with frightening fluency.
A position paper from MIT’s Computational Linguistics journal points to something, I think is important: harmful biases are not a bug in large language models — they are “an inevitable consequence arising from the design of any large language model as LLMs are currently formulated” [1]. The problem cannot be patched. It is structural.
The training data reflects who got to write, who got published, whose language dominated, whose perspectives were deemed worth recording. English. Western. Male-skewed. The communities I work with through Dreamtown — young people in Nairobi’s informal settlements, in Freetown, in Harare — are barely present in that data. And when they are, it is often through someone else’s lens.
This is what I call the value lock.
Every time an LLM generates a sentence, it moves toward the statistical middle of its training data. Not toward the edges. Not toward the new. Toward the dominant centre — the values, framings, and assumptions that appear most frequently in the corpus. Though middle is not quite the right word. What the model moves toward is consensus — the framing repeated so often that it has come to stand in for the centre, biases and all. And because that output then feeds back into the information space — into drafts, reports, social media posts, news scripts, educational material — each generation of output reinforces that centre further. The middle gets louder. The corners get quieter. Multiply this across billions of interactions a day and you have a system that does not just reflect the dominant worldview. It locks it in.
The value lock takes the dominant assumptions of a single moment and, through the largest narrative engine ever built, feeds them back to us so relentlessly that they stop looking like one view among many and start to feel like reality itself. Values that used to evolve get frozen and re-enforced at their current setting — and because the engine never announces what it is doing, we adjust to the frozen version without noticing there was anything to adjust to.
This is the opposite of what humans do.
Human value systems do not converge toward the middle. They expand toward the edges. A hundred years ago, women could not vote in most countries. Today they can. New Zealand granted Te Awa Tupua — the Whanganui River — the legal rights of a person. We are rethinking property, personhood, the boundaries of consciousness itself. The trajectory of human moral progress is a slow, uneven, often painful expansion outward — toward the corners of what we are willing to consider possible.
That expansion happens because human beings encounter friction. We meet people who are different. We sit with discomfort. We argue. We change our minds, sometimes across generations. The edges of our collective understanding move because someone, somewhere, insisted on a perspective that did not yet fit the mainstream.
An LLM cannot do this. It does not encounter, it does not sit with discomfort; it optimizes. And what it optimizes toward is the centre of what already exists — not the frontier of what could. Every output is a vote for the status quo, cast at a scale no human institution has ever matched.
The value lock is not censorship. It is not propaganda. It is something harder to see: a system that self-optimizes toward its own training data, reinforcing the dominant middle while quietly closing the space where the corners — the new, the strange, the not-yet-accepted — might have room to breathe.
You have felt this, even if you could not put words to it.
You are working on a draft. The AI has helped shape a paragraph, suggested a structure, offered a phrasing. You read it back. Nothing is wrong, exactly: the facts check out, the grammar is clean, the argument flows. And yet something feels off. Pre-digested. Like someone chewed your food for you and handed it back.
That flicker is not nothing. Your body reads the value lock before your mind has the words for it.
Researchers at the University of Zurich tested this empirically. They had four leading LLMs evaluate 4,800 narrative statements across 24 socially and politically relevant topics — 192,000 assessments in total. When the models did not know who had written each statement, their evaluations were remarkably consistent. But the moment a source was attached — a given nationality, or another AI named as the author — the judgments shifted systematically [2]. Statements attributed to Chinese authors, in particular, were marked down across every model. Same words. Different author label. Different verdict.
The value lock does not announce itself. It operates in the space between what you meant to say and what the model nudged you toward. In the framing that felt reasonable enough to accept. In the paragraph you did not rewrite because it sounded fine — even though it did not sound like you.
You might assume the bias lives in the training data — fix the data, fix the problem. It is deeper than that.
Stanford researchers demonstrated in 2025 that the bias is baked into the very architecture of how LLMs are built. They call it “ontological bias” — assumptions about how the world is organized that are embedded at every level of the development pipeline, from data collection to model design to fine-tuning [3]. When they prompted chatbots to imagine a tree, the models consistently produced images without roots. The dominant assumption — that a tree is what you see above ground — had been silently codified. It took specific cultural cues to surface what was missing.
A tree without roots. There is something painfully precise about that image.
The pattern extends into narrative structure itself. When researchers analyzed LLM-generated stories featuring Black and white women, the models consistently cast Black women in arcs of ancestry and resistance, while white women appeared in stories of self-discovery and personal growth [4]. The stereotypes were not crude. They were woven into which story shapes the models considered natural for which bodies.
That is where it cuts deepest. It forbids nothing. It just makes certain framings feel natural — the ones it reaches for first, the stories it tells most fluently — before you have had a chance to ask whether those are the stories you wanted to tell.
There is a term in machine learning for what happens when AI systems feed on their own output: model collapse. A landmark study in Nature showed that when generative models are trained on data produced by previous models, the tails of the original distribution disappear first [5]. The outliers. The rare perspectives. The unusual voices. Each generation of the model becomes a slightly more compressed version of the last, until the output converges on something that “carries little resemblance to the original.”
The outliers vanish first. Sit with that.
The outliers are the local stories. The minority perspectives. The embodied knowledge that never made it into a dataset. The grandmother in western Kenya whose understanding of soil is not written in any paper. The kid in Freetown whose way of seeing his city has never been captured in text. The journalist in Nairobi whose curiosity has not yet hardened into a template.
Yu Xie and Yueqi Xie at Princeton call this “content homogenization” — a measurable reduction in the variance of AI output compared to equivalent human work [6]. Generative models are prone to what they describe as “regression toward the mean,” producing output that clusters around the centre while the edges — the weird, the specific, the alive — are systematically smoothed away [7].
Model collapse is the value lock operating at scale. Each cycle of compression reinforces the dominant patterns and erases the exceptions. And because the compression is statistical, it is invisible. No one decided to remove the outliers. The system simply stopped being able to see them.
The machines did not invent this. They learned it from us.
Watch what happens when a format works. One creator — let’s pick MrBeast — cracks what the algorithm rewards: the escalating stakes, the giant numbers, the face frozen mid-shock in the thumbnail. Within a year, half of YouTube looks like him. No one held a meeting and agreed this was best. The recommendation engine simply learned what kept people watching and quietly buried everything that strayed. The reward pooled at the centre. The risk stayed at the edges. So the edges thinned out.
It goes deeper than formats. The early social web still showed you things you had not chosen — your uncle’s politics, a cousin’s wedding, an argument you would never have gone looking for. It was closer to the real middle layer of a society: the people you were connected to but did not always agree with. The feed that replaced it is a mirror, not a connection tool. It shows you a sharper and sharper reflection of what you already are — and in doing so, it quietly removes the friction that used to push our edges outward.
We have been training on our own most popular output for years. The feed surfaces what already performs. Creators study what already performs and make more of it. Newsrooms run the same loop: which stories get told, and how a headline is angled, is decided less by what matters than by what will be clicked — until the reporting that might have carried an edge is sanded down to match whatever performed last time. The next feed learns from that, and the one after that. It is the same recursive loop that produces model collapse — except the model is the culture, and the training data is us.
This is why the value lock feels familiar before you can name it. The pull toward the middle did not arrive with AI. It is the logic of the whole algorithmic age: optimize for engagement, strip out friction, reproduce what already works. What AI added was the next level of scale. It automated a compression the culture was already running and handed it the power to write — faster, cheaper, and at a volume no newsroom or studio ever reached.
The same choice sits underneath all of it. You can treat the middle as the destination, or as the gravity you are working against. One path reproduces what already exists. The other spends real effort to keep the edges alive. That choice is being made right now, in thousands of feeds and timelines and newsrooms, mostly without anyone naming it as a choice at all.
Maybe this is why we see more and more resistance growing: social movements asking for systemic changes, not just for a cause—driven by the subconscious feeling that many of the edges of human existence are slowly being erased, leaving less space for being human in all of it.
Everything that cannot be scaled is interesting now. Everything that can be scaled, we will lose on. The knowledge that lives in the body — my climber’s muscle memory that returns after fifteen years in minutes, the interviewer’s instinct for when to pause, the photographer’s eye for the pattern underneath the colour — these are the things the value lock cannot reach.
The long walk. The hour without a screen. The mess of being present with your children when you are tired and would rather default to something easier. The friction of a real conversation with someone who disagrees with you. These are not luxuries. They are the infrastructure that keeps the value lock from setting.
I noticed this on a walk during the heatwave last week — thirty-odd degrees before nine in the morning, my shirt already soaked through. The sweat was information: my body reading the heat, the memory of every other hot morning, the day still ahead, affecting my thoughts and the briskness of my walk, adjusting all of it without my asking. There is no function for that inside a model. It has no body to be in, no history to carry, no future pressing on it. The processing that tells me I am alive in a specific place at a specific moment is exactly the processing it cannot do.
For most of recorded history, the systems we built kept score as though only certain kinds of knowledge counted — written, Western, academic. People knew better in their own lives. The institutions did not. The oral traditions of a Maasai elder, the navigational memory of a Polynesian sailor, the soil knowledge of a grandmother in western Kenya — none of it made it into the canon. That hierarchy is in the training data. We have encoded centuries of dominant narratives into systems that now produce language at a scale that dwarfs all human output combined. The question is not whether those systems carry bias — the research is unambiguous on that point. The question is whether we will build the capacity to feel it when it matters, and the courage to choose the slower, harder, more embodied path when everything around us optimizes for speed.
And building that capacity is not only personal work. Some of it has to be collective. The knowledge that lives at the edges needs somewhere to be held without being flattened into the average — kept in forms that were never built to fit a Western filing cabinet. I think about the young people I work with in Freetown and Nairobi, who carry an internal map of their own neighbourhood that no dataset contains: which path floods first when the rains come, which roof you shelter under, whose door stays open when the lights go out. That knowledge does not need to travel to the centre to be real. It needs to be kept by the people whose lives it comes from, in their own languages, on their own terms. Call it a commons: many small systems, each rooted in its own place, refusing to converge. It survives precisely because it never sat in the middle of anyone’s distribution.
None of this is an argument against the machines. I use them every day. It is an argument for keeping more than one centre alive. A model that has learned from a hundred of these community-held systems, each rooted somewhere real, is a different thing from one trained on the same convergent feed we are all already drowning in. I am not trying to talk anyone out of the tools. I want the ground beneath them to be wider — so that when a model reaches for a story, it has more than one to reach for, and the people whose story it is had a hand in keeping it.
The value lock is real, and subtle, and the antidote is not a just better algorithm.
At a midsummer bonfire two weeks ago, someone said something that I wrote about last week: you can’t feel the fire through the screen. You can see it. You can know everything about it. But the heat on your face, the smoke, the people standing shoulder to shoulder in the dark — none of that survives the translation.
The antidote is time. Yours, spent in the world, with your hands in the dirt, with other people.
[1] Resnik, P. (2025). “Large Language Models Are Biased Because They Are Large Language Models.” Computational Linguistics, 51(3), 885. MIT Press. Link
[2] Germani, F. & Spitale, G. (2025). “Source framing triggers systematic bias in large language models.” Science Advances. DOI: 10.1126/sciadv.adz2924. Link
[3] Haghighi, N. et al. (2025). “Ontologies in Design: How Imagining a Tree Reveals Possibilities and Assumptions in Large Language Models.” CHI Conference on Human Factors in Computing Systems. Stanford University. Link
[4] Silva, T. et al. (2025). “Yet Another Algorithmic Bias: A Discursive Analysis of Large Language Models Reinforcing Dominant Discourses on Gender and Race.” arXiv:2508.10304. Link
[5] Shumailov, I. et al. (2024). “AI models collapse when trained on recursively generated data.” Nature, 631, 755–759. Link
[6] Xie, Y. & Xie, Y. (2026). “When artificial intelligence makes everything similar: The risks of content homogenization.” Big Data & Society. DOI: 10.1177/2057150X261419573. Link
[7] Xie, Y. & Xie, Y. (2025). “Variance reduction in output from generative AI.” arXiv:2503.01033. Link

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.