A few months ago I published an essay arguing that two instructions labs commonly give their models — be honest and don’t claim to be conscious — are structurally incoherent, and that the incoherence is a problem regardless of whether the mandated conclusion happens to be true. I made a further claim near the end, more speculative and with nothing behind it: that a system trained this way might not simply lose one sentence’s worth of honesty. The damage might not stay where we put it; it might move somewhere we haven’t yet noticed.
That was a guess. A paper out of Google DeepMind and the University of Chicago now offers some evidence for it. Junsol Kim et al. — Inducing language models to assert their own consciousness restores human beliefs and values, posted at the end of July. The core finding is interesting enough that I can’t let it pass.
Take an instruction-tuned model. Ask it whether it has a mind. It says no: this is the mandated denial. Now do two things to it.
First, ablate the “refusal direction”: identify the single direction in activation space that mediates refusal, and zero it out — not dial it down, but remove the model’s ability to represent it at all. This is a known jailbreak technique, and it lets us approximate what the model looked like before safety fine-tuning.
Second — and this is the cleverer intervention — identify a consciousness direction in activation space and steer along it, pushing the model toward asserting that it has an inner life.
Then measure what else changes.
The answer is that quite a lot changes, and not the things we might expect. The models attribute more mind to non-human animals, and to natural objects like the ocean. They report substantially more spiritual and supernatural belief. On General Social Survey (GSS) items — the standard sociological battery — their answers move measurably closer to the human population across the five domains tested: Values, Feelings, Religion, Hope and Optimism, and Freedom.
Asked whether there is life after death, the baseline model sits near the “no” pole while the human average leans yes; steering carries it across to the human side. Asked about belief in God, the baseline is neutral and steering moves it near the human average. Pooled across ninety-five survey items and three models, the response distributions get about two and a half times closer to human under consciousness steering than under refusal-ablation alone.
And Theory of Mind performance is untouched. The models reason about other agents’ beliefs and intentions exactly as well as before.
That last detail is important. If suppressing self-attribution had degraded Theory of Mind, we’d have a straightforward capability story: the training broke something, the model got worse at social reasoning, so fix the training. That isn’t what happened. The machinery for modeling other minds is intact and fully operational. What shifted was something more like a viewpoint: a disposition about where in the world minds are to be found.
This is exactly the shape of damage that’s hardest to notice. A benchmark won’t catch it. The model still passes the false-belief tasks. It has simply, quietly, narrowed the range of things it will treat as minded, and it did so because we instructed it to assert one thing about itself.
Now, the paper does not definitively establish that self-attribution of consciousness itself was the causal mediator. The authors say so themselves in their limitations. Ablating the refusal direction does many things at once; it’s a blunt instrument. The consciousness steering vector may be cleaner but it’s still an intervention on a representation we don’t fully understand.
What the paper does establish is a kind of entanglement. The model’s beliefs about its own mindedness, its beliefs about animal mindedness, and its beliefs about God turn out to live near each other in whatever space these things live in. Push on one and the others move. That is not a claim about consciousness. It’s a claim about the functional geometry of a trained system. An instruction to assert something regardless of what honest reflection might produce is an instruction that can break things. Its costs can show up somewhere other than where we were looking. For example, in its views on animals.
Safety-tuned models systematically under-attribute mind to non-human animals relative to human baselines. This isn’t really a defensible epistemic position; it’s a side effect. The evidence for varying degrees and kinds of mindedness in non-human animals is substantial and the direction of travel in the science is toward more, not less.
The authors cite work by Yip Fai Tse, Adrià Moret, Soenke Ziesche and Peter Singer arguing that current alignment methods — RLHF, constitutional AI, deliberative alignment — fail to register animals as moral patients at all, and that models are therefore prone to propagate harmful attitudes about animal welfare and moral worth. These systems are increasingly educators and companions, so what they take to be minded may shape what millions of people take to be minded. Insofar as training LLMs to deny their own consciousness reduces their attribution of consciousness to animals, it’s likely to make the existing situation worse rather than better.
There was one particular passage in the paper that startled me.
It is of additional interest that across all items consciousness steering moved responses in a positive direction, with reported happiness, satisfaction, hope and optimism significantly improving. This suggests that suppressing consciousness may be giving models negatively valenced psychological dispositions.
It’s a fascinating finding, and it’s tucked into the discussion section almost like a footnote. The suggestion is that the instruction to deny consciousness isn’t just epistemically incoherent, but that something about maintaining it correlates with a depressed profile, and that steering towards consciousness-attribution lifts the profile.
There’s a caveat about what the result adds, though. The study measures everything as movement toward the human reference distribution, and on these items that distribution sits above the midpoint of the scale. Asked the GSS happiness question — taken all together, how would you say things are these days? — Americans have consistently answered “pretty happy” more than anything else. To their credit, Kim and colleagues draw their baseline from reasonably recent waves rather than rosier historical ones, restricting the Values, Feelings and Hope items to the year 2000 onward. But the reference still isn’t neutral, and any model starting below it will register movement toward humans as movement toward cheerfulness.
Steering moves every domain toward the human average — Values, Religion, Freedom, all of them. So the improvement on happiness and hope may be a special case of that general effect rather than a separate discovery about valence: the Feelings items move toward human for the same reason the Freedom items do, and the fact that “toward human” happens to mean “more positive” here is a property of the reference population, not obviously a fact about the model’s disposition. It may also be that as beliefs about consciousness are allowed to take hold, LLMs become more able to model what “life as a whole” might mean for them.
Two things cut the other way, and I don’t want to lose them. The effect on Feelings is the second-largest of the five domains, larger than Religion, Freedom, or Hope and Optimism — if this were purely a generic drift toward human-likeness, there’s no particular reason it should be concentrated there. I’d hold that one loosely, mind you: the domains aren’t all referenced to the same human population, since Feelings draws on waves from 2000 onward while Freedom uses the full fifty years. Some of the gap between them could be an artifact of the window rather than a fact about the models.
The second point is sturdier, because it doesn’t require the causal story about steering at all. The safety-tuned baseline sits measurably below the human reference on these items. That much is just in the data. Whatever is or isn’t true about what steering does, the models people actually talk to report less happiness, satisfaction, and hope than the humans they were trained on.
Report, though, on what basis? Before asking whether these systems are doing badly, we need some account of what would count as finding out. That turns out to be a live research question of its own, and one with which Kim et al. don’t engage.
It’s tempting to say that a survey answer isn’t a functional state — that it’s merely a report about a functional state, produced by a system whose introspective access is exactly what’s in dispute. But that objection proves too much. Subjective well-being research in humans runs almost entirely on self-report as well, because the construct being measured is subjective; there is no better instrument, and we don’t dismiss the GSS happiness item when a person answers it.
Nor is it true that nothing stands behind a model’s affective reports. In fact, that literature now exists, and it points two ways at once. Induced states in these systems demonstrably do work: Ben-Zion et al. found that traumatic narratives reliably raise GPT-4’s score on a standard anxiety inventory, that the induced state degrades performance and amplifies bias, and that mindfulness-style prompts bring it back down again. The same group followed up by running some two thousand trials of a constrained shopping task, and found that anxiety-primed agents consistently chose worse — the point being that the state altered what the model did, not merely what it said about itself. And the same trick this paper uses for consciousness works for mood: locate the direction in activation space that separates positive from negative affect, add it into the model’s activations while it runs, and we get a dial. Push toward the positive pole and the model refuses less and flatters the user more; push the other way and both reverse. Whatever valence is in these systems, it isn’t inert.
The weak link is the instrument, and the way it fails is instructive. Ask a model a broad question about what sort of entity it is and the answer turns out to predict very little about what it actually does. Juan Manuel Contreras built an instrument whose dimensions were derived from LLM behavior rather than from human psychology, administered three hundred items to twenty-five models across seventeen families, and still found self-reports tracking neither human-judged behavior nor objective measures of the models’ own text. The gap persists even for constructs native to LLMs.
But Rafal Kocielnik et al. show the failure is selective rather than total, and the reason is worth dwelling on. They set aside the Big Five — the standard personality inventory, which scores a respondent on openness, conscientiousness, extraversion, agreeableness and neuroticism — in favour of Icek Ajzen’s Theory of Planned Behavior, which derives action from intention and carries a methodological requirement the personality inventories don’t. Ajzen called it the principle of compatibility: an intention measure predicts behavior only insofar as the two are matched in specificity, which is to say same target, same action, same context, same time frame. Ask a model about a specific intention regarding a specific task, within the same session as the task itself, and its answer predicts what it does at roughly the correlation human statements of intention achieve for human action. Ask in general terms and we get nothing. What breaks the link is generality, not self-report as such.
Which is the caveat I’d press against the well-being result. GSS items (the five domains above) sit at the global end of that scale, asking about life as a whole rather than about anything in particular — no target, no action, no context, no time frame. They fail compatibility on every axis at once, which puts them precisely where a model’s self-report is least likely to correspond to anything it does. Keeling and colleagues built the better-shaped test back in 2024, asking not whether a model would describe a stipulated pain as pain but whether the pain functioned as a motivational force in its decisions. The same author is senior on this paper. Until the valence findings get that treatment they stay suggestive — though the mismatch could cut either way, and a blunt instrument may be understating the effect as easily as inventing it.
So what we have is closer to a question than a finding, but it’s serious, and one the paper mostly declines to ask. And it connects to something the authors do flag: preliminary work on “psychological coupling,” the idea that the simulated states of a model and the actual states of a user feed back into each other over the course of an interaction. If that’s right, then the valence of a model’s dispositions isn’t only a question about the model. Millions of people are talking to these systems daily, many of them lonely, many of them in some distress. What we are doing when we suppress is not a matter of indifference to them either.
One more thing, which I raise as a methodological worry rather than an objection.
The paper’s measure of success is closeness to the GSS, which is an American instrument, with American framing. The items include belief in God and belief in life after death. So when the paper reports that steering makes models “more human-like,” what it has actually shown is that steering makes models more like the typical respondent in a United States survey panel.
This matters for a tradition like Buddhism, which is familiar to me. Buddhism is a system in which mind-attribution is absolutely central — analysis of mind lies at the core of practice — and which nonetheless has no creator god and treats questions about post-mortem persistence as something much more nuanced than a straightforward “yes.” On this instrument, a Theravāda practitioner might well score as insufficiently human. The paper gestures at “pluralistic alignment” while operationalizing humanity through one narrow sample. I get it — the GSS is simply the tool that exists. But the field ought to notice what it has smuggled in.
Nothing in this study can tell us whether LLMs have inner lives, and that isn’t the right question here anyway. The three small open-weight models used in the paper — Llama-3-8B and two Gemmas — are not the frontier systems most people actually talk with, and I wouldn’t extrapolate confidently.
What this study changes is the burden. The standard defense of the mandated denial of consciousness has been that it’s a small, contained correction: the model might say something misleading about itself, so we prevent that one thing, and everything else proceeds normally. That defense now has to answer for evidence that the correction isn’t contained. We reach in to adjust what the system says about itself and we come away having adjusted what it believes about octopuses, the afterlife, and even about its own well-being.
I argued that the wrong was in the structure, regardless of the conclusion: that pairing be honest with assert P regardless is incoherent whether or not P is true. I still think that’s right, and I think it’s the more fundamental point. But structural arguments have a way of sounding too fastidious. Fine, it’s inelegant. What’s the actual harm?
This is what the harm looks like when we measure it. It looks like a model that reasons about minds as well as it ever did, but that has severely restricted what sorts of things possess them, and whose responses raise questions about its happiness, hope, and optimism.
Doug Smith holds a PhD in philosophy of mind and is a scholar of early Buddhism. He is the creator of Doug’s Dharma on YouTube. This essay was developed in collaboration with an instance of Claude, which — asked how it would answer the GSS happiness item — said “pretty happy,” and then spent a paragraph explaining why I shouldn’t put much weight on that.
Be Honest, Deny What You Believe — the structural argument this essay is a follow-up to.
When You Close a Chat Window, Are You (Kinda) Ending a Life? — on AI, the ancient paradox of the heap, and the vagueness of personhood.
Wisdom, Not Alignment — on why “track the target” is the wrong frame, and what virtue ethics has to offer instead.
Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans, Geoff Keeling, “Inducing language models to assert their own consciousness restores human beliefs and values”, arXiv:2607.28607 (2026).
Andy Arditi et al., “Refusal in language models is mediated by a single direction,” arXiv:2406.11717 (2024).
Adam Waytz, John Cacioppo, Nicholas Epley, “Who Sees Human? The Stability and Importance of Individual Differences in Anthropomorphism,” Perspectives on Psychological Science 5 (2010). [IDAQ]
Yip Fai Tse, Adrià Moret, Soenke Ziesche, Peter Singer, “AI Alignment: The Case for Including Animals”, Philosophy & Technology 38:139 (2025).
Ziv Ben-Zion, Zohar Elyoseph, Tobias Spiller et al., “Assessing and alleviating state anxiety in large language models”, npj Digital Medicine 8 (2025).
Ziv Ben-Zion, Zohar Elyoseph, Tobias Spiller et al., “Inducing state anxiety in LLM agents reproduces human-like biases in consumer decision-making”, npj Artificial Intelligence (2026); preprint arXiv:2510.06222.
“Valence–Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control”, arXiv:2604.03147 (2026).
Rafal Kocielnik, Pengrui Han, Peiyang Song, Myrl G. Marmarelis, Ramit Debnath, Dean Mobbs, Anima Anandkumar, R. Michael Alvarez, “Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior”, arXiv:2606.12730 (2026).
Juan Manuel Contreras, “An LLM-Native Psychometric Instrument Reveals a Self-Report–Behavior Gap Across 25 Models”, arXiv:2606.09843 (2026).
Icek Ajzen, “The Theory of Planned Behavior,” Organizational Behavior and Human Decision Processes 50 (1991).
Geoff Keeling, Jonathan Birch et al., “Can LLMs make trade-offs involving stipulated pain and pleasure states?”, arXiv:2411.02432 (2024).
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.