RSS Amplifier

Derek’s Substack · May 18, 2025

We are more likely to over-attribute digital consciousness

0
Sign in to vote or save

Derek Shiller · Derek’s Substack

The welfare of possible future digital beings is worth taking seriously. The progress we’ve seen in AI is bewildering and only likely to speed up. We don’t understand the systems we’ve built very well and there is little reason to think we’ll be less uncertain about future systems. AIs can already engage with human beings in ways that are basically indistinguishable from other humans. It is hard to provide clear non-question-begging explanations of why we should be confident that they lack anything worth calling a mind.

People like Rob Long, Jeff Sebo, and Kyle Fish have pushed for greater consideration of the needs of potential digital minds. They think that we ought to do more than engage in speculation or general hand-wringing: we need to have practical suggestions that can be implemented soon that will safeguard the possible well-being of AIs.

At the moment, I’m with those who think that over-attribution of minds is a much bigger concern than under-attribution. In other words, we’re likely to think we see consciousness or other morally relevant states where they don’t exist. If we’re not careful, practical near-term safeguards will encourage us to misread the cues for consciousness. It’s easy to see how a desire for implementable policies could, in the present environment, force us to look to text output as the gold standard of our assessments. Credulity now could be bad for digital beings in the long run if it distorts our thinking (or that of the public). Once we treat an entity as a moral subject — and particularly if we give it legal rights or political power — it may be hard to retract it without decisive evidence. And it may be difficult to develop more nuanced theories if they preclude entities that we have already decided merit protection.

There are a few reasons for thinking that in the short term, over-attribution is a much bigger risk than under-attribution. This post discusses some reasons why. The dangers of over-attribution are often downplayed or glossed as a matter of wasted resources. That framing undersells the issues, but I’ll save that discussion for another time.

We naturally see agency in the world and attribute chance events to the deliberate work of minds. In the past, people saw spirits or gods behind good or bad luck. We blamed crop failures on curses. We treated astronomical events as moral verdicts or warnings for the future, delivered to us by some celestial intelligence. This tendency has been attributed to an overactive theory of mind module that helps us understand our interactions with each other. We read thoughts and beliefs easily into observable patterns that they might help explain. AI will allow for mental interpretations, and so we’re likely to see them as having minds.

We’re not used to withholding attributions from things that act like us. There is a history of denials of moral relevance that have primed us to resist conservativism. Blake Lemoine cited this past in his case: “I think every person is entitled to representation. And I’d like to highlight something. The entire argument that goes, ‘It sounds like a person but it’s not a real person’ has been used many times in human history. It’s not new. And it never goes well. And I have yet to hear a single reason why this situation is any different than any of the prior ones.” It is easy to see how this argument will have influence.

Our willingness to see artificial minds is also well-attested by science fiction. We’re primed to see robot characters as genuine persons with real interests. It is a common trope of science fiction that robots can have feelings, and when they act like people, the reader is almost always invited to think of them as such. These feelings are often wrongfully dismissed by callous characters and championed by the protagonists. Readers aren’t normally invited to assume that the human-like robots display some clever engineering trick that allows them to mimic us. It is assumed that they act like us because they are like us.

We use psychological language when describing the text of current AI language models. We talk about AIs being deceptive or sycophantic as if they have communicative intents. We treat the text produced with their help as if it gives us important insights into what they think or what they want. We infer that they are more or less confident in the things they say and expect that if we ask them to explain what they are doing, they will give us a genuine reply.

This is not normally justified by the underlying setup. LLMs reflect both sides of the conversation equally, and there is no more reason to think that they identify with the assistant half. We read the ‘assistant’ part as if it alone represented their mental lives, rather than a part of a script they are mimicking. From the inside, it is indistinguishable to them whether they are receiving text or producing it. The fact that we read one side as representing their mind just shows how susceptible we are to the power of mental explanations. We don’t see any need to explain why the text we input is what it is, so we ignore any of the equivalent processing the LLM does for it.

You might think that we should be motivated to extend or withhold attributions on the basis of well-supported theories. We’re not likely to over-attribute unless our best theories do, and if they do, what right do we have to ignore them?

I’m deeply skeptical that theories can bear the weight we want to put on them. There are nearly as many theories of consciousness as there are theorists, and some can be used to make predictions about the possibility of consciousness in artificial systems. No theory is particularly well supported by the evidence and no theory has found the favor of a majority of the field. The most-celebrated theories in their classic formulations also don’t have direct applications to arbitrary artificial minds. They identify different things in humans that are said to be critical, but abstracting and generalizing so they can be applied more broadly requires further assumptions.

The empirical evidence is too coarse-grained: it is natural to want to look at what special thing is happening inside human brains that makes some states conscious and others not, but there will be many different ways of identifying the differences such that various systems will and will not count as conscious. Insofar as we want to say a certain system is conscious, we can probably find a theory consistent with the evidence that makes it so.

In general, people have more confidence about the things they think are or are not conscious than the theories that could justify those verdicts. They are likely to go with their guts to assess their theories than go with their theories against their gut feelings about attributions.

Suppose that a theory of consciousness predicted that dogs are not conscious. Should we still take it seriously? Some theories clearly predict that dogs are conscious. Some that they are not. Many leave it somewhat uncertain. People often react to these predictions by discounting the theories that predict canine unconsciousness, or by favoring variations that allow for canine consciousness. Why? I take it that there is some pressure to account for what the pre-theoretically think, and our pre-theoretical convictions are rather strong.

Most people are confident that dogs (and other pets) are conscious. Very few people can spell out a compelling argument why this should be. The natural assumption we might fall back to is that like causes produce like effects. However, we know that evolution may have crafted the same behavioral patterns in very different ways. If we were to discover that animal brains functioned in a completely different sort of way than human brains, would we still be licensed to infer that they are conscious from their behavior?

Very few people can explain the similarity of the relevant neural mechanisms (the role of the cortex in humans and dogs, the length of time we’ve spent on separate evolutionary branches, etc.) Instead, we see some important and suggestive similarities and take their behavior (and perhaps our relationships with them) as sufficient evidence. It is easy to see a mind behind their actions. We know what we would experience in their place. That seems to be enough.

Similar considerations apply to hypothetical aliens. What would we say about a biological extraterrestrial creature that had fundamentally different brain structures and processes. There might be abstract reasons to doubt that our best theories apply to them, but I can’t imagine that would hold anyone back from attributions of mental states, if they acted just like us. We take behavior to be a guide to feelings absent any clear reasons to expect it to correlate with specific structures in our brains.

The point is, few theories of consciousness are well-enough developed and clearly enough applied to decide one way or the other for dogs (or aliens, or AIs) to make a confident prediction. We’ve got vague guesses. Those shouldn’t license deep confidence.

If and when we develop relationships with AI, we should expect more of the same. If people have meaningful relationships with AI of the sort they have with their pets or their friends they will be inclined to judge them to be conscious, particularly if they act much like us and don’t give off any clear artificial tics.

AI companies can control model behavior, and through that, public perception. If they want, they can play up or down those aspects that make us more inclined to view the thing we’re interacting with as a fellow mind. They can change whether and when the systems are spontaneous, the kinds of mistakes they make, and so forth. They can give them coherent personas and memories or have them act like a more sophisticated Google search or autocomplete.

To some extent, the companies (OpenAI, Anthropic, Google) have crafted experiences that are suggestive. The interface with which we use a chatbot looks like the same interface we talk to each other. We give the model names and have the name represent the character in the dialogue the model constructs so that the Claude model predicts the ‘ai assistant’ whose words it generates will self-identify as Claude. (OpenAI, at least, has tried showing multiple generations: the kind of presentation that disrupts the anthropomorphizing presentation.

Interacting with a base model is rather different from interacting with a chat model through its interface. It is hard to experience a base model as a person. We are influenced by such experiences by certain AI traits that have little to do with mentality. The more AI systems remind us of computers, the less we are likely to attribute them minds. The fact that ChatGPT doesn’t initiate conversations seems like a critical barrier to granting them personhood. They talk to us about what we are interested in and may express desires but never follow up.

Other tics play important roles. The fact that they are servile, that their memories are poor, that they occasionally make mistakes that seem to us to be deeply stupid (such as the ‘strawberry’ test). If you were to get rid of these things, I think people would be much more disposed to see the AI systems as being substantial entities and be more prepared to judge them to be conscious. These things seem to have nothing to do with consciousness. Rather, they help play into other categories we already use.

Researchers at the companies may also have much better access to the systems, giving them additional authority in any pronouncements. Sophisticated language models currently are quite like one another, all being minor variations of the transformer architecture. That could change, particularly if we see an AI research explosion and any novel architectures might be treated like trade secrets. In that case, the only people with the relevant information about the models may be people with obligations to follow their companies’ marketing.

Overall, this might be likely to push in the direction of downplaying consciousness. So far, we’ve seen somewhat mixed signals. But it is easy to imagine that there would be incentives to sell lifelike companions and that could pressure certain competitors to build AIs that encourage this reading.

I’m somewhat skeptical that the major AI companies will continue to offer the best social chatbots. The AI technology for this purpose is already pretty achievable, so having the best-and-most-cutting-edge AI is unnecessary for a good product. The thing that sets competitors apart will be engagingness and cost-per-interaction. Cost may itself be negligible for most purposes in a few years. It isn’t clear why OpenAI or Anthropic or Google should have the edge in terms of selling social chatbots. And if there is fierce competition from smaller companies, it is conceivable that some of them will angle to build systems that are convincingly ‘real’.

I think that consciousness is unlikely to come along for the ride with intelligence or agency. I’m partly influenced by looking at the mechanisms underlying current LLMs and partly influenced by major theories of consciousness.

Current LLMs appear to produce intelligent behavior in a way that is very different from how we do it. We are organisms that need to maintain our bodies in a hostile world, and complex human intelligence is built out of the sensory and motor capacities of our simpler ancestors. LLMs generally don’t have very sophisticated sensory capacities and certainly aren’t built up from primitive sensory or motor foundations. They are kind of an add-on and the possibilities of a labelled training regimes make this viable. Given that they can do the same things we can do — math, poetry, etc. — with very different approaches, we should expect that they can do them without consciousness and that they can do those things even better without gaining conscious experiences.

Major theories of consciousness often posit something that seems to me to be unnecessary for the sophisticated behavior we want in our LLMs. If current LLMs can do what they do without a global workspace, why think that adding a workspace is the ideal way to get the most out of them? In response, you might think that the fact that we do things a certain way is evidence that it is the best way for it to be done. But we can really infer — at most — that the way we do things is the best way for us to do them given the constraints we face. AI face such different constraints and we shouldn’t expect them to do things the way we do.

Or consider attention schema theory. Is it helpful for an LLM to track its own attention? Why? There might be some use cases, but I submit most use-cases make it unnecessary. In general, LLM systems function well enough as is without doing this. Given that they don’t do it, it isn’t clear what future use case would necessitate it. Do they need to track their own attention in order to be good at math, or write bug-free code, or make a coherent movie? If we think so, we owe some story as to why.

If it is generally unnecessary to include a capability, it is inefficient to add in. If systems don’t need global workspaces or attention schemas, it would be a mistake to add them. And so I think it is likely that future AI systems won’t have them, or most of the other things that we might think are an architectural foundation for conscious experience.

One lesson from the success of LLMs is just how much is possible without consciousness. Many experts, even the most credulous, think that current LLMs are unlikely to be conscious. They can do all the things they can do because of their remarkable reasoning abilities. Those abilities do not require consciousness.

There are things that they can’t do that are barriers to wider adoption. They struggle with complex contexts or performing series of tasks to achieve goals in the real world. They have limits in their creativity and reasoning abilities.

AI companies are working on addressing these issues. How likely is it that the solution involves inducing consciousness? Unlikely, I think. For one, many of the things we think of as conscious can’t do these things either. Dogs are not obviously any better at reasoning or contextual thinking than cutting-edge LLMs. If they are conscious, they didn’t develop consciousness to that end and it hasn’t helped them in that way. Just as we got intelligence and sophisticated reasoning without consciousness, there is no obvious reasoning we can’t go a few steps further with more clever RAG frameworks or better RL training regimes. If the path we pursue is incremental, there is no clear reason that it should invoke consciousness.

On the other hand, consciousness might require changes or additions that are inefficient. Insofar as consciousness involves some computational structure, imposing that structure on top of one that already works is likely to just make things less efficient. If we don’t need a global workspace or an attention schema to do the things LLMs are tasked with doing, companies would have little incentive to include them but to able to use it in marketing.

Companies may accept some choices not driven by performance if they are striving to produce conscious systems. However, even so, we shouldn’t expect them to strive to hit a very high bar. We don’t know what it takes to be conscious: we could try to include some of the things that go in human brains, but we won’t know if we’ve succeeded. The public, I expect, will believe more or less what they want to believe.

No posts

Read the original on transitionalforms.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.