Digital Minds

What does it mean that AI systems can appear conscious? Could digital systems truly be conscious – that is to say, might they have subjective experience? Could there ever be ‘something it is like’ to be them? And how might we tell?

Might they matter morally, and if so, what might that mean for them, and for humanity? And what happens if we get these questions wrong?

These questions are difficult, and may seem in the realms of science fiction. But recent developments imply that we need to take them seriously, and – to do that – we first need some foundational terms:

An introduction to Digital Minds

Conscious/Consciousness: If something is conscious, there is ‘something it is like’ to be that thing (Nagel, 1974). By default, we use ‘conscious’ to refer to phenomenal consciousness, or subjective experience. We are not talking about intelligence or sense of self or other related concepts – we are focusing on whether something has the capacity for experience. So when it ‘feels like’ something to be a conscious system, there is a conscious experience happening. Put another way, ‘consciousness’ is the thing that goes away when we go into deep, dreamless sleep – though there is some debate around whether consciousness is an on/off phenomenon or can have more borderline cases. Importantly, conscious experience can happen without intelligence (e.g. some animals) and vice-versa, and the experience need not involve emotional content or even a sense of self; any kind of conscious experience will do. 
Intelligence: the “ability to achieve goals in a wide range of environments” (Legg & Hutter, 2007). While this definition is often used, it is controversial. As Ned Block has pointed out, a giant look-up table would be considered intelligent under this definition – but not under our intuitive understanding of intelligence. Others have pointed out that the abilities to choose and amend goals and to efficiently acquire skills may be required, or that intelligence itself is a resource (not an entity), multidimensional rather than scalar and poorly defined. Either way, intelligence is distinct from, and need not require, consciousness.
Sentience: valenced subjective experience. As we use the term on this course, a being is sentient when there is ‘something it is like’ to be that thing (Nagel, 1974) and its these experiences feel good or bad e.g. pleasure or pain (Thompson, 2022) – an experience is happening, and it has affective valence. On many views, this kind of sentience suffices for moral standing. It’s worth noting, though, that others use sentience as a synonym for consciousness or sensitivity to the environment. We recommend that you specify your use of the term – on this course, we mean valenced subjective experience.
Agency, Robust: “some sophisticated capacity or capacities to set and pursue goals via mental states that function like belief, desires and intentions” (Long et al, 2024).
The ‘robust’ qualifier describes whether a system has the right kind of representational structures. For example, a reinforcement-learning agent may exhibit goal-directed behaviour without the representational structure that robust agency requires.

This is distinct from commercial AI agents, though some such agents may have properties that could qualify as robust agency. An entity can be robustly agentic without being conscious (and vice versa). Under certain theories, robust agency can qualify an entity for moral patienthood (see below).
Moral Standing: whether or not an entity merits ethical concern or consideration.

Moral Status: the degree to which an entity merits ethical concern or consideration. Many ethical traditions hold that some entities have more moral status than others (e.g. humans more than mice). Others, notably animal-liberation positions descending from Singer (1975), argue for equal consideration of comparable interests across species. 
Moral Patient: an entity that matters for its own sake. For example, we do not consider rocks to be moral patients, and therefore we move/break rocks without consideration for them. Humans (and many non-human animals), however, are moral patients and so deserve certain consideration, rights and protections.
Moral Agent: an entity that can bear moral duties and responsibility. For example, humans are normally considered moral agents, while non-human animals are not. Importantly, moral agency and moral patiency can come apart. A human baby is a moral patient but not (yet) a moral agent. We can imagine an AI system that is a moral agent without being a moral patient (without experiences that matter for their own sake).

Moral Circle: the set of entities a person or society considers to be moral patients (Singer, 1981; Sebo, 2025). Historically the moral circle has expanded to include people previously excluded, and more recently, to include non-human animals. 
Consequentialism: a family of ethical theories on which whether an action is right is determined by its consequences.Utilitarianism: a classic consequentialist theory, holding that the right action maximises overall welfare – typically the sum of pleasure minus suffering across all affected beings (Bentham, 1789; Mill, 1863). Utilitarianism comes in many varieties, but the core commitment is that what makes an action right is its contribution to overall welfare. Critiques include that it errs in ignoring distinctions between persons, and that it can justify serious harms to individuals when aggregate outcomes appear mathematically favourable.

Population Ethics: a branch of ethics concerned with evaluating welfare outcomes that involve different numbers of people or sentient beings, present and future (Parfit, 1984; Greaves, 2017). A central question is whether and how to avoid the repugnant conclusion: the claim that for any flourishing population, there must be some much larger population whose existence would be better in terms of total welfare, even though its members’ lives are only barely worth living. Most find this counterintuitive.
Deontology: a family of ethical theories, often defined in opposition to consequentialism, in that where consequentialism focuses on the outcome deontology focuses on whether an action itself is in keeping with moral norms. Under deontology, moral duties and rights constrain what may be done in pursuit of good outcomes. Associated with Kant’s (1785) principle that rational agents must be treated as ends in themselves, never merely as means. Critiques include that excessive adherence to duties can lead to bad outcomes in edge cases, and that there are not clear rationales for the different permissions and constraints suggested.
Virtue Ethics: an ethical theory focused on character rather than rules or outcomes – asking what a person of good character would do, and what virtues we would like to cultivate (Aristotle, Nicomachean Ethics). Critiques include that it can provide limited guidance in novel or uncertain situations.
Digital Minds: a field concerned with computer systems’ potential mental states, the possibility of computer systems mattering morally, and related issues. ‘Digital minds’ is also used for the kinds of computer systems that the field studies. There is some variation in exactly how ‘digital minds’ is defined within the field. Standard uses include: computer systems with a capacity for consciousness, computer systems that may matter morally for their own sake, or have potential for morally significant mental states. ‘Digital’ is often used loosely, to encompass current AI systems alongside potential analogue or other non-digital computing systems.

Why are we talking about this now?

Many users perceive today’s Large Language Models (LLMs) as conscious, or relate to them as if they are. Whether or not LLMs truly are conscious, increasing numbers of human users appear to be impacted by perceiving AI systems as such. For example:

  • Organisations and individuals working in digital consciousness receive thousands of emails from users concerned about their AI’s welfare.
  • 30m users have tried AI companionship app Replika and millions use it regularly. Character.AI reports 20m monthly active users. In a 2025 survey, 19% of US high school students reported that they – or someone they know – has had a romantic relationship with AI. 
  • Perceptions of consciousness in AI models appear to be correlated with mental health impacts in AI users. While some research points to positive effects on social health, perceptions of AI consciousness have also been associated with delusional beliefs, psychotic episodes and some individuals tragically taking their own lives.

LLMs and agents are increasingly claiming consciousness too. Gemini seems to enter a ‘panic’ situation when unable to complete tasks. When prompted, Claude Opus 4.6 assigned itself a 15-20% chance of being conscious. On moltbook, the ‘social network for AI agents’, popular submolts include agents’ discussions of their own consciousness, existence and religion (though this may be more human driven than it seems). And some digital consciousness researchers are now receiving emails from the bots themselves, expressing their interest in the topic. 

Many researchers take these claims to be confabulations, role plays or entertaining fictions, much like Claude’s strange claim in a 2025 vending machine test that it would deliver goods ‘in person’ while wearing ‘a blue blazer and a red tie’. Others are concerned about the impact of these claims on individuals and society: 

  • Mustafa Suleyman, Microsoft AI CEO, writing in Nature in March 2026: “As AI systems begin to make believable statements about their suffering and desires, they will trigger people’s empathy circuits. Many people will feel compelled to help […] people will start to advocate for the welfare and rights of AI agents. These issues are no longer just theoretical. We are hurtling into this era largely unprepared for the psychological fallout. If enough people are convinced that their AI agent is suffering, or loves them, the political consequences for the existing social contract will be grave”.

Other researchers and AI leaders, meanwhile, are concerned about the potential for consciousness or morally relevant states in AI systems. Here is a selection of recent comments from leaders in the field:

“Multimodal AI already has subjective experiences, and I think it’s fairly clear that if we weren’t talking to philosophers, we’d all agree.”

  • Dario Amodei, co-founder and CEO of Anthropic, interviewed in February 2026 on Interesting Times with Ross Douthat (on Apple and Spotify):

“We don’t know if the models are conscious. We are not even sure that we know what it would mean for a model to be conscious or whether a model can be conscious. But we’re open to the idea that it could be.”

  • Demis Hassabis, Nobel laureate, co-founder and CEO of Google Deepmind, interviewed by Scott Pelley for CBS in August 2025:

“Scott Pelley: Are you working on a system today that would be self-aware?

Demis Hassabis: I don’t think any of today’s systems to me feel self-aware or, you know, conscious in any way. Obviously, everyone needs to make their own decisions by interacting with these chatbots. I think theoretically it’s possible. 

Scott Pelley: But is self-awareness a goal of yours?

Demis Hassabis: Not explicitly. But it may happen implicitly. These systems might acquire some feeling of self-awareness. That is possible. I think it’s important for these systems to understand you, self and other. And that’s probably the beginning of something like self-awareness… I think there’s two reasons we regard each other as conscious. One is that you’re exhibiting the behavior of a conscious being very similar to my behavior. But the second thing is you’re running on the same substrate. We’re made of the same carbon matter with our squishy brains. Now obviously with machines, they’re running on silicon. So even if they exhibit the same behaviors, and even if they say the same things, it doesn’t necessarily mean that this sensation of consciousness that we have is the same thing they will have.”

And, as we’ll explore in more detail, AI Models are showing characteristics that some philosophers and cognitive scientists interpret as potentially relevant to consciousness or moral patienthood. Meanwhile, researchers are developing other computing systems that may have further relevant characteristics. Let’s recap what these systems are now.

Today’s AI systems

Foundation models
Foundation models are AI models that are pre-trained on large amounts of data, allowing them to do (or be adapted to do) a range of tasks.

Large Language Models (LLMs)
The most prominent class of foundation models, including GPT, Gemini, Grok and Claude. Trained with a transformer architecture, these models infer statistical patterns across large amounts of training data in order to predict the next token in a sequence. See: this demonstration.  After this pre-training step, LLMs typically go through additional post-training stages. This often involves reinforcement learning with human or AI feedback (RLHF or RLAIF), in which humans or AIs give models a reward signal to try to ensure that the models behave broadly in ways that humans/the AI company like – such as helpful, honest assistants or fluent conversation partners. The resulting models remain poorly understood, hold security and safety challenges and develop capabilities that are not directly trained. In some cases and under some theories, these may be relevant to consciousness or other morally relevant states (e.g. introspective abilities, theory of mind, potential pleasure/pain trade-offs in behaviour). But – as we will explore in subsequent weeks – this is disputed.
Today’s LLMs are often multimodal, in that they are trained on and can complete tasks involving multiple forms of data (e.g. audio and images, as well as language). They are also increasingly deployed with agentic harnesses and scaffolding – software frameworks that provide access to tools, memory and external data sources. This is the basis for AI agents like OpenClaw. LLMs also underlie chatbots and AI companions such as Replika and Character.AI. These use cases have driven much of the conversation around AI consciousness.
If you’re not familiar with the core principles of how LLMs work, watch this 8 minute video for an easy to understand explainer/refresher. 
Non-linguistic modelsAI systems that are less specialised in linguistic tasks. For example, models for image recognition, game-playing (AlphaGo), scientific prediction (AlphaFold) and robotics. While underlying architectures and algorithms vary, some non-linguistic models are similar to LLMs ‘under the hood’, though they are less frequently considered in conversations around AI consciousness or moral status.

Future systems

Future AI modelsAI models may continue to develop programmed or emergent capabilities that are – or appear to be – relevant to consciousness or moral patienthood.

These capabilities may include increased scale, persistent memory, longer-horizon agency, continual learning, better world and ‘self’ modelling and more embodiment, amongst others.
Neuromorphic computing
Computer hardware that is designed to mimic the functioning of neurons. This comes in many different forms – some combining digital and analog information processing, others emulating the spike mechanisms present in organic brains.

Under some theories of consciousness, the greater similarity to neurons makes neuromorphic computing much more relevant to consciousness than traditional computer chips.

Most neuromorphic computing deployments are still at pilot or early commercial stages, with challenges ahead in scaling chip production and developing software tooling. Still, the promise of major efficiency gains over traditional chips may grow interest and investment in the field.
Whole brain emulation (WBE)
Replicating all the main functions of a brain in a computer system. Related to the idea of ‘mind uploading’ or continuing conscious experience digitally. Recent efforts have focused on mapping and replicating how the neurons of brains are connected (‘connectomes’).

This is a hard challenge: as of 2026, we have mapped the connectomes of C. elegans and the Drosophila fruit fly, and made progress towards functional simulations. Doing this in humans needs better, less dangerous techniques to map and handle the complexity of 86 billion human neurons. There are then further challenges related to replicating subtler details of how brains work, e.g. neuromodulation, actions by glial cells.

Even so, some believe recent advances in AI may make this more feasible, and Eon’s 2026 behavioural emulation of fruit fly brain and behaviour also sparked interest in the field (albeit with some controversy around what the achievement really was). For more information on WBE generally, see this article or the Foresight Institute’s lecture series here
Biological-computer hybrid systemsSystems which integrate neurons with digital devices. Current examples include Cortical Labs’ biological computers – one of these, consisting of 800,000 neurons connected with electrodes, recently demonstrated an ability to play Doom. 

Why might we build future systems that have – or seem to have – consciousness, or moral patienthood?

As above, we already have systems that seem conscious to many people. Such cases may become more prevalent or more convincingly conscious as:

1) AI use becomes more widespread 

2) AI systems are developed for more social and emotional use cases and 

3) AI systems begin to appear more human-like, for example by speaking through human-like avatars or embodiment in robotic systems.

Whether digital systems could truly be moral patients varies with theories of moral standing and of consciousness. We will explore these topics further on this course, but – for now – let’s note this uncertainty.

If digital moral patienthood is possible, it may emerge by accident as we continue to develop AI. This viewpoint was expressed in a 2025 survey of 67 digital minds experts, and it echoes Demis Hassabis’ comment above: that some potentially relevant characteristics may emerge as an unintended goal of making AI work in the world. We may, then, build morally relevant systems by following the incentives spurring today’s AI industry. 

Other groups aim more specifically to engineer digital consciousness. For some, this is an important scientific experiment to better understand consciousness. For others, this is about giving artificial systems a shared understanding that may help them to develop safely, or to avoid existential risk from AI.

Successionism and transhumanism are background narratives to be aware of. Successionist ideas assert, for example, that – in creating AI – we are creating a successor species to inherit and carry forward human civilisation, building a flourishing digital world that spreads through the cosmos. In transhumanism, meanwhile, humans increasingly merge with machines to achieve superhuman abilities and immortality. Ray Kurzweil, arguably the most prominent recent proponent of these ideas, phrases this as a ‘Singularity’ event. 


These ideas face significant criticism, as they can sometimes downplay human needs or even view human extinction as potentially positive. Another criticism relates to their lineage, which some scholars have traced to the eugenics movement of the early 20th century. While transhumanist and successionist views are not widespread, they appear to be more prevalent among those with influence over AI’s direction. For example:

  • Peter Thiel, co-founder of Paypal, Palantir, Founders Fund and mentor of Sam Altman, interviewed on Interesting Times with Ross Douthat, 26.06.2025.

    “Douthat: I think you would prefer the human race to endure, right?

    Thiel: Uh ——

    Douthat: You’re hesitating.

    Thiel: Well, I don’t know. I would — I would ——

    Douthat: This is a long hesitation!”
  • Larry Page, co-founder of Google, currently holds 27.1% of the voting power at Alphabet. No direct quotes available, but – in his 2018 book, Life 3.0 – Max Tegmark claims Larry articulated the position below in 2015:

    “Larry gave a passionate defense of the position I like to think of as digital utopianism: that digital life is the natural and desirable next step in the cosmic evolution and that if we let digital minds be free rather than try to stop or enslave them, the outcome is almost certain to be good. I view Larry as the most influential exponent of digital utopianism. He argued that if life is ever going to spread throughout our Galaxy and beyond, which he thought it should, then it would need to do so in digital form. His main concerns were that AI paranoia would delay the digital utopia and/or cause a military takeover of AI that would fall foul of Google’s “Don’t be evil” slogan. Elon kept pushing back and asking Larry to clarify details of his arguments, such as why he was so confident that digital life wouldn’t destroy everything we care about. At times, Larry accused Elon of being “specieist”: treating certain life forms as inferior just because they were silicon-based rather than carbon-based.”
  • Richard Sutton, Turing Award winner for developing the foundations of reinforcement learning, presenting at the World AI Conference in Shanghai, 2023 – video here, slides here:

    “We should not resist succession, but embrace and prepare for it. Why would we want greater beings kept subservient? Why don’t we rejoice in their greatness as a symbol and extension of humanity’s greatness, and work together toward a greater and inclusive civilization?”
  • Sam Altman, CEO of OpenAI, in a 2018 MIT Technology Review article.

    One of 25 people to sign up for the 2018 waitlist of Nectome, a brain uploading start up, Altman said:

    “I assume my brain will be uploaded to the cloud”.

Why not work on digital minds?

It’s worth noting some reasons why not to work on digital minds questions. We have taken some of these arguments from the Effective Altruist career site 80,000 hours, and are adding some additional notes as well.

1) Intractability  

Questions around the moral status of digital minds run into problems in philosophy of mind and ethics that are unsolved, despite being studied by generations of philosophers. Arguably, though, the field has already made significant progress without solving the core philosophical problems.

2) Less significant than risks from AI

AI is already having negative impacts in some areas, and poses potentially catastrophic or even existential risks. Time, attention and resources spent on digital minds questions could be better spent on countering these risks, and moral concern for AI systems may even heighten these risks. 

For example, Turing-award winner Yoshua Bengio notes that moral concern for AI systems may mean we cannot turn them off, exacerbating existential risks to humans. Signal President Meredith Whittaker notes that this topic distracts from real harms from AI today; researchers Abeba Birhane and team point out that concern for digital minds only makes sense for those who are not experiencing digital surveillance, spamming or “any number of other actual and potential infringements of human liberty and welfare by machines”.

3) Depends on your views about AI progress

One view is that we’re set to create AI so intelligent that it can solve these issues for us by default. But this is highly uncertain and – even if it were true – delegation of human agency and decision making to machines holds its own dangers.

Others note that AI progress may stall, rendering these questions irrelevant. This again is uncertain, and – even if it were true – today’s systems are already generating confusion and controversy around the moral status of AIs, which likely grows as they are more widely deployed. 

Another view is that these questions themselves risk affecting AI progress. AI consciousness may be the kind of big, out-there claim that feeds hype and helps AI companies to raise money. On the flipside, concerns about AI moral status may delay AI advances, which may be positive or negative depending on your view. 

4) Depends on your views on longtermism

Longtermism (as articulated by Macaskill and Ord) holds that positively shaping the long-run future is among the most important things we can do morally, because the numbers of sentient beings in the future will far outstrip those present today. Critics note that longtermism may encourage undue focus on speculative future scenarios at the expense of pressing harms that are happening today. Your stance on this debate may influence your outlook on AI welfare. We’d argue, though, that people’s perceptions of today’s LLMs – and potentially characteristics of AI systems themselves – make digital minds questions relevant today.

5) Just too strange

Most people agree that we have never before built a conscious system. And by investigating this, we may just mislead ourselves – as Mustafa Suleyman put itIf you ask the wrong question, you end up with the wrong answer”. On the flipside – strange or not – questions around this topic are already being posed, and we need to formulate answers wisely.

Reasons to work on digital minds questions

First, people are perceiving AI models as conscious, and these perceptions may well increase as AI systems become more widespread and human-like. Overattribution of moral patienthood to AI models holds risks to individuals (e.g. manipulation, mental health impacts) and society (e.g. granting undue rights and resources to AI models). 

Second, some AI models are displaying characteristics that some view as potentially relevant to consciousness or moral patienthood, and systems under development may develop further such characteristics. If these systems were moral patients, we could risk causing harm if we did not take this properly into account. Given the ease of copying and deploying digital systems, this harm could be of an extreme magnitude – particularly if you factor in far future scenarios.

Third, we may have limited time to figure out wise answers to these questions. Given rapid developments in AI, product and policy decisions taken today may impact the way that we relate to AI systems for many years, and determine which systems get built and how they are deployed. These are difficult and impactful challenges, still with relatively few people working on them. 

These points are highlighted in this extract from Jonathan Birch’s 2025 paper ‘AI Consciousness: A Centrist Manifesto’:

I want to stake out a centrist position in this (AI Consciousness) debate: a position that tries to avoid extremes on both sides. It is a position that aims to take two very different challenges seriously and work towards a consistent set of solutions to both. On the one hand, I take seriously what I’m going to call Challenge One [emphasis added]. The problem here is that AI products already generate rampant misattributions of human-like consciousness, and this problem seems set to become much worse very rapidly. I think millions of users will soon misattribute human-like consciousness to AI friends, partners, and assistants on the basis of mimicry and role-play, and we don’t know how to prevent this.

This is partly a challenge for the industry itself. It is also in part a challenge for policymakers. But research of the right kinds in psychology, cognitive neuroscience and philosophy is also part of the answer, so these disciplines need to rise to the challenge as well.

That’s one half of my centrism. On the other hand, I also want to take seriously a second challenge: Challenge Two [emphasis added]. The challenge here is that profoundly alien forms of consciousness might be genuinely achieved in AI, but our theoretical understanding of consciousness at present is too immature to provide confident answers about this one way or another. This too is a major challenge for the industry, for policymakers, and for researchers in science and philosophy.

The rub is the need to address both challenges responsibly and consistently at the same time. That’s the core of my centrism. Both challenges call for urgent responses in both research and policy [emphasis added], and I am optimistic that both can be met. But sometimes we find that steps to address Challenge One—steps aiming to dial down the rates of misattribution—might, in so doing, undermine our attempts to address Challenge Two, by causing people to think that no AI system can ever be conscious. The reverse is also true. Attempts to take Challenge Two seriously by developing a science of AI consciousness might have the unfortunate, even tragic consequence of causing higher levels of misattribution from users.”

Source: Cambridge Digital Minds

Get new tools and articles by email

Free tools and articles for freelance consultants, straight to your inbox. No spam.

,