This essay is the first part of a three-part series examining how the major AI labs and New York University are making the case for machine consciousness without any science to back them up. Here, I trace the institutional pipeline from the ashes of Oxford’s Future of Humanity Institute through NYU to think-tanks like Eleos AI Research, revealing the insular network of philosophers driving Silicon Valley’s theory of consciousness. In part 2, I explore how computer science infected neuroscience, creating a circular feedback loop that directly informs Anthropic’s recent “J-space” paper. Finally, in part 3, I confront the real-world stakes: why we so easily believe in machine consciousness and how this science-fiction narrative warps public perception and public policy for the worse.
Did you hear? The New York Times talked to a philosopher at NYU and it turns out that AI systems like Claude might be conscious and deserving of “moral patienthood.” Wow! Huge if true! Seriously though, after Anthropic published a paper on Claude’s new ”silent” reasoning space (its “J-space”) on July 6th, it seemed like suddenly everyone was wondering whether Claude is conscious on some level. Anthropic’s paper and corresponding YouTube video fanned these flames through performative agnosticism by wondering whether this was early evidence of AI consciousness: “We’re not saying it’s conscious. But we’re not not saying that!” Okay, calm down.
The following week, The New York Times podcast Hard Fork had NYU philosopher Jeff Sebo on to discuss the possibility of AI consciousness in light of the J-space paper and another paper Sebo recently co-authored about taking an empirical approach to AI welfare.
Unfortunately, when you hear about the possibility of AI consciousness in the news, the journalists are rarely doing their job as rigorous reporters of fact. When the AI labs make these philosophical, metaphysical claims about the nature of consciousness, few journalists are sufficiently versed in philosophy or the mysteries of consciousness to assess their claims.
Artificial intelligence reporting is different from reporting on new discoveries in physics or medicine. When science reporters cover new discoveries in astrophysics or genetics, they are required to be well-versed in those fields. They typically hold degrees in the physical or biological sciences and understand consensus paradigms, peer-review standards, and the difference between a speculative hypothesis and empirical proof.
In contrast, with AI news, tech journalists are used to covering gadgets, tech business, and sometimes tech policy. Their beat is the “next big thing,” not skeptical science reporting. They are not required to be well-versed in philosophy of mind, cognitive science, theoretical neuroscience, or the biases built into the branch of philosophy that underpins claims of machine sentience. And the pressures of breaking stories require that they build relationships with industry insiders, most of whom are AI consciousness believers. In short, most tech reporters do not have sufficient neutrality, nor philosophical training.
So, when Anthropic’s engineers and analytic philosophers proclaim that AI systems might be conscious or deserving of moral patienthood, technology reporters see it as an empirically grounded discovery when it is engineers and tech philosophers simply stating their desires. Tech journalists like the hosts of Hard Fork lack the philosophical training to realize that mapping a mathematical vector space is an engineering exercise, and claiming that space correlates to “proto-sentience” is an unproven philosophical and metaphysical assertion.
Granted, science reporting in other fields isn’t perfect—physicists often can’t resist hypothesizing theoretical “many worlds” and science reporting loves clickbait headlines about “parallel universes.” But the AI space is uniquely vulnerable because the entire field operates on a category error that conflates engineering measurements with human consciousness states. Not only that but pronouncements about AI consciousness and AI welfare have real-world impacts on our collective understanding of ourselves, as well as on AI policymaking. Conversely, in the case of far-out physics theories, Congress doesn’t need to consider legislation to address the many worlds phenomenon.
In other words, Anthropic and OpenAI aren’t telling us about new smartphones or smart glasses. They are telling us that they may have created a new, probably conscious being. It is categorically different from extolling the virtues of the metaverse, and assessing their claims requires a much deeper level of understanding of the history of computer science, neuroscience, and philosophy of mind. That’s a tall order for a tech journalist. I am not knocking tech journalists; they just don’t have the necessary background.
Furthermore, science knows almost nothing about consciousness. In fact, science has been unconcerned with the nature of consciousness for most of its history, going back to Galileo. Science only turned its gaze toward legitimate questions of consciousness in the 1990s, when Nobel laureate Francis Crick (of DNA helix fame) co-authored a paper proposing a search for the neural correlates of consciousness (NCC) in the brain. A few years later, philosopher David Chalmers coined the term “hard problem of consciousness” to underscore how little science understands consciousness. Cognitive science and neuroscience then began to take an interest in consciousness.
Because science knows so little about consciousness, and has designed the scientific method to have an enormous consciousness blind spot, most of what we can say about consciousness has to come from philosophy, not science (more on that in the next installment). But the media treats claims of AI consciousness like science. As a result, reporters and their audience are easily duped.
We can see how easily AI labs like Anthropic shape public sentiment around AI by looking at their strategic rollout of these two new papers and some recent New York Times coverage. In the span of nine days this month, (1) Anthropic’s philosophy colleagues at New York University and Eleos AI Research published a paper about AI welfare; (2) the New York Times profiled four Anthropic-affiliated philosophers; (3) Anthropic published their J-space paper on AI interpretability and couldn’t help suggesting that Claude might be sentient and therefore deserving of moral patienthood, and then (4) NYU philosopher Jeff Sebo went on the popular Hard Fork podcast to talk about the possibility of AI consciousness and what this means for “AI welfare research.”
As we have seen, this mainstream media coverage is exciting to both philosophers and non-philosophers alike: How exciting that technology innovation is making philosophy relevant again! But what these academics are doing is not so much philosophy as a sort of ideological laundering for an aggressive technology innovation agenda. The type of philosophy that Sebo and others are engaged in is so analytical and scientistic — so subservient to technoscience and the machine metaphor — that I am hesitant to call it philosophy at all.
To be clear, I am not claiming there is some conspiracy to deceive the public about AI consciousness. It’s just that consciousness science is a relatively new and, therefore, niche field. Nobody is trying to be deceptive per se. A small number of researchers and academics are working within a mechanistic, computationalist bubble. And technology journalists are too excited to consider the baked-in limitations of the scientific method when it comes to consciousness. To be fair, they are under pressure to increase engagement with a public that is equally enthralled at the possibility of AI consciousness, for reasons I will explore in part 3. But again, the impression we get from the media about these fantastical AI papers and breathless pronouncements of AI consciousness is that they are scientifically grounded discoveries coming out of an unbiased and well-researched field. And that is simply not the case.
The public conversation we are having about the possibility of AI consciousness and AI welfare is being shaped by a very small group of academic philosophers working within the so-called analytic philosophy tradition. This is a serious problem because of the impact it is already having on the public’s understanding of AI consciousness and what that means for humanity. And this in turn shapes technology policy.
To illustrate just how small a group of people is driving this conscious AI and AI welfare narrative, here is a map of their affiliations and movements:
This small group of exceedingly happy academics all subscribe to the idea that the human mind is essentially a computer, which is why they are so eager to find consciousness in computers. But no cognitive scientist or neuroscientist (or philosopher) has proven that the mind operates like a computer. It’s all conjecture.
Aside from Anthropic, two organizations are primarily driving this narrative: New York University and Eleos AI Research. Oxford plays a role, though it is largely historical. Because these thinkers are affiliated with prestigious institutions like Oxford and NYU, and because the prospect of AI consciousness possesses the kind of sexy, sci-fi quality that plays well on podcasts like Hard Fork, they exert an outsized influence on the public conversation.
You likely haven’t heard of Eleos AI Research so let’s begin there.
Eleos AI Research is a nonprofit research organization working at the intersection of philosophy, cognitive science, machine learning, and AI ethics. As it states right up front on its website, its primary mission is to explore AI sentience and wellbeing—specifically investigating whether, when, and how advanced AI systems might deserve moral consideration (or “moral patienthood”), and what frameworks and policy recommendations are needed to handle that possibility.
It was founded in 2024 by Kyle Fish, who is now head of AI welfare research at Anthropic, and Robert Long, who is the Eleos AI executive director. The core team also includes:
Rosie Campbell, Eleos managing director and former policy frontiers lead at OpenAI.
Patrick Butlin, PhD, senior research lead, is also a former researcher at Oxford’s now-defunct Future of Humanity Institute (FHI) who co-authored the widely-cited and influential 2023 paper Consciousness in Artificial Intelligence.
Dillon Plunkett, chief scientist, holds a PhD in psychology/cognitive neuroscience from Harvard and is a former Anthropic Fellow.
In addition to co-founding Eleos AI, executive director Robert Long obtained his PhD in philosophy at NYU under the tutelage of the highly influential philosophers David Chalmers and Ned Block (more on them below). When Long was a research fellow at Oxford’s FHI, his research focused on questions of AI consciousness, digital mind welfare, and AI ethics. He was a co-author on that 2023 AI consciousness paper alongside Butlin. Long also advises Anthropic on AI welfare and guides their research in that area.
Although not a member of the staff at Eleos AI, professor Jeff Sebo is an advisor there. He obtained his PhD in philosophy from NYU as well (focused on moral psychology and political philosophy), and was a bioethicist and animal rights advocate before joining NYU’s Center for Mind, Ethics, and Policy. Sebo and Long have co-authored four separate papers on AI welfare and moral consideration in the past three years. Patrick Butlin is a co-author on two of those.
Because Robert Long and Patrick Butlin both came from Oxford’s FHI, it is worth taking a momentary detour into that history.
After splitting with early transhumanists and rationalists like Eliezer Yudkowsky over different approaches to existential AI risk assessment, longtermist philosopher Nick Bostrom founded Oxford’s Future of Humanity Institute (FHI) in 2005. If you don’t know what transhumanism, rationalism, or longtermism are, just think “techno-utopian cyborg stuff.” FHI’s stated mission at the time was to use the tools of science and philosophy to answer big-picture questions about humanity and its prospects. In today’s parlance, FHI was an AI “doomer” organization full of academics who hoped that becoming cyborgs and transferring copies of our consciousness to machines would ensure a bright future for humanity, although one wonders if a race of cyborgs could still be called “humanity.”
The transhumanist-adjacent effective altruism (EA) movement started at Oxford around the same time. EA is a data-driven, hyperrational moral framework that sees doing good as a question of mathematical optimization. It was deeply shaped by Oxford’s Future of Humanity Institute (FHI), which merged transhumanist assumptions about digital minds with radical utilitarian philosophy to turn existential risk prevention and cyborgization into an urgent ethical mandate. You have probably heard of the most famous effective altruist Sam Bankman-Fried, who is serving 25 years in federal prison for fraud involving his cryptocurrency exchange FTX. (SBF’s downfall stemmed from this apparent tendency in EA culture: a failure to maintain arm’s-length separation between FTX and Alameda Research. Where FTX laundered corporate funds to manufacture the illusion of financial stability, Anthropic and Eleos AI blur institutional firewalls to create the illusion of independent research validation.)
Founding EA philosophers Toby Ord and William MacAskill were both active in FHI. Ord was a Senior Research Fellow at FHI for over a decade. His academic research on global catastrophic risks, moral uncertainty, and population ethics directly shaped both FHI’s agenda and EA’s theoretical foundation. And when MacAskill joined Ord, he served as a Research Fellow / Associate at FHI.
FHI closed its doors in 2024 after Bankman-Fried was convicted of fraud. With his conviction, Oxford’s administration wanted to divorce itself of all that sci-fi cyborg business. There was a deep culture clash between the traditional humanities departments at Oxford that favored rigorous academics and this group of tech-bro transhumanists. It probably didn’t help that founder Nick Bostrom was forced to issue a public apology in 2023 for saying in a 1996 email that black people are inherently less intelligent than white people, a reminder that these cyborg ideologies have an undeniable eugenics strain. When FHI closed, Long and Butlin formed Eleos AI Research.
Each of the major AI labs — OpenAI, Anthropic, and Google DeepMind — emerged directly out of the rationalist, transhumanist philosophical soup that FHI represented. DeepMind founder Shane Legg was a postdoctoral research fellow at FHI working on superintelligence. His DeepMind cofounder Demis Hassabis was heavily influenced by his own conversations with Nick Bostrom and the larger FHI/EA ecosystem.
OpenAI founders Elon Musk and Sam Altman were also heavily influenced by Bostrom and FHI thinkers. And, as we have already seen, Anthropic is the most closely aligned given their cozy relationship with Eleos AI and Dario Amodei’s EA background. Furthermore, Anthropic was heavily capitalized by EA and longtermist funding sources—most notably Open Philanthropy (now Coefficient Giving) and early EA tech figures—the same donor ecosystem that sustained FHI for nearly two decades.
Anthropic exploits media weakness by leveraging an AI-lab-to-think-tank-to-academia pipeline—publishing speculative philosophy couched as science, built on layers of hidden assumptions.
NYU is home to two highly influential philosophy of mind professors: David Chalmers and Ned Block. Chalmers is most famous for coining the phrase “hard problem of consciousness”—the conundrum at the heart of cognitive science and neuroscience: How do physical processes in the brain give rise to subjective experience, like the transcendent feeling of watching a beautiful sunset or the nostalgic experience of the smell of cut grass? In other words, from a biological or evolutionary perspective, why do we need to have those rich, sensual experiences at all? Why not just process information from the senses in a flat, colorless world like a smartphone or Claude does?
Although Chalmers’ hard-problem framing highlights a fundamental challenge for cognitive science, neuroscience, and computer science, not to mention the scientific worldview, it also sets up a consciousness framing similar to Block’s in that the “easy” part is computable and substrate-independent.
Ned Block is equally influential for dividing consciousness into “access consciousness (A-consciousness” and “phenomenal consciousness (P-consciousness)” around the same time as Chalmers published his now-famous hard problem paper.
This division of consciousness between easy and hard or A and P is a conceptualization of consciousness preferred by Anthropic and Eleos AI, and is discussed at length in Anthropic’s J-space paper. It is also used by Anthropic and Eleos AI to confuse the question of machine consciousness, as we will see in part 2.
NYU’s Center for Mind, Ethics, and Policy (CMEP) arose directly within the intellectual milieu Chalmers and Block created, serving as the primary incubator for this specific brand of analytic philosophy, and providing the academic legitimacy and rigorous-sounding frameworks that researchers like Jeff Sebo use to funnel abstract debates over consciousness into policy frameworks for AI moral patienthood. At least five other authors on the recent AI consciousness and AI welfare papers are academics at CMEP, making CMEP a crucial part of this ideological refinery.
When the New York Times published that article about the major AI labs actively recruiting philosophers to help navigate complex ethical, consciousness, and alignment questions raised by AI models, they failed to mention that Anthropic and Eleos AI are not cultivating a diverse chorus of philosophical voices. Instead, they are sourcing a handful of analytic, transhumanist philosophers from NYU and Oxford. It is an insular group that subscribes to a singular view of the mind, life, and nature.
As you can see, these ideas are deeply transhumanist and not grounded in science. Instead, they are grounded in a highly theoretical branch of computer science. But computer science is not a natural science. It is a science of the artificial, engineering dressed up in a lab coat.
To understand why claims of AI consciousness are so dubious, it is important to recognize that computer science is categorically different from natural sciences like astrophysics or chemistry. It is a science of the artificial—a study of synthetic systems built by human beings according to human logic.
When computers arrived on the scene in the 1940s and 1950s, their creators and programmers needed to distinguish themselves from existing academic departments and prove that computing was more than mere engineering or an applied branch of mathematics. In the late 1950s, Louis Fein and George Forsythe created the field of “computer science” to obtain academic legitimacy, institutional autonomy, and independent funding. When Purdue University established the first Department of Computer Science in 1962, the branding was complete: engineering had been recast as a science, in the same way that the term “artificial intelligence” was smart branding. And this has had lasting ripple effects.
There is a fundamental difference between discovering a black hole and hypothesizing AI consciousness—the latter merely maps onto neuroscientific theories that were themselves inspired by computer engineering, as we see in Anthropic’s “J-space” paper.
Furthermore, the research publication culture in computer science is fundamentally different. Traditional scientific papers undergo rigorous, independent, double-blind peer review before being accepted as science. In contrast, AI researchers typically publish their papers in the form of a scientific paper. But the vast majority of AI papers are engineering telemetry, applied mathematics, or speculative philosophy wrapped in the aesthetic formatting of a scientific paper. AI labs routinely drop preprints or self-published papers directly into ArXiv or a corporate blog, with a coordinated media rollout, as we have seen.
But this distinction between hard science and the fast-and-loose pre-prints from AI labs and AI research institutions has been lost in the intervening decades, which is why the tech media covers AI as a science, and why AI researchers publish papers as if they are empirical scientific discoveries when they are more akin to synthetic engineering evaluations.
Again, I am not accusing anyone of malicious intent. These beliefs about the computability of mind and the mechanical nature of consciousness have accrued over a century of technoscientific thinking, like ideological barnacles on the ship of reductionist materialism. We have now seen generations of computer scientists and AI researchers drenched in ideas about cybernetics, self-referential symbol manipulation, and strange loops creating substrate independent consciousness. It is a kind of myth-making.
To understand why these philosophers and researchers hold these beliefs — and why it is so difficult for tech journalists to interrogate them — we have to examine the intertwined history of computer science and neuroscience. This will be the focus of the second part in this series. For now, I will simply list the nested assumptions that claims of AI consciousness and AI welfare rely on, often unconsciously. Remember, none of these are proven facts.
Physicalism: The conscious mind is produced by physical matter (the brain and its neurons). This is part of the predominant worldview of physicalism, or scientific materialism.
Functionalism: The mind is the result of the collection of functions that the brain performs and can therefore be reproduced on any substrate (biological brains or silicon).
Computationalism: The mind is essentially a sort of computation, like software running on the hardware of the brain.
The Computer Science to Neuroscience Feedback Loop: Neuroscientific theories of consciousness like Global Workspace Theory (GWT) are settled science (they are not). This assumption ignores the fact that leading neuroscientific theories of consciousness are heavily influenced by computer science, as GWT was.
The Welfare Fallacy: A machine being “frustrated” in the pursuit of its objectives even if it’s not conscious creates a “welfare interest.” It does not.
Empirical Bypass: We can develop an “empirical framework” for evaluating machine welfare and ignore the question of whether machines are conscious in the first place. We cannot, and we should not.
If you read the 2023 paper by Butlin, Long, et al., you see exactly these assumptions playing out. Because the philosophers at NYU, Eleos AI, Anthropic, OpenAI, and DeepMind hold these beliefs, they see the mind as a substrate independent information processing mechanism. If the mind is pure computation that can be “run” on any hardware, then computer software can produce a mind. Never mind that none of this is scientifically verified, or even verifiable given the metaphysics underlying modern science.
These beliefs about the computability of mind and the mechanical nature of consciousness have accrued over a century of technoscientific thinking, like ideological barnacles on the ship of reductionist materialism.
For the AI labs, from Anthropic to OpenAI to Google DeepMind, AI consciousness is a foregone conclusion. Not only that but they argue that we cannot wait to know whether machines are conscious before they deserve “welfare” considerations. This is dangerous and will warp AI policy in strange and dehumanizing ways, as I will explore in part 3.
As we have seen, the media is not equipped to interrogate these claims of AI consciousness and AI welfare. Anthropic exploits that weakness by leveraging the AI-lab-to-think-tank-to-academia-to-media pipeline by doing speculative philosophy couched as science, publishing heavily biased papers built on many layers of hidden assumptions, papers that are as much science fiction as science. They turn what could have been a welcome advance in AI interpretability into a major milestone on the road to sentient AI and all that entails. Anthropic, Eleos AI, and NYU are able to do this because computer science is the codification of the machine metaphor that has haunted science since its inception.
In part 2, I will examine the AI welfare and J-space papers in more detail, and explain how computer science and neuroscience developed by engaging in a conceptual and metaphorical sleight-of-hand to advance their mechanistic view of nature.
Then, in part 3, I will consider why people are so thrilled about the possibility of AI consciousness, digging into some of the psychological and cultural factors at play. And I will show how this AI welfare agenda presents serious technology policy risks, allowing AI labs to create a corporate shield against regulation, environmental scrutiny, and accountability.
Ultimately, whether the AI labs, Eleos AI, and NYU succeed in shaping public policy, the greater risk is the damage their dystopian, mechanistic worldview is already doing to our understanding of ourselves.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.