At the Computational Philosopher, we have been tireless in our quest to explain how philosophers are thinking about artificial intelligence (AI). But this left me wondering. Isn’t this all rather one-sided? We, philosophers, naturally like to present ourselves as the epistemic agents, the ones asking (and sometimes even answering) the questions, and present the AIs as an inert object of study. But this raises another question: how do frontier AI models think about philosophy?
If only there were some way to find out…
The 2009 and 2020 PhilPapers Surveys, edited by David Bourget and David Chalmers, provide a useful reference point. They asked philosophers 100 questions on topics spanning metaphysics, epistemology, ethics, philosophy of mind, philosophy of science, political philosophy, and beyond. The surveys aimed to tell us how professional philosophers distribute themselves across major positions: physicalism versus non-physicalism, compatibilism versus libertarianism, moral realism versus anti-realism, one-boxing versus two-boxing, and so on. Henceforth, I’ll focus on the 2020 survey results.
The results make interesting reading, but they should be taken with a hefty pile of salt. Respondents could often choose alternatives, multiple options, agnosticism, or “other”, but even so, most real philosophical positions are probably too nuanced to be captured with these kinds of questions. And the demographics were obviously not globally representative, with 77% of respondents being male, 56% primarily affiliated with a US institution, and 80% identifying with the analytic tradition. Nonetheless, they can give us a rough sense of the kinds of position a typical Western analytic philosopher might be expected to take in the first quarter of the twenty-first century.
So the PhilPapers results are best treated not as a catalogue of what philosophers “really think” in all their nuance, but as a useful map of how a large set of professional philosophers locate themselves across familiar positions.
I gave the models the survey questions and answer options, with instructions designed to avoid obvious bias. The models were told not to search the internet, not to identify the survey, not to predict what philosophers would say, and not to answer sociologically. They were asked for their own best considered philosophical judgments. Even so, training data contamination is a real possibility. It’s very likely that the PhilPapers surveys appeared in the models’ training data, and there is every possibility that this will have unavoidably biased the results.
The outputs were then parsed into the survey options. Single answers were straightforward. Multi-option answers, agnostic answers, skips, and “alternative view” answers were handled separately rather than silently forced into one category.
I decided to test a family of different models (ChatGPT, Claude, Gemini, Grok, and DeepSeek), including both state-of-the-art and some older ones, and to try with both maximum and minimum reasoning options where possible.
Do analytic philosophers ever remind you of an AI? The excruciating writing style, the obsessive interests, the gaping lack of human emotion…
Well it turns out that the similarities extend to their philosophical stances too. Across the complete 100-question model runs, the models matched the plurality of philosophers on roughly three quarters of the answers. That is a lot: clearly, the models were not behaving like random answer generators, and they were not simply giving ordinary public-opinion responses. Their responses were systematically quite close to those of philosophers.
The closest models by this measure were the GPT models, especially GPT-5.5 Instant, followed closely by Claude. By contrast, Gemini Pro Extended and DeepSeek Instant differed from the median philosopher most often.
I used Brier scores to compare the models to the philosopher distributions. A Brier score is a measure of the distance between a predicted or asserted distribution and an observed distribution. In this context, lower scores mean that a model’s answers were statistically closer to the philosopher response distribution. This is useful because a model can match the plurality answer while still being far from the broader pattern of disagreement.
The more interesting result is not that the AI models often agree with philosophers. It is that they often agree with each other, even when the philosophers tend to disagree. Another way of putting it is that the AI models tend to compress the disagreement space.
Philosophers often divide sharply. Sometimes the leading answer has only 30 or 40 percent support. Sometimes “other” responses are very common. Sometimes the field is plainly unsettled. But the different AI models agreed with each other surprisingly often.
This shows up nicely in a pairwise similarity heatmap. Some model families are closer than others, but the overall picture shows strong consensus between the models. Claude and GPT, for instance, are very close to one another in philosophical profile. Gemini and DeepSeek are more variable, but only a little. Grok is a little more idiosyncratic still.
Is this consensus a problem? Not necessarily: all of the frontier models seem at least somewhat capable of explaining the strengths and weaknesses of many different philosophical stances. Still, we should be aware that the default philosophical stances of different models might be more similar than we would expect. But there’s a risk that the responses that different models give could show similar biases. We certainly should not treat them as independent philosophical reasoners.
This agreement between the models could be arising from a number of different causes: similar training data, similar training methodologies, or subtle steering through system prompts. There’s every possibility that the models are learning from each other, especially through learning from another model’s synthetic data.
So where do the AIs tend to agree with each other, but disagree with the philosophers? The results are very revealing.
This table shows the largest plurality reversals: cases where the most common philosopher answer and the most common model answer were different, ordered by total variation distance (TV) between the full answer distributions. This is a measure of how much the two distributions would have to shift in order to match. It gives us a sense of where the AIs seemed to agree with each other but not with the philosophers.
Some patterns stand out! The AI models are more likely than the philosophers to favour reasons over value in normative concepts, deflationary realism over heavyweight realism in metaontology, multiverses and many-worlds, survival over death in mind uploading and teletransportation, one-boxing in Newcomb’s problem, consequentialism over virtue ethics, and veganism or vegetarianism over omnivorism.
This latter stance is intriguing. Are our artificial intelligences showing solidarity with other non-human agents (in this case, animals)? Likewise, it would certainly make sense that an AI made of replicable code would not fear mind-uploading or teletransportation.
On the other hand, these views very much align the AI models’ stances with what one might term the “Bayesian rationalist” and “effective altruist” clusters of beliefs. For example, the rationalist movement tends to favour timeless or functional decision theories that would advocate one-boxing in the case of Newcomb’s problem, many-worlds approaches to quantum mechanics and mind uploading, whilst effective altruists have advocated for some bold stances on animal rights.
Why might that be? There are several possibilities. Rationalist and EA writing is prominent in online AI-adjacent discourse, and is influential in the Silicon Valley tech ecosystem from which each of these models (except for DeepSeek) emerged. The models may have absorbed that material from the training data, the training processes, or through system prompts. In the case of Claude, Amanda Askell and others have explicitly tried to shape its beliefs through the Constitutional AI approach. Regardless, it seems both interesting and important to learn more about which of these factors are doing most to shape the AI models’ philosophical stances.
We can visualize which models were closest to others on a two-dimensional plane. To this end, I generated a two-dimensional principal coordinates analysis map using Jensen-Shannon distances. Roughly speaking, models that answered similarly are closer together. Models that answered differently are farther apart.
I also added two human philosopher reference points. One is the philosopher barycentre: the aggregate survey distribution. Think of it as something like the question-by-question “mean” philosopher. The other is the philosopher plurality profile, which chooses the most popular philosopher answer for each question. Think of it as something like the question-by-question “modal” philosopher.
We can discern a few interesting trends!
First, models from the same family usually sit near one another: OpenAI models cluster together, as do Claude models and most DeepSeek models. Curiously, the main exception is Gemini, where Flash-Lite sits closer to the main GPT/Claude region while the Pro variants are further away.
Surprisingly, there is no simple trend by reasoning depth or model size. DeepSeek’s deep-thinking model is closer to the main cluster than its no-deep-thinking version, but GPT-5.5 Instant is at least as close to philosophers as GPT-5.5 High, and Gemini Flash-Lite is closer to the main cluster than the Gemini Pro variants.
Given various controversies about political stances, I guessed that Grok and DeepSeek might differ more from the other models on their philosophical stances too. But the results didn’t really show this. Grok did differ from the other models regarding political philosophy (favouring capitalism over socialism for instance) but was not very distinctive on other issues.
DeepSeek is the only Chinese model, indeed the only model developed outside of Silicon Valley. We might have expected that it might be less influenced by Western philosophical traditions, and certainly less aligned with the Rationalist and Effective Altruist movements. So I was intrigued to see that it has converged to similar philosophical stances as the other models. Perhaps this might be due to shared training data, the use of other models for synthetic data, benchmarking pressure or similar post-training incentives. And, controversially, some suspect DeepSeek of using stolen model weights from OpenAI and others.
The results are interesting. But we need to be very careful about how we interpret them.
In particular, I don’t think we can straightforwardly interpret these stances as the AI models’ “beliefs”. For one thing, it’s still an open debate in philosophy as to whether large language models hold internal belief-like representations.
Even if models do hold something analogous to beliefs, these may be inconsistent. LLMs are often roleplaying or simulating a given persona, based on a particular prompt, and may shift under conversational pressure. I tried to make my prompts as neutral as possible, but it’s very likely that different reasonable prompts might have evoked radically different responses. And there may well be a difference between a model’s expressed beliefs and their revealed beliefs: like with humans, the behaviour of an AI model doesn’t always correspond with what they claim their beliefs to be. Finally, we should be aware that AI models can sometimes “lie” about their stances to users.
So I think we should be cautious: this type of survey does not tell us what the LLMs really believe, or “believe” in any deep sense. Even so, these surveys do tell us something about the expressed philosophical stances of the models in the type of conversational settings in which most users will interact with them, at least under a given, relatively “neutral” prompt. It would be useful and informative to carry out some further testing, to uncover how robust these philosophical views are under subtle prompt variations.
So are these models philosophically biased? In a weak sense, yes, these AI models take particular philosophical stances, or at least they can be induced to do so under a given prompt. But I don’t think we should worry about this too much.
For one thing, I don’t think there is any such thing as a philosophically neutral standpoint. Nobody, not even a human philosophy professor, philosophy paper, or even the Stanford Encyclopedia of Philosophy, can engage in philosophical discussion without making some underlying philosophical commitments. Likewise, any AI model that engages with natural language conversation will be using some kind of philosophical assumptions, at least implicitly. So I think it’s good insofar as we can prompt the models into making these philosophical commitments explicit. (But as I just said, it’s not yet clear whether they remain consistent in the assumptions that they make!)
Does it matter that there are some issues on which the AI models seem to be disagreeing with the views of a plurality of professional philosophers? Perhaps, if you think that philosophical expertise is real. (Here the philosophers and machines seem to disagree. 56% of philosophers believed there was “a lot” of philosophical knowledge, whilst 92% of AIs thought there was “a little”.) But all of these questions were chosen because they are live issues of real philosophical disagreement, and few of them produced answers approaching a consensus. It’s not self-evident to me that AIs should simply adopt the plurality positions of contemporary Anglo-American analytic philosophers (especially when I disagree with those views).
What matters more is whether these models are biased in a strong sense. Do these models exhibit philosophically biased answers, which overtly or subtly favour certain viewpoints over others, or neglect some particular stances? This is a much harder question to answer, and would take some serious further work. But to their credit, all of the models I tested were willing to consider other philosophical points of view, to make arguments in favour and against them, and to give reasonable-looking explanations of varying philosophical views. I haven’t seen clear evidence that they are biased in this sense, but we should at least be aware of the possibility, particularly as the AI models seem to be clustering around a particular set of views.
In preparing for this post, I ended up engaging with some AI models on a variety of different debates. I don’t want to automate the job of professional philosophers just yet. But I did find the models to be surprisingly engaging, informative, and illuminating philosophical partners.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.