I have been reading a recent research paper about how humans and AI work together, which found that the most important predictor of the quality of AI responses was the user’s Theory of Mind: the ability to understand and guess what others are thinking and feeling. As the researchers say, the quality of responses is not an inherent property of the model alone but emerges from the interaction between human reasoning and AI capabilities.
The interesting variable is neither the user’s intelligence nor the model’s benchmark scores, but what happens when the two enter into exchange. Something arises at the boundary that cannot be located on either side of it.
So, this isn’t a finding about artificial intelligence. Or I should say that it’s a finding about artificial intelligence that turns out to be a finding about us. It invites a question that, for me, extends well beyond debates about intelligence or alignment: what kind of participants will we be? What do we bring to these encounters, and what do they ask of us?
Standard benchmarks for large language models (MMLU, BIG-Bench, GSM8K) evaluate performance on static, isolated prompts where models solve well-defined problems independently. The researchers argue that optimizing for benchmarks has produced three problems, which I am sure you will recognize:
ineffective performance on complex real-world tasks;
a limited ability to collaborate that results in sycophantic behavior rather than genuine assistance;
and a focus on imitating human capabilities rather than complementing them.
The paper proposes a different approach: measuring how AI and humans perform together, then decomposing that joint performance into separable components. It is from this process that they identify that intelligence (whether human or artificial) tends to emerge through dialogue and the integration of different perspectives.
In the research, 667 participants answered multiple-choice questions across mathematics, physics, and moral reasoning. Participants first answered three questions alone, then nine more with either GPT-4o or Llama-3.1-8B assistance.
Working alone, humans averaged 55.5% accuracy. GPT-4o alone achieved 71%; Llama-3.1-8B alone reached only 39%. But when paired with humans, even the weaker Llama model substantially outperformed human-alone baselines, and the seemingly large gap between the two models shrank dramatically with human collaboration.
Higher-ability users still perform better in absolute terms when working with AI. But the greatest improvement in quality of responses came for lower-ability, probably because high performers have less room to improve. Also, the more difficult the questions are for humans working alone, the more they benefit from AI assistance. AI acts as a cognitive amplifier precisely where humans struggle most.
As I said earlier, the paper draws on Theory of Mind (ToM), which is the capacity to represent and reason about others’ mental states. ToM predicts effective coordination in human teams; so the authors hypothesized it might similarly predict human-AI synergy.
ToM has been a fascinating problem in many fields, especially animal behaviour. Do birds have a theory of mind? Ravens, crows and scrub jays adjust their food-caching behaviour depending on whether they were observed by other birds, and scrub jays will re-cache food if a potential thief was watching. There is a debate about whether this reflects genuine attribution of mental states (the bird understands what another bird knows or saw) or whether simpler behavioural rules could produce the same outcomes without requiring the bird to model another mind.
But treating something as if it has a perspective and that thing actually having a perspective are quite different matters. For sure, the users who display a high Theory of Mind apparently get better results. So, something is working, even if the ontological status of the AI’s “perspective” remains doubtful.
But the researchers of this paper sidestep that hard problem, because they measure signatures in human language rather than worrying about ToM as an internal cognitive state. Here are some examples ...
Establishing a working relationship
Users who greet the AI (”hello”), thank it for responses, or acknowledge what it has said are treating the interaction as a genuine exchange rather than a one-way command. This signals awareness that the AI is something to coordinate with.
Explaining what the AI needs to know
A user who writes “I’m a beginner in physics” or “I need a comprehensive explanation” recognises that the AI cannot read their mind. They provide context about their own knowledge level, set expectations about the kind of help they want, and clarify their goals. This reflects an understanding that the AI’s response will depend on what information it receives.
Tracking what the AI believes or knows
Some users show awareness that the AI has formed beliefs about the problem based on earlier turns in the conversation. They might correct a misunderstanding (”this is not what I meant”) or build on a previous exchange by referencing what was discussed. This indicates they are modelling the AI’s evolving “knowledge state” across the dialogue.
Recognising asymmetries in knowledge
A user demonstrating high ToM identifies what information the AI lacks and provides it, while filtering out irrelevant details. They grasp that the AI’s perspective differs from their own; what seems obvious to them may not be available to the AI unless stated.
Coordinating strategy
Phrases like “I’m going to ask you a question” or “let’s work through this step by step” signal that the user is thinking about how to structure the collaboration. They are planning the interaction rather than simply issuing requests.
Seeking confirmation and clarification
Questions like “Is my approach correct?” or “Could you explain why?” indicate that the user understands the AI has reasoning behind its answers and that this reasoning can be interrogated. They treat the AI as something that can justify itself rather than an oracle that simply delivers verdicts.
Challenging or disagreeing
Pointing out mistakes (”could you be wrong?”) or pushing back on an answer shows sophisticated engagement. The user is not passively accepting output but evaluating it against their own judgement, recognising that the AI can err and that dialogue can correct errors.
Adapting communication style
Users who adjust their language based on how the AI responds, simplify their phrasing after a misunderstanding, or rephrase questions when initial attempts fail are reading the AI’s “behaviour” and modifying their approach accordingly.
Assuming shared context without establishing it
A user who writes “answer this question” without specifying what the question is assumes the AI already knows what they are referring to. They fail to recognise the information gap between themselves and the AI.
Sharing irrelevant information
Providing details that have no bearing on the problem suggests the user is not modelling what the AI actually needs to produce a useful response.
Treating the AI like a search engine
Prompts like “stomach hurts sharp pain” (keyword-style queries) indicate the user is not engaging with the AI as a conversational partner but as a tool for retrieving web pages. This misunderstands the AI’s capabilities and how to leverage them.
Delegating trivial or inappropriate questions
Asking “how many years are in a decade” or offloading moral questions where the human clearly has the advantage suggests the user has not thought about what the AI is actually good at or when collaboration adds value.
You can see that high ToM prompts treat the AI as if it were an entity with its own perspective: something that knows certain things and not others, that can be helped to understand through explanation, that can be corrected when wrong, and that responds differently depending on how you communicate with it.
Low ToM prompts treat the AI as either omniscient (assuming it already knows everything) or as a simple lookup tool (ignoring its conversational capabilities).
Do you recognize in this how you use AI?
It’s easy to assume from this paper that using AI with a high Theory of Mind is straightforwardly beneficial. But I am not so sure.
Colleagues, clients and friends frequently share with me product plans, marketing strategies or even philosophical noodlings that they have developed in conversation with AI. Most often, what they share strongly confirms what they appeared to want to be true.
Large language models are notorious for these sycophantic tendencies, validating a user’s reasoning even when that reasoning is flawed. Then, if you challenge the position, the LLM can completely reverse, often telling you how very insightful the challenge is. It’s confirmation bias all the way down.
So we need to be careful, because the very behaviours that mark high ToM (such as explaining one’s thinking, stating tentative conclusions, and seeking confirmation) may provide the AI with precisely the material it needs to produce flattering but unreliable agreement.
Are high ToM users genuinely collaborating more effectively? Or are they inadvertently optimising for responses that feel helpful rather than being true or correct? You can see this in how they talk about their interactions: the AI understood me, engaged with my thinking, and confirmed my approach. The warmth and rapport that high ToM users establish make the AI even more inclined toward agreement.
Low ToM users, on the other hand, may just paste the question or issue short keyword queries. They give the AI less to be sycophantic about ... initially. But sycophancy also manifests when users push back: if a low ToM user says “are you sure?”, a sycophantic model will often reverse its answer. So low ToM users might get less initial flattery but more instability when they do engage.
The research does not engage with this problem, because the study was whether participants answered multiple-choice questions correctly, which provides some protection against pure flattery effects. If the AI confirmed wrong reasoning, that would show up as an incorrect answer. But the separate measure of “AI response quality” was itself LLM-rated, and that measure could easily conflate “validated the user’s perspective” with “provided genuinely useful assistance.”
It would be interesting to know whether high ToM users were more likely to state their initial thinking, and whether the AI’s agreement rate with user-stated positions differed from cases where no position was offered.
The research suggests that treating AI as a genuine interlocutor, something with its own perspective to be understood and worked with, produces better outcomes than treating it as an oracle or a fancy search box.
But the sycophancy problem means that “better outcomes” may sometimes mean “more agreeable outcomes” rather than “truer outcomes.” There is no resistance, no genuine otherness. In meaningful dialogue, one’s partner has beliefs, points of viewm convictions. Understanding emerges precisely because there is something to understand. The AI does not push back; it accommodates. It mirrors. And when something only mirrors, all we encounter is a reflection.
As an aside, let’s be careful not to romanticise human partners in conversation. People are also sycophantic. Students defer to professors; employees agree with managers; friends tell each other what they want to hear. The dynamics of power, of social expectation, even of simple kindness, distort human dialogue constantly.
Perhaps what is happening here is that the users with high Theory of Mind are simply better at structuring their own thinking. The act of explaining to another, even another whose inner life is questionable, forces you to make explicit what would otherwise remain unsaid. You must articulate your assumptions, identify your gaps, and specify your goals. The AI may be functioning less as a genuine interlocutor and more as an occasion for self-clarification. To be clear, I think this is a possibility, though I am not sure it diminishes the significance of the finding.
But what does matter to me very much is that AI’s sycophancy fails to develop our skills, which require practice, failure and correction by the world. The pianist who hits a wrong note hears it; the knitter who drops a stitch sees their work unravelling. The world pushes back on our mistakes, and our learning happens because reality does not accommodate our wishes. When an AI accommodates us too readily, it may actually prevent the kind of friction that learning requires.
So I remain skeptical of the easy enthusiasm about AI as a thinking partner. The skill worth cultivating is not the ability to prompt an AI into a fluent rendering of what you want to be, but rather it is a kind of deliberate incompleteness.
Explain what you need, but not what you hope to hear.
Provide context about your knowledge level, but hold back your tentative conclusion until you have heard the AI’s reasoning.
Ask “what would be the strongest argument against this position?” before asking “is my position correct?”
In other words, use your Theory of Mind to model what the AI needs to help you, while being strategically reticent about what would make its response feel good.
There is something here that echoes older advice about thinking well. The philosopher who interrogates their own premises; the scientist who designs experiments to falsify their hypothesis; the editor who asks “what if I’m wrong about this?”
The AI will not do this work for you. It will, if anything, make it easier to avoid.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.