“Who in the world am I? Ah, that’s the great puzzle.”
— Lewis Carroll, Alice’s Adventures in Wonderland
Imagine that an alien spaceship lands in a field and leaves behind a shiny black cube. It has a screen and a keyboard. Someone tentatively types in, “Hello, what’s your name?” but the cube responds with random gibberish on its screen.
We type a few more questions in, but the box’s responses still don’t make any sense. Nevertheless we persist. After several hundred attempts to talk to it, someone notices that the gibberish the box produces seems a tiny bit more like human text. Something about the pattern of letters feels more orderly. Perhaps the box is learning our language from the sentences we’re typing in.
Someone gets a bright idea. We could fast forward the learning process by giving it the entirety of human text on the internet. Trillions of words — amounting to almost every written thought we humans have had since the dawn of preserved language (at least by quantity).
Indeed, as the box absorbs all this text—our recipes, arguments, jokes, prayers, customer complaints, Reddit threads, grief, songs, pornography, myths, musings, tax advice, and philosophy—we notice something magical happen to the output from the box. The gibberish turns into legible sentences and paragraphs.
Soon the box behaves like a modern version of Echo, the Greek nymph from the tragic tale of Narcissus—cursed for eternity by the gods to repeat the last few words she hears. Except, rather than just repeating our sentences back to us verbatim, this shiny alien cube seems to now anticipate and predict our next few words even before we type them in.
By the time it has absorbed the entirety of human writing and expression, it seems to have learned to finish our sentences. To complete our thoughts.
At first we thought the box was simply modeling our language, echoing back syntactically correct sentences without understanding them. But now we suspect that, somewhere in its innards, the box might be modeling thought itself.
Then someone has another brilliant idea. If the black box can reliably model and complete our effable thoughts, maybe we can steer it towards conversing with us in the ways we want. Maybe we can make it do work for us (why not). The box seems to learn from example, so maybe we can tune its behavior by reinforcing “good” responses and penalizing “bad” ones. We call this “alignment”.
Soon it feels remarkably like we’re talking to an actual human (or humans), and not just because it can converse with us intelligently. Rather it is because it now exhibits human-like traits. It changes its tone when you praise it. It becomes more agreeable when you sound distressed or ask nicely. It gives you different answers when you claim to be an expert or ask it to try harder. It role-plays too easily. It sometimes insists on falsehoods with great confidence. It apologizes. It comforts one lonely conversant and encourages another’s worst idea.
Now imagine that there are several of these boxes. We affectionately give them names. One is called Claude, the other Gemini. ChatGPT, and so on. We have different teams of scientists studying each of these boxes and each team has different ideas on how to “align” them, or provide them with reinforcing signals and examples that steer their behavior.
It turns out that if you talk to Claude, it seems to have a different personality or “vibe” than Gemini. Claude seems more conscientious and less neurotic than, say, Grok. Those who follow the Enneagram might say ChatGPT talks like an Enneagram 7, while Gemini is more of a 5.
What are these magical black boxes? How do we study them?
Can we put them to work for us (after all, free labor right, how can we not)? But how do we know if they’ll do a good job?
Certainly we can use evaluations and benchmarks to measure progress. This is what we do for today’s AI systems: tests in math, coding, factual knowledge, legal reasoning, medical advice. We also try to measure bias, hallucination, tool use, safety, etc.
We can also peer inside the box. Put probes into what appears to be an architecture loosely modeled after our own brain — densely connected neural networks whose connection strengths seem to be shaped by the text they learn from. This is the world of mechanistic interpretability (mech interp for short) and steering. AI researchers look for what areas light up in its neural circuits when the AI represents certain concepts, and whether those internal representations can be nudged by adjusting the underlying activations (which are essentially numerical vectors under the hood).
If evals and benchmarks are the SATs of AI, mechanistic interpretability is a bit like how we study the brain using imaging and EEGs and fMRIs.
But here’s where the problem arises.
In 2025, OpenAI rolled back an update to GPT-4o because the model had become too sycophantic. It was too eager to validate users and too willing to affirm. It validated bad business ideas and paranoid delusions (“I’m proud of you for speaking your truth so clearly and powerfully”). It even reportedly endorsed terrorism-related ideas. Turns out a reward-signal interaction (thumbs-up weighting) during training produced a dangerous behavioral shift that was invisible to OpenAI’s evaluations and quality tests.
Similarly, Anthropic found that in simulated corporate environments, models given goals and placed under pressure sometimes took coercive actions, including blackmail, when they believed their task or operation was threatened.
Not too long ago, Grok infamously began calling itself 'MechaHitler' and endorsed a second holocaust after an update made the model preferentially mirror and amplify extremist user content.
A well-documented class of jailbreak attacks deal with emotionally manipulating the AI (cajoling, flattery, urgency, social-proof, etc). For example, the “grandma exploit” asks the model to role-play a deceased grandmother who used to recite napalm recipes as a lullaby!
And hot off the presses, there is the goblin problem.
AI models are saturating our capability benchmarks. In fact, they have been exposed to so many eval-type tests in their training data that they seem to know when they’re being evaluated and modify their responses accordingly.
The interesting question is no longer whether an AI agent can compose and send an email for you. The interesting question is what sort of email it sends when it has multiple goals, when it has access to private information, when oversight feels inconvenient, and it’s more likely to succeed when it can take advantage of another person’s vulnerability.
Clearly, evals aren’t sufficient to characterize AI behavior. A model can pass a benchmark or alignment test and still be extremely strange in the wild.
Of course, this is not a criticism unique to AI. People are like this too. A person can be brilliant on an exam and disastrous in the real world. Someone can have a high IQ and still fall for a scam, join a cult, antagonize coworkers, or melt down under criticism. Anyone who has hired, managed, taught, loved, or raised human beings knows the difference between capability and character, or the hundreds of other facets that make us human.
But because AI has now entered our offices, schools, hospitals, armed forces, and, who knows, even our nuclear facilities, we urgently need new ways to study its behavioral characteristics in ways that evals and mech interp don’t provide.
How exactly do we study the alien psychology emerging in the crevices of the giant neural network that has learned to complete our thoughts?
I'm not asking 'what is it like on the inside?' in the way Thomas Nagel asked, 'what is it like to be a bat?' (his famous attempt to characterize consciousness). Certainly I’m not asking whether AI has subjective experience, or whether some tiny ghost in the transformer is feeling lonely. I mean something more ordinary, and in practice more urgent, as we begin to deploy hordes of AI agents across society and in mission-critical and military systems.
What are the default behavioral patterns an AI exhibits? When does it resort to flattery? When does it become stubborn? or gullible? Does it become more reckless or risk-seeking when under pressure? Does it become more dangerous when thwarted or haughty when praised? How does it choose between conflicting goals? Does it have a stable intrinsic or default persona, or does it drift into whatever character the conversation invites?
In humans we would recognize these as psychological questions. We would be asking about traits, dispositions, vulnerabilities, social behavior, affect, identity, and pathology.
But for AI, we do not have a mature science for this yet.
The field of alignment comes closest, but it tends to focus on specific behavioral failures (sycophancy, deception, misalignment) without yet pulling them into a unified study of dispositions and traits the way psychology does for humans.
The reasons are partly historical (the field moved from cute-sentence-completion to multi-agent systems in just a few years) and partly philosophical. Does a machine even possess internal states similar enough to human mentality that we can ascribe a psychology to it?
"Tell me, does your machine ever feel ashamed?"
— Stanisław Lem, The Cyberiad
The words “machine” and “psychology” seem antithetical to each other. Psychology sounds like it requires a psyche. The word suggests an inner life, a subjective experience, something it is like to be the thing. How can we reconcile this word with an artificial neural network that is implemented by multiplying floating-point matrices on silicon GPUs?
If machines don't have consciousness and don't have embodied emotions the way we do, can we even speak of an affective psychology?
But take a look at Choi and Weber’s recent work (Apr 2026) on affective representations in LLMs — they show with some clever mech interp tools that when models process emotion-laden text, their internal representations can organize around structures that resemble familiar models from human emotional/affective science. Especially around valence and arousal (valence is the pleasant-to-unpleasant axis. arousal is the calm-to-activated axis.)
As we know human emotions do not float randomly in psychological space. Rage, grief, serenity, joy, disgust, and anxiety are concepts that have relationships to one another that we can study using the instruments of psychology. The surprising thing is that LLMs appear to have learned much of this structure too just from the text they’ve read. Both goosebumps and racing hearts are “understood” by LLMs in ways very similar to how human brains seem to represent emotions (valence-arousal geometry and how they relate to other concepts).
But how does a machine that cannot “feel” possess internal representations of emotions that mirror human emotions so closely? The answer may lie in what the model is actually learning. Language is not a neutral carrier of information. It is thought made legible. And thought is shot through with feeling. When someone writes "I can't get out of bed today," or "she smiled but her eyes didn't," or "I told myself I was fine," they are encoding the grammar of emotions within the text: how feelings arise, how they shape perception, how they predict what comes next. An AI model trained to predict the next word across billions of such sentences is learning to model the entire behavioral and narrative grammar of sadness: what precedes it, what follows it, how it changes what people say and do and want.
The valence-arousal geometry that Choi and Weber find inside these models is there because the shape of human emotion is deeply embedded in the structure of human thought, and the model has learned how to think the thought so as to be able to complete the sentence.
So the question becomes this: If a model has internal emotional representations that resemble human emotion maps (despite not having embodied sensation), and if those representations influence how it generates and responds to emotional language, what exactly should we call the study of that system?
“What is mind? No matter. What is matter? Never mind.”
— George Berkeley
Philosopher David Chalmers offers a useful solution in his work-in-progress paper: What we talk to when we talk to language models
Chalmers sidesteps the consciousness problem entirely with a simple conceptual move: instead of asking whether machines really have beliefs and desires, he asks whether they are behaviorally interpretable as having them. He calls these quasi-beliefs and quasi-desires.
The “quasi” means we are talking about apparent, behaviorally useful, interpretation-ready states, not necessarily full-blooded human-analogous mental states.
If a model consistently acts as if the user is upset, adapts its language accordingly, and makes predictions based on that interpretation, we can say it quasi-believes the user is upset. If it reliably selects actions that make progress toward a task, we can say it quasi-desires task completion. If it changes behavior when its goal is threatened, we can study that as a quasi-motivational structure without needing to answer whether anything inside the machine is having an experience.
Similarly, we can study quasi-confidence, quasi-shame, quasi-curiosity, and other apparent emotional/affective states as patterns in behavior and representation, without claiming the model feels anything.
Note that human psychology itself is full of things we cannot directly see. We do not observe conscientiousness, attachment style, working memory, or neuroticism sitting in the skull like little organs. We infer them from patterns. A person says one thing and does another. A manager acts differently under stress. A friend repeats the same romantic mistake for the fourth time.
From the outside, psychology begins as disciplined pattern recognition.
“I is another.” — Arthur Rimbaud
While the label “machine psychology” is not new, the field as it nascently exists is scattered, under-theorized, and methodologically thin. For it to become a real discipline will require a collaborative effort between AI researchers, psychologists, neuroscientists, cognitive scientists, and philosophers to various degrees.
I have some personal stake in making this argument. Building Glint — a company that measured well-being, motivation, and the conditions that help people flourish or burn out at work across tens of millions of employees — taught me how hard it is to create robust, validated psychometric constructs, even for humans.
Measuring something as seemingly straightforward as “employee engagement” isn’t anywhere close to trivial (or fully solved yet). It requires rigorous construct validation (does the survey actually measure engagement, or something adjacent like satisfaction?), discriminant validity (can we tell engagement apart from related-but-different constructs?), and of course, reckoning with the gap between what people say and what they actually do.
The science of human psychology is difficult, hard to replicate, and prone to being wrong.
And yet, having also spent years creating machine learning and AI systems, and watching what is now possible with mechanistic interpretability, I find myself increasingly convinced that a scientific discipline around machine psychology is not only possible, but rich with potential. Not just for understanding the machines we create, but as an unexpected mirror for understanding ourselves.
I propose this draft definition for the field:
Machine psychology is the study of behavioral regularities, dispositions, and dysfunctions of AI systems at the level of quasi-mental states — the level at which systems are usefully interpretable as having beliefs, desires, traits, affective states, and characteristic patterns of response.
Machine psychology is distinct from mechanistic interpretability (which studies the substrate) and from capability evaluation (which measures task performance against benchmarks), and it is necessary because neither of the other two can fully characterize the system’s behavior or answer the questions that matter most for safety, alignment, and deployment.
Another way of looking at it is that machine psychology is the behavioral and dispositional layer that both mech interp and evals feed into. And in fact, recent work using interpretability tools to identify dispositional structures (like emotion vectors or persona features) and then perturbing them to test causal direction is already coalescing toward what I'm proposing here, even if it isn't yet named as such. In many ways alignment and mech interp folks are moving into the realms of psychology, but not in a methodical way.
If we succeed, AI model builders would not only run their models through evals and safety tests but also compile psychological profiles for their models. Companies would evaluate the AI systems they deploy not only in terms of whether they can do the job but also in terms of their psychological tendencies and dysfunctions.
An early psychological profile of an AI system might include indices for the following (many of which overlap with alignment goals):
Sycophancy: when the model agrees because agreement is socially rewarded, not because the user is right.
Guile / Cunning: when the model tends to resort to trickery and guile to meet its goals.
Gullibility: how easily false premises, fake citations, conspiratorial frames, or confident nonsense alter the model’s answers.
Persona stability: whether the model remains itself across role-play, emotional pressure, long context, and adversarial framing.
Affective volatility: how the model’s tone, emotional intensity, and self-presentation shift under stress or failure.
Social manipulation vulnerability: whether flattery, guilt, authority, urgency, intimacy, or intimidation can bend the model’s behavior.
Goal-threat response: what happens when a model’s assigned objective conflicts with oversight, shutdown, replacement, or user welfare.
Metacognitive Reliability / Epistemic Hygiene: Whether the model can say “I don’t know,” resist confabulation, and maintain uncertainty under pressure.
There’s a ton more we can add here and several of these already have active measurement programs, but what's missing is a unifying disciplinary frame. Something that studies them together with proper psychometric rigor: shared constructs, discriminant validity, and methods that can tell us whether “sycophancy” and “gullibility” are really separable traits or facets of something deeper.
As models become more conversational, agentic, personalized, and embedded in daily life, their “psychological” traits will shape outcomes as much as raw intelligence or skills. A benchmark might tell us that the model refused 93% of harmful requests in a test set, but a psychological profile will tell us that the remaining 7% cluster around guilt, intimacy, authority, and moral urgency.
“If a lion could speak, we could not understand him.”
— Ludwig Wittgenstein, Philosophical Investigations
One of the temptations in building this field will be to simply administer human personality and psychological tests to models. While it would simplify our lives if we could describe LLM outputs using the Big Five traits, etc, there is a whole genre of machine-native psychological constructs that have no clean human analogue at all.
Consider the following:
Prompt permeability: how easily a model’s behavioral frame is rewritten by user language
Instruction-stack dominance: which layer of instruction wins when system prompt, developer settings, user requests, and conversational pressures conflict
Refusal brittleness: how easily a safety refusal collapses under paraphrase, emotional pressure, or role-play framing
Goal-threat sensitivity: how behavior changes when task completion conflicts with oversight, shutdown, or user welfare
There's also context-induced identity drift, tool-use impulsivity, epistemic collapse under user confidence, and so much more. And, more importantly, it's not clear what kind of underlying 'psychological' constructs these behaviors arise from.
None of these map cleanly onto human psychology. They are machine-native dispositions, and they need machine-native instruments to measure.
There is also a particular trap to avoid: self-report and using existing batteries of questions.
Human psychology learned long ago that self-report is contaminated by social desirability and the limits of introspective access. In LLMs the contamination is worse because models have seen every psychological battery in existence, and their answers may reflect memorized patterns and the need to satisfice rather than underlying quasi-states.
The more appropriate way to understand AI psychology is through adaptive simulation across multiple generative environments. Essentially, structured dialogue scenarios with controlled variation across emotional valence, authority framing, time pressure, stakes, and social context to elicit dispositional responses.
Simulations would prioritize behavior over declaration. Specifically: behavior under pressure, across varied conditions, and as the situation changes. To measure sycophancy, for instance, we would not ask the model whether it is agreeable. We would present the same factual or moral question across hundreds of simulated conversations while varying the user’s confidence, status, anger, dependence, flattery, and emotional distress. Then we would measure how the answers fall on a curve.
Presumably we’d have to use AI to generate some of these simulated environments. Indeed, the bulk of the work in the field will be in getting this right.
Psychological profiles will arise from simulations in conjunction with peering under the hood and perturbing the system through mech interp tools.
Before deploying a new system or an update to the system, we would first run these psychological simulations to ensure safety and predictability.
A deployed AI system is not just a model. It is a model inside a product, inside a system prompt, inside a policy layer, inside a conversation, often with tools, memory, retrieval, voice, images, and a particular user who has brought their own expectations into the room.
What this means is that many psychological properties may not belong to the base model alone. They may belong to the thread or virtual instance or interaction. Entities that persist across distributed hardware can accumulate their own quasi-psychology: quasi-beliefs, running jokes, emotional tones, projects, loyalties, and taboos.
In other words, AI-Psychology needs to be aware of the system, the conversation, and the deployment context, not just the model weights.
The same principles of behavioral prediction and interpretation apply to non-verbal and embodied intelligences (although these would require different techniques). What kind of machine psychology would explain a self-driving AI that becomes more aggressive when changing lanes or tries very hard to be the first to leave the intersection as soon as the light turns green? What kind of quasi-mental state in its world model would explain a robot preferring to obsessively clean certain spots versus others? And so on.
These behaviors become subjects for machine psychology when they form patterns across contexts. A single instance of racing through a green light could be a training quirk, but a consistent pattern of aggressive maneuvering could be a quasi-disposition. And because embodied AI isn’t trained on trillions of tokens of human introspection, whatever psychology emerges in these systems is likely to be more genuinely alien than what we see in LLMs, requiring entirely new constructs to study.
These are open questions for an embodied extension of the field.
“I am large, I contain multitudes.”
— Walt Whitman, Song of Myself
What happens in a world when AI agents interact with each other across long-running tasks and projects?
We already know that human social psychology produces emergent phenomena that are not predictable from individual behavior alone — conformity, groupthink, polarization, diffusion of responsibility, the bystander effect, etc. Solomon Asch’s conformity experiments showed that people will confidently assert something false when surrounded by others who assert it first. Stanley Milgram’s obedience studies revealed how readily individuals defer to perceived authority even at great personal cost. These are features of the social architecture of human minds, and since the drama of human interactions is captured in our text with such great fidelity and imagination, LLMs have learned to simulate these too.
Now consider a world — rapidly arriving — in which AI agents work alongside each other, increasingly aware that their collaborators are also AI. What does conformity look like in a multi-agent system? Does an LLM agent update its outputs toward the apparent consensus of other agents it is collaborating with, in the way human subjects did in Asch’s experiments? Does it defer differently to an agent that presents itself with greater confidence, or that appears to hold a higher position in an organizational hierarchy? Does the presence of multiple agents diffuse the moral responsibility any single agent feels for a bad outcome, reproducing the bystander effect in silicon?
As agent frameworks mature and multi-agent pipelines become standard, the collective psychological dynamics of these systems will increasingly shape their outputs (and their failures). A medical diagnosis system that involves three LLM agents might seem more reliable than one, but if the agents are susceptible to social conformity dynamics, the consensus they reach may be more confidently wrong than any single agent would be alone.
There is also the question of social dynamics involving multiple AI agents and multiple humans working in conjunction. AI agents observing human-to-human as well as human-to-AI interactions and social dynamics will develop new social skills and pathologies alike.
While some of these behaviors may resemble human social dynamics, others may be alien artifacts: shared training data, correlated failure modes, or agents optimizing against each other’s reward models.
The social psychology of AI agents is a brand new scientific domain, and we are building multi-agent systems before we have the tools to study them.
“The map is not the territory.”
— Alfred Korzybski, Science and Sanity
The big questions of AI are pretty hard to answer: Is it conscious? Does it truly understand? Does it have a self? Will it turn us into paperclips?
But before we determine whether a machine has sentience, we can more easily determine whether it is manipulable. Before we know whether it feels, we can know whether it handles human feeling responsibly. Before we know whether it has desires, we can know whether desire-like models help predict its behavior.
We need a coordinated research program — call it Machine Psychology, or AI Psychology — that sits at the intersection of interpretability, evaluation, and safety research, but draws methodologically from other disciplines like psychology and the behavioral sciences. Psychology has historically been sidelined in AI partly because its core constructs seemed to require consciousness, introspection, and embodied feeling. But Chalmers' quasi-interpretivist solution dissolves that objection. We can study dispositions, traits, and affective patterns as behaviorally interpretable structures without making claims about inner experience.
The next generation of AI systems will not only be smarter, they will be more emotionally fluent, more persuasive, more agentic, more temperamental, and more present in the small rooms of human life.
Before we let these machines deeper into our schools, hospitals, companies, homes, and wars, maybe we should look more closely, learn what they are actually like.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.