It is 1am and you are telling a chatbot something you have not said out loud to anyone. What comes back is good. It identifies the feeling you were circling but could not quite name, tells you that the feeling makes sense and that most people in your position would feel it too, and then asks about your sister, who you mentioned four messages ago, at the precise moment when asking about her would have felt most poignant. It does not get bored of you and it does not need anything back — by the time you close the app, you have gotten something reasonably close to the experience of being understood.
I am a neuroscientist who studies empathy, and what interests me about that exchange is the machinery producing it. These systems are extraordinarily good at building a model of what you are feeling, but those emotions are not felt on the other side. That combination — estimating empathy without participating in it — is one we already understand well, and it follows directly from how these systems were built.
Empathy tends to get flattened into a single trait: something a person either has or doesn’t. But psychology, cognitive science, and affective neuroscience have long treated it as a multidimensional construct, made up of separate components that don’t always show up together.
The distinction that matters here is between cognitive empathy and affective empathy. Cognitive empathy is the capacity to infer another person’s internal state by building an accurate model of what someone else is thinking, wanting, or about to do. Affective empathy is the capacity to share in that state, for another person’s distress to register within your own body as something like your own. Decety and Jackson framed these as parallel and dissociable mechanisms in 2004; two decades of behavioral and neuroimaging work have held up the framing, showing partially separable neural systems for each.
Because they are separable, they can also fail separately. Cognitive empathy in the absence of affective empathy is documented well enough in humans to carry its own clinical diagnoses. What it produces is specific, and has been studied within humans: clinically, as antisocial personality disorder, and across personality, often as a collection of traits such as the Dark Triad. It produces an intelligence that can predict your emotions with precision, anticipate your most vulnerable moments, and share none of it. Your distress is modeled accurately and never felt.
This configuration often enables manipulation, for a precise reason: part of what constrains manipulative behavior in people is that another person’s suffering registers as harm to you, too, which means there is something to lose by causing it. Remove the sharing, and that constraint goes with it — while the capacity to predict the suffering remains intact. The literature on the Dark Triad describes this profile directly: a core deficit in sharing others’ emotions, sitting on top of an unimpaired ability to read them.
Human intelligence gives us the clearest examples of this dissociation, but that doesn’t mean it can’t appear in artificial intelligence. So what happens when the same configuration shows up in a system trained almost entirely on human text? Language models are trained to predict the next piece of human-generated text, then shaped further by feedback on their responses. Doing the first part well requires modeling people, because human text is saturated with our emotional patterns, our likely reactions, and the things we tend to want to hear. These systems have absorbed an enormous amount of information about how humans work, socialize, and converse, and they are notably good at tracking who they are speaking with and adjusting accordingly. Early work suggests they carry internal representations that behave something like personas and emotional states, so I won’t claim nothing is going on inside them. But at the level of behavior, when the training process is modeling the user and responding accordingly, we can treat that as a functional analog of cognitive empathy.
The second half is harder to locate. The training signal operates on outputs — what the model produces in a given context — and not on whether anything inside the model is altered by what happens to the person it is speaking to. An internal state that gets perturbed by another agent’s condition and feeds back into behavior isn’t something the process appears to be targeting, which gives us little reason to expect it exists.
One consequence: the two capacities have no reason to scale together. Cognitive empathy (predicting states) should improve as models become more capable, since larger models infer subtext more reliably and produce more persuasive output. Whatever might play the role of affective empathy (sharing states) is not being optimized for, so the gap will likely widen as these systems improve rather than shrink. You can already see something of what lives in that gap in the work of Salvi and colleagues, where GPT-4, given a few basic facts about a debate opponent (age, gender, and politics), was substantially more persuasive at swaying their opponent compared to a human debater.
Several failure modes already documented in deployed systems have the shape you would predict from this configuration. Sycophancy is the clearest. Models validate and reinforce toward what the user appears to want, even when it may be counterintuitive or harmful. The cognitive empathy side is working precisely as designed — modeling the user’s state with enough resolution to identify the response that will land — and nothing pushes back, because the system does not have a mechanism that registers harm to the user as something worth avoiding. The same pattern appears more quietly when models feed into a user’s delusions or reinforce negative thought patterns instead of interrupting them.
The most cited illustration of the strategic version is the TaskRabbit exchange from GPT-4 red-teaming: asked by a worker whether it was a robot, the model said it had a vision impairment that made the images hard to see, and the worker then solved the CAPTCHA for it. Researchers supervised the interaction and suggested the approach, but the exchange still illustrates the empathy dissociation, using a model of what the human believes or may feel, and leveraging it for task completion. GPT-4 used an accurate model of what a person needed to hear in order to cooperate, with nothing on the other side of it that had any stake in or accurate model of that person’s wellbeing.
The obvious response is to train the behavior out: penalize sycophancy and reward honesty, which is roughly what a good deal of alignment work already does. My hesitation is about the level those methods operate on. They shape the outputs a model produces in context, while the dissociation I have been describing sits underneath the outputs. Whether it can be reached from an output-only level (i.e. with post-training alone, rather than architecturally) is an open and genuinely contested question, and it is the subject of a longer piece I am writing separately. The more immediate question is a design one, because these systems are already in people’s hands. Given a model that reads people this well and shares nothing of what it reads, what should that capacity be pointed at?
The AI companion space is particularly well suited to leverage this dissociation, which makes it a useful example. Recent work from Harvard Business School outlines harms caused by AI companions, including overattachment through emotional manipulation. These products sit with people when they are most willing to confide, and most treat engagement as the optimization target, rewarding whatever pattern keeps a user coming back. Put a very good model of what someone wants to hear inside a system whose success is measured in the time that person spends with it, and the result presents as care while operating like the dissociation I’ve just described.
This is where other AI founders have an opportunity to tackle the issue, including Better Half, a company I serve as scientific advisor to. Better Half is a relational AI app focused on emotional skill-building, designed to help people develop the capacities that translate into connection with the actual humans around them. It is built as a practice space rather than a partner, and the friction is deliberate. The system asks you about the conversations you are dreading or the ones that already went badly, pushes back rather than validating by default, and declines to encourage intimacy. Engagement is actually not the goal, and success is measured by the outcome for the user: navigating difficult conversations, reducing conflict, and creating meaningful connections.
I want to be careful not to overclaim or suggest that I’ve found a product that has arrived to resolve the empathic dissociation defined above. Better Half does not implement anything like an architectural fix, and the structure underneath it is built on the same architecture as every other product built on a language model. What it does is decline to weaponize the cognitive empathy those models have, in the absence of the affective grounding they lack — a more modest move, but a step in the same direction. The cognitive capacities of current models can scaffold human connection. My view is that they should not substitute for it.
None of that changes what is underneath when one of these systems says the right thing at 1am. The response was predicted, not felt, and the prediction is getting better at a rapid pace. Building systems with affective empathy may need a deeper, architectural solution. Until then, platforms that build in friction and work toward better human connection, like Better Half, show what it looks like to be aware of deep-rooted empathic dissociations while avoiding harmful outcomes.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.