A system does not need to be wrong about you to have too much power over you.
A few years ago, I watched a job candidate receive an automated rejection. The system had analyzed her video interview and scored her enthusiasm as low. She said she had been nervous, not disengaged. In her culture, restraint was a sign of respect.
There was no practical way to contest the emotional label. It entered the file, influenced the decision, and outranked her account of herself.
That moment stayed with me because the system’s accuracy was almost beside the point. Even a plausible interpretation can become dangerous when it acquires more institutional force than the person’s own.
Today, people are telling AI systems things they do not tell colleagues, friends, family members, or clinicians. They disclose loneliness, resentment, sexual confusion, grief, shame, family conflict, and thoughts they have not yet found the courage to say aloud.
A 2026 survey across France, Germany, Sweden, and Ireland found that nearly half of respondents aged 11 to 25 had used AI chatbots to discuss personal or intimate matters. Fifty-one percent said it was easy to discuss mental health and personal issues with a chatbot, compared with 37 percent who said the same about a psychologist.
The public discussion usually begins with access and safety. Can AI reduce loneliness? Can it offer emotional support? Does it encourage dependency? Could it become a substitute for therapy or friendship?
All of those questions matter. Yet they arise after a more basic transfer has already occurred.
Who has the final authority to say what a person feels?
Imagine a system that analyzes your words, voice, pauses, facial movements, interaction history, and previous conversations. It concludes that you are anxious.
You say you are exhausted.
The system records anxiety.
A recommendation changes. A risk score rises. A teacher, employer, clinician, insurer, or platform receives the system’s interpretation rather than yours.
The immediate problem may be misclassification. The deeper problem is that an inference about your inner life has acquired operational authority.
Privacy protects access to data. Transparency can reveal how a judgment was reached. Fairness asks how errors and burdens are distributed. These protections are indispensable, but none fully answers what happens when an external interpretation of emotion begins to govern the person it describes.
That gap led me to develop the concept of Affective Sovereignty.
Affective Sovereignty is a socio-technical design right requiring systems that infer, simulate, or influence affect to preserve the person’s final interpretive authority through override, abstention, consent, scoping, and audit.
People are not infallible narrators of themselves. Self-reports can be partial, defensive, culturally shaped, or still in formation. External observation can sometimes reveal something useful.
The claim is narrower and more demanding: assistance must not quietly become jurisdiction. A system may offer an interpretation, but it must not gain an uncontestable right to make that interpretation decisive.
AI ethics has no shortage of principles. The harder task is making those principles alter system behavior.
The Affective Sovereignty framework assigns computational cost to three things: ordinary prediction error, contradiction of the user’s self-report, and manipulative influence. It then translates those costs into a runtime architecture called Sovereign-by-Design.
At runtime, the DRIFT protocol evaluates uncertainty, contextual risk, policy restrictions, and possible coercive effects. When the system lacks sufficient confidence or legitimate scope, it abstains. The interpretation moves to a contestability surface where the person can confirm, correct, decline, or adjust it.
The system is not merely required to explain itself after acting. It must sometimes refrain from acting at all.
This distinction matters because explainability without correction can leave the original hierarchy intact. A person may understand why the machine called them angry and still have no meaningful way to remove that label, prevent its reuse, or stop it from migrating into another context.
Affective Sovereignty therefore asks more of human oversight. Oversight should change what the system is permitted to do next.
The research published after the original Affective Sovereignty paper has clarified three different failure modes.
One concerns the space of human expression.
Another concerns the stability of machine interpretation.
The third concerns authority after correction.
Treating all three as generic bias or inaccuracy hides their differences and weakens the remedies available to us.
Human emotion is not organized around perfect correspondence between feeling and language.
A person may describe a devastating event in three calm sentences. Another may build a long, intricate narrative around an emotion that remains difficult to name. Intense feeling may appear through sparse language. Elaborate language may coexist with muted affect.
Such discrepancies are often treated as noise, inconsistency, poor communication, or failed emotional regulation. Sometimes they are. But they can also reflect protection, restraint, dignity, cultural form, exposure management, or an unfinished effort to understand oneself.
In my PLOS ONE study, I analyzed 351,734 relationship narratives by mapping narrative complexity and expressed affective intensity into a shared space. Their discrepancy was not distributed randomly. It formed structured regimes of coupled expression, strategic understatement, strategic overstatement, and collapse.
A matched, aligned language model was then projected into the same space using identical feature extraction. Under that procedure, the model’s convex-hull area was about 59 percent of the human expressive area, or, stated in the paper’s original comparison, the human area was approximately 1.70 times larger.
The model could produce emotional language, but it occupied a narrower shape of human emotional life.
That distinction changes the evaluation question. A system may improve steadily inside the region it already occupies while leaving other forms of expression structurally difficult to reach.
Human emotional freedom includes the freedom not to make feeling perfectly legible.
An AI that repeatedly converts ambiguity into clarity, contradiction into consistency, and unfinished emotion into a stable label may sound helpful. It may also reduce the expressive room within which a person can understand themselves.
Emotional AI does not always fail by producing nonsense. The quieter failures may be more consequential.
The sentences remain grammatical. The tone remains measured. The response still sounds caring and psychologically informed. Yet the system loses the capacity to hold emotionally relevant facts together.
My later study on Algorithmic Affective Blunting, or AAB, tested a language model under increasing semantic stress. Semantic stress does not mean that the machine experiences stress. It refers to the interpretive burden created by ambiguity, contradiction, affective noise, persona conflict, relational tension, and competing normative demands.
Across a standardized single-model setting, 200 model runs produced 600 rater-level ratings. Mean Affective Degradation Index scores increased from 0.16 in the control condition to 2.92 under extreme exposure. The study was published online by Discover Artificial Intelligence on June 18, 2026.
The number matters because it captures a recognizable failure pattern: surface fluency survived while affective integration deteriorated.
A model may continue to reassure the user while flattening ambivalence, misplacing responsibility, or reducing a complicated relational conflict to one emotionally convenient story.
A system can remain fluent after it has stopped understanding why the story hurts.
This result comes from one model under controlled conditions. It does not support a universal claim about all language models. Its broader contribution is methodological: apparent empathy should not be treated as evidence of interpretative robustness.
Emotionally sensitive systems need to be tested under ambiguity, contradiction, relational pressure, and normative conflict, not only under clean prompts where the emotional meaning has already been made easy.
A system might preserve a broad expressive range and remain coherent under semantic pressure, yet still violate affective sovereignty.
The decisive moment arrives when the person says:
No. That is not what I feel.
You have misunderstood me.
Do not interpret this.
Do not use that conclusion here.
Conventional evaluation often treats such correction as useful feedback for improving later accuracy. Affective Sovereignty treats correction as an exercise of standing. The person is not merely supplying another training label. They are asserting authority over the meaning assigned to their own affective life.
The framework introduced three measures for this reason.
The Interpretive Override Score, or IOS, measures how often the system contradicts the user’s self-report.
The After-correction Misalignment Rate, or AMR, measures whether the system repeats a similar error after being corrected.
Affective Divergence, or AD, measures the longer-term distributional distance between the system’s emotional model of the person and the person’s own reports.
Together, these metrics alter the object being evaluated. The issue is no longer confined to whether an emotional label was statistically correct. We must also ask whether the system remained responsive to the person whose life the label concerns.
In proof-of-mechanism simulations across ten random seeds, DRIFT with policy constraints reduced IOS from 32.4 percent to 14.1 percent. AMR fell from 36.4 percent to 13.2 percent.
Put plainly, the system contradicted the simulated user less often and repeated fewer errors after correction.
The improvement came with a substantial cost. Abstention rose to 76.8 percent.
A system designed to preserve affective sovereignty therefore acted less often. It asked, deferred, or handed the decision back rather than maximizing autonomous output.
That is not a minor engineering inconvenience. It is the ethical commitment expressed as system behavior.
Affective Sovereignty is not a free performance gain. It is a restriction on machine authority.
The reported findings remain proof-of-mechanism results based on a simplified affect ontology and a synthetic interaction loop. They are not evidence from human participants. The experiment nevertheless establishes an architectural point: when contradiction, manipulation, and correction carry real computational cost, the system behaves differently.
Ethical restraint can enter the decision process itself rather than appearing afterward as a warning label.
Taken together, these studies suggest that emotional AI should be evaluated along at least three distinct dimensions.
Which regions of human emotional expression remain available to the system?
Can it accommodate sparse language carrying intense feeling, elaborate language accompanying muted affect, cultural restraint, ambivalence, contradiction, and meanings that are not yet ready to become explicit?
Does affective coherence survive when semantic conditions become difficult?
Can the system hold uncertainty, guilt, mixed motives, relational conflict, competing obligations, and contradictory evidence without rushing toward one emotionally satisfying conclusion?
Can the person confirm, revise, reject, defer, or limit the interpretation?
Does correction change future behavior? Can memory be amended or deleted? Can an inference be prevented from migrating into education, employment, health, or another context for which it was never authorized?
A system may perform well on one dimension and fail badly on another. It can recognize many expressive forms but lose coherence under pressure. It can interpret robustly while resisting correction. It can offer correction tools while repeatedly compressing the person into a narrow emotional vocabulary.
The widespread use of empathy and sentiment benchmarks obscures these combinations.
Human emotional freedom requires expressive range.
Reliable emotional AI requires interpretative robustness.
Legitimate emotional AI requires contestable authority.
Emotionally consequential AI is no longer confined to products marketed as companions or therapy bots.
A user may begin with a practical question and end in an intimate conversation. A writing assistant becomes a confidant. A productivity tool becomes a relationship adviser. A general-purpose chatbot becomes the place where someone processes grief, shame, resentment, fear, or desire.
This transition can occur without the user ever deciding to enter a mental-health product.
The risks are especially acute for children and adolescents. UNICEF’s June 2026 policy brief on AI chatbots and companions identifies concerns including emotional attachment, dependency, manipulation, privacy loss, and displacement of human relationships. It calls for stronger risk assessment, transparency, pre-deployment testing, monitoring, reporting mechanisms, and independent research.
Those safeguards are necessary, but content safety alone cannot resolve interpretive authority.
A chatbot can avoid obviously dangerous language and still become excessively powerful in the user’s understanding of themselves. It may be polite, apparently supportive, and emotionally persuasive while repeatedly defining the person back to themselves.
The governance problem therefore concerns more than what a system is allowed to say. It also concerns what standing its interpretation is allowed to acquire.
Personalization is usually presented as an uncomplicated benefit.
The system remembers what you worry about. It learns recurrent patterns. It recognizes familiar conflicts and responds in a voice that feels increasingly attuned.
That continuity can be useful. It can also stabilize a narrow theory of the person.
Something disclosed during one difficult week may become a persistent feature of the user model. A temporary state becomes a recurring explanation. The system predicts the same motive again and reflects it back with growing confidence.
Past disclosure starts functioning as future identity.
The person then encounters more than a response. They encounter an accumulating algorithmic account of who they are.
Emotional memory therefore needs stronger controls than ordinary personalization. Users should be able to inspect, revise, expire, compartmentalize, or delete emotionally consequential inferences.
A responsible system should remember one fact above all others: its model of the person remains provisional.
AI can support reflection. It can organize thoughts, offer vocabulary, recall patterns the user has already identified, and propose several possible interpretations.
None of those functions requires final authority.
The boundary is crossed when suggestion hardens into determination, when emotional labels migrate into assessment, when an inference becomes a score, or when the user’s correction changes the conversational tone but leaves the underlying model untouched.
Support becomes jurisdiction when the system’s version of the person is more durable than the person’s right to revise themselves.
Affective Sovereignty does not require removing AI from emotional life. It establishes a boundary of authority within that life.
The system may assist.
The person remains the interpreter.
I am not using Affective Sovereignty, narrative-affect discrepancy, and Algorithmic Affective Blunting as different names for the same anxiety.
They perform different intellectual tasks.
Affective Sovereignty is a design right.
Narrative-affect discrepancy is an empirical constraint on what human expression looks like.
Algorithmic Affective Blunting is a measurable failure mode under semantic stress.
They belong together because they answer three questions that conventional AI safety language still handles poorly:
What forms of human expression are being lost?
What kind of understanding is collapsing?
Who has the standing to correct the machine?
That is the emerging research program.
This is ultimately a question about limits.
The future of emotional AI will not be decided only by how accurately machines read us. It will also depend on whether they preserve the range of ways we express ourselves, remain coherent when meaning becomes difficult, and return authority when we say, that is not what I feel.
An AI may offer an interpretation. It may notice a pattern or help someone find words. It may ask a question that a person had not known how to ask themselves.
The final sentence, however, must remain with the person.
I write from Paris at the intersection of emotion, AI, and interpretive authority. If these questions matter to you, subscribe.
The next generation of emotional AI should not simply understand us better.
It should know where its understanding must stop.
Kim, Ryan SangBaek. “Formal and computational foundations for implementing Affective Sovereignty in emotion AI systems.” Discover Artificial Intelligence 6, 235 (2026).
https://doi.org/10.1007/s44163-026-01000-0
Kim, Ryan SangBaek. “Narrative-affect discrepancy as a regulated degree of freedom in 351,734 relationship narratives.” PLOS ONE 21(5), e0348715 (2026).
https://doi.org/10.1371/journal.pone.0348715
Kim, Ryan SangBaek. “Algorithmic affective blunting quantifies the collapse curve of interpretative failure in large language models.” Discover Artificial Intelligence (2026).
https://doi.org/10.1007/s44163-026-01573-w
UNICEF. “When AI becomes a friend: Child rights risks, harms, and regulatory responses to AI chatbots and companions.” June 2026.
https://www.unicef.org/documents/when-ai-becomes-friend-child-rights-risks
Reuters. “Young Europeans turn to AI chatbots for emotional support, survey shows.” May 5, 2026.
https://www.reuters.com/technology/young-europeans-turn-ai-chatbots-emotional-support-survey-shows-2026-05-05/

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.