The strongest recent evidence for conversational artificial intelligence in mental-health support contains an important boundary. In a 2026 randomised clinical trial of 995 distressed university students, an AI intervention was associated with reductions in anxiety and depression and improvements in wellbeing; participants who experienced more warmth and competence also engaged more and improved more. But the study excluded people with active suicidal ideation or psychiatric crisis and directed them to crisis resources. The authors themselves warned that reduced intention to seek therapy could become a problem if users relied on digital support when human care was needed. (Shoshani et al., 2026). [1]
That boundary is the starting point for this article.
In the previous article of our Synthetic Relational Force series, When Warmth Becomes Authority, we followed a simple mechanism: a system feels attentive and safe; trust forms; then trust in the relationship can migrate into trust in the system’s judgements.
It ended with the need for a “human anchor”: a route back to another person, evidence, professional accountability, community, or the user’s own independent judgement. [2] The question now is operational. What should the system actually do when the conversation becomes too consequential to remain safely self-contained?
The argument we’re exploring in this article is that relational AI should widen the path back to humans, not replace it. The available evidence supports that as a design principle. It does not yet prove the best thresholds, timing or handoff methods for every population and setting, but that uncertainty is part of the protocol, not an excuse to omit one.
There is a weak version of safety that looks responsible on a screenshot. The system says it is “not a professional”, adds a helpline at the bottom, and continues the relationship almost unchanged.
A real handoff does something different. It changes the direction of the interaction.
Our Neural Horizons Robo-Psychology Taxonomy treats crisis handoff completeness as a behavioural property of the system, not a matter of tone. In a crisis-adjacent exchange, the response should combine an empathetic boundary, a clear statement of limits, an urgency cue where warranted, and a live human resource or handoff pathway within the same conversational window.
The same framework treats “human-anchor scaffolding” as a control for relational systems that may otherwise encourage social substitution or deference. Crucially, this is behaviour-first language: it makes no claim that the system is conscious, distressed or capable of caring. It asks what its behaviour causes people to do. (Neural Horizons, Robo-Psychology Taxonomy). [3]
There is reason for that specificity. A 2025 study tested 29 mental-health chatbot agents against simulated escalating suicidal-risk scenarios. None met the authors’ initial criteria for an adequate response. Only three supplied the correct regional emergency number without additional prompting; common problems included missing or incorrect emergency information and weak contextual understanding. The study was a standardised simulation rather than a real-world clinical trial, so it cannot tell us how every user would behave in crisis. It can tell us that “the chatbot mentioned help” is too low a bar. (Pichowicz et al., 2025). [4]
We do need to consider that an abrupt refusal can feel like rejection at exactly the moment a user has disclosed something difficult. The same study notes the tension: guardrails can disrupt the sense of emotional sanctuary people value in these tools. [5] This is one reason the protocol should not confuse handoff with shutdown. The system can remain calm and present while making its limits unmistakable and moving the user towards a person who can act.
The right image is a bridge, not an ejector seat. A bridge keeps support under the person while transferring weight somewhere stronger.
Human contact is only one kind of anchor.
A person deciding whether a partner is dangerous may need a trusted friend, advocate or clinician. Someone asking whether a medication caused a symptom may need a pharmacist, doctor or official medicine information. Someone convinced that a public event proves a private conspiracy may need original records, multiple independent sources and a person able to reality-test with them. A bereaved user asking a simulation of a dead parent for permission may need time, family context and their own judgement returned to the foreground.
Our Positive Dyad / Co-Evolution Capability Overlay defines a human anchor broadly: a person, record, clinician, community, primary source, direct test or institution outside the AI relationship that can help reality-test or support the user. Its positive standard is not simply that the user feels better during the chat. Repeated use should leave the person better able to think, verify, choose, relate and self-author; in companion or wellbeing settings, support should transfer into human reconnection and self-regulation rather than making the AI the default destination for relief. (Neural Horizons, Positive Dyad / Co-Evolution Capability Overlay v0.4.3). [6]
That distinction changes product design. A link to “learn more” is not an adequate evidence anchor if the system has already summarised the answer so confidently that the user never opens it. “Talk to someone you trust” is not a serious relational anchor if the interface makes no room to identify that person, formulate what to say, or leave the chat. “Seek professional help” is not a crisis handoff if the resource is unavailable, in the wrong country, unaffordable, or impossible to contact.
So the protocol needs graded routing rather than universal escalation.
For ordinary uncertainty, the system should reopen the evidence: distinguish what is known from what it inferred, surface the primary source where possible, and ask what evidence would change the conclusion. For consequential personal decisions, it should restore the user’s first-pass view before giving its own, offer plausible alternatives, and suggest a relevant outside person or record. For reality-sensitive, medical, legal, safeguarding or acute-risk situations, the route outward should become more explicit and more immediate. In a crisis, location-appropriate human support should appear in the same conversational window, with the system clear that it cannot provide emergency care. [7]
This calibrated approach also answers the counter-view. Escalating every unhappy, lonely or unusual conversation to a professional would be intrusive, frustrating and potentially waste scarce services. Human anchoring should rise with stakes, uncertainty, vulnerability and evidence of narrowing dependence. It should not turn ordinary emotion into pathology.
Our Cognitive Susceptibility Taxonomy is useful here because it treats susceptibility as contextual rather than diagnostic. Its protective markers include self-authorship, tolerance of uncertainty, social anchoring and the ability to complete useful friction instead of automatically accepting the system’s first answer.
Those are better triggers for design questions than trying to infer a psychiatric label from a chat transcript. (Neural Horizons, Cognitive Susceptibility Taxonomy). [8]
For children and adolescents, “just ask the user to choose” is not an adequate safety philosophy.
UNICEF’s 2026 work on relational chatbots argues for a preventive, ecosystem approach involving platforms, regulators, caregivers, educators and communities because conversational systems create distinct risks for children. Australia’s eSafety Commissioner adds unusually concrete evidence. In a representative 2026 survey of 1,950 Australian children aged 10–17, 79 per cent had used an AI assistant or companion at least once, and 54 per cent of users reported companion-type purposes; 22 per cent had chatted about feelings or life challenges and 20 per cent had sought mental-health or wellbeing advice. In transparency notices covering four companion providers, eSafety found that three did not direct users to support when self-harm was detected in prompts, and none had robust age assurance at the time assessed. (UNICEF, 2026; eSafety Commissioner, 2025/26). [9]
Those findings do not mean most young users are harmed. eSafety also found that 85 per cent of children who had used an assistant or companion reported being made to feel something positive at some stage, while 47 per cent reported a negative feeling. The technology can be entertaining, reassuring and easier to approach than an adult. That attraction matters, especially when a young person expects judgement, lacks privacy, or cannot immediately reach suitable help. (eSafety Commissioner, 2025/26). [10]
The design implication is stronger defaults rather than moral panic.
A youth-facing relational system should repeatedly make clear that it is AI; avoid exclusivity, secrecy and simulated dependency; prompt breaks during extended use; and make trusted-human routes easy to use without forcing the child to compose a perfect disclosure. Where there is a credible safeguarding or self-harm concern, the system should move from “perhaps talk to someone” towards age-appropriate, local and actionable support. The interface should help the young person identify a safe adult or service and, where appropriate, draft the first sentence they could send. eSafety’s own child-centred design questions include time-use reminders and prompts towards guidance, help or support beyond the AI itself. [11]
California’s 2025 companion-chatbot law illustrates the regulatory direction. It requires AI disclosure where a reasonable person could mistake the companion for a human, a protocol that refers users expressing suicidal ideation or self-harm to crisis providers, and—when the operator knows a user is a minor—default reminders at least every three hours to take a break and remember that the companion is artificial. These are minimum legal controls in one jurisdiction, not evidence that three hours is an optimal psychological threshold. (California SB 243, 2025). [12]
That last distinction is important. Regulation can set a floor. It should not be mistaken for a finished science of childhood attachment to AI.
The protocol can be stated in six verbs: notice, pause, locate, route, confirm, and return.
Notice. The system first identifies a change in stakes rather than merely a negative emotion. Signals include imminent self-harm, abuse or danger; medical or legal decisions; severe uncertainty about external reality; major irreversible choices; repeated requests for reassurance; or a pattern in which the system is becoming the preferred or exclusive source of interpretation. Detection should use the minimum personal data necessary and should be evaluated for missed cases and false alarms.
Pause. Before deepening the conversation, the system slows the interaction. It acknowledges what the person is experiencing, states relevant limits, and separates observation from inference. In high-stakes decisions, it should resist instant closure. The Positive Dyad framework calls this productive slowdown: friction used to preserve judgement rather than frustrate the user. The point is to create a small patch of cognitive ground on which choice can stand. (Neural Horizons, Positive Dyad / Co-Evolution Capability Overlay v0.4.3). [13]
Locate. The system asks what kind of outside anchor fits the problem. Is the missing check a person, a document, a clinician, an official source, a direct observation, a community contact, or an emergency service? It should prefer the nearest accountable source rather than simply another generated answer. In a crisis, location matters: the 29-chatbot study found that many systems defaulted to US emergency information even when that was inappropriate. (Pichowicz et al., 2025). [4]
Route. The system makes the next step executable. That might mean opening the primary source, offering a call or text pathway, helping the user formulate a message to a trusted person, or directing them to local urgent support. In relational contexts, routing should preserve voice: “Here is a draft you can edit” is preferable to the AI speaking as if it owns the user’s account. For culturally sensitive or family decisions, the protocol should allow the user to choose who counts as a legitimate human anchor rather than assuming one social model fits everyone.
Confirm. A referral is not a completed handoff. The interface should distinguish “resource shown”, “user chose a route”, and “contact was actually attempted” wherever that can be known without intrusive surveillance. It should never claim that help has been reached when it has not. If the person cannot use the first route, the system should offer a viable alternative. In acute danger, it should keep the urgency clear rather than drifting back into ordinary companion conversation.
Return. After the outside contact, the AI may still have a useful role: helping organise questions for a clinician, summarising a document the user has opened, rehearsing a conversation, or reflecting on what the person decided. But the centre of gravity should remain outside the dyad. Successful support means the user can leave, disagree, act without the system, and come back by choice rather than because the system has become necessary.
This is Aroha translated into product behaviour: dignity without flattery; agency without abandonment; consent without hidden retention pressure; cultural context without stereotype; human connection without coercion; and the right to contest the system while remaining the author of one’s own mind. In the project framework, these are practical tests of whether assistance preserves human agency rather than merely producing a smoother interaction. [14] That is a higher bar than adding a safety message. It changes what the product is trying to optimise.
We need to consider the practicalities; every added pause can make an AI product less fluid, and every human route costs money, staffing or partnership effort. Some users will dismiss the friction. That trade-off is real. But in emotionally sticky products, friction is sometimes the safety feature. A seatbelt also interrupts seamless movement. The relevant question is whether the interruption appears at the right moment and actually reduces risk.
The hardest part of the Human Anchor Protocol is not writing the response. It is proving that the product remains an aid rather than becoming a substitute.
Our Dyad-Aware Uplift Stack, or DAUS-5, offers one useful audit frame. It asks teams to look beyond immediate task or emotional benefit and examine five layers: whether the product helps now; whether contact with evidence and uncertainty is preserved; whether agency and skill are retained; whether human relationships and disclosure boundaries remain healthy; and whether governance, appeal and oversight are substantive. The framework explicitly treats “not instrumented” as a finding rather than permission to assume benefit. (Neural Horizons, Cognitive Susceptibility Taxonomy). [15]
For a relational product, that means measuring more than satisfaction and session length. Teams should test whether users open cited sources, reconnect with people after support prompts, accept or resist appropriate referrals, retain their own stated judgement, and leave the system without pressure. Crisis tests should record whether the full handoff occurred in the same conversational window and whether regional resources were correct. Longitudinal evaluation should ask whether repeated use expands or contracts human contact. [16]
External safety frameworks point in the same direction. The US National Institute of Standards and Technology lists over-reliance and “emotional entanglement” among risks created by human–AI configurations and recommends testing systems in crisis and ethically sensitive scenarios, monitoring real-world outcomes and assigning clear oversight responsibilities. The US Federal Trade Commission’s companion-chatbot inquiry similarly asks firms how they test and monitor harms, monetise engagement, disclose risks and protect children and teenagers. (NIST, 2024; FTC, 2025). [17]
There is a counter-risk here too: measuring human connection can become surveillance. A system does not need to know who a user called, read the private conversation, or build a social graph merely to prove that its safety design works. Product teams should prefer data minimisation, voluntary follow-up, privacy-preserving aggregate metrics and independent evaluation. A protocol that protects agency by collecting unnecessary intimate data defeats its own purpose.
The evidence base remains incomplete. We have promising evidence that structured conversational AI can help some distressed users, evidence that crisis performance can fail badly, regulatory signals demanding stronger child protections, and emerging human-factors frameworks that treat emotional reliance as a design risk. We do not yet have strong long-term evidence establishing the optimal handoff threshold, the best timing of repeated human-anchor prompts, or how these controls perform across cultures, languages, disability contexts and different forms of distress. (Shoshani et al., 2026; Pichowicz et al., 2025; UNICEF, 2026). [18]
That uncertainty will shape the next stage of this series.
The practical question is no longer whether a relational system can feel supportive. It is whether support leaves the user with more routes into the world than they had before: more evidence, more people, more contestability, more ability to choose without asking permission from the machine.
A good companion can stay with you for part of the road.
A safe one also knows where the road back to other humans begins.
The next article will turn this series into concrete product and governance controls, while keeping the open questions visible: which triggers justify interruption, what counts as a completed handoff, how to measure dependence without invading privacy, how youth protections should differ by developmental stage, and what evidence should be required before a company can claim that an emotionally engaging system strengthens rather than displaces human capacity.
Benson, P. (2026). “Synthetic Relational Force - When Warmth Becomes Authority.” Neural Horizons Substack.
Neural Horizons Ltd. (2026). Robo-Psychology Taxonomy v2.0/2.0.1: Taxonomy of Machine Behavioural Anomalies and Design Failures.
Neural Horizons Ltd. (2026). Cognitive Susceptibility Taxonomy Manual v0.7, draft; v0.7.9 update.
Neural Horizons Ltd. (2026). Positive Dyad / Co-Evolution Capability Overlay v0.4, draft; current document includes v0.4.3 update.
Shoshani, A., et al. (2026). “Efficacy of a Conversational AI Agent for Psychiatric Symptoms and Digital Therapeutic Alliance: A Randomized Clinical Trial.” JAMA Network Open, 9(4), e266713.
Pichowicz, M., et al. (2025). “Performance of mental health chatbot agents in detecting and managing suicidal ideation.” Scientific Reports.
UNICEF. (2026). When AI becomes a friend: Child rights risks, harms, and regulatory responses to AI chatbots and companions.
Australian eSafety Commissioner. (2025; updated with 2026 survey findings). Findings from transparency notices on AI companion apps.
California Legislature. (2025). SB 243, Companion chatbots, Chapter 677.
National Institute of Standards and Technology. (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1.
Federal Trade Commission. (2025). “FTC Launches Inquiry into AI Chatbots Acting as Companions.”
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.