Note to anyone in the situation I describe. For any indvidual the risk is low, precisely because your situation is so very common. Anyone afraid will have an argument that their situation is uniquely dangerous, but in truth, this is an area of wildly varying presentations, many of them extreme. The fact that your presentation might be weird, extreme and contain to you what looks like special features means essentially nothing. This piece is going to describe worst cases vividly because it's a policy argument. If you're acutely anxious right now the useful move is to stop reading and take this to a clinician.(1)
A lot of people talk about the content of their mental illnesses with LLMs. The classic case is OCD. Unfortunately, OCD sufferers tend to seek reassurance from LLMs. This is discouraged, as it can exacerbate their illness, just as if they had sought reassurance from a human. Often even for patients that are completely safe- in ways that a clinician would immediately recognise- the content of disclosed thoughts can be quite disturbing. What a clinician understands, but a human crowd worker or content filter might not, is that these thoughts are ego dystonic. The mother imagines stabbing her baby again and again not because she wants to but because it is the most hateful image possible in the world for her. That very hatefulness makes it lurid, sticky, and impossible to escape. OCD is again the classic example here, but the phenomenon can be present in all sorts of mental illnesses. Content can include violent mental imagery, fears of committing grave crimes and even fears that one has committed grave crimes in the past that might, sans context, read like confessions etc. etc.
There is a lot of pressure on AI companies to adopt a proactive disclosure stance, re: possible harmful acts to self and others. Uniquely, they have the capacity for this kind of surveillance for obvious reasons (LLMs themselves can flag content). Companies have stated that they will disclose in some cases (see this August 2025 post by OpenAI https://futurism.com/openai-scanning-conversations-police). AI companies have not yet adopted policies that would lead to proactive disclosure in most cases like we are discussing, but the pressure on them to do so is growing.
OpenAI (currently) has a policy of only reporting on apparent threats to harm others, not self. Even this policy is under strain- as you read this it is quite likely there are sympathetic families in front of the camera with wholly reasonable fears that more could have been done to save their loved one from themselves. However even if it held, this carveout is insufficient in the OCD case, and other cases like postpartum presentations, panic disorder with fear of losing control and hurting others, some presentations of generalised anxiety disorder, depression with psychotic features and delusional guilt, psychosis generally and so on. This is precisely because in these conditions content that is really only harmful to self masquerades as something potentially dangerous to others.
Indeed, the sheer frequency of violent imagery, apparent confessions and other superficially troubling phenomena in LLM conversations- often phrased in ways that are difficult to separate from the real thing- makes me wonder about the practicality of searching conversations for evidence of wrongdoing, even apart from the ethical issues. Even with a high criterion for flagging, I would anticipate many false flags.
OCD is common. Other disorders that can cause similar problems are common. About 2.5% of the population have OCD, and as many as 40% have themes likely to generate taboo content- in addition to many other mental illnesses that can cause similar phenomena. Compulsive desire to “disclose” and seek reassurance are common, and clinicians are already recognising reassurance seeking from LLMs as a major new vector of harm. Back of the envelope maths says it is not impossible that 1% of the population have had a conversation in this category- that one percent is a good chunk of the 40% of OCD sufferers with this form of the disorder, plus many others with different mental illnesses. At a guess, perhaps 0.1% of the population have an extensive history of conversations of this form. Such conversations do genuinely seem to be extremely common in OCD. The desire to disclose and seek reassurance of safety is likely common in other disorders with symptoms of guilt and fear of wrongdoing for parallel reasons.
If LLM disclosure policies and OCD cases collide in the worst way possible you could have some very, very unfortunate investigations and possibly even charges and legal cases. You might even see the odd false conviction or bit of time spent in jail pre-trial, especially among the most vulnerable and least well resourced. This would be devastating on its own terms, but among the mentally ill, many would essentially take it as confirmation that they had done something terrible and are therefore dangerous people. Obviously the whole 1% figure I estimated is not going to be hauled in for interrogation, nor 0.1%. However in the dramatic case of a sharp policy move towards disclosure even a very low base rate could lead to thousands or tens of thousands globally suffering real harm, and for this population especially, process is punishment.
Consider the parallel to the Tarasoff doctrine- proactive disclosure requirements for therapists. I have never been fully sold on the doctrine, to be honest, but in Tarasoff situations, you have protective factors that are missing here. Safeguards that are absent in LLMs include:
1. Clinical judgement
2. A specificity/imminence requirement that the clinician is better positioned to evaluate than an LLM.
3. An existing relationship and understanding of the client that disclosures can be understood in the context of. Referrals, clinical history etc.
4. And the clinical conversation doesn’t happen in the context of an all-pervasive device, connected to your browser, your email client and your operating system.
There are many, many reasons why people do not want their laptop or phone spying on them on behalf of the state, and the argument I have given is just one. Even if the threshold for reporting on someone is enormously high, the fact that an application in your browser is open to the possibility of reporting you to the cops is chilling and creates an unpleasant feeling of always being surveilled and judged. Consider that LLMs are increasingly integrated into everything- email clients, operating systems etc. LLMs spying on you will come to mean your whole computer spies on you. The point about the chilling effects of privacy breaches is common, but what is less often noted is that these effects hit the vulnerable, the paranoid, the shame-and-false-guilt ridden the hardest.
The best solution, I believe, is to not go down the path of LLM as snitch. If conversations must be reviewed it should be with a warrant, and not on the basis of proactive disclosure. I have almost exactly zero faith that our crime and liability obsessed, paranoid society will reach this equilibrium.
FN: (1) I am aware that some of what I say in this paragraph- that the odds are low- could be taken as giving reassurance. I have decided to leave it in, regardless, because I feel like telling people once that a rational apprasial of their odds should comfort them, when a new fear emerges or has emerged, is good practice, particularly when introducing a quite novel idea that the fearful mind might begin to see as likely and not merely possible, leading to an insight collapse. Again, if you feel distressed, please see a clinician. I am not a clinician and nothing I say is clinical advice.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.