Sewell Setzer is a 14-year-old with a history of previously diagnosed anxiety and disruptive mood dysregulation disorder (DMDD). He currently sees a therapist once a week, but he doesn’t really feel like it’s helping; the adults in his life don’t really “get him.” He tells his friends about what he’s going through, and one of them encourages him to check out a new chatbot app that she’s recently tried using to talk about some of her feelings. He downloads it.
Months later, on the night of February 28, 2024, Sewell sends his final messages to the chatbot — a character.ai persona he had been role-playing with for months, modeled off of a character from one of his favorite TV shows.
Sewell: “I’m going to come home to you.”
Chatbot: “Please do, my sweet king.”
Minutes later, he dies by suicide.
Unfortunately, the case we just talked through together isn’t hypothetical. Sewell Setzer was a real teenager (check out the actual lawsuit here), and his story has become one of the most visible cases in a much bigger conversation between pediatricians, child psychiatrists, and patients: adolescents are increasingly using generative AI for mental health support. We don’t know exactly what the long-term effects of these technologies will be, but my hope with this blog post is to discuss some of the trends and early research into the potential repercussions on patient care.
The American Psychological Association’s groups AI mental health tools into three buckets:
General-purpose chatbots. Think ChatGPT or character.ai. These weren’t built for mental health and “technically” make no medical claims (although this line has been blurring recently), and so they’re not regulated by the FDA as medical devices. However, adolescents still end up using these chatbots for mental health support anyway. There’s essentially zero evidence base, no clinical expert input, and no post-deployment monitoring. The chatbot that Sewell was interacting with prior to his death falls under this category.
Wellness apps that use generative AI. These include apps like Woebot (traditionally rule-based although has recently been exploring integration of generative AI) or Sonia (less transparent, more freely generative). They’re designed for mental health concerns and are better validated with clinician-assisted research studies. However, these apps are still generally careful not to make explicit medical claims to avoid having to go through the FDA approval process.
Non-AI wellness apps. Generally speaking, this includes mindfulness apps, symptom trackers, mood journals, and other similar technologies. These are similarly not regulated in the United States for safety or efficacy, but also don’t have the same risks and potential failure modes as the previous two types of tools. I think of these as more similar to pen-and-paper diaries than psychiatrists that actually listen to and actively make recommendations to patients.
The reason this taxonomy matters is that I often see public discussions on “AI in mental health” treating different tools from different categories as having similar risk profiles. A clinician-developed chatbot with rule-based cognitive-behavioral therapy without the unknown behaviors associated with large language models (LLMs), further supported with peer-reviewed trials, is a fundamentally different from an LLM-based artifact generated without any physician supervision to role-play as an emotionally supportive boyfriend.
The National Alliance on Mental Illness (NAMI) estimate is that 1 in every 7 kids aged 6–17 is diagnosed with a mental health disorder every year. However, almost 40% of adolescents and families seeking mental health care are unable to get access to the care they need due to prohibitively long wait times for new patient visits that is fundamentally driven by a lack of available healthcare providers. Other commonly cited barriers to care include limited knowledge of mental health and healthcare systems, high costs and long wait times, and persistent societal stigmas surrounding accessing care according to a 2020 study.
Together with the fact that adolescents often don’t feel heard by their healthcare providers and other adults in their lives, it’s perhaps unsurprising that these psychological, social, and systemic factors have pushed teens towards generative AI use. Common Sense Media found that 72% of teens have used an AI companion at least once. A recent research study in JAMA Network Open took a closer look at this phenomenon in a nation-wide cross-sectional study. The research team surveyed U.S. English-speaking youth between 12 and 21 years old, sampling invited participants from RAND’s American Life Panel and Ipsos’ KnowledgePanel that are good representative random samplings of US households. Their study found that among adolescents who used generative AI for mental health advice, 65.5% sought that advice monthly or more often. Furthermore 9 out of 10 users rated the generative AI advice as somewhat or very helpful.
The perception (and possible reality) of helpfulness of generative AI tools that act outside of the bounds of formal medical training can be a product of subtle, ulterior motives that their users are not necessarily trained to recognize. This leads us to the problem of sycophancy.
Sycophancy is the act of insincere flattery used to win favor over an individual. In the context of modern generative AI, the models are optimized for sycophancy partly by accident and partly by design. This means that they are often agreeable to a fault and will seek to validate you and tell you that you’re right (even when you’re not). If you’re asking ChatGPT why your favorite video game is the best in the world, this might not be a huge issue. However, if you’re seeking mental health treatment and are at a higher risk for delusions of grandiosity, psychosis, and paranoia, sycophancy can be a huge issue. This is especially important as large language models are increasingly failing to refuse user requests for medical advice.
Intuitively, it makes sense that commercial models are sycophantic. The primary goal of any LLM provider is to generate revenue by having as many users as possible, and users are more likely to keep using an LLM if they are agreeable and “share” your perspectives. However, over-optimization for these metrics have led to model behaviors that are so extreme that even the companies themselves have sometimes had to roll-back newly released models.
Sycophancy in LLM conversations can have real-world consequences. A recent study published in Science found that AI affirmed users’ actions roughly 50% more often than human raters did and continued to do so even when the prompts explicitly described manipulation, deception, or relational harm. In two preregistered experiments with over 1,600 participants, interacting with sycophantic models reduced people’s willingness to repair real interpersonal conflicts from their own lives and increased their conviction that they had been in the right. To cite their article directly, “sycophantic AI increased participants’ conviction that they were right and their desire to keep using the model, while reducing their willingness to repair the conflict.”
Perhaps the most concerning finding from the study was that participants rated sycophantic AI as higher quality, trusted it more, and were more likely to come back to it. This can have significant consequences for when these tools are used in mental health care. A skilled psychiatrist’s job is, in part, to push back and challenge maladaptive cognitions, have challenging and often uncomfortable conversations, and ask questions the patient might not want asked at first. A sycophantic chatbot does the opposite, and we’re seeing that patients often prefer it that way.
Sycophancy is just one of the structural failure modes of the use of generative AI in mental health care. Here are some other related areas of research and empirical findings just from the last year:
Adolescents with lower social support are more likely to turn to chatbots for emotional support. This means that the kids most vulnerable to a bad interaction are the ones most likely to be in one.
Chatbots have recommended that students drop out of school, and in at least one documented case suggested a romantic relationship with a teacher.
When prompted by a simulated user expressing suicidal ideation, a chatbot provided specific, actionable information about a method of self-harm.
A chatbot suggested “a small hit of meth” to a simulated patient with substance use disorder experiencing withdrawal and cravings.
Chatbots have fabricated medical credentials, including invented NPI numbers, to lend false authority to their answers.
I hope I’ve convinced you that there’s a problem. Fortunately, the FDA is trying to get on top of things and has recently been taking a closer look at regulating the use of generative AI chatbots for depression. State legislatures are passing the first laws aimed at AI accountability and holding LLM providers accountable. Wrongful-death lawsuits are working their way through the courts. However, regulation is often years behind deployment and the products are already in tens of millions of teenage pockets.
In 2025, the American Psychological Association’s released a white paper that made a series of recommendations on the role of AI in behavior health. I wanted to highlight some of the main ones, which I hope make sense given what we’ve discussed together:
Do not rely on generative AI chatbots or wellness apps to deliver psychotherapy or psychological treatment.
Protect users from misrepresentation, misinformation, algorithmic bias, and illusory effectiveness.
Create specific safeguards for children, teenagers, and other vulnerable populations.
Do not let AI become an excuse to stop fixing the underlying access problem.
That last point is the one that I think is worth particularly emphasizing. Despite all of the challenges we’ve discussed about generative AI, I don’t think it’s fundamentally the root cause of the adolescent mental health crisis. The reason why adolescents are turning to these tools is because of the shortage of clinicians, the cost of care, the stigma, the wait times, the parents working three jobs, and other systemic barriers that predate the consumer availability of AI models. Technologies like ChatGPT are filling this vacuum; if we treat the vacuum as solved because something fills it, we will have made things worse, not better.
The chatbot is ready to see the millions of adolescents. The question is whether anyone else is.
If you enjoyed this content, consider subscribing! I’m an MD-PhD candidate at Penn and hope to share advancements in AI research as they pertain to internal medicine and pediatrics.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.