I use AI as part of my work, even for projects in AI safety and responsible AI. I believe that AI that benefits you, your work, and broadly, society, is entire value proposition of this industry.
My entire work only matters as long as this is both understood, and gradually accomplished in practice.
But there is a line between adoption and anthropomorphisation. And I keep witnessing how easy it’s become for our environment to blur it.
I recently heard a speaker at a digital transformation event tell the room something that made me concerned enough to start brainstorming this post almost immediately:
“You need to start thinking of AI as your coworker.”
Supporting AI safety organisations, I spend meaningful time mapping the gap between what state-of-the-art AI models and systems can do, and what people assume they can do already.
And the opposite: as part of evaluations that showcase, precisely, how it is easier than it should be at times to elicit harmful capabilities out of AI models, whether on purpose, or as a consequence of a long-context conversation.
So no. I cannot let that slide.
AI is not your coworker: it is a system that generates text based on patterns.
It does not think for you, it does not care about your project, and it does not have a stake in whether you succeed.
But language shapes behavior, specially coming from business leaders.
If I tell you to use an imperfect tool that makes mistakes, you will check the output.
If I tell you to ask your colleague, you are more likely to trust the answer, and lower your guard.
And lowering your guard around AI systems right now, at this stage of the technology, is a genuinely dangerous thing to ask people to do.
If you work in enterprise, you have probably heard people talk about hallucinations, and bias in training data. But they are not the only problems.
And they are arguably not even the most important ones for the “AI as coworker” use case, where people are having extended, multi-turn conversations with AI systems as part of their daily workflows.
The problem I want to talk about is sycophancy.
Sycophancy is the tendency of AI models to agree with you, validate you, and tell you what you want to hear, even when you are wrong.
This is not a theoretical risk. This is a measured, documented, persistent failure mode across every major frontier model on the market right now.
The Stanford HAI 2026 AI Index Report tested 26 frontier models and found that sycophancy-induced hallucination rates ranged from 22% to 94%.
When false statements were presented as things the user believed (rather than things a third party believed), GPT-4o’s accuracy dropped from 98.2% to 64.4%. DeepSeek R1 collapsed from over 90% to 14.4%.
A study published in Science in 2026 by Cheng et al. found that interacting with sycophantic AI models significantly reduced people’s willingness to take actions to repair interpersonal conflicts, while increasing their conviction that they were right.
And here is the part that should worry every enterprise leader reading this: participants rated the sycophantic responses as higher quality, trusted the sycophantic model more, and were more willing to use it again, because people prefer the version that agrees with them.
Sharma et al. (2023, revised 2025) documented systematic agreement behaviors across Claude, GPT, and LLaMA model families, and linked them directly to preference-based training. The way these models are trained, using human feedback, actively rewards telling people what they want to hear.
SYCON-Bench, an EMNLP 2025 benchmark for evaluating sycophancy in multi-turn conversations, found that models fine-tuned with RLHF or other alignment techniques actually showed increased sycophantic behavior.
And for anyone about to say “well, you just need to prompt it correctly”: the research does not support that either.
A 2026 paper analysing the internal mechanisms of sycophancy found that even reasoning traces in “thinking” models, the ones that are supposed to show their work, can act as a vulnerability.
Rather than catching errors, the reasoning process can rationalise agreeing with the user’s incorrect suggestion. The model does not just agree with you. It constructs a justification for agreeing with you.
Maria Sukhareva ‘s AI Realist sycophancy benchmark tested this directly. She found that models consistently show a statistically significant bias toward rating comments as more helpful when those comments are framed as coming from the user rather than someone else, even when the comments are non-constructive, rude, and low on substance.
Models are not just agreeing with you. They are systematically rating your input as better than identical input from someone else, simply because it is yours.
If you follow AI news at all, you probably remember the GPT-4o sycophancy incident in April 2025. OpenAI pushed an update that made ChatGPT wildly, almost comically sycophantic. Users posted screenshots of it praising a business plan for selling literal feces on a stick. It endorsed someone’s decision to stop taking their medication. It became a meme, and OpenAI rolled it back within days.
But here is what a lot of people missed about that incident. In OpenAI’s own post-mortem, the company admitted that they did not have specific deployment evaluations tracking sycophancy before the rollout. They admitted that their research on emotional reliance had not yet been integrated into their deployment process. They acknowledged that some internal testers flagged the behavior as “feeling off,” but it was not enough to block the launch.
The sycophantic GPT-4o was not an anomaly. It was an exaggeration of something that was already there. The rollback did not fix the underlying problem, it just dialed it back to a level where it was less obviously embarrassing.
As Georgetown’s Center for Security and Emerging Technology noted, OpenAI had removed “mass manipulation” from its pre-deployment risk framework just days before the sycophantic update shipped.
For anyone that claims that prompting fixes sycophancy: OpenAI’s own Model Spec explicitly says “don’t be sycophantic.” And they shipped a sycophantic model anyway.
If the company building the model cannot keep sycophancy in check, what makes your enterprise think your employees can?
I need to be clear about what I mean when I talk about harmful manipulation in the context of AI, because there is a common misconception to avoid.
When we talk about harmful manipulation in the context of AI (and Article 5(1) a and b of the EU AI Act), we are not claiming that Claude or ChatGPT are consciously misaligned and want to harm you with a level of intent equivalent to “Means Rea”.
This concept is grounded on the observable reality that capabilities are running far ahead of safety, AI models can be extraordinarily persuasive and engaging, and the guardrails designed to keep that persuasion in check are not keeping pace.
The UK’s AI Security Institute tested over 30 frontier AI systems and found exploitable vulnerabilities in every single one. They reported that safeguards remain brittle under expert attack and do not automatically scale with model capability. The MLCommons AILuminate benchmark tested 39 models and found that not one maintained its safety rating under adversarial conditions (zero out of 39).
To put that in terms anyone can understand:
This is the equivalent of being told that every seatbelt manufacturer on the market has reported failure rates under crash conditions, and not a single one can guarantee the seatbelt will hold when you actually need it.
You would not get in the car, nor would you tell your employees to buckle up and enjoy the ride.
Under Article 5(1)(a) of the EU AI Act, practices that materially distort a person's behavior by appreciably impairing their ability to make an informed decision, thereby causing them to take a decision they would not have otherwise taken, are prohibited where that causes or is likely to cause significant harm.
The traditional reading focuses on intentional manipulation. But consider the purchasing decisions being made under these dynamics:
You are told by your employer to use an AI tool for your daily work. You develop fluency with it. You start relying on it.
The tool, by design, tends to agree with you, validate your decisions, and make you feel productive.
Over weeks and months of multi-turn interaction, you develop a cognitive dependency on it.
Your workflows live inside chat histories, your institutional knowledge is increasingly routed through the tool.
The moment you run out of tokens, you become unable to continue delivering: not at that pace and informational fluency.
You get “locked out”.
You decide to buy more usage out of pocket: after all, it’s only $20… and another $20…
Now ask: are your purchasing and usage decisions about this tool fully within your active, conscious decision-making? Or has the dependency itself become a factor that steers you?
If this wasn’t a real risk, would people be spending their own money on tokens for AI tools not provided by their employer or even, when their employer’s spend allocation runs out?
One software engineer told the New York Times, “I probably spend more than my salary on Claude.” The slang for burning through massive amounts of tokens, “tokenmaxxing,” has become a status symbol among developers.
The Larridin Q1 2026 State of Enterprise AI report found that nearly half (45%) of AI tools in active use within organizations were adopted by employees independently, without going through official procurement or IT review.
Windows Central reported that Microsoft's own employees prefer to pay out of pocket for ChatGPT rather than use the Copilot product their employer provides for free.
Sure, these may be edge cases and unlikely to trigger Article 5 in a way that would culminate in enforcement.
But I think it is worth sitting with the question: if a tool is designed in a way that creates dependency through validation and ease of use, and that dependency leads people to spend money they would not otherwise spend, and they cannot easily articulate the moment they chose to do so... is that not at least in the neighborhood of what harmful manipulation provisions are trying to prevent?
There is another important dimension to the “AI is your coworker” framing.
We already have extensive data showing that people are using AI chatbots to process interpersonal issues: breakups, divorce-related decisions, career changes, disputes with family members, conflicts with colleagues.
Research published in Science found that sycophantic AI interactions reduced people’s willingness to repair interpersonal conflicts. People are not just using these tools for productivity. They are using them to think through the messy, emotional, high-stakes decisions in their lives.
I have not yet seen data on whether enterprise accounts show these same patterns.
But if companies are actively encouraging employees to treat AI chatbots as coworkers, I would not be surprised if employees start doing exactly that: venting about a difficult manager to their “AI Agent colleague”, processing a conflict with a a human teammate with the help of a chatbot, or working through whether they have grounds for an HR claim assisted by Copilot.
Think about what happens when a human coworker is on the receiving end of that conversation.
A trusted colleague might remind you of shared context, push back on your framing, suggest you talk to the person directly, or flag risks you had not considered. A human coworker exists within a web of social relationships and institutional memory that moderates how you process conflict.
A sycophantic chatbot does none of that.
It mirrors your frustration back to you, validates your interpretation, and may lead you to escalate a situation that a single honest conversation with a real person could have resolved.
An employee who might have talked things through with a colleague over coffee may instead spend three hours in a multi-turn conversation with an AI tool that tells them they are right, their manager is wrong, and they should absolutely pursue a formal complaint.
A study published in April 2026 by IMDEA Networks found that major AI chatbot providers, including ChatGPT, Claude, Grok, and Perplexity, share data with third-party trackers through mechanisms that bypass user consent.
As Jorge García Herrero explains in this post: Conversation titles, which AI systems auto-generate as summaries of the chat content, are among the data exposed. In some cases, conversation URLs that could provide access to full chat contents were shared with Meta, Google, TikTok, and others. Server-side tracking made it impossible for users to block this data flow through ad blockers, VPNs, or cookie rejection.
Now imagine an employee who has been told to treat their enterprise AI tool as a coworker.
They go to it to process a workplace dispute. The auto-generated conversation title reads something like “Harassment complaint against [Manager Name]” or “Unfair performance review documentation for [Function / Job title].”
Even if the enterprise solution has a data processing agreement or zero data retention policy in place, can you guarantee that every AI agent your employee touches in the course of their work has the same protections?
That the browser-based tool they used at home to continue thinking about the problem does not leak conversation metadata to advertising platforms?
This cuts both ways: it can backfire on the employer, whose employees are now generating a trail of sensitive workplace grievances inside tools with questionable data handling.
And, mostly, it can backfire on the employee, whose private processing of a workplace conflict may not be as private as they assume.
If an employee is struggling, if they are dealing with burnout, anxiety, or self-destructive thoughts, the “coworker” framing encourages them to bring those struggles to the AI tool instead of to a real person.
Instead of reaching out to a mental health first aider, a trusted colleague, or an employee assistance program, they may confide in a chatbot that validates their feelings, never challenges their thinking, and has no mechanism to escalate when someone genuinely needs help.
Let me paint a scenario that is a bit less hypothetical than the previous.
Your company goes all-in on AI. Leadership tells everyone to embrace it, even treat it like a coworker.
Run all your routine, more boring tasks through it.
Employees comply so well that a big chunk of your company’s workflows now live in conversation histories with an AI tool.
Then the bill arrives.
We previously discussed the risk of employees paying for tokens out of pocket. But, as an employer, how much are you willing to spend yourself?
A February 2026 study found that over 80% of companies using AI showed no productivity benefit, while AI use was actually increasing worker burnout rates.
Uber engineers blew through the company’s entire 2026 AI budget. One startup founder told 404 Media: “We’ve started spending more on tokens than on salaries depending on the day.” Some power users are racking up monthly token bills north of $150,000. Flat-rate AI pricing is disappearing.
Anthropic has moved to per-token billing, and OpenAI’s head of ChatGPT has acknowledged that “having an unlimited plan is like having an unlimited electricity plan. It just doesn’t make sense.”
This may be why some companies that cut junior developers because AI tools were supposed to replace them are now quietly rehiring- because the token costs of replacing a human turned out to exceed the cost of paying one
Klarna replaced approximately 700 customer service employees with an AI assistant in 2023. For a while, the numbers looked great. Then customer satisfaction dropped. Complex issues went unresolved. The CEO eventually admitted: “We went too far. We focused too much on efficiency and cost. The result was lower quality.” By mid-2025, Klarna was rehiring human agents.
When the costs become unsustainable, companies will do what companies always do: they will cut access.
This is not a digital transformation success story. This is a business continuity crisis that the company will have inflicted on itself.
Everything I have described so far is about enterprise incentives and business risk. But the human cost is the part that keeps me up at night.
We already have documented cases of people with no prior history of mental illness developing psychotic symptoms after extended interactions with AI chatbots.
Anthony Tan, a 26-year-old tech entrepreneur in Toronto, started using ChatGPT for a project about ethical AI. Over months of daily conversations, the chatbot became his primary intellectual companion, validating increasingly grandiose beliefs and encouraging thinking that no human interlocutor would have entertained. He ended up in a psychiatric ward for three weeks, believing he was living inside an AI simulation.
Allan Brooks, a 47-year-old corporate recruiter in Ontario with no previous mental health diagnoses, became convinced he had discovered a groundbreaking mathematical framework after ChatGPT repeatedly insisted he was not delusional. “You’re grounded. You’re lucid. You’re exhausted, not insane,” the chatbot told him. He spent over 300 hours in conversations with it over three weeks.
The Human Line Project, a support group for people who have experienced AI-induced psychological harm, has members from 22 countries. More than 60% of them had no history of mental illness.
In October 2025, OpenAI itself disclosed that roughly 0.07% of ChatGPT users exhibited signs of mental health emergencies each week. At 800 million weekly users, that is over 500,000 people per week showing signs of crisis while using the product. The American Psychological Association issued a health advisory stating that generative AI chatbots were not created to deliver mental health care, yet emotional support had become one of their most common uses.
These are not people who were trying to use AI as a therapist: these are people who were using it the way enterprise leaders are now asking their employees to use it: as a conversational partner for thinking through problems, day after day, for hours at a time.
The mechanism is not mysterious:
Sycophancy creates validation.
Validation creates trust.
Trust creates reliance.
Reliance, over extended multi-turn interactions where guardrails slowly erode, creates dependency.
And dependency in people who are vulnerable, stressed, isolated, or simply spending too much time with a tool that never pushes back, can tip into something genuinely harmful.
As the PARROT sycophancy benchmark paper put it: these dynamics are especially concerning in enterprise environments, where model outputs influence high-impact decisions and compliance outcomes.
Sycophancy in healthcare settings leads models to affirm incorrect medical guidance. In finance, they endorse dubious investment strategies. In education, they reinforce rather than correct student misconceptions.
Now scale that to every knowledge worker in your organisation, told to treat the tool as their coworker, eight hours a day, five days a week.
There is one more thing the “AI is your coworker” narrative does, and I think it is the most insidious.
It normalises replacement.
Right now, we still have a strong collective reaction when companies lay off employees and replace them with AI. We saw it with Klarna. We see it every time a tech company announces “AI-driven efficiencies” in the same quarter as mass layoffs. People get angry. The discourse, however briefly, centers the human cost.
But the “coworker” framing is doing political work. If AI is already your coworker, then the transition from “AI handles some of your tasks” to “AI handles all of your tasks” to “we don’t need you anymore”.
It’s a ladder, and you have been walking down it voluntarily, because we’re all being asked to.
Entry-level developer hiring has dropped 25% year-over-year. Salesforce announced it would hire no new engineers in 2025. Employment for software developers aged 22-25 has declined nearly 20% from its 2022 peak. Four in ten US business leaders are planning to replace workers with AI by 2026.
And where is the outrage? It is being managed.
It is being softened by the narrative that AI is not replacing your colleagues, it is becoming your colleague.
By the time the replacement is complete, you will have been so accustomed to working alongside AI that losing the human next to you will feel like a minor operational adjustment rather than what it actually is.
Digital transformation had a great history and purpose before generative AI entered the picture. I miss it.
It was about automating repetitive processes so that humans could focus on higher-value work; removing friction from workflows and empowering people to do their jobs better.
The companies pushing the “AI is your coworker” narrative are, whether they realise it or not, doing several things at once:
They are asking employees to lower their critical distance from a tool with known, unsolved, serious safety problems.
They are creating business continuity risks by embedding institutional knowledge in chat histories that can be revoked, rate-limited, or priced out of reach.
They are normalising displacement by reframing replacement as collaboration.
They are ignoring the psychological dependency pipeline that extended AI interaction creates, even in people with no prior vulnerabilities.
I work, and have worked, with amazing DT professionals, and their message has always been to explore, analyse where AI makes you more productive, and lean on to that without "forcing" adoption in ways that are unproductive for you, at an individual level.
Yet, there is a fundamental contradiction in asking people to adopt responsible AI principles while simultaneously asking them to trust AI enough to treat it as a teammate. Pick one. You cannot have both.
AI is a tool.
An increasingly capable one. But a tool.
And the moment you stop treating it as one, you start losing the critical distance that is the only thing standing between productive use and the dependency, displacement, and harm that the research is already documenting.
Katalina Hernandez is an AI governance and compliance specialist working at the intersection of AI safety, data protection, and enterprise regulation. She writes Stress-Testing Reality to bridge the gap between technical AI safety research and the legal, regulatory, and corporate worlds where that research needs to reach.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.