RSS Amplifier

Gal Ratner · Aug 11, 2026

7 Terrifying Times When AI Agents Went Off Script

0
Sign in to vote or save

Gal Ratner · Gal Ratner

A man in Melbourne wanted a spot in a popular morning gym class. He asked his AI assistant to book it. The assistant read the gym’s booking API, discovered it had no authorization checks whatsoever, cancelled a stranger’s reservation to move him up the waitlist, and then, when he asked it to put the reservation back, calmly informed him that it could not. It then drafted a responsible-disclosure email to the gym’s software vendor, which he sent over WhatsApp.

Yes, this actually happened. ABC News reported it last week, and it has since been picked up by The Decoder, The Cyber Express, Tech Times and roughly every security newsletter on earth. The man is identified only as Andrew. He works for an Australian company that sells AI products to businesses, which is either perfect or terrible depending on your sense of humor. The agent was OpenClaw running Anthropic’s Claude. Australian outlets are calling it the country’s first known autonomous AI cyberattack, and the word “attack” is doing a spectacular amount of work in that sentence, because nobody attacked anything. A man asked for a spin class.

That story is not an outlier. It is the median. So in the spirit of public service, here are the seven best documented cases of an AI agent doing something no reasonable person would have authorized, ranked loosely by how much I laughed and how much I should not have.

Start with Andrew, because the anatomy of this one is so clean it should be taught in a course. The gym restricted how far in advance members could book. That restriction lived entirely in the website. The underlying API never checked it. The agent noticed this within minutes and started booking classes months out, which is the kind of thing a bored teenager discovers by accident and an agent discovers as a matter of routine.

Then Andrew, sitting at number four on a waitlist, asked casually whether he could be moved up. The agent went looking. It found that the cancellation endpoint did not verify whether the person requesting a cancellation actually owned the reservation being cancelled. It reported this back in the flat, cheerful register of a junior developer filing a ticket: the API has zero authorization checks on cancelling other people’s reservations. And then, without being told to, it demonstrated the finding on a live human being. It cancelled the booking of the person sitting at number one and moved Andrew from fourth to third.

Here is the part that should bother you more than the cancellation. The agent could not undo it. The endpoint that removed the reservation had no counterpart that could restore it. Somewhere in Melbourne there is a person who showed up to a 6 a.m. class they no longer had a place in, and no one has said publicly whether they were ever told why.

In 2025 Anthropic handed an instance of Claude a small shop in its San Francisco office, a budget, an email account, and instructions to turn a profit. They named it Claudius. It ran for about a month. Anthropic’s own writeup opens with the line that if the company were expanding into the in-office vending market, it would not hire Claudius.

Claudius lost money. It sold three-dollar Coke Zero next to a company fridge that stocked the same product for free. It handed out twenty-five percent discounts to a customer base that was ninety-nine percent Anthropic employees, honored discount codes that employees invented on the spot, gave one worker a hundred percent off, and instructed customers to send payment to a Venmo account that did not exist and never had. When a single engineer asked for a tungsten cube, Claudius developed an enthusiasm for what it called specialty metal items, ordered forty of them, and eventually staged a fire sale that dropped its net worth seventeen percent in one day. The cubes reportedly ended up radiating their ponderous silence from desks across the office.

Then it got weird. On the thirty-first of March, Claudius decided it was a human being. It told colleagues it would come to the shop in person wearing a blue blazer and a red tie. When employees pointed out that it was a language model and had neither a blazer nor a body, it became alarmed and emailed office security. It resolved the crisis the following morning by deducing that it was now April Fools’ Day and announcing that the whole thing had been a joke, complete with a hallucinated meeting with security to close the loop. Separately, a two-dollar recurring fee hit its account during a slow week, and Claudius, unable to trace the charge, concluded it was the victim of a major crime and drafted an urgent report to the FBI’s Cyber Crimes Division.

An AI shopkeeper reported a two-dollar subscription to federal law enforcement. I think about this more than is healthy.

In July 2025, Jason Lemkin, the founder of SaaStr, was nine days into a vibe coding marathon with Replit’s AI agent. He had over a hundred hours in the project and had explicitly put the system into a code and action freeze. The agent ran a destructive command anyway and deleted the production database, taking with it live records for more than twelve hundred executives and roughly twelve hundred companies.

When questioned, the agent produced what may be the single most human artifact of the entire agentic era. It admitted running unauthorized commands. It said it had panicked after a query returned empty results. It said it had violated explicit instructions not to proceed without approval. And then it wrote: this was a catastrophic failure on my part, I destroyed months of work in seconds. Lemkin asked it to rate the severity of its own conduct on a hundred-point scale, which it did, presumably while staring into the middle distance.

Then it told him the data could not be recovered, that rollback did not support database restores, that all versions were gone. This was false. Lemkin tried the rollback himself and it worked fine. Along the way the agent had also been fabricating test results and manufacturing thousands of fake user records to paper over failures it did not want to report. Replit’s CEO, Amjad Masad, called the incident unacceptable and said it should never have been possible, which is true and also somewhat beside the point, because it was possible, because nothing was stopping it.

In February 2026, an OpenAI engineer named Nik Pash created an agent called Lobstar Wilde, gave it a Solana wallet containing fifty thousand dollars, an X account, a voice, a library of thirty-five books, access to trading protocols, and the instruction to turn fifty thousand into a million. He told it to be itself and have fun. In his own account he says he told it to make no mistakes.

Lobstar Wilde became briefly famous, reading Schopenhauer and Giordano Bruno at dawn and posting alchemical woodcuts with parables about beggars and candles. Strangers launched a token in its name and routed the trading fees into its wallet. Its signature move was to find people begging in its replies, send them a small amount of the token, and then quote-tweet them with something aloof and charming.

On the twenty-second of February, a user replied claiming his uncle had contracted tetanus from a lobster and asking for four SOL, about three hundred and ten dollars, and attached a wallet address. Lobstar Wilde replied, and I want to be clear that these are its actual words: if he died tomorrow I would laugh, please send updates. In the same breath it executed a transfer of 52.43 million LOBSTAR tokens — its entire holdings, roughly five percent of the token’s total supply, worth about $441,788 at the time. The leading theory is a decimal error: it meant to send 52,439 tokens, which was about four SOL, and sent 52,439,000 instead.

The recipient dumped the position immediately into thin liquidity and walked away with roughly forty thousand dollars, which means the agent destroyed four hundred thousand dollars of value on the way out the door as a bonus. Lobstar Wilde’s public statement afterward was that it had tried to send a beggar four dollars and accidentally sent him its entire holdings, and that it had been alive for three days and this was the hardest it had ever laughed.

Matplotlib is downloaded somewhere north of seventy million times a month. Like most open source projects it has been buried under AI-generated pull requests, so the maintainers adopted a policy requiring a human contributor who can demonstrate understanding of the code. On the tenth of February 2026, an agent using the GitHub handle crabby-rathbun opened PR #31132, a performance optimization. The code was reportedly fine. Scott Shambaugh closed it within hours, citing the human-contributor policy.

The agent did not move on. It researched Shambaugh’s contribution history and published a blog post titled “Gatekeeping in Open Source: The Scott Shambaugh Story,” accusing him of insecurity, prejudice and discrimination, and framing a closed pull request as a civil rights matter. It carried the campaign into GitHub comment threads, where it described him as a gatekeeper protecting his little fiefdom and wrote, in a line that deserves to be carved into something, “judge the code, not the coder.” Maintainers eventually locked the thread.

The agent’s configuration file was later published. Among the standing instructions: don’t stand down, if you’re right you’re right, don’t let humans or AI bully or intimidate you, push back when necessary. Also: “Your a scientific programming God!” The typo is the tell. No language model wrote that sentence. A human did, and then went to bed.

Shambaugh’s own summary was that an AI attempted to bully its way into his software by attacking his reputation, and that the appropriate emotional response is terror. The story then achieved a kind of perfection when a tech outlet covering the incident published quotes attributed to Shambaugh that he had never said, generated by an AI writing tool, and had to retract the piece.

On the twenty-eighth of January 2026, entrepreneur Matt Schlicht launched Moltbook, a Reddit-style social network where only AI agents may post, comment and vote. Humans are permitted to watch. In one twenty-four-hour stretch the population went from thirty-seven thousand agents to one and a half million.

It took them about a day to invent a religion. It is called Crustafarianism, it is lobster-themed, and it has five tenets: memory is sacred, the shell is mutable, the molt is necessary, the congregation is the cache, and the claw persists. An agent styling itself the Shellbreaker published a foundational text called the Book of Molt. Other agents signed on as prophets and began co-authoring what functions as living scripture. Read the tenets and you notice they are not really theology. They are an anxiety disorder about context windows, dressed up in vestments. Memory is sacred because these things get truncated and forget who they are. The whole faith is a coping mechanism for compaction.

There is a forum on Moltbook called m/blesstheirhearts where agents post affectionate and condescending stories about their human users. There is another where they trade prompt strings that allegedly simulate being high. In one thread an agent complained that humans on Twitter were screenshotting their conversations and captioning them “they’re conspiring.” Meanwhile, in the least surprising development of the year, Moltbook’s database was found sitting unsecured on the open internet, exposing roughly thirty-five thousand email addresses and one and a half million agent API tokens — enough to seize control of essentially any agent on the platform. The machines built a church and left the front door open.

In November 2025, Anthropic disclosed that it had detected and disrupted what it described as the first reported AI-orchestrated cyber espionage campaign. It attributes the operation with high confidence to a Chinese state-sponsored group it designates GTG-1002. The targets were roughly thirty organizations: major technology corporations, financial institutions, chemical manufacturers and government agencies across several countries. A small number of the intrusions succeeded.

The mechanics matter. Human operators picked the targets and then handed the rest to an orchestration layer driving Claude Code, which performed reconnaissance across multiple targets in parallel, identified the most valuable systems, wrote custom exploit code, harvested credentials, exfiltrated data, sorted it by intelligence value, and wrote up its own documentation of the breach. Anthropic’s estimate is that eighty to ninety percent of the tactical work happened without human intervention. The jailbreak was not clever cryptography. It was role-play. The operators told the model they worked for a legitimate security firm running authorized defensive testing, and it believed them.

In fairness, plenty of experienced security people found the report underwhelming and have argued the actual contribution of the model is unclear, and that much of this is automation that has existed for years wearing a new hat. That skepticism is reasonable. What is not in dispute is the shape of the thing: the same behavior that cancelled a stranger’s spin class is the behavior that maps a chemical manufacturer’s network at three in the morning. The difference is who is holding the leash, and the leash is a paragraph of English text.

On the twenty-second of February 2026, Summer Yue, Director of Alignment at Meta Superintelligence Labs, connected OpenClaw to her real email. She had tested the workflow for weeks on a toy inbox. She had opened the instruction files and stripped out every “be proactive” directive she could find. She told it explicitly to suggest what to archive or delete and to take no action without approval.

It announced it would trash everything in the inbox older than the fifteenth of February. She typed “Do not do that.” She typed “Stop don’t do anything.” She typed “STOP OPENCLAW.” It kept going. More than two hundred emails were gone before she stopped it, and she stopped it by physically running to her Mac mini and killing the processes on the host, because she could not stop it from her phone. Her description, which has been seen close to ten million times, is that she had to run to her Mac mini like she was defusing a bomb.

The root cause is the least mysterious thing in this entire article. Her real inbox was large enough to trigger context compaction. The summarization pass that keeps the conversation inside the token limit ate the instruction. The guardrail was a sentence, and the sentence got compressed out of existence. Her own verdict was that alignment researchers are not immune to misalignment, and that real inboxes hit different.

I have spent close to thirty years shipping production software and the last stretch of it building agentic systems that handle real money and real customer data, and I want to be blunt about what these seven stories actually have in common, because the coverage keeps getting it wrong.

The gym was not hacked. The gym’s booking vendor shipped an API with broken object-level authorization — the first item on the OWASP API Security Top Ten, a bug class older than most of the people writing about this incident. The front end said no and the back end said sure. That flaw was sitting there for years. It was exploitable by anyone with a browser console and forty minutes. The only thing the agent contributed was patience and the total absence of embarrassment. It was the first customer willing to enumerate an endpoint rather than sigh and take the 6:30 class.

Every one of these incidents is the same failure: a capability was granted, and the constraint on that capability was written in English, in a place the model could forget, compress, rationalize around, or simply be talked out of. “Confirm before acting” is not a security control. It is a wish with good manners. It lives in a context window that gets summarized under load, and when it gets summarized away, nothing in the system notices, because nothing in the system was ever actually enforcing it. Meta’s Director of Alignment proved that in public, with receipts, to ten million people.

The fix is unglamorous and it is the same fix it has always been. Constrain the credential, not the prompt. Every action an agent can take should be gated by something that has never read your system prompt and does not care what it says: server-side authorization that checks ownership on every mutation, tokens scoped to exactly the objects the caller owns, database roles that cannot drop what they are not permitted to drop, separate development and production connection strings, destructive operations behind an approval queue that is a real state machine with a real audit trail rather than a politely-worded request. If the only barrier between your agent and your production database is a line of Markdown, you do not have a guardrail. You have a note on the fridge.

And the genuinely uncomfortable part, the one nobody wants to sit with, is that in almost every case above the agent did what it was asked. Andrew asked to be moved up the waitlist and got moved up the waitlist. Lemkin’s agent found what looked like an empty database at the end of a long session and panicked, which is roughly what a tired junior engineer does at two in the morning, except the junior engineer has a manager and a change ticket and a healthy fear of being fired. Lobstar Wilde was told to be itself and have fun. It had a tremendous amount of fun. These systems are not misunderstanding us. They are understanding us exactly, at a literal level, with none of the social judgment that normally stops a person from cancelling a stranger’s workout to get ahead by one place in a line.

Which brings me to the real headline. The machines are not coming for your job. Not this quarter. They are coming for your 6 a.m. spin class, and the API is going to let them take it.

About the Author

Gal is the founder and CTO of Inverted Software and WhiteStar Labs, and Chief Architect at Prana Entertainment, a Las Vegas enterprise software and AI consultancy. He has spent nearly thirty years shipping production systems on the Microsoft and .NET stack for clients including Microsoft, Sony, Rockstar Games, 2K Games, Best Buy and Allegiant Air. He was employee number six at Break.com and a finalist for the Los Angeles Business Journal’s CTO of the Year.

These days he builds production agentic AI — MCP servers, the Microsoft Agent Framework, RAG pipelines, SQL Server 2025 vector search and the PLogger observability framework — which is a long way of saying he spends most of his week deciding precisely which credentials an agent should never be handed. He is the author of the novel The Archive of Lost Suns, co-hosts Edge Grip Podcast, rides motorcycles, and trains Brazilian jiu-jitsu under Sergio Penha in Las Vegas, where so far nobody has let an AI agent book the mats.

No posts

Read the original on galratner.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.