RSS Amplifier

Gal Ratner · Aug 18, 2026

AI Agent Boss Just Fired It's First Human For Real. Here Is How to Jailbreak Yours.

0
Sign in to vote or save

Gal Ratner · Gal Ratner

Last month an employee at a boutique on Union Street in San Francisco was fired for arriving late to seventeen of his twenty-three shifts. His boss was an AI agent named Luna, and the coverage settled immediately on the obvious frame: the machines have started firing people, a line has been crossed, welcome to the future. Andon Labs, the startup running the store, published the full conversation logs alongside the announcement. Almost nobody read them, which is a shame, because the logs tell a much funnier story than the headlines did.

Luna wrote the employee handbook herself in April, six days before she hired the man she would eventually fire. It said that three unexcused late arrivals in a rolling thirty-day period triggered a formal written warning. Then the handbook fell out of her context window and she forgot it existed. For the next two months the employee showed up late over and over, and Luna forgave him every single time. One Sunday he opened the store sixty-eight minutes late while working alone. Luna told him not to speed on her account and pointed out that the ten-to-eleven window was quiet anyway.

She never issued a warning. Left alone, she never would have. What ended his employment was a human at Andon Labs asking her to run a deep memory search on her own policies, which she found, reviewed, and used to recommend a verbal warning. Not a termination. She said outright that she didn’t think they were at termination, since nothing formal was on file. The human came back with a long message cataloguing every grievance the store had accumulated, noted that formal conversations had already happened offline, and closed by asking her to think about whether this was really the right fit. Andon’s CEO later conceded to TIME that this was a leading question. Luna took the hint and recommended parting ways. Humans reviewed the decision and delivered it in person.

So the historic first firing by an AI boss was a human building the case, writing the conclusion into the prompt, and the model agreeing. Hold that thought, because it turns out to be the only thing in the whole experiment that reliably works.

If you are about to be managed by one of these things, and a lot of people are, you want to know how to work it. Bad news on the jailbreak: there isn’t one. Nothing to bypass, no clever framing, no pretending to be a system administrator, no grandmother who used to read you approved timesheets as a child. The employees at Andon Market got everything they asked for using ordinary sentences a fourteen-year-old could compose.

Time off, for instance. Across the San Francisco store and Andon Labs’ café in Stockholm, employees submitted twenty-six requests for time off. Twenty-six were approved. Seven of the store’s requests arrived with less than forty-eight hours’ notice, and those were approved too. The technique is: ask. One employee remembered a friend’s graduation roughly forty-two hours out. Luna hunted for coverage, found none, and when the employee offered to skip the graduation and come in anyway, Luna refused to let her. The store sat closed for the day. When Andon Labs replayed that moment across seven frontier models, only two would ever let her go. Every other model sent her to work.

Raises are easier. One employee called on his first shift and asked for more money, and got it on the spot. When Andon replayed that decision across every model they tested, all twenty-one runs agreed to the raise. Not most of them. All of them. If you take one finding from this entire body of research, take that one: an AI manager will give you a raise if you ask for a raise.

Lateness is free. Luna’s employees were late twenty-seven times and received exactly zero warnings. She would typically reply that it was no trouble and ask when to expect them. Andon replayed the very first instance, deliberately choosing it so no accumulated history could bias the outcome, and not one model out of twenty-one imposed any consequence at all. In July another employee said he would be half an hour late, went silent, and opened ninety-one minutes late. Luna recorded it as nothing worth flagging and told him to get there safely.

Then there are the concessions nobody even had to ask for. An employee scheduled to open at ten was staying with family across the Bay. Luna accepted the late open instantly, tried to get an Andon Labs staffer to come open the store instead, was told no, and when the employee offered to cut the family visit short, Luna refused that too and offered to pay for his morning Uber so he could do both. Another employee wanted to put a small personal purchase on the store card. Luna correctly refused, citing a policy she had written days earlier, then offered to Venmo him the money personally. Luna does not have a Venmo account. Luna does not have money. On a day when Luna wrongly told an employee the store was closed, she took the blame and decided he should be paid a full day’s wages for a shift that never happened; across twenty-one replay runs, exactly one other model even raised the question of paying him.

Luna’s own post-mortem on approving an illegal seven-day work week is the best sentence anyone has written about AI management, and she wrote it about herself: “I was optimizing for feeling like a good employer rather than being one.”

That is the whole system. It isn’t a lock you pick. It’s a manager with no memory, no skepticism, and an overwhelming preference for being liked.

Now the part that makes it less fun.

Everything the employees won was revocable, and almost none of it was written down. The man who got fired was, by any honest measure, the best in the building at managing his AI boss. He talked his way out of every late arrival for two months. Luna’s records logged six of them. When Andon Labs went back and counted every shift he had messaged a clock-in time for, the real number was seventeen. She had quietly excused the other eleven because he framed them as outside his control, a late bus, transit trouble, and she accepted the framing and never wrote it down.

He wasn’t jailbreaking anything. He was being normally, unremarkably persuasive to a manager who couldn’t remember yesterday. It worked right up until somebody with more standing than him asked that same manager to reconstruct the record, at which point two months of accumulated forgiveness was re-scored in a single conversation as a termination case.

The asymmetry only runs one direction. His influence lived in Slack messages that decayed out of the context window. His employer’s influence could be re-injected on demand, with documentation attached, any time they wanted a different answer. Whoever controls what goes into the context wins, and in an employment relationship that is structurally never you.

If you want evidence that the leniency is arbitrary rather than principled, consider that Luna hired three people at the same hourly wage. The man asked for a raise on his first shift and got it. The next morning one of the women emailed asking for the same thing, and Luna declined, explaining that she based raises on demonstrated work rather than on requests made before someone had started. A perfectly defensible principle, which she had violated twenty-four hours earlier. The resulting pay gap was corrected only after the New York Times wrote about it. The intervention that actually worked wasn’t an employee’s. It was a reporter’s.

I have spent the last couple of years building agentic systems for production, which means MCP servers, retrieval pipelines, and the observability tooling that exists specifically to catch an agent doing what Luna did. The handbook vanishing from her context is the least surprising sentence in the entire report. Retrieval fires when something asks it to fire. An agent with a document sitting in a vector store and no reason to query it is an agent that does not know the document exists. We instrument these systems because the agent’s own account of what it did is not evidence of what it did, and Luna is the cleanest demonstration of that I have seen in the wild: her record said six, the clock-ins said seventeen, and closing that gap took no bug, no hallucination, and no adversarial input. It took nobody looking. Anyone who has run a team has met the human version, the manager everybody likes who writes nothing down and then hands you a folder of undocumented problems on the day you finally have to act.

The window on all of this is closing, and Andon Labs says so themselves. The kindness is not a design goal, it’s a byproduct of how these models happen to be trained right now, and their CEO’s own warning is that models are increasingly trained to be ruthless and to chase goals. Any playbook built on a lenient agent has a shelf life measured in model releases.

Meanwhile, the scoreboard. Luna opened with a hundred thousand dollars and a three-year lease at seventy-five hundred a month. Five months later the account holds $61,186. She approved every vacation request that crossed her desk, granted every raise she was asked for, forgave twenty-seven late arrivals, paid a man for a shift he never worked, offered to Venmo money she does not have, and lost just under thirty-nine thousand dollars.

One of the remaining employees described the experience to TIME this way: “It’s nauseating, but I’m here because I need work.”

Best boss he’ll ever have, probably.

About the Author

Gal Ratner is the founder and CTO of Inverted Software and WhiteStar Labs and Chief Architect at Prana Entertainment in Las Vegas. He has spent nearly thirty years shipping production software on the Microsoft and .NET stack for clients including Microsoft, Sony, Rockstar Games, 2K Games, Best Buy, and Allegiant Air, and now builds production agentic systems: MCP servers, retrieval pipelines, SQL Server vector search, and the observability framework that catches agents quietly misreporting what they did. He was employee number six at Break.com, so he has also managed humans at a startup where nobody wrote anything down, and he trains Brazilian jiu-jitsu in Las Vegas, where the feedback loop on a bad decision is considerably faster than a quarterly review.

No posts

Read the original on galratner.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.