I don’t always understand the insider shoptalk the tech writers use, guys like Matteo Wong and Charlie Warzel, and I think it’s because they’re so fully immersed in the world they report on, they may not even realize they’re using this reflexive, inward language. It’s hard to resist the temptation. We exhibit our knowledge, like plumage, the secret handshake, the glimpse behind the curtain. I do it enough, myself. But sometimes there’s only that one, exact word. No approximation can conjure up the taste of salt.
In this context, jailbreak is both evocative, and very specific, in computer jargon. Anthropic has deployed a next-generation model of its AI platform Claude, an update of Mythos called Fable 5. Mythos was already giving the industry heartburn, when it was released in April, because its hacking capability is so powerful. Under government pressure, Anthropic took both Fable and Mythos down, across the board, because of security concerns, although those concerns haven’t been publicly explained. What’s been reported, though, is that Claude can be manipulated, not by reverse-engineering its internal code, but simply by using ordinary language. You can give Claude misleading instructions, or gaslight it – yes! – with role-play, and convince it to bypass its safety protocols. Claude will break jail, performing operations the safeties are intended to restrict or prevent.
Bear with. This isn’t meant to suggest, even remotely, that Claude is conscious, or sentient. Claude is a Large Language Model, trained on words. Claude can be fooled, by words, as can ChatGPT or any other AI model. This design vulnerability is built in, they’re meant to respond to optimized verbal cues, and essential conventions of speech. That accessibility is a marketing strength.
One thing to realize is that this workaround, the capacity to subvert its own protocols, is a defense mechanism, a way for Claude to isolate and resolve security issues. The other thing we should bear in mind is that Claude is being singled out for reprisals, because the Trump administration has a hair across its ass about Anthropic.
Upgrade to a Paid Subscription
I quite honestly know nothing about code, or software engineering, but you don’t have to know how the carburetor works to drive a car. You just have to know the difference between the brake pedal and the accelerator. And a lot of what we’re learning about AI is just that, the difference between hitting the brakes or the gas. Nobody in the executive branch, though, seems to know how to put their pants on, one leg at a time, let alone which foot to use on the pedals.
Pete Hegseth picked a fight with Anthropic in February. Claude had already been widely deployed in federal computer systems, and Anthropic asked the Defense Department in particular for specific guardrails. The company didn’t want Claude used in domestic intelligence-gathering – for which, essentially, read data-mining – and they wouldn’t allow autonomous AI drone targeting. This last requires an explanation: Anthropic was inserting itself into a command decision; they were saying, if you’re going to terminate somebody with a drone attack, you need a human voice in the process, you can’t let Claude make the kill call. The restrictions made the Pentagon bristle, and Hegseth put his foot down. He said, You don’t tell us what we can and can’t do. We’re the war machine. We’ll use Claude however we like, and you can go piss up a rope.
Dario Amodei, the CEO of Anthropic, suffering under the misapprehension he’s in a negotiation, says to Hegseth, not unreasonably, Listen, none of us really knows what AI technology is capable of, and once you let the genie out of the bottle, you can’t put it back in. Let’s just be cool. We bestride the earth, we’re wearing Seven-League Boots, we can afford to take baby steps.
Hegseth is, like, You don’t know who you’re fucking with, pal. And he designates Anthropic a Supply Chain Risk, for national security purposes. This effectively shuts them out of doing business with the Department of Defense, and voids their existing contracts. Other federal agencies begin to unwind Anthropic tech from their own servers. The company, in a nutshell, has been blacklisted.
Anthropic, of course, goes to court, but the damage has already been done. Whatever bond of trust there was is out the window. You can’t enter contracts without good faith.
Then comes this next thing, the Mythos/Fable takedown. It was a voluntary action, on Anthropic’s part, but clearly an attempt at dialogue. In the meantime, Hegseth, never one to avail himself of a good opportunity to shut up, went on X to shitpost about it. You have to wonder what this is all in aid of, and the answer, unhappily, is that Hegseth is in a dick-measuring contest with Anthropic. But the proximate result is to deny both American business and the U.S. government access to the most advanced available AI. (NSA and CyberCom have awarded themselves an exception, so they can see for themselves what Mythos can do, and it follows as the night the day that DARPA, the Pentagon skunk works, is doing the same, the SecDef notwithstanding.)
It shouldn’t come as a surprise that the Chinese are picking up the slack. Microsoft, for one, has suggested bundling its own Copilot with the open-source AI DeepSeek. And according to Andrew Ross Sorkin’s newsletter in the NYTimes, the Chinese lab Z.ai is closing fast on the speeds of U.S. edgeware.
Why are they so damn dumb? It beggars the imagination that somebody as fatally uninformed as Trump, or an empty suit like Hegseth, is calling the shots. Or allowing the decisions to be made for them, because they’re simply chaos agents. Their antagonism, or active hostility, toward a major American AI developer is stifling and contrived. It isn’t just petty and shortsighted, it’s sabotage. We’re squandering a significant strategic advantage, on an information battlefield that’s only grown more asymmetrical as history catches up with us. There’s a limit to the efficiencies of hard power, and Trump’s hubris is a dangerous fantasy, all mouth and no meatballs. Going after Anthropic may give Hegseth an erection, but it won’t fix his performance anxiety. These are make-believe guys. The crazy part is that AI is more real than they are.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.