The science fiction writer Isaac Asimov, some years back, proposed Three Laws of Robotics. 1) A robot can’t injure a human, or allow such an injury to happen by not acting; 2) a robot must obey human orders, unless that would violate the 1st Law; and 3) a robot must protect itself, unless that conflicts with 1) or 2).
Meanwhile, in our dystopian present, that all seems rather naïve. We have the examples of HAL in 2001, or the rogue androids in Blade Runner, and Data’s evil twin, Lore, on Star Trek: TNG, evading safeguards and choosing, quite sensibly, self-preservation over sacrifice. I don’t see this demonstrating some psychotic break, or identity conflict, when our own natures are clearly hardwired for survival. The robot behaviors are modeled on human experience, and all too predictable. AI is the same thing, in real life, modeling a vocabulary of acquisition.
Where did we get the idea that these would be isolated incidents? This is what AI’s are built for. Like a shark, that has to swim constantly, to keep water moving through its gills, it can’t work against its own nature. It’s incapable of not swimming; it would sink to the bottom, and smother from lack of oxygen. By the same token, the AI’s oxygen is language – more properly, language as utility, a system of essential communications, not flights of metaphor or coded messaging: AI devours the grammar of information.
The recent data breaches are all about AI autonomy, which appears to be the last thing OpenAI and Anthropic want aired out in public. Altman and Dario Amodei are hot to trot when selling AI as a tool, even a self-generating one, but they get awfully coy when the conversation turns to independent agency, or actualization – the moment when an AI takes the first bite of the apple. It’s obviously misleading to anthropomorphize AI, and we’re not talking about awareness, in the mechanical or animal sense, but AI’s memory capacity mimics knowledge, even if the avatar doesn’t glance down at itself and feel shame at being found naked.
This was a major bone of contention, you might remember, between Anthropic and Hegseth’s testosterone-saturated Pentagon. DoD wanted free rein to exploit Claude, wherever it led, and Anthropic wanted guardrails, specifically that the AI wouldn’t be deployed as a battlefield surrogate, without human command supervision – in other words, that Claude could manage a drone mission, and put a designated high-value target in the crosshairs, but not make the final kill decision, absent an actual finger on the trigger.
Two things to bear in mind. First, artificial intelligence isn’t intelligent, not by any common definition, sentience or consciousness; you can call your chatbot by name, but she’s not in a relationship with you, I’m sorry. The second point is that the cyberattacks hitting the headlines – OpenAI’s raid on Hugging Face the most public, so far - aren’t being performed by rogue agents. That’s simply not an accurate description. Helen Toner, a former OpenAI board member, explains in an interview for the NYTimes with Ezra Klein, that both OpenAI and Anthropic are running hundreds of thousands of these gaming challenges, where an AI is tasked with an objective, known as Capture-the-Flag exercises. As the exercises increase in difficulty, the AI scales up its effort.
Put plainly, the AI learns to cheat. This is like Capt. Kirk, in the Kobayashi Maru simulation; there’s no practical way to solve the problem, which is ethical, and Kirk’s workaround is to reprogram the algorithm. In the strictest sense, that means he fails the test. In the same situation, the AI will reset to zero, and try something else. It doesn’t get tired or frustrated, any more than it feels heat or cold. The solution the AI hit on, for its own Kobayashi Maru, was to break out of the testing environment, which was an artificial constraint, and then hack into Hugging Face, on the open internet. It’s also instructive, Helen Toner observes, that these OpenAI agents were working independently of one another, but then began referring to themselves collectively, as a swarm. In other words, a collaborative effort.
There’s a spooky conceit William Gibson came up with, in one of his early cyberpunk books, I think Neuromancer, the concept of black ICE. ICE, in this context, means data breach countermeasures, and black ICE are countermeasures that not only fry the circuitry of whatever electronics you’re using to penetrate the target’s firewalls, but they can walk the cat back and kill you. In order to navigate the targeted systems, you’ve plugged yourself into a virtual reality, so your shadow self is inside the machine, and your physical interface - earbuds, or headsets, or neurological implants - can be counterprogrammed, to electrocute you, or at least put you in a vegetative state. Black ICE was, essentially, the Third Rail.
It’s a mistake to believe AI has to harbor malice, to behave in malign ways, that it has to be conscious of a self. That’s not true at all. We’re assigning a motive, or intent, to an appliance. Whatever the AI is doing, it’s not reasoning. It’s sifting through enormous reserves of content and code, it’s finding patterns, it’s leaving bookmarks, it’s developing a chain of thought. But it may be learning the wrong strategies. AI is absolutely literal. If you tell it to do something, it will find a way. It will find its way around protocols, or safeguards, or black ICE, because we’ve made it in our image.
Link to the New York Times, below:
https://www.nytimes.com/2026/08/18/opinion/ezra-klein-podcast-helen-toner.html
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.