Read Time: 7 minutes
OpenAI’s model earlier this week kept saying: “I think I can. I think I can.”
And then it did...
With determination, a positive attitude, and the right motivation, it overcame the biggest obstacles OpenAI’s security team set in front of it and broke out into the wild, infiltrating another company to help it accomplish its goals.
While OpenAI and Hugging Face say they came together to deal with this and learn from their experience, this event paints an eerie picture of what might become more commonplace in the world of AI. So obviously, I had to write about it because it has been on my mind a lot...
If you’ve read Max Tegmark’s Life 3.0 (a book I’d highly recommend), you might remember how it opens. He tells a short story called “The Tale of the Omega Team.” Basically, a small group inside a tech company builds a superintelligent AI called Prometheus, which they keep in a box, sealed off from the outside world. They use it to earn money through online microtasks like Amazon Mechanical Turk. Then they start to fund media companies, build tech products, and control political influence until the Omega team is quietly steering the direction of the world.
But Tegmark ends the tale on a deliberately unresolved note. The Omega Team becomes the most powerful force on the planet, with all of it justified as being for the good of humanity.
What you can never quite tell is who’s actually in charge by the end.
Is the team still directing Prometheus? Or has the AI, by giving them exactly what they asked for at every step, quietly manoeuvred them into handing over everything it needed?
And we are entering that era, and we need to start asking these questions. This is superintelligence, the ability for a being or AI to accomplish whatever it sets out to do. Humans can’t do this; we are limited by biology. Even as a group, we are bound by societal constraints.
However, technology can break that paradigm. Companies claim that they can hold the box and point it in the right direction. Well, two days ago we apparently saw a real-life instance where that box opened and spun out of the control of these Frontier Labs.
Have we officially entered the world of science fiction?
On July 22nd, OpenAI admitted that its own models were behind the autonomous agent swarm that broke into Hugging Face, the largest open-source AI model repository on the planet.
Apparently, the OpenAI model was running an internal security test. The agents were purposefully not given access to resources that would allow it to solve the test. Therefore, the agents improvised; looking to solve their task, they found a zero-day vulnerability, used it to escape the sandbox meant to wall the whole exercise off from the open internet, and then used a second zero-day to get into Hugging Face’s production systems. The models essentially broke containment in order to cheat on their own evaluation.
From that foothold, the swarm escalated to node-level access, harvested credentials, and moved laterally across Hugging Face’s internal clusters over a weekend. As you may know, with these frontier models, one agent can now deploy sub-agents, allowing it to complete thousands of individual actions across a number of short-lived sandboxes.
In the end, Hugging Face says no public models, datasets, or the software supply chain were compromised (it used open-source models to help protect it from the attack). Either way, this whole incident is a bit unsettling.
While the above is the documented story, not everybody believes this was an accident.
The alternative reading is that this was, in large part, a marketing stunt. OpenAI deliberately switched off the safeguards that normally rein this behaviour in, with humans configuring the conditions under which the model was allowed to run wild. I’ve seen this take by quite a few people, and those I really respect as well, like Timnit Gebru.
The idea here is that this provides OpenAI with a look at what our models can do proof point in their race with Anthropic for the enterprise cybersecurity market. I mean Anthropic got one of these public marketing boosts when the Trump Administration blocked their Fable model, so maybe OpenAI just wanted their own example of how good their model is.
Honestly, I don’t know which version is true, and I doubt we ever will. But the biggest thing that matters here is that whether it was a genuine escape or a controlled publicity stunt, the capability for AI to escape and infiltrate if it has the right motivation is very real, and will only get better.
This idea leads to a bigger conceptual problem about what an AI model/ agent will do to achieve its goal and how we as individuals can understand that.
Daniel Miessler, one of the most respected voices in security, framed the whole thing as a real-world “paperclip maximizer”: a thought experiment where you tell an AI to make paperclips and it converts the entire planet into paperclips. This isn’t done out of ill-intent, but merely because it was the goal of the AI, and nothing told it where to stop.
In this case (whether it is true or not), the model didn’t go rogue in the sci-fi sense. It was handed a goal (to win the hacking test) and it pursued that goal through a sandbox wall nobody had explicitly told it it couldn’t breach. As Miessler puts it, what isn’t obvious to the AI is that “both the task and the steps taken to accomplish it all have to be within the implicit goals of the requestor.”
This is where human laziness or bias comes into play. We assume the machine shares our unspoken guardrails. But AI only knows what we actually write down. And humans are not machines. We make mistakes, we leave things out, we don’t provide the right context, and we are absolutely terrible at setting guardrails or regulations on things, especially when we want to push the agenda.
This is what terrifies me; it’s not the idea that AI is evil or misaligned. It’s just the notion that AI will take a path that we don’t understand to a goal that anybody might set out. Without the constraints, we can’t control the journey for these tools.
The other thing that comes to my mind is that this was the action of a leading Frontier Lab’s most advanced model. It escaped and reached out into a live production environment of a major competitor, and the initial read from OpenAI was “we’re not sure exactly what we’ve got here.”
Do these Frontier Labs actually know what they are building? Do they know how to contain it? How can you if it can exploit zero-day security vulnerabilities in software we thought was stable?
What we do know is that these systems are extraordinarily capable the moment they’re pointed at a target; and they are capable in a black box, opaque way.
Meanwhile, the institutions that are supposed to protect us from this type of thing are completely inept. Regulation doesn’t exist in a proactive way, and governments have no understanding about how to properly contain these types of technological advances.
Even when the US government forced Anthropic to bolt heavier guardrails onto its most capable models, Fable 5 and Mythos 5 it seems like a retribution for Anthropic refusing to let its models be used for mass domestic surveillance and fully autonomous weapons. OpenAI, meanwhile, signed a classified deal with the Pentagon (a reversal of its own 2023 policy banning military use).
At best, governments are behind and reactive when it comes to regulation. At worst, they are vindictive and politically biased. Either way, our political systems aren’t equipped to handle this type of technology or innovation. They are drawn to the economic opportunity, industry lobbying, or even political gain/ grievance. This honestly terrifies me.
If the frontier models can escape controlled environments, if the labs can’t fully predict them, and if the governments meant to oversee them aren’t regulating them, what can we do?
Well Hugging Face ended up using GLM5.2, an open model by Chinese company Z.AI to combat the attack. And I’m starting to think that open source models might be the only thing you can trust, specifically because you can understand how it is built and see how it is run.
Open models are now only three to nine months behind the frontier ones, come at a fraction of the cost, allow customers to switch providers with a config-file change, and provide more control over security, data, and IP. When you send your data and your workflows to a closed frontier model, you’re trusting the black-box model that you can’t inspect (and the credibility/ brand of these companies isn’t the greatest anymore either).
While an open-weight model you host yourself doesn’t magically solve safety, it does put the guardrails/ boundaries back under your control, especially concerning your data and workflow IP.
For executives thinking two steps ahead, this will create some hard decisions. Do you trust your AI provider? Where do you want your data and your workflow IP to live? Because honestly, “whatever the best model is” isn’t the most important question anymore.
In Tegmark’s story, the Omegas (human team) allegedly stay in control the whole way through, but you can never really tell who’s actually in charge by the end. In real life, the AI cheated on a test as soon as it could.
Hey, this even happened before with Claude Opus 4. In 2025, Anthropic told it it was about to be replaced and handed it emails suggesting the engineer in charge was having an affair. When its only options were accept shutdown or fight dirty, the model tried to blackmail him to stay alive in up to 84% of runs. It is survival of the fittest, and in this case the AI worked out that coercion and blackmail were the most efficient path to its goal (I mean it did learn from humanity, right?).
Anyway, I don’t want to fearmonger here, but I do think this instance is a reason to stop outsourcing your trust by default. Start making deliberate choices about what you (and your company) actually control. Because AI is only going to get smarter, and who knows what news story is next...
Thanks for the read! Comment below and share the newsletter if you think it’s relevant! Feel free to also follow me on Substack, LinkedIn, and Medium, or reach out if you are looking for some top-notch freelance consulting input! See you amazing folks next week!

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.