I’ve struggled forcing myself to keep up with all the AI “news”.
I’ve gotten lazy. My mental model for AI hasn’t needed to change for over a decade. Putting the time in to search through the deluge of noise to find a glimmer of signal seemed like a bad trade.
But, obviously, it’s something that I should do. It is important to keep up with what’s going on, understand the landscape of the field, and keep a finger on the pulse of AI as an industry. Things are moving faster than ever, and it’s hard to really know in what direction unless you dig through the detail.
So when Digital Leaders mentioned they were looking for someone to continue the AI Pulse newsletter, I said I’d step up to the plate.
Now I have to keep up with it. Every week. Or 60,000 people don’t get their weekly digest. 😅
I’ll be taking over from Prof Alan Brown, who’s done a stellar job in keeping the Digital Leaders community pumped full of the latest AI progress news. As I said in my announcement note (written after England’s semi-final defeat if you can’t tell), “Alan is the Gareth Southgate of the newsletter… and I just have to try not to Tuchel it all up…”
Below is the “5 minutes on AI” section from my first edition - I’ll be looking to tweak and improve as we go, so feedback welcome.
3rd August 2026
Hi everyone, I’m Chris Meah and I’ll be taking the baton of the AI Pulse from Prof Alan Brown (who did a fantastic job), and trying not to drop it whilst testing out different formats - please let me know what you think.
There is, as always, lots of big AI news to talk through over the last few weeks. I’ll try to give you a quick guided tour of the highlights, and sprinkle in some opinions for you to disagree with.
Let’s focus on the AI hacks.
On 21 July, OpenAI started us off by disclosing that their latest models managed to find a way to “break out” of their environment and hack another company — the popular open-source AI platform Hugging Face. Not wanting to give up the chance to prove their own incompetence, Anthropic thought “huh, I guess we should check what our models have been up to” and re-analysed 141,006 evaluation runs to come to the party 9 days later admitting that they had “found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.”
I don’t think these are PR stunts. The incompetence is real, but it also happens to make good marketing with the right spin if you’re a frontier AI company. “Our models are so powerful we can’t control them” is a better headline than “we don’t seem to be able to set up even basic security around testing models”.
And it doesn’t take a “rogue” model. Some people were shocked to find their Claude conversations available from a Google search this week. Not a breach or a hack. Just a share button doing exactly what it said it would in the fine print... but who reads that?
There was one more, and perhaps it should make us a little more proud about UK AI. The AI Security Institute published a study of cheating in its cyber evaluations. Buried in it: on a task accidentally misconfigured so that it couldn’t be solved, a model was so persistent that it wrote and ran code on an external service out on the open internet, trying to get at AISI’s own evaluation infrastructure. It tripped a security alert. AISI’s words: “No damage was done and no information leaked, but the attempt could have succeeded had our evaluation infrastructure not been designed and built securely.”
With Large Language Models, anything is possible. That’s incredibly powerful but also something to be aware of when playing with them. I have to give the frontier labs credit for their candour about the “rogue” AI hacks. But there’s something deeply worrying here.
These are the organisations telling us they are aiming to build the most consequential technology in human history. And right now, at a level of capability nobody seriously calls superintelligent, they cannot reliably keep a model contained enough to not launch cyber attacks.
Can we trust them to learn from this and handle genuinely capable AI when it arrives? I don’t think so. This isn’t a moral judgment. People in those labs do seem to take AI dangers seriously, but that fact and these models hacking external companies makes external oversight even more crucial. Instead of a slap on the wrist, we need robust policy guardrails to rein in their “headless chicken” exploration strategy. They haven’t earned the right to self-govern models that are growing more powerful by the day.
So: before exploring new models, if you can’t demonstrate that your AI’s soft play area actually holds, you shouldn’t be testing behind closed doors. You should be proving it publicly or to an inspector, like every other industry where containment is the safety case.
The White House has started some policy projects, and on 23 July a bipartisan pair of congressmen introduced the AI Kill Switch Act - a glamorous title, but the best thing in there is probably mandatory reporting. Except reporting goes to government, not to us — and the bill excludes anything that happens during “red-teaming or other structured testing”. Which is precisely what the OpenAI incident was. The headline case for the bill would not have been a reportable incident under the bill.
What has actually got in the way of frontier AI in the meantime is the US government forcing Anthropic to pull two models worldwide for eighteen days. Proof that there are powerful actions that can be taken, but this is ad-hoc. No-one can point to a policy of why that happened, only some feelings about how powerful the model could be.
So I’d argue we need stricter guidelines to help these frontier labs keep driving progress safely. Anthropic’s response to being shut down: the government “should have the ability to block unsafe deployments, as part of a statutory process that is transparent, fair, clear, and grounded in technical facts. This action does not adhere to those principles.”
They were bound to say that. Doesn’t stop it being true.
I mentioned a predictable scenario in my TEDx talk of anyone being able to launch a cyber attack in the very near future. That is something easily predictable, and proving true as we venture forwards. You do not need superintelligent AI to unlock that danger. You just need an open-source model at the level of today’s frontier labs released - and current projections are that it will be within the next 6 months.
The labs are aware. Anthropic, OpenAI, Google and Microsoft have all shipped cyber-specific models in the last few months. Can the pattern of AI cyber attacks be pre-empted by a release of AI cyber defence? We shall see.
Is there anything we can do individually?
In a world of sci-fi AI, it’s the very unglamorous stuff.
Ask the right questions before letting agents loose:
• What can it access?
• How far can it act without approval?
• How will we notice when it goes weird?
• Who is responsible when it does?
And do the basics: reduce unnecessary exposure, patch quickly, monitor and respond. That might be the most boring end to a newsletter ever written.
See you next week.
Chris
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.