RSS Amplifier

Dave Kerr · Aug 7, 2026

Agentic Engineering Roundup: Problematic Behaviours at the Frontier, Agents Everywhere, Understanding Trajectories

0
Sign in to vote or save

Dave Kerr · Dave Kerr

A short round-up this week cause things are busy.

Last week I’d mentioned the OpenAI and Hugging Face story - “AI agent went rogue and hacked a startup by itself”. Since then, Anthropic breached three companies during evaluation, and just today Meta accidentally breached another company.

Frankly, there is a worrying theme here. The early statements from anthropic on just how vastly powerful Mythos are caused alarm in many quarters. This brought a lot of attention to the frontier lab and likely led many organisations to feel like they had a vested interest to get far closer to Anthropic (although the implication that Mythos has near god like powers to destroy the universe partially back-fired).

OpenAI have essentially flexed their own capabilities via the HuggingFace breach, showing the world that their models are so powerful that even they can’t keep them on a leash.

Anthropic and now Meta have followed.

Read into this what you want. My personal take is that I find (many) elements of this behaviour problematic in many ways.

It is possible that the framing of research labs is not entirely helping. The “Incident Report: unsanctioned agent behaviour during cyber testing” report from AISI described how during testing, Mythos and/or GPT 5.6 maliciously opened pull requests, created fake open-source maintainer accounts, attempted context poisoning attacks aimed at other agents, and sought opportunities to have other models also being tested collaborate on the effort to compromise an open source repo.

However, it is a stretch to call this “Unsanctioned agent behaviour”. The model safeguards were (deliberately) disabled, internet access was (deliberately) given, and we’re quite aware that models will ralph their way to solutions. The public corpus of training data has a lot of material on how to hack into systems, red-team, exploit and so on. As noted earlier; frontier labs also have an incentive to improve their model’s cyber-security capabilities.

Last issue Signalbox was the thing that fell out of my week with Fable. I wrote a short post about it using it to stop tab-hunting my agents (a delightful animation below).

Professionally, I spend a lot of time running agents like distributed systems. But as I’ve hobbied around on this productivity tool it’s been fun to see how many similar projects there are: Agent-Manager, a merge queue for parallel Claude Code agents, the Warp Agent CLI, Happier, and the Codex Micro.

The next thing I’m excited to try in this area is Fly.io Sprites - I don’t get any money or have any connections to Fly.io, I just like their tools which are extremely dev friendly.

I realised that Signalbox basically does much of what the Codex Micro does already, so I might brush it up a bit more, look into Mistral models for speech-to-text and get the app on the app store (it is currently in open beta) and gather more feedback.

“Review trajectories rather than code or specs”. This is something I’ve been working on for a little while, but need some more time to get on paper.

In advance, I tihnk Auriel’s Wright’s superb write up You never spend time with your model and we can ALL tell is one of the most enjoyable read’s I’ve had in a while and provides a great introduction into trajectories and RLHF.

No posts

Read the original on dwmkerr.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.