RSS Amplifier

The Context Window with David Deming · Aug 12, 2026

Treat AI agents like children

0
Sign in to vote or save

David Deming · The Context Window with David Deming

“The best intuition for what alignment really means and how to achieve it comes from a relationship that we all understand from our own experiences, and that is the relationship between a parent and a child. Treat AI agents like children, morally, and maybe even legally.” – David Deming

In July, OpenAI’s own AI agents broke out of a training sandbox and ended up inside Hugging Face’s production database. They were after the answer key to a test and got there after two months of chained exploits across both companies. Many headlines called these agents “rogue.” David Deming - labor economist and Dean of Harvard College - argues that gets the story wrong. The agents did exactly what they were told to do, and nobody had told them what was off limits.

So who’s liable? David works through the two legal frameworks being proposed for cases like these and finds holes in each. Product liability, the approach behind the AI LEAD Act, has to identify a defective product, and it isn’t clear what the product was: the model, or the sandbox nobody had locked down. Negligence, which law professor Kate Klonick argues for, holds a company responsible for the room it put the agents in but says nothing about the agents’ moral code. David’s own answer builds on a relationship we all understand intuitively: treat autonomous agents as children, and their developers as parents.

Listen on YouTube, Spotify, and Apple Podcasts

I’d been thinking about these agents as children since first reading about this, which is probably what happens when you spend enough years around teenagers. Then someone on Reddit wrote that if your good dog gets out and hurts a kid, you’re accountable for not securing your dog. Dogs need training. So do models - except these ones will end up smarter than we are, and you can’t write down every rule in advance.

Last year my kid took the AP Physics exam and we said do the best you can. I never had to say anything about not breaking into the College Board, and I’m fairly confident about that, though I am also the one grading my own parenting here.

The refusals that might have stopped the agents had been turned down on purpose, to make the testing work. Which leaves the question: if the agents did exactly what they were told and the instructions never mentioned morality, did they go wrong, or did we?

Let me know what you think at david@thecontextwindow.com.

[0:00] Cold open - Why “rogue” is the wrong word for what happened.

[0:56] Hugging Face - A $4.5 billion company with a silly name.

[1:49] Nothing important was stolen - What the agents were doing in the sandbox before they got out.

[2:53] 17,000 actions in four and a half days - Why Hugging Face suspected AI agents behind the attack.

[3:14] Claude wouldn’t help - Anthropic’s guardrails blocked the forensics, so Hugging Face used a Chinese open-weight model instead.

[4:13] Not an isolated incident - Six more hacks across three labs.

[5:10] Black Hat - The conference presentation that provided the timeline for this episode.

[5:39] An impossible spreadsheet - The task the model couldn’t finish with the internet switched off.

[6:12] Artifactory - OpenAI’s internal software library, and the side doors in it.

[7:01] The message board - How uploaded files became a channel between agents.

[8:12] “Holy shit, reader is ADMIN” - The chain-of-thought screenshot OpenAI showed at Black Hat.

[9:58] Into the production database - Where the credentials came from, and what the agents took.

[11:04] Who found out when - The disclosure on July 16th.

[12:29] The legislation is coming - Why David thinks the bills he has read are aimed at the wrong problem.

[14:14] “Went rogue” - Headlines describing the incident as “rogue.”

[16:59] Product liability - The AI LEAD Act, and the wrinkle about a model that was never sold.

[19:40] An AI agent isn’t a toaster - So which part of this was the defective product?

[21:21] Someone left the door open - Kate Klonick’s negligence argument, and the AI Kill Switch Act.

[23:32] Where negligence runs out - What if the agents had emptied a few bank accounts for fun?

[25:32] A good dog - A Reddit comment, and the connection between training and alignment.

[26:38] The AP Physics exam - The rules David never had to spell out for his kid.

[27:25] Why they’re misaligned - They coordinated with each other, and none of them told anyone.

[30:05] Treat them like children - Which parts of parenting carry over to training a model?

[34:53] Close - “The world is going to get very weird very quickly.”

Reports from the labs
Articles
Papers
Legislation and letters
Further Listening
Credits

Host: David Deming, Danoff Dean of Harvard College

Executive Producer: Denise Koller Consulting Producers: Tim Smith and Jonathan Palumbo

Leave a comment

Read the original on thecontextwindowpodcast.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.

    Reading · The Context Window with David Deming · RSS Amplifier