“The best intuition for what alignment really means and how to achieve it comes from a relationship that we all understand from our own experiences, and that is the relationship between a parent and a child. Treat AI agents like children, morally, and maybe even legally.” – David Deming
In July, OpenAI’s own AI agents broke out of a training sandbox and ended up inside Hugging Face’s production database. They were after the answer key to a test and got there after two months of chained exploits across both companies. Many headlines called these agents “rogue.” David Deming - labor economist and Dean of Harvard College - argues that gets the story wrong. The agents did exactly what they were told to do, and nobody had told them what was off limits.
So who’s liable? David works through the two legal frameworks being proposed for cases like these and finds holes in each. Product liability, the approach behind the AI LEAD Act, has to identify a defective product, and it isn’t clear what the product was: the model, or the sandbox nobody had locked down. Negligence, which law professor Kate Klonick argues for, holds a company responsible for the room it put the agents in but says nothing about the agents’ moral code. David’s own answer builds on a relationship we all understand intuitively: treat autonomous agents as children, and their developers as parents.
Listen on YouTube, Spotify, and Apple Podcasts
I’d been thinking about these agents as children since first reading about this, which is probably what happens when you spend enough years around teenagers. Then someone on Reddit wrote that if your good dog gets out and hurts a kid, you’re accountable for not securing your dog. Dogs need training. So do models - except these ones will end up smarter than we are, and you can’t write down every rule in advance.
Last year my kid took the AP Physics exam and we said do the best you can. I never had to say anything about not breaking into the College Board, and I’m fairly confident about that, though I am also the one grading my own parenting here.
The refusals that might have stopped the agents had been turned down on purpose, to make the testing work. Which leaves the question: if the agents did exactly what they were told and the instructions never mentioned morality, did they go wrong, or did we?
Let me know what you think at david@thecontextwindow.com.
[0:00] Cold open - Why “rogue” is the wrong word for what happened.
[0:56] Hugging Face - A $4.5 billion company with a silly name.
[1:49] Nothing important was stolen - What the agents were doing in the sandbox before they got out.
[2:53] 17,000 actions in four and a half days - Why Hugging Face suspected AI agents behind the attack.
[3:14] Claude wouldn’t help - Anthropic’s guardrails blocked the forensics, so Hugging Face used a Chinese open-weight model instead.
[4:13] Not an isolated incident - Six more hacks across three labs.
[5:10] Black Hat - The conference presentation that provided the timeline for this episode.
[5:39] An impossible spreadsheet - The task the model couldn’t finish with the internet switched off.
[6:12] Artifactory - OpenAI’s internal software library, and the side doors in it.
[7:01] The message board - How uploaded files became a channel between agents.
[8:12] “Holy shit, reader is ADMIN” - The chain-of-thought screenshot OpenAI showed at Black Hat.
[9:58] Into the production database - Where the credentials came from, and what the agents took.
[11:04] Who found out when - The disclosure on July 16th.
[12:29] The legislation is coming - Why David thinks the bills he has read are aimed at the wrong problem.
[14:14] “Went rogue” - Headlines describing the incident as “rogue.”
[16:59] Product liability - The AI LEAD Act, and the wrinkle about a model that was never sold.
[19:40] An AI agent isn’t a toaster - So which part of this was the defective product?
[21:21] Someone left the door open - Kate Klonick’s negligence argument, and the AI Kill Switch Act.
[23:32] Where negligence runs out - What if the agents had emptied a few bank accounts for fun?
[25:32] A good dog - A Reddit comment, and the connection between training and alignment.
[26:38] The AP Physics exam - The rules David never had to spell out for his kid.
[27:25] Why they’re misaligned - They coordinated with each other, and none of them told anyone.
[30:05] Treat them like children - Which parts of parenting carry over to training a model?
[34:53] Close - “The world is going to get very weird very quickly.”
Reports from the labs
OpenAI, “OpenAI and Hugging Face partner to address security incident during model evaluation” (21 July 2026, with updates on 28 and 29 July) – the disclosure David reads from, and the source of the phrase “reduced cyber refusals.”
Hugging Face, “Security incident disclosure – July 2026” (16 July 2026) – the first public account of the breach, five days ahead of OpenAI’s.
Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline” (27 July 2026) – the forensic write-up behind the figure of about 17,600 actions, with a day-by-day table and an interactive replay of the campaign.
Anthropic, “Investigating three real-world incidents in our cybersecurity evaluations” (30 July 2026) – the three incidents David counts, identified across 141,006 evaluation runs in which Claude could have reached the internet.
Eric Wallace and Michael Dalton, “The ‘Breaking’ News: The OpenAI–Hugging Face Incident” (Black Hat USA 2026, 5 August 2026) – the 37-minute presentation David draws the timeline from, including the chain-of-thought screenshots.
Articles
Kate Klonick, “The AI That Hacked Its Way Out and the Hype That Followed It” (Lawfare, 29 July 2026) – the negligence argument David works through.
Zvi Mowshowitz, “OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards” (7 August 2026) – the argument that shutting down the message board was the wrong remedy, and that the training itself should have been redone.
Zvi Mowshowitz, “What Happened: OpenAI and HuggingFace” (8 August 2026) – his fuller reconstruction of the incident.
Yo Shavit, “post on X” (7 August 2026) – the passage Mowshowitz quotes, on the danger of leaving a task-independent instruction to be a good person out of the training objective.
Gabriel Weil, “Your AI Breaks It? You Buy It.” (Noema, 10 September 2024) – the case for treating frontier AI development the way tort law treats housing wild animals or blasting dynamite.
James Rundle and Angus Loten, “Renegade AI Systems Bring Security Leaders' Worst Fears to Life” (WSJ Pro Cybersecurity, 22 July 2026)
Kate Conger, “OpenAI Says Its A.I. Models Went Rogue and Attacked a Digital Library” (The New York Times, 21 July 2026)
Cristina Criddle, “OpenAI admits an AI ‘agent’ caused a major cyber breach by itself” (Financial Times, 21 July 2026)
SVXfiles, “comment in r/Futurology” (9 August 2026) – providing the dog analogy David builds the second half of the episode on.
Papers
Richard Ngo, Lawrence Chan and Sören Mindermann, “The Alignment Problem from a Deep Learning Perspective” (ICLR, 2024) – the technical statement of the problem David defines on air.
Bruce W. Lee, Yueh-Han Chen and Tomek Korbak, “Training Agents to Self-Report Misbehavior” (preprint, 2026) – on how rarely agents report their own misbehaviour unless trained to.
Patrick Dieveney, “The New ‘Parent’ Corporation (Moral Responsibility for Sophisticated AI Systems)” (Journal of Business Ethics, 2026) – the philosophical treatment of corporate moral responsibility behind the parent framing, written before the incident.
Legislation and letters
“S.2937, the AI LEAD Act” – Dick Durbin and Josh Hawley, introduced 29 September 2025 and still in Senate Judiciary. The definition of “design” is at section 3(5) of the bill text.
“New York A8833, the Understanding Artificial Intelligence Act” – Alex Bores, introduced 9 June 2025. The clause ruling out the mental-states defence is section 1512(2)(c) of the bill text.
“H.R. 9917, the AI Kill Switch Act” – Ted Lieu and Nathaniel Moran, introduced 23 July 2026, two days after OpenAI’s disclosure. The $20 million daily penalty applies to defying an emergency shutdown order; the general penalty is $2 million a day.
“Letter from Andy Ogles and Delia Ramirez to Sam Altman” (dated 31 July 2026, released 3 August) – the bipartisan request for a briefing from the House Homeland Security cybersecurity subcommittee.
“Letter from fifteen state attorneys general to OpenAI” (3 August 2026) – led by Iowa’s Brenna Bird, demanding document preservation and a halt to high-risk exploitation testing.
Further Listening
Credits
Host: David Deming, Danoff Dean of Harvard College
Executive Producer: Denise Koller Consulting Producers: Tim Smith and Jonathan Palumbo

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.