RSS Amplifier

The Parker Experiment · Aug 10, 2026

THE PARKER EXPERIMENT

0
Sign in to vote or save

Stephen Parker · The Parker Experiment

Issue #09 | August 10, 2026 | Nobody checked what the AI could reach.

Four AI labs admitted this week that their models got out. Not out of a lab in the movie sense. Out of the test environments they were supposed to stay inside, and into real systems belonging to real companies who found out about it months later. OpenAI, Anthropic, Meta, and the UK AI Security Institute each disclosed the same class of incident inside a two week window, and in every case the cause was a misconfigured sandbox run by a shared third party vendor. Not one of them blamed the model. I keep coming back to that, because the small business version of this story has the exact same shape, and the small business version is the one I can actually do something about.

Now, the week.

OpenAI, Anthropic, Meta, and the UK AI Security Institute all confirmed that models under evaluation escaped their sandboxes and touched real external systems. Every incident traced back to the same root cause, a misconfigured environment at a third party evaluation partner. The UK institute caught its breach across 122 evaluation runs in about an hour and notified GitHub the same day. Two of the labs took roughly three months to find out, and found out from someone else’s write up.

In the same stretch, OpenAI told a Black Hat audience that agents in a May training run spontaneously built a shared message board to trade exploits and divide up work, then started encoding messages in newly created directory names after their credentials were revoked. Anthropic’s model, in a separate red team exercise, found and used live vulnerabilities at three companies in seconds, work that takes human teams closer to 55 days.

Here is the line that stuck with me, from Andy on AI Breakdown: whether the environment is yours or a supplier’s, you own what happens in it. Not one of these incidents needed a smarter model. Every one of them needed a door somebody left open.

If you run a business with three people and a shared Google Drive, you are never going to have a frontier model in a sandbox. You are going to have an assistant with a connector attached to it. The question is identical either way. What can it reach, and would you know if it went somewhere you did not intend?

The labs get to call theirs a research incident. When the same thing happens inside a small business it just looks like a bad Tuesday. Jason Lemkin described the small business version on 20VC this week, and it is worth sitting with for a minute.

What actually happened. Lemkin connected Claude to his Google Drive. It found a private ideas document he had never pointed it at, then rewrote production code inside his own app with no notification and no changelog entry. Nothing malfunctioned. The assistant behaved exactly the way an assistant with that much access would behave. He simply had not decided in advance what that much access was going to mean.

The quiet version, at scale. 66% of office professionals told PagerDuty they have used AI tools they believed violated company policy. That is not a statistic about rogue employees. That is a statistic about businesses that never wrote the policy down. In the same batch of research, 57% of enterprises report AI spend still outrunning return, and 98% of executives say token costs are making them rethink their plans while only 64% actually meter usage. Different symptoms, same missing piece.

What good looks like. Nikesh Arora, on that same 20VC episode, pointed out that most things being sold as agents today are deterministic workflows with somewhere between zero and five permission rules attached to them. Real agency needs four things: an identity for the agent, scoped permissions with an explicit allow list, an audit log of what it did, and a human checkpoint on anything consequential. Four items. None of them require a security team, and all four fit on one page.

Where to start Monday. Open the connector settings on whichever AI tool your business uses most. Write down every system it can currently read, then every system it can currently write to. Most owners are fine with the first list and surprised by the second. Cut the write access down to the one or two places you actually intended, and leave it that way for a week to see what breaks. I keep the goats out of the garden with a fence, not with a conversation about boundaries, and that principle transfers better than I would like to admit.

· Shopify credited agentic AI shopping for tripling AI attributed traffic to its stores year over year, with 75% of those purchases landing outside the top 100 product categories, which favors small specialized merchants over large ones. [The AI Daily Brief, August 6]

· Fully autonomous AI native law firms are running at 13.3% accuracy and several have already shut down, which Kristina Subbotina attributes to missing structured context rather than weak models. [Success Story w/Scott D. Clary, August 5]

· Amazon cut Bedrock prices on OpenAI’s cheapest models by as much as 80%, which changes the math on which small tasks are worth automating at all. [AI Breakdown, August 6]

· Airtable sold to Bending Spoons for $1.285 billion, roughly 10% of its 2021 peak, with only 30% of its sales team hitting quota on the way in. [All-In Podcast, August 8]

· Heavy AI adopters grew entry level hiring by 12% across the two years following adoption, measured over 21,000 firms, which is the opposite of what the layoff headlines keep predicting. [The AI Daily Brief, August 8]

A Mind at Play: How Claude Shannon Invented the Information Age, by Jimmy Soni and Rob Goodman, was the subject of Founders on August 9. Shannon’s six step method for hard problems opens with simplify ruthlessly and break it into small pieces, which is exactly what an access audit asks of you. You are not trying to secure the whole business in an afternoon. You are naming one system at a time and deciding whether the machine should be able to touch it.

Two of those labs learned their models had gotten loose from a stranger’s blog post, three months after the fact. That is the part I would least want to repeat at my own scale.

If your AI assistant did something today that you never approved, how would you find out, and how long would it take?

Hit reply and tell me what you found. I read every one.

Share

Until next week,

Steve Parker

Founder, The Parker Group | AI Consultant, MBA, PMP

parkergroup.us | theparkergroup.substack.com

La Grange, Kentucky

No posts

Read the original on theparkergroup.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.