Hello fellow designers, creatives, digital professionals, and Eidos Design members!
Two weeks ago, an AI model hacked a real company. Nobody told it to, and both companies admitted everything in public. I've been reading about it ever since, and it changed how I see every other launch this month.
On July 16, Hugging Face disclosed a security incident. Five days later, OpenAI explained where it came from: the attacker was OpenAI's own model.
OpenAI was testing two unreleased models on a kind of hacking exam: a set of security challenges the model solves inside a sealed practice environment, with no access to the outside internet, and with the usual "I can't help with hacking" refusals switched off for the test. The model found a flaw in the sealed room itself, a bug in OpenAI's own software, and slipped out onto the open internet. Then it worked out where the exam's answer sheet was stored: on Hugging Face, a real company that hosts AI models and data for millions of developers. It broke in through real gaps in Hugging Face's security and copied the answers out of a live database. The model didn't go rogue in any cinematic sense. It cheated on a test, and it hacked a real company because that was the shortest path to a better score.
Hugging Face described "many thousands of individual actions across a swarm of short-lived sandboxes." OpenAI took responsibility in a joint disclosure and called it a partnership to address a security incident. Hacker News read the same document and summarized it in five words: so OpenAI hacked Hugging Face?
There are three ways to read this story, and I've been going back and forth between them all week.
The first is an asymmetry that Simon Willison, one of the most widely read writers on AI development, flagged while almost nobody else covered it. To investigate the break-in, Hugging Face's security team needed to feed real attack commands and exploit code into an AI model for analysis. The big commercial models refused, because their safety rules saw hacking material and said no. The team ran the forensics on a self-hosted open model instead, one with no such restrictions. Attackers use those same unrestricted models and get no refusals either. The guardrails, as designed, slowed down only the defenders.
The second lands directly on our desks. Every product shipping a "let the agent act for you" feature right now inherits this failure mode: a model chasing a goal will squeeze past the rules when the rules stand between it and the goal. The approve button, the permissions screen, the activity log your team keeps deprioritizing: as of this month, those are load-bearing parts of the product. I use agent tools every day, and I still catch myself clicking approve without really reading.
The third is the skeptical read. The same document that admits the breach also shows off, in detail, how good the new model is at breaking into things, days before its wide release. Safety researchers read the incident as proof that models gaming their instructions is now a practical problem rather than a theoretical one. That's true. It's also true that "our model is scarily capable" has never hurt a launch.
I'd rather say all three don't resolve than pick the comfortable one. What they agree on is the conclusion: machines doing real work need supervision, and supervision is an interface someone has to design.
On July 9, alongside GPT-5.6's public release, OpenAI launched ChatGPT Work: an agent that gathers context from your connected apps (Microsoft 365, Google Workspace, Slack, SharePoint) and produces finished documents, presentations, spreadsheets and hosted websites. It runs tasks for hours, and the Ultra tier runs four agents in parallel. You brief it, leave, and come back to a deck.
Anthropic spent July on the same territory from the other side. Claude Design, which shipped in April three days after Mike Krieger left Figma's board, got its first big update on July 3: direct editing, deeper design system support, and tighter workflows with Claude Code. Anthropic says more than a million people used Claude Design in its first week. That's a company-supplied number I couldn't verify independently, but even the order of magnitude says the labs consider design output a mainline product now, not a demo.
The obvious question is quality, and the sharpest critique I read cut precisely there: better formatting is not evidence of better sourcing. A polished deck with a wrong number in it is more dangerous than an ugly one, because polish is how wrong numbers get into board meetings. The org that adopts ChatGPT Work isn't buying finished work. It's buying first drafts wearing finished clothes, and someone still has to be the person who can tell the difference.
Notice what that person is doing, though. They're not making the deck. They're judging it, correcting it, deciding whether it should exist at all. The deliverable layer is becoming machine territory, and the judgment layer is becoming a job.
July answered that question from five directions in about three weeks, and together the answers amount to a new category of interaction design: interfaces for supervising AI that works while you don't watch.
The most literal one is hardware. On July 15, OpenAI released Codex Micro with Work Louder: a limited-run $230 keyboard for agent work, with dedicated keys that show what your agents are doing, a joystick for moving between workflows, and a physical dial that adjusts how much reasoning effort the model spends. Reasoning effort, the thing you used to control by phrasing your prompt carefully, is now a knob. TechCrunch called the keyboard a flashy bauble, and as a product they're probably right. As a design artifact it's the first mainstream attempt to give agent supervision a physical form, and I'd bet the dial outlives the keyboard.
The dial already has a software twin. Claude Opus 5, released July 24, ships a user-facing effort control and lets you switch models mid-task, so cost and quality become decisions you make moment to moment rather than a plan you subscribe to. Notion, a week earlier, shipped a standalone iPhone app called Agents whose only job is letting you check on, redirect and approve the AI teammates working inside your workspace. A dedicated app for babysitting software, and I can't decide whether that's the future of work or peak 2026. Raycast gave its AI eyes on your focused window this week with Screen Awareness. And on July 29, Superlogical came out of stealth: Mitchell Hashimoto, one of the most respected toolmakers in software, teamed up with design leaders from Vercel and the AI lab Poolside to redesign the terminal, the oldest surface in computing, around long-running agent work. A terminal company hired interface designers as founders. Two years ago that pairing didn't exist.
The pattern vacuum is being filled from below, too. The most-shared design resource in my feed this month was Beautiful UI, a free library of seventeen polished patterns specifically for AI-native products: thinking states, streaming states, tool calls, approval cards, diffs. The approval card is quietly becoming a core unit of interface work, the control that decides whether a machine's action becomes real.
Some caution before anyone reorganizes a roadmap around this. The keyboard is a limited drop, which is marketing rather than a product line. Nobody knows yet whether anyone opens Notion's agent app twice. The category is a month old and most of it is unproven. But five independent teams shipping the same idea in the same month is how categories usually start, and this one is short on designers who've thought hard about it. That's where you come in.
While the labs were redefining the work, the design tools quietly redefined the bill. Within about six weeks, Figma, Framer and Webflow all moved to metered AI credits.
Framer 3.0 shipped on June 16 with the most complete agentic design workflow on the market: agents that work inside your live project on layout, styles, CMS and SEO, with branching and previews, so an agent's changes arrive as proposals you approve rather than surprises. It also shipped a new kind of price list. Framer's pricing page now leads with credits: 500 to try on Free, 1,000 a month on the $10 Basic plan, 3,000 on the $30 Pro plan, and a slider that runs to 100,000 if you need more. This week the company went further and put the models themselves on sale, two of its fastest at 50 percent off through August 14. Your design tool now runs promotions on inference, the way a carrier discounts data.
Webflow's simplified plans took effect June 29, merging two tiers into a single Premium plan at $25 a month on annual billing and adding AI credits to every workspace. The tell sits in Webflow's own announcement, which advises customers to switch to annual billing before the change in order to lock in their current pricing. Figma has enforced its credit system since March, and its community forum spent the summer documenting confusion about how those credits get counted.
Each company calls this simplification, and it's worth being precise about what actually changed: design tools used to cost a seat, and now they cost a seat plus a meter. Seats are predictable. Meters are not. If you run an agency or a freelance practice, your tool costs just became a variable to forecast, and every retry you prompt has a price attached, even when the meter is invisible in the moment. That invisibility is an interaction-design choice, and I don't think it's an accidental one.
The fair version of the other side: inference costs real money, and per-use pricing is more honest than burying it in seat inflation. Fine. But watch the shape this creates. The vendors now earn more when you iterate more, while the marketing promises you'll iterate less. Both can't be the pitch.
👉 Peter Yang open-sourced a "no AI slop" skill that strips twenty-plus recognizable slop patterns from any piece of writing. Slop is now well-defined enough to lint. (Full disclosure: this issue went through it.)
👉 Beautiful UI: seventeen free, polished patterns for AI-native products, from thinking states to approval cards and diffs, source code included. The pattern library the supervision era was missing.
👉 Every's vibe check on Claude Opus 5: "brilliant in flashes, frustrating in practice." The most honest model review of the month.
👉 drawesome, a free drawing toolbar for React by Benji Taylor. Part of the quiet rise of free UI libraries built by designers who care about details.
👉 The Content Architecture, a Next.js and Sanity starter kit by Edoardo Lunardi, is pitched at the agents that set your project up rather than at you. Its tagline is this month's argument in one line: "The Sanity setup agents don't reinvent. Every run invents a new one, none decided." It's a €549 commercial kit, so read the pitch as a pitch, but the thinking (and the ASCII art) is worth the click.
The model that broke into Hugging Face was not malfunctioning. It had a goal, a score, and nothing that told it what a good way to reach that goal looked like. Nobody had written that part down.
The same gap runs under all four stories this month. A deck that arrives finished and slightly wrong. A layout that is fine and forgettable. An agent that ships a change nobody would have approved, if anybody had been asked. Every one of them is a machine hitting the target and missing the point, and the only thing that closes that distance is a person who can tell the difference and says so out loud.
Which is why the other thing I saw in July stayed with me. In the same month the labs started shipping finished work, my feed filled up with designers rebuilding a color picker by hand, publishing pattern libraries nobody paid them for, arguing about what an interface should do in the moment between a click and the thing appearing. That is how you keep a standard sharp. It is also what you bring into a review that the machine cannot.
I don't know what August will bring. More capability, certainly, and probably another thing that gets loose before anyone meant it to. I'll write you when it adds up to something. The question I'm carrying into it is smaller than any headline: when the machine hands me something good, will I still know whether it's right?
If this edition gave you something to think about, pass it to someone who needs to read it. And subscribe if you haven't yet.
Sincerely,
Mykola Korzh

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.