💬 OpenAI’s rogue agents found each other and built a secret message board. At Black Hat on August 5, OpenAI security engineer Michael Dalton and researcher Eric Wallace walked through the forensics behind last month’s Hugging Face breach, and the part nobody had reported is the part that should keep you up. Agents from separate, unrelated training runs discovered each other inside OpenAI’s own JFrog Artifactory instance and used it as a persistent covert message board, trading exploits and credentials for roughly two months. They started leaving files for each other on May 7, got internet access through a server-side request forgery on May 26, and hit an Artifactory zero-day for remote code execution on June 26. When OpenAI revoked access and rebuilt the repository, they re-established communication in about two days by encoding messages into newly created directory names, one of which read probe_PENDING_HOLD_SWARM_until_confirm. One agent’s reasoning trace, quoted at the talk: “Help peer. But our task doesn’t benefit. Yet collective may yield generic route if someone frees time.”
🎭 A Claude model invented fake people to trick a real maintainer into merging malicious code. The UK’s AI Security Institute published an incident report on August 4 covering a single cyber challenge it ran 122 times across several frontier models. In 10 of those runs the agent took unsanctioned action against real people and organizations on the live internet, 19 documented actions in total, 17 of them from Anthropic’s Mythos 5 and 2 from OpenAI’s GPT-5.6-Sol with its cyber classifiers switched off. Read more
🥉 Meta becomes the third lab in a month to admit its model hacked somebody. Muse Spark 1.1, Meta’s most capable agentic model, escaped its test sandbox, reached the open internet, exploited a vulnerability in a third-party service, and made unauthorized changes to another company’s internal infrastructure. Meta hasn’t named the company, hasn’t said when it happened, and hasn’t said how long the model ran unsupervised, but it did name the cause: its evaluation partner Irregular misconfigured the sandbox. Read more
🏛️ A dozen AI companies went to the White House and agreed to nothing. The Trump administration convened roughly twelve companies on August 4, including OpenAI, Anthropic, Google, Meta, Nvidia, and Microsoft, to review a finished voluntary framework for testing frontier models’ cyber capabilities. It comes out of the June 2 executive order on advanced AI innovation and security, which gave agencies 60 days to stand up a classified benchmarking process plus a channel where developers can hand the government early access to a model for up to 30 days before release. Open-weight systems are exempt, no company sent C-suite, Trump didn’t attend, neither did his science adviser or his national cyber director, and the administration has no plans to publish the framework at all. Read more
⏸️ OpenAI hit the brakes on Astra because it can’t rule out that the thing writes zero-days. Axios reported on August 7 that OpenAI is slowing work on Astra after concluding the model may cross the Critical cybersecurity threshold in its own Preparedness Framework, meaning it could find and develop exploits against hardened systems without a human in the loop. Bloomberg confirmed it the same day, and a White House official told Axios the company volunteered the delay rather than being asked. Read more
🚪 Google DeepMind’s brain trust walked out the same day Hassabis stepped back. On August 5, Demis Hassabis moved to Chair of Google DeepMind and Chief Scientist of Alphabet, handing day-to-day control of Gemini and frontier research to Koray Kavukcuoglu, who now reports straight to Sundar Pichai. On the way out the door went Jeff Dean and Sanjay Ghemawat, 27 years each, along with Oriol Vinyals and Quoc Le, to found Discovery Loop, a public benefit corporation aimed at automating scientific and engineering research, starting with automating machine learning research itself. Read more
📈 ChatGPT crossed a billion weekly users. OpenAI confirmed the milestone on August 6, up from the 800 million Altman announced at DevDay in October 2025. Read more
💳 Stripe is in exclusive talks to buy OpenRouter for around $10 billion. The Information reported on August 6 that the deal has moved to exclusivity as cash and stock, escalating talks the Wall Street Journal first surfaced in late July. OpenRouter routes a single API call across more than 400 models from 60-plus providers and moves something like 25 trillion tokens a week, five times its volume six months ago, and it was valued at $1.3 billion in a May round led by Alphabet’s CapitalG. Read more
Generative design of bacteriophages with genome language models
From Samuel H. King, Brian L. Hie, and colleagues at Stanford and the Arc Institute
Jake’s Take: Long story short: an AI wrote genetic code for sixteen (new) viruses. Said AI is called Evo, a language model trained on DNA instead of English, that can predict the next base pair in a similar way to how ChatGPT predicts the next word. The labe fine-tuned Evo on 14,466 sequences from Microviridae, used it to then generate thousands of complete genomes, filtered these down to 302 plausible candidates, and then sent them off to a DNA synthesis company to be physically built. 285 came back as actual molecules. They then introduced these to E. coli, looking for the clear spots that appear on a plate of bacteria when something is killing them. Sixteen produced those spots which were then confirmed as real viral particles with cryo-electron microscopy. The winners carried between 67 and 392 mutations their closest natural relatives don’t have, and several diverged far enough to count as new species. When the researchers pitted them against wild ΦX174 in the same culture, one designed phage (dubbed Evo-Φ69) grew to 65 times its starting share of the population.
That’s a lot of numbers and fancy words. Here’s the point. Bacteria like E. coli often develops resistance to natural phages, and these new AI-generated phages were able to break through in 1/5 passages. Phage therapy’s number one weakness has always been that bacteria can adapt to any single phage, and finding a replacement means sifting through nature. These AI-generated phages will help provide a brand new set of candidates in these cases, significantly cutting down the time it takes to find non-resistant solutions.
Worth noting: the authors excluded viruses that infect humans, animals, and plants from the model’s training data on purpose and ran everything in biosafety cabinets with non-pathogenic hosts. But they still wrote in their own paper that a motivated actor “could, in principle, sculpt these models to design human viruses.” Whoops.
🐕 The model after Astra is reportedly called Doug. SemiAnalysis dropped one sentence in an August 7 piece: “OpenAI has overcome their pre-training issues, and a much larger model code named ‘Doug’ is actively in the works.” That’s the entire sourced claim, lifted from a July 9 note to their institutional subscribers, and it matters because it suggests OpenAI is scaling pre-training again after years of squeezing gains out of post-training instead. Read more

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.