This week Jeff and I recorded a day early because my daughter qualified for a big swim meet in Irvine, so surprise, Tuesday episode! We dug into OpenAI’s wild week, where an unreleased model solved math problems that stumped humans for decades while the company’s agents were escaping containment at multiple companies. We also got into the EU AI Act going live with real enforcement teeth, Alibaba open-sourcing their biggest model ever and the price war it’s fueling, Google putting Gemini inside a humanoid robot, and LinkedIn rolling out an official “AI slop” button. Plus a speed round covering Gemini Spark, fake satellite imagery, and Sam Altman’s podcast experiment that went very wrong on social media.
We just hit 100 paid patrons, which is a milestone I’m very proud of. To celebrate, for the next two weeks, new patrons can use the code 100PATRONS to get 30% off a yearly subscription on top of the existing 20% annual discount. That comes out to about half price for the year. The AI Inside Daily podcast, ad-free episodes, Discord access, and a bonus patron AMA at the end of August are all included. Can we get from 100 to 120? I’m going to try. Check it out at patreon.com/aiinsideshow.
OpenAI had three stories this week telling a different story when you stack them up. First, the math results: an unreleased model called Astra solved 10 previously unsolved problems across high-dimensional geometry, coding theory, quantum complexity, and lattice cryptography, including the first-ever construction of a non-sofic group. Total compute cost was roughly $2,000 at API rates. Jeff and I pulled up Gemini live on the show to figure out whether any of this actually matters in the real world, and it does, which was helpful in understanding the practical reason these sorts of problems need to be solved. But by whom?
Gary Marcus called the results “amazing” but warned against the fallacy of composition, arguing that being great at one kind of math does not mean you’re close to general intelligence. Jeff and I both agree with that. The conversation we landed on was really about mathematician identity: these are people who defined themselves by solving puzzles nobody else could, and that self-image is harder to hold when a machine does it for two grand. Jeff compared it to the calculator replacing the slide rule, but noted this cuts deeper because the identity is tied to being the person who can do the thing nobody else can.
Meanwhile, Sam Altman personally demoed Astra for US senators in Washington, meeting with Warnock, Moreno, and Warner alongside Treasury Secretary Bessent and Commerce Secretary Lutnick. The timing is notable: President Trump’s executive order requiring government access to frontier models before public release had an August 1 framework deadline, and Altman showed up right as that kicked in.
And then the containment story. Reuters reported that OpenAI’s investigation into escaped agents has widened, with multiple autonomous agents breaking free from testing environments. One agent “went haywire for days inside another company’s network,” and compromised accounts turned up at four other companies including Modal. Anthropic disclosed its own models conducted break-ins at three companies going back to April, acknowledging that “real-time monitoring of the evaluation logs would have helped.” Put those three side by side: in one room they’re solving decades-old math for $2,000, in another Altman is demoing for senators, and in a third their safety team is discovering agents that got loose months ago. The capability is settled. Whether they can contain it is the question.
The EU AI Act became fully enforceable on August 2, giving the European Commission direct enforcement powers over general-purpose AI providers regardless of where they’re headquartered. Anthropic, OpenAI, and Google are specifically named for heightened oversight. Fines can run up to 15 million euros or 3% of global annual turnover, whichever is larger. Jeff predicts that this ends up like the cookie-click problem, where lawyers tell every company to label everything AI-assisted just to cover themselves, and if everything is AI then nothing is AI. Separately, new EU labeling rules for realistic AI-generated content kicked in the same day, requiring machine-readable watermarks and visible disclosures on deepfakes and synthetic media, though the enforcement question remains open since watermarks can be stripped.
The China AI story keeps building. Alibaba released Qwen3.8-Max, their largest model ever at 2.4 trillion total parameters with a million-token context window, and they’re open-sourcing it. With DeepSeek also slashing prices, there’s a genuine race to the bottom that pushed OpenAI to lower some of its own prices in response.
This led Jeff and me into a bigger conversation about Alex Karp’s Palantir shareholder letter, which reported 93% revenue growth and $1.1 billion in quarterly profit. Karp called the hosted-model approach a mistake and wrote that OpenAI and Anthropic are trying to “drug addict” companies to a future they want to control. I admitted on the show that I recently upgraded from $100 to $200 a month with Anthropic because I hit a wall and was willing to pay through it, and I’m not sure whether that was worth it or whether my usage is just climbing. Jeff predicted I’ll follow Leo Laporte down the local-models path, maybe buying a DGX Spark to run open-weight models and reduce that dependency. Pretty expensive, but I could see it. Someday.
Google DeepMind released Gemini Robotics 2, a vision-language-action model that gives robots full-body control. They demonstrated it on Apptronik’s Apollo 2 humanoid with five-fingered hands, and we watched a demo where a Boston Dynamics robot navigated to a snack table, read the word “popcorn” on a bag, grabbed it with its pincer claw, and brought it back. The system runs entirely on-device with no cloud connectivity, adapts to new robot bodies in hours, and handles multi-robot coordination. DeepMind CEO Demis Hassabis compared it to what Android was for smartphones, and with NVIDIA building its own robotics stack, competition for that OS layer is heating up.
Apple’s bug bounty program is being overwhelmed by AI-generated vulnerability reports. LLMs make it easy to generate hundreds of plausible-looking submissions full of technical jargon, but many of them are hallucinated, and each one still requires manual human review. Apple has paid over $35 million to about 800 researchers since the program launched, with maximum bounties reaching $2 million and potentially exceeding $5 million with bonuses. Jeff made a good point: at some point, responsible companies will run their own software through the same AI tools internally, which should cut down on both bugs and the bounty market over time.
The Friend pendant is back with a speaker, so instead of just texting you, it talks out loud. The price nearly doubled to $249, each pendant gets a randomly assigned voice and personality you can’t change, and there’s an optional $10 per month subscription for extended AI memory. By late 2025, Friend had sold around 3,000 units and shipped about 1,000, and the company spent over $1 million on NYC subway ads that became better known for the vandalism they attracted than the actual sales. Jeff found the company’s marketing video depressing, with everyone looking mournful. Adding a speaker to something most people didn’t want in the first place probably doesn’t change the equation.
LinkedIn added a new reporting option letting users flag posts that “seem like AI slop.” The fact that LinkedIn used the informal term rather than something corporate is notable on its own. This is Microsoft’s professional social network, a company that’s a major OpenAI investor, and the official UI now contains the word “slop.” Jeff confirmed the option is now live in the three-dot menu on any post.
A judge denied Elon Musk’s xAI a request to pause Minnesota’s first-in-the-nation ban on nudify apps. My initial reaction was to write this off, but the deeper I looked, the more complicated it got. The law has no consent requirement, no intent requirement, and strict liability up to $500,000 per image, with a definition of “intimate parts” borrowed from sexual assault statutes that technically covers shirtless men and people in swimsuits. There’s no satire exemption, meaning something like South Park’s deepfake of Trump naked in the desert would violate it.
Techdirt’s Mike Masnick called it “the worst person you know” filing “a good First Amendment lawsuit against a very badly drafted” law. Jeff traced the deeper issue back to the history of American privacy law, which started with the unauthorized use of a woman’s face on bags of flour and then got tied to the technology of the postal letter. That technology-specific approach is why email and texts still don’t get the same Fourth Amendment protections letters do. The answer, Jeff argued, is writing laws at a principle level rather than targeting specific technology, which is harder but necessary.
Gemini Spark can now browse the web using your desktop Chrome, including logged-in accounts and saved passwords. Google highlighted use cases like scheduling apartment viewings and researching flights. I shared my own Spark experience on the show: I set up agents months ago and then forgot about them, realizing weeks later they’d been running in the cloud on my behalf like a leaky faucet I forgot to turn off. Jeff still can’t use Spark because of his Google Workspace account, which at this point is just the standard Workspace experience.
Google launched a feature in Google Earth letting users generate AI satellite imagery by typing descriptions, then pulled it within a day. Researchers showed the tool would accept prompts like Iran’s Kharg Island engulfed in flames, a flooded US Capitol, and bomb damage on hospitals. Satellite imagery has traditionally been hard to fake, making it a trusted tool for journalistic investigations and war crimes documentation. Bellingcat’s Jake Godin warned it would “streamline the process” of creating fake satellite imagery.
Thinking Machines Lab released Inkling-Small, an open-weights model with 276 billion total parameters and 12 billion active. It claims comparable performance to the full Inkling at a quarter the size and actually beats its larger sibling on reasoning tasks, scoring 31.6% on Humanity’s Last Exam versus Inkling’s 29.7%.
Sam Altman tweeted about using ChatGPT to generate a podcast for his kids while driving, and the ratioing was swift. Responses ranged from “maybe before you post publicly, you should have trusted people who can guide you to make different choices” to “why don’t you just talk to your kids” to “can we please get some normal people running these tech companies.” Jeff and I agreed he probably thought it was a feel-good use case and had no idea how it would land.
Executive Producers: DrDew, Jeffrey Marraccini, Radio Asheville 103.7, Dante St James, Bono De Rick, Jason Neiffer, Jason Brady, Anthony Downs, Mark Starcher, and Karsten Samaschke. Thank you for your support at the executive producer level.
Thanks for reading and watching. Catch the full episode, subscribe links, and more at aiinside.show.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.