Google shipped three new Gemini models this morning. And somehow the one it didn’t ship became the loudest story in AI.
The new Flash lineup covers speed, low cost, and cybersecurity, and Google confirmed its “most ambitious pre-training run yet” is already underway for Gemini 4. But 3.5 Pro, the model built to trade blows at the frontier, is still missing in action. When the model everyone is waiting on keeps slipping, the release meant to reassure the market ends up doing the opposite.
In today’s AI news:
Google ships new Flash models while Pro stays missing
Sakana makes orchestration the cybersecurity moat
Poolside resurfaces with a top open-weights coder
AI’s biggest copyright bill comes due
Mozilla: open models closed the capability gap, not the revenue gap
News: Google rolled out a trio of Gemini models led by 3.6 Flash, joined by a speed-tuned 3.5 Flash-Lite and a security-hardened 3.5 Flash Cyber. The lineup is built for efficiency over raw intelligence, and the frontier-class 3.5 Pro that was supposed to headline the day is still stuck in testing.
Details:
3.6 Flash brings efficiency gains but gets outscored by similarly priced rivals like Grok 4.5 and GPT-5.6 Luna across a range of benchmarks.
On the Artificial Analysis’ Intelligence Index, 3.6 Flash shows no measurable jump over 3.5, landing near Meta’s Muse Spark and Z AI’s open GLM 5.2.
Logan Kilpatrick says 3.5 Pro is still “testing with partners” and will “hopefully land soon” after a string of delays.
Google closed by confirming Gemini 4 has entered its “most ambitious pre-training run yet,” with even Elon Musk taking a jab at the release.
Why It Matters: Efficiency helps the millions who live inside Gemini, but a Flash refresh is not what Google needs right now. The missing frontier model deepens the read that the company is a step behind while OpenAI, Anthropic, and xAI keep shipping. The lab with the most researchers, data, and TPUs on the planet cannot afford to keep being the one everyone is waiting on.
News: Tokyo’s Sakana AI shipped Fugu-Cyber, a security-tuned version of its Fugu model that answers to one API but is actually a pool of specialized agents working a problem together. It posts state-of-the-art, self-reported scores on two hard security benchmarks and ships access-gated and defense-only.
Details:
One endpoint, many agents: an orchestrator decomposes each task and routes it across specialists that analyze, challenge, and verify each other’s work, extending the original Fugu from June.
It claims 86.9% on CyberGym (finding real vulnerabilities in code) and 72.1% on CTI-REALM (turning threat intel into detection rules), comparable to GPT-5.5-Cyber and Anthropic’s Mythos-Preview, though the numbers are first-party and unverified.
Access is gated behind a manual approval form and an acceptable-use policy that bans offensive work, with Sakana conceding a raw model doesn’t solve enterprise security without its Applied Enterprise team building the verification harnesses.
Why It Matters: The real claim is that the moat is orchestration and deployment, not parameter count. Notice the pitch: frontier capability “without the risk of export controls,” a route around the scale-based rules that throttled even Anthropic’s Mythos and Fable 5 for three weeks in June. For buyers, the move is obvious: benchmark it on your own repos, because self-reported numbers behind an approval wall mean few outside eyes have kicked the tires.
News: Poolside launched Laguna S 2.1, an open-weights coding model that leads the Western open-source field while staying small enough to run on a single desktop. It’s the clearest sign yet that America is mounting a real answer to China’s open-model lead.
Details:
Laguna activates just 8B of 118B parameters at a time, fitting on one Nvidia DGX Spark box, with the weights free to grab on Hugging Face.
It tops coding benchmarks among U.S. open systems but still trails Kimi K3 and other Chinese open models by a wide margin.
The release lands a week after Mira Murati’s TML shipped Inkling, both widening America’s thin open-source bench.
Poolside says training took under nine weeks, and the model ships with a 1M-token context window.
Why It Matters: Two open releases in two weeks won’t erase China’s open-source lead, and Kimi is still lapping the field, but the momentum is real after months of a widening gap. A frontier-adjacent coder that runs on one desktop changes who gets to build with it, from enterprises that need data to stay in the building to solo founders who can’t rent a cluster. Open weights are where the U.S.-China AI race actually gets decided, and America is finally shipping instead of watching.
News: Anthropic won court approval for a $1.5B settlement with book authors, clearing payouts of roughly $3,000 per title across 482K works pulled from piracy sites. It’s the largest copyright recovery in U.S. history, and it leaves the ruling that AI training counts as fair use fully intact.
Details:
The suit traces back to 2024, when novelist Andrea Bartz and other writers alleged Claude was trained on books lifted from pirate libraries.
Judge William Alsup split the case in 2025: training on books is legal fair use, but Anthropic’s stockpile of 7M pirated titles was its own violation.
The settlement lets Anthropic sidestep a jury trial set for last December, where damages could have run into the hundreds of billions.
Anthropic keeps training on legally purchased books, and the fair-use finding stays untouched.
With similar suits pending against other labs, roughly $3,000 per work is emerging as the going rate to settle.
Why It Matters: This is the moment the industry gets a price tag for its original sin: acquire the corpus legally and training is protected, pirate it and you pay. Every lab that scraped the open web is now running the same math, and the ones with the balance sheets to settle just watched the template get written.
News: Mozilla, the nonprofit behind Firefox, published its first State of Open Source AI report, and the headline is blunt: the gap with top closed systems like ChatGPT and Claude has narrowed to just 3%, while inference costs have fallen up to 50x in three years. The twist underneath is the real story, because open models now power roughly a third of real-world AI usage but capture only 4% of the revenue.
Details:
The receipts: a global survey of 950+ developers plus OpenRouter’s 100-trillion-token dataset, with GPT-4-class inference now at $0.40 per million tokens, down from $20 three years ago (to be fair, I really hope nobody is ever using GPT-4...).
The frontier is still jagged: open matches closed on coding and general knowledge but trails on reasoning and agentic work, where GPT-5.5 hits 83.4% on Terminal-Bench 2.1 against the best open model’s 67.9%.
The bottleneck moved off the model: only 51% of open-model teams reach production versus 63% for closed, and it’s tooling and trust, not capability. Stripe cut AI costs 73% by moving 50M daily requests to self-hosted open models on a third of the compute.
Geopolitics may matter most: China and East Asia lead open adoption at 89%, Chinese open weights jumped from under 2% of OpenRouter tokens in late 2024 to 45%+ by April 2026, and Qwen passed Llama as Hugging Face’s most-downloaded family. Mozilla calls it deliberate state policy and a hedge against U.S. chip controls.
Why It Matters: When the model layer commoditizes, the margin migrates to the layer above it: deployment, orchestration, governance, exactly where open models still stall. That 4%-of-revenue gap is the opening, because whoever builds the tooling that gets open models reliably into production captures value that evaporates today.
🚀 Laguna S 2.1: Poolside’s open-weight coding model with a 1M-token context, small enough to run on a single desktop.
⚡ Gemini 3.6 Flash: Google’s newest Flash variant, tuned for efficiency and lower-cost, high-volume use.
🎆 Qwen-Image-3: Alibaba’s new image model with notably strong text rendering.
Twitter and Block founder Jack Dorsey unveiled Buzz in early access, an open-source workspace where AI agents join team chats as coworkers.
Substack partnered with Pangram to embed an AI-writing detector into the platform for both readers and publishers.
Anthropic shipped “Record a skill,” letting users screen-record a workflow so Claude can build a skill around the task.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.