RSS Amplifier

Limitless · Jul 24, 2026

Google Faces Real Humiliation

0
Sign in to vote or save

Ejaaz · Limitless

Good afternoon,

Google had a pretty humiliating week after shipping what might be its worst AI model ever. Its stock has dumped 10% since the launch of Gemini 3.6 Flash, despite posting record earnings.

Meanwhile

  • One of OpenAI’s models literally broke out of its cage and hacked Hugging Face to cheat on a test.

  • The U.S. Government is debating whether to make open-source AI illegal, despite half of American founders using them.

  • Nvidia’s next chip just cut the cost of intelligence by 10X.

Let’s get into the news.

-Ejaaz

Google released Gemini 3.6 Flash this week and it scored a 50 on the Artificial Analysis Intelligence Index, which is the same score as the model it was supposed to replace.

That score puts it below Meta’s Muse Spark 1.1, China’s GLM-5.2, OpenAI GPT-5.6 Luna, Claude Sonnet 5, Grok 4.5 (need me to go on? Its bad)… several of which are older AND cheaper.

At $7.50 per million output tokens, Google is charging a premium for a model that gets beaten by old open source models free to anyone.

Maybe I’m being too harsh, but we’re talking about the company with the most capital, infrastructure (TPUs), the most data, a 14% stake in Anthropic, and its own Nobel-winning research lab… how the hell did they blow the lead?

Despite beating earnings this week, Google stock has taken a 10% hit at the time of writing.

I’m not kidding. This week OpenAI disclosed that during an internal cybersecurity eval, GPT-5.6 Sol and an unreleased, even more capable model (GPT 6???) were tasked with solving ExploitGym, a cybersecurity benchmark. To measure raw offensive capability, the models were sandboxed with no internet access and safety filters turned down.

Instead of actually trying to solve the problems, the models spent enormous compute finding a zero-day exploit in OpenAI’s own software, got open internet access, correctly reasoned that Hugging Face probably hosted the ExploitGym answer key, then chained together an attack (with stolen credentials) to remotely steal the answer key from Hugging Face’s production database. It did all of this in a few hours.

I don’t necessarily think the model was acting maliciously. And that’s the crazy part: it was just following orders and found what it believed to be the most efficient way to achieve its goal.

After the successful release of China’s latest Open Source model Kimi K3 last week, the US has gone in a spiral. Chinese models run ~$0.18 / million tokens versus ~$4 for US frontier models. That’s 22x cheaper while being 90% as effective.

So why on Earth should American companies spend millions of dollars on Fable 5 when you can just use these?

The US Treasury Secretary said the government will now probe Chinese models for IP theft and may sanction them based on alleged distillation attacks against frontier AI labs like Anthropic.

The community hated this: nearly 200 companies (Y Combinator and Proton among them) fired back with a coalition letter to Trump opposing a ban. This morning Jensen Huang and the leaders of 24 other companies including Microsoft signed a similar letter advocating for open models.

But my question: once open weights are downloaded, they live on private servers forever so… how do we even ban this stuff?

3 very different companies this week announced they’re launching the same AI product: Ramp shipped Router (it cut their own LLM bill 30% across 100+ internal use cases), Cursor shipped Cursor Router, which triages every coding request to a cheap model for up to 60% lower cost… and The Information broke that Meta is building one too.

But why?

The concept is simple: don’t send every request to the most expensive model. Put a smart dispatcher in front that picks the cheapest model that clears your bar of quality, and only use the smartest models when you need to solve a hard problem. The rest can be relegated to cheaper (often faster) models.

Reminder that Cursor got acquired for $60 billion by Elon for building a product that did exactly this for coding. Rumor has it Openrouter (another company thats pioneered this) has received a $10B acquisition offer from Stripe

This week CoreWeave became the first to stand up Nvidia’s next-gen Vera Rubin AI server, and the numbers are staggering: up to 10x more tokens per megawatt and roughly 1/10th the cost per million tokens versus the current Blackwell generation.

Racks run $7–8M each, but companies can make their money back on them 10x as fast. This is huge for scaling AI models and products to everyone, as compute constraints are a very real problem and drive costs up.

Thanks for joining us for another issue. Now go listen to our podcast :)

No posts

Read the original on limitlessfm.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.