RSS Amplifier

The Change Constant · Aug 23, 2026

Weekly Wrap Sheet (08/21/2026): Costs, Control & Corpora

0
Sign in to vote or save

Saanya Ojha · The Change Constant

  • AI competition has entered its Costco era. Frontier intelligence still matters, but the bigger commercial prize is making yesterday’s frontier cheap enough to run everywhere.

  • The AI policy fight is really a trust problem. One camp fears regulation will entrench a handful of labs; the other fears unconstrained diffusion will distribute catastrophic capability faster than institutions can react. Everyone agrees concentrated power is dangerous. They disagree on where concentration becomes least tolerable.

  • The easy training data is running out. Now labs are moving into progressively stranger and more expensive sources: rare books, enterprise exhaust, expert reasoning, and eventually the physical world.

AI competition has entered its Costco era. The throughline across the three most important recent model releases was cost.

xAI launched Grok 4.6 at $2 per million input tokens, $0.50 for cached input, and $6 per million output tokens below 200K prompt tokens, positioning it as a frontier model for coding, agentic tasks.

Google launched Gemini 3.7 Flash as its “most intelligent workhorse model” and priced it at $0.75 per million input tokens and $3.75 per million output tokens - half the original price of Gemini 3.6 Flash, a model that was barely 3 weeks old.

OpenAI pushed the argument furthest. It cut the price of GPT-5.6 Luna, its high-volume model, by 80%, from $1/$6 to $0.20/$1.20. Luna delivers performance comparable to models that were frontier-class a year ago at roughly six cents on the dollar per task - and at 9x the speed. It now powers unlimited text chats for free users.

This market direction makes sense. Pushing the absolute performance frontier is increasingly expensive, technically difficult, prone to diminishing returns, and involves many meetings with regulators. Pushing down the cost curve, by contrast, expands the addressable market every time you do it.

The frontier still matters enormously. If you are proving a theorem, discovering a vulnerability, or reasoning about protein structures, you probably want every available unit of intelligence. But most economic activity is considerably more mundane. Extract this invoice. Write this SQL query. Update this CRM record.

You don’t need Einstein to rename a file. And there are vastly more files to rename than theorems to prove.

Software that runs continuously has to be economical. The more ambitious the AI vision becomes - billions of autonomous agents performing trillions of tasks - the less sustainable it is to use the smartest model for every step.

Over time we’re going to move to a model org chart. The frontier model becomes senior mgmt: make the plan, resolve ambiguity, handle key decisions. Cheap models become the workforce: execute the plan, call tools, transform data, check work.

A nod also to the accelerating speed of intelligence commoditization. Luna operates at last year's frontier but at a fraction of the task cost. Google replaced a three-week-old Flash model with one that is simultaneously better and half as expensive. We are seeing rapid intelligence deflation. This is how technology becomes infrastructure.

Falling prices rarely mean falling spend. Usually, they mean exploding consumption. When bandwidth got cheaper, we invented YouTube. When computers got cheaper, we put one in every pocket. When storage got cheaper, we stopped deleting photos. The same thing is likely to happen with intelligence.

At $20 per task, you automate only your most valuable workflows. At $2, you automate far more. At $0.20, you embed everywhere. And when the price approaches zero, entirely new behaviors emerge.

Twitter is usually optimized for heat, not light: hot takes, performative certainty, and the occasional ratio. But at its best, it can do something rare: expose the underlying structure of a disagreement.

This weekend, a remarkably substantive exchange between Gavin Baker, Sholto Douglas, Dario Amodei and David Sacks underscored the central question at stake in the AI policy debate: Is AI too powerful to distribute or too powerful to centralize?

The Gavin / Sacks / Zuckerberg view:
AI is enormously powerful, therefore the worst possible outcome is concentrating control over it in a handful of labs and govt agencies. Their preferred safety mechanism is pluralism: more models, more companies, more open weights, more independent actors.
The intuition is almost Madisonian: ambition must counteract ambition. If intelligence becomes the most important economic input in the world, you do not want a small committee - however well-intentioned - deciding who is allowed to build it, deploy it, or access it. Competition becomes a form of checks and balances.
Their biggest concern with regulation is that it has a tendency to turn into an incumbent moat. The frontier lab with a $100B balance sheet and an army of lawyers can survive an elaborate compliance regime. Who does not? Three college dropouts with a GPU cluster and a term sheet. That is the regulatory-capture argument: rules written in the name of safety can unintentionally freeze the competitive hierarchy.

The Dario / Anthropic view:
AI is already structurally centralizing because frontier intelligence requires extraordinary amounts of compute, chips and capital. Making model weights open doesn’t magically democratize a world where a handful of actors own the infra. The barriers to entry are already arriving in shipping containers full of Nvidia GPUs. So the answer cannot simply be “more diffusion.”
Dario’s preferred model is asymmetric regulation: put the strongest constraints on the most powerful frontier labs, exempt smaller players, test dangerous capabilities before deployment, and use institutions to constrain corporate power.
While the other camp trusts competition to constrain power, Dario trusts competition plus institutions to constrain power. Markets need referees because eventually some players become bigger than the field.

This is the AI policy debate in summary:
The optimists fear that safety becomes an excuse to centralize power.
The safety camp fears that decentralization becomes an excuse to ignore risk.

Both fears are credible.

The most important question is which institution you trust least with superhuman intelligence: the market, the state, or a small set of frontier labs?

Most of the AI policy debate is an elaborate argument over which failure mode scares you most.

Remember when the AI industry’s answer to “where will the data come from?” was basically: the internet. It was enormous, free-ish, continuously expanding, and conveniently stored in formats computers could read.

AI labs inherited the accidental product of several decades in which humanity digitized civilization without realizing it was assembling a training corpus. We wrote websites, uploaded photographs, published code, digitized encyclopedias and put videos of every conceivable human activity online. Then the models ate it.

Now the easy calories are running out: publishers are putting up tollbooths, copyright litigation has made indiscriminate scraping more expensive, and the web itself is increasingly filling with AI slop. So the industry is moving down the data supply curve. And things are getting weird.

This week offered a particularly strange example. Google bid $10m for the internal business data of bankrupt Spirit Airlines to use for AI training. The assets include 100m emails and 500m Microsoft Teams messages, along with documents, spreadsheets, calendars and other business records.

Google is picking through the carcass of a bankrupt airline because buried inside are hundreds of millions of examples of humans coordinating, deciding, escalating, correcting and getting things done.There is perhaps no more fitting epitaph for the knowledge worker: in the end, we return to dust, while our Teams history gets fed to an LLM.

Spirit is only one manifestation of a much broader scramble. Silicon Valley has also rediscovered books. Booksellers in the UK and Ireland are reporting bizarre bulk orders for obscure and out-of-print titles, from old agricultural texts to vintage racing biographies. The labs have been buying and destructively scanning physical books.

A forgotten book sitting in a secondhand shop contains something surprisingly scarce in 2026: a large block of coherent, professionally edited, definitely human-generated text that may never have appeared on the internet.

But digging through bankrupt companies and secondhand bookstores only gets you so far. The more direct solution to scarce data is simply to manufacture it. That has created a booming industry around supplying frontier labs with increasingly exotic forms of training signal. Companies like Mercor, Turing, Snorkel, Surge, Handshake, AfterQuery etc are connecting labs directly to experts to generate high quality training data.

The AI industry’s data strategy is starting to look like a hierarchy:
Internet → licensed content → offline archives → enterprise workflows → expert reasoning → interactive environments → real-world behavior

Each step moves us toward data that is harder to acquire, harder to standardize and dramatically more expensive. AI's hunt for human exhaust is getting ever weirder.

Read the original on saanyaojha.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.