RSS Amplifier

Good AI's newsletter · Jul 20, 2026

The $0.50 Barrel of Intelligence

0
Sign in to vote or save

Darwin Ling · Good AI's newsletter

This week on CNBC, Chamath Palihapitiya priced intelligence like crude oil: a barrel of WTI runs about $81, and a barrel of intelligence — one million tokens — runs anywhere from fifty dollars down to fifty cents, depending on whose pump you pull up to.

I verified the current rate card against vendor pages this week (full table and methodology in the appendix — note these are list prices per token, not cost per completed task, which varies further because models differ two- to three-fold in how many tokens they burn to finish the same job), and the shape of it is simple: the frontier tier — Anthropic’s and OpenAI’s flagships — charges a steep premium over everyone else — roughly 11x over the best open-weight barrel, 57x if you reach for the cheapest production-grade one. Below that tier sits a crowded, fiercely repricing middle: Google, xAI’s newest models, and the Chinese open-weight labs (GLM, Kimi, DeepSeek), all clustered within a few dollars of each other and falling. And here’s the part that matters: for most enterprise workloads, the cheap barrels are now good enough that the premium is hard to justify. Chamath’s own assessment of whether the non-frontier models can handle most use cases: “a screaming yes.” The exception he carved out — narrow, high-stakes work like cybersecurity, where the frontier’s edge genuinely pays — is real, but it’s the exception. The broad middle of enterprise AI work is repricing toward the floor.

Update — July 20, 2026. Three days after this essay's rate card was pulled, Moonshot released Kimi K3 — an open-weight model listed at $15 per million output tokens: level with the closed mid-tier (GPT-5.6 Terra, Claude Sonnet 4.6) and roughly 17x DeepSeek. The chart above and the appendix table are updated to include it. K3 changes one number in this essay: the frontier-to-best-open-weight spread quoted below compresses from ~11x (vs. GLM 5.2, on the July 13 card) to ~3x — not because open weights got cheap, but because the newest one is priced like an incumbent. "Open" and "cheap" were never the same axis; K3 is the proof. Within days of launch, Moonshot paused new subscriptions on GPU capacity — the constraint layer arriving on schedule. Weights land July 27; a follow-up note then.

A 57x spread for a commodity input, with quality converging from below, is what a market looks like mid-commoditization. The model is becoming an input, not the product — and the scarce asset is shifting to the system that decides which model to use, proves it can be trusted, and captures the value created. The rest of this essay is about that shift: who’s driving it, and how to build on the right side of it.

In this country, at every single enterprise I deal with, these people are livid. They’re like, I am paying for tokens that create no value. These people are stealing the weights in alpha of my business, and they are creating a wealth tax.

Alex Karp on CNBC, July 1

Which brings us to the most-watched AI television of the summer. On July 1, Alex Karp went on CNBC and torched the frontier labs’ business model. Enterprises, he said, are chillaxing — burning budget on tokens that create no measurable value, while harboring a suspicion they can’t verify: that their data, their prompts, their alpha — the proprietary edge that makes a business win — may be feeding the very models they’re renting. Karp never quite asserted that the labs do this — he was careful to frame it as the questions every client must be able to ask: who owns the data, where is it cached, are the prompts secure. The labs’ enterprise terms say customer data isn’t trained on by default; his point is that inside a closed deployment, you’re taking their word for it. The written version of the indictment came the day before, in Palantir’s sovereignty doctrine, which coined tokenmaxxing for the spending pattern: organizations incentivized to consume tokens, mistaking activity for progress, paying what Karp called on air a wealth tax on American business. The CEOs he deals with, he claimed, are livid in private and silent in public.

The sharpest version of his argument fits in two lines, and it’s worth separating from the theatrics: token price is what you pay. True cost is what you pay plus the value of whatever might leak. Karp’s on-air phrasing was that true cost is what you make minus what you lose — “the value of your business.” The rate card measures the first number. His entire pitch is about the second, unpriced one.

Why silent? Because until recently, the frustration had nowhere to go. You don’t attack a supplier you can’t replace, and you don’t admit your flagship AI initiative has no ROI while your own stock is long the AI story. So the anger travels privately — and the financial reckoning is still ahead. Most CFOs, as Chamath observed in his follow-on interview, don’t yet know how much tokenmaxxing lives inside their own organizations. They’ll find out the way CFOs find out everything: on an earnings call, tracing an OpEx miss back to the $50 barrels their teams were burning where fifty-cent barrels would have done.

Coverage called Karp’s appearance a crash-out. Almost everyone missed the sequencing: on Monday, Palantir and NVIDIA announced a partnership to deploy NVIDIA’s open-weight Nemotron models — as independently evaluated (by Artificial Analysis) to be good enough for the broad middle of enterprise work — inside Palantir’s sovereign stack. On Tuesday, Palantir posted a nine-point written doctrine on AI sovereignty — point four: “Controlling your weights is controlling your fate.” On Wednesday, Karp went on television. Deal, doctrine, distribution. That’s not a meltdown; that’s a launch.

And the launch is aimed precisely at the two problems the rant named. Governance: the Palantir–NVIDIA stack lets an enterprise or agency run open-weight models it actually controls — its compute, its deployment, its weights — inside Palantir’s application layer, with the audit and control surface regulated buyers require. Data retention: closed-model vendors do offer no-training commitments, zero-retention modes, and dedicated deployments — the difference is that with owned weights and owned infrastructure, control is architectural rather than contractual. You’re not verifying a promise; there’s no promise to verify. Karp conceded as much on air: a closed frontier model can pass the trust test too, if its vendor will answer the hard questions. The partnership’s pitch is simply that with open weights, no one has to ask the vendor.

Strip the theatrics, and the strategy is elegant. NVIDIA profits when the model layer commoditizes completely — every enterprise that self-hosts open weights buys GPU capacity, owned or rented, and the order lands with NVIDIA either way. Palantir profits the same way — every increment of anxiety about the model layer raises the value of the trust layer it sells. Karp said the quiet part on air: in Palantir’s own financials, the money is made in the application layer and in compute. Not the model. And none of this is new for him — back in 2024, Karp was already saying the market’s value would go to “chips and what we call the Ontology,” and we wrote it up at the time. July wasn’t a change of view. It was the 2024 sentence, productized — with NVIDIA now formally attached to the “chips” half. The layer below the model and the layer above it just allied, and the instrument is a free model. Karp, to his credit, flagged his own conflict mid-rant — his claim, he admitted, was “obviously slightly true, but slightly self-centered.” Both halves are accurate.

One asterisk, and it’s load-bearing: the premise that an open model plus a great application layer reaches frontier outcomes is asserted, not demonstrated. Nemotron is not frontier-parity today. If the premise proves out, this essay describes the next two years. If it doesn’t, “sovereign” quietly becomes “sovereign but second-tier” — a different product. Watch this.

If Karp is right that the model is just an input, the obvious next question is who captures the value when the input commoditizes. Follow the economics one step past the rate card, and the answer emerges. If price commoditizes and the cheap models are good enough for most work — but not all of it — then for the broad middle of enterprise workloads, buying a model stops making sense. What a rational buyer acquires instead is an allocation: the high-volume, well-scoped work routed to the fifty-cent barrels; the judgment-heavy work that genuinely needs the frontier routed there; everything behind a switch that can be repointed as the leaderboard moves. Model choice is becoming a procurement and routing problem. (The math of routing well — cost per completed task, not price per token — is in the appendix; the punchline is that the differences are even larger than the rate card suggests.)

Which means the scarce, compounding asset is no longer the model. It’s the layer that decides: the router, the evaluation harness, the governance wrapper, the ontology — Palantir’s term for the data foundation that catalogs a business’s operations so an LLM can be used, refined, and imposed on the enterprise (our 2024 explainer covers how it works and why it’s the differentiation). The labs themselves are confirming this with their feet — OpenAI ships Codex, Anthropic ships Claude Code, xAI ships Grok Build — because they can see the naked model becoming substitutable underneath them. For the evaluable middle of enterprise work, the model has stopped being the product. The model plus the layer that directs and governs it is the product.

Every business built on this layer shares one structural property worth naming, because it’s the general test — the one line I’d ask every founder to write on the wall:

Does cheaper AI below you make you stronger or weaker?

Palantir gets stronger — every increment of model-layer anxiety raises the value of its trust layer. NVIDIA gets stronger — every self-hosted open model is a GPU order, because the cheap barrel doesn’t eliminate cost so much as move it from the API line to the infrastructure line, which is exactly where NVIDIA wants it. The labs’ own harnesses exist because the naked model gets weaker. If cheaper AI below you is a tailwind, you’re in the constraint layer. If it’s a headwind, you’re in the commodity.

This is the thesis I’ve been writing under for a year — value migrates to the constraints: power, capacity, and trust. The token price collapse isn’t an exception to it. It’s the mechanism.

If you’re building right now — especially in a regulated or trust-sensitive vertical — the pricing collapse is either your margin expansion or your extinction event, and the difference is architectural. Five things:

Architect for model-switching from day one. The switch is leverage even if you never flip it. Every negotiation with a model vendor now happens against your credible ability to leave — an alternative that didn’t exist eighteen months ago. Abstract the model behind an interface; route by workload; re-evaluate quarterly.

Price on value delivered, never on token pass-through. If you bill per token, the deflation curve is your customer’s windfall. If you bill per outcome — per matter, per claim, per completed workflow — every point of token-cost decline flows to your margin. The best-positioned businesses for the next three years are the ones whose pricing was never denominated in tokens at all.

Run the tailwind test on your own moat. If your differentiation is access to a model, you’re a thin wrapper and the race to zero is already priced in. If it’s workflow depth, proprietary data rights, domain governance, or regulatory trust — things that get more valuable as the intelligence beneath them gets cheaper — commoditization is working for you.

Treat trust as product, not compliance overhead. Provenance, auditability, data handling, deployment control, chain of custody — in regulated industries these aren’t checkboxes; they’re the line items customers actually pay for. The market just spent a month demonstrating that trust is the contested, monetizable layer. Build it as a feature with a price, not a policy with a PDF.

Route by cost-per-completed-task, not by brand. Benchmark your actual workloads: the well-scoped, high-volume work probably belongs on the cheap barrels; the judgment-heavy work may justify the frontier. The teams that instrument this well will run AI cost structures their competitors can’t explain.

I put predictions in these essays so you can score them. Three for this one — revisit this post in January 2027.

Watch frontier labs’ enterprise contract terms over the next two quarters, not their revenue. Commoditization arrives as concessions — on pricing, data retention, IP protection, deployment flexibility — before it ever shows up as churn. The concessions are the repricing. (They won’t be announced plainly; they’ll be dressed as “vendor flexibility” and “customer-friendly terms.” Translate accordingly.)

Watch for the first public-company earnings miss traced to ungoverned token spend. Chamath’s prediction, and I think he’s right. The first EPS miss attributed to an AI OpEx line will mark the moment enterprise buying discipline actually turns — and the moment routing stops being an optimization and becomes a control function.

Watch whether open-plus-application-layer actually reaches frontier outcomes. This is Karp’s unproven premise and the hinge of the whole sovereign-AI trade. If Nemotron-class models inside serious harnesses close the gap on real enterprise workloads within a year, the trust layer captures the market. If the gap persists where it matters, the frontier keeps its pricing power longer than this essay implies.

The model layer is doing exactly what commodity layers do: getting cheaper, better, and interchangeable. That’s not the industry failing. That’s the industry working — and the value, as always, is flowing to the constraints.

List prices, output tokens per million, pulled from vendor pricing pages the week of July 13, 2026 (links at each vendor name in the published version). Two definitional notes: these are hosted-API output prices — input tokens bill separately and lower, and self-hosting an open-weight model is a different cost structure entirely (GPUs, ops, latency, reliability), which can land above or below hosted pricing depending on scale and utilization. “Open weights” means the model can be self-hosted and controlled; it does not automatically mean cheaper all-in.

Four notes for anyone doing the math on their own workloads:

Like-for-like spread: frontier flagship to best Chinese open-weight model is ~11x on output (vs. GLM 5.2; Kimi K3's July 17 listing compresses this to ~3x — see update); to the cheapest production-grade barrel, ~57x. Cached input bills at roughly 10% of list across most providers, which lowers everyone’s effective price but preserves the spread.

Open-weight vendors can’t enforce a price floor once the weights are public. GLM 5.2 lists at $4.40 output on Zhipu’s own API; the identical model through third-party aggregators runs about a third of that. When weights are public, anyone with GPUs can undercut the creator. Meta drew the logical conclusion and skipped the fight: it wound down its first-party Llama API, released Llama 5’s weights openly, and monetizes through its ecosystem instead.

Price per token isn’t cost per task. Models differ enormously in how many tokens they burn to finish a job — the newest OpenAI flagship solves standard benchmark tasks in roughly half the tokens of Anthropic’s models, which means a higher list price can still be a lower cost per completed task. The metric that matters: price × tokens-per-task ÷ quality. Third-party evaluators (Artificial Analysis and others) now publish the inputs; run it on your own workload mix.

The premium buys the last few points. On standardized coding benchmarks, a cluster of models — including open-weight ones — sits within half a point of each other at prices from $2.40 to $12, while the top two frontier models buy roughly eight more points at $25–50. Whether your workload lives inside those eight points is the entire procurement question.

Related: The Week the Frontier Went Two Directions at Once — on GLM-5.2, the Fable 5 export-control whiplash, and why neither the closed model nor the open one solves the trust problem.

No posts

Read the original on goodai.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.