RSS Amplifier

Hans Royal · Jul 2, 2026

The "Token Price War"

0
Sign in to vote or save

Hans · Hans Royal

The most watched number in AI right now is the price of a token, and it’s definitely falling rapidly. As predicted, this number drops more and more as the commoditization of certain compute continues.

OpenAI’s GPT-5.6 Luna is supposed to cost $1 per million input tokens. DeepSeek V4 Flash runs for cents. Open-weight models like GLM-5.2, Qwen 3.5, and Llama 4 cost nothing at all to license; you download the weights, load them onto your own hardware, and the marginal cost of a token approaches zero after you have paid for the hardware to run it.

So there’s a bear case for AI electricity demand to think about here. If tokens are getting cheaper, the argument goes, then the economics that make AI workloads price-inelastic to electricity cost must be weakening. The Compute Heat Rate1, which measures the maximum electricity price a data center operator can rationally sustain, should be compressing. The grid should be getting relief. That would be great if true!

It’s not that simple. Two denominators, one moving in each direction. The price of a token and the revenue generated per megawatt-hour of electricity consumed are not the same number. Per-token pricing is falling, but the revenue per MWh of AI compute is not.

When a token gets cheaper, more tokens get consumed. That is textbook Jevons: efficiency lowers unit cost, unit cost lowers the barrier to adoption, adoption increases total consumption. But the mix of what gets consumed also changes, and the mix shift matters more than the volume for the grid.

The value migrates up-stack. Cheap commodity inference draws more users and more volume. That volume is good business for the model providers, but it runs on the lowest-revenue-per-MWh tier of the workload stack. The high-value work, the work that actually sets the ceiling, is moving in the opposite direction, which I’ve written about multiple times before. You can check it out here.

Anthropic’s Fable 5 prices output tokens at $50 per million, exactly double the prior frontier. It ships with an “ultra” reasoning mode that coordinates sub-agents across extended task chains, burning through token budgets that would have been unthinkable a year ago. OpenAI’s GPT-5.6 Sol, released five days later, prices at $5/$30 and introduces its own ultra mode with the same architecture: coordinated sub-agents, long-horizon execution, massive token throughput per task. The per-token price at the frontier rising.

Meanwhile, the commodity tier is getting cheaper, which means more people use it, which means more GPUs are running, which means more electricity is consumed. The grid does not distinguish between a $1 token and a $50 token. It sees a megawatt-hour.

The Compute Heat Rate is calculated from revenue per MWh, not revenue per token. When the Q2 2026 CHR index moved to approximately $8,000/MWh, roughly 160 times the gas heat rate, that increase was driven by the frontier and agentic tiers repricing upward, not by the commodity tier. The consumption-weighted blend is dominated by the workloads where operators capture the most economic value per unit of electricity. A 10x compression at the commodity end barely moves the blend, because commodity inference is a small share of total electricity consumption relative to the high-throughput, high-value tiers that run around the clock.

This is the inversion the token price chart obscures. The most-cited number in the industry (cost per token, falling) and the most consequential number for the grid (revenue per MWh, rising) are moving in opposite directions. If you are building a forward electricity price curve, a PPA pricing model, or a capacity market reform proposal, the token chart is the wrong input.

Then the government actually made it worse. On June 26, the Trump administration asked OpenAI to limit the release of GPT-5.6 to a small group of government-approved partners. This followed the administration’s export control order on Anthropic’s Fable 5 and Mythos models earlier in June, which forced Anthropic to pull access for foreign nationals entirely. For the first time, the two most powerful AI model families in the world were behind government firewalls, available only to vetted domestic partners on a case-by-case basis.

The stated reason is cybersecurity. My take is that the structural consequence is a massive acceleration of self-hosted open-weight deployment, in every jurisdiction that suddenly cannot rely on American API access for business continuity.

Every one of those sovereign deployments is a data center. Every one draws electricity, and every one runs workloads whose economic value per MWh is governed by the same CHR dynamics. The government gating of frontier models does not reduce the total quantity of AI compute, it just redistributes it geographically across more grids, in more countries, with less coordination and less visibility.

There’s an interesting paradox with self-hosting. Take the bear case to its logical extreme. Imagine a data center that pays nothing for its model. It downloads open weights, runs them on owned hardware, and pays zero dollars per token to any API provider. Its token price is literally zero.

Does its electricity price tolerance/CHR collapse? Not at all. It may actually increase.

Think of it this way: the operator paying zero for the model still captures the full economic value of the work the model performs. It essentially cuts out the middleman. The self-hosting operator keeps 100% of the economic output per megawatt-hour instead of sharing it with a model provider, which means the rational willingness to pay for the electricity that produces that output goes up, not down. For example: a biotech lab researching new drugs on owned hardware and open-source tokens gets the economic benefit of the new drug without paying for tokens. So the MWh ‘value’ of the electricity to create that new drug actually was much higher.

The CHR framework already accounts for this. The enterprise agentic tier in the published workload table is valued on the economic output of the work, not on a token price. Self-hosting is simply the regime where every tier defaults to that valuation method.

In summary: the token price war makes AI cheaper to use. Cheaper to use means more people use it. More people using it means more data centers. Government restrictions on frontier models mean more of those data centers are being built in more countries on more grids. And the highest-value work, the work that actually sets the price tolerance ceiling, is getting more expensive per token, not less, because the tasks that justify frontier pricing are getting more complex and more compute-intensive with every model generation.

The Compute Heat Rate is compounding. Tokenomics is only part of the story.

---

Hans Royal is the originator of the Compute Heat Rate™ (CHR) framework. All views are his own and do not represent those of any employer or affiliated organization.

1

Royal, Hans, The Compute Heat Rate: Quantifying AI-Driven Electricity Price Tolerance
and Its Implications for Wholesale Market Repricing (February 28, 2026).
Available at SSRN: http://dx.doi.org/10.2139/ssrn.6322318

No posts

Read the original on computeheatrate.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.