Every year or so, something triggers a referendum on whether AI is a bubble. Last year it was the infamous MIT study reporting that 95% of corporate genAI pilots had delivered no measurable return. A statistic that launched a thousand “it’s all hype” takes before getting picked apart on methodology, from its absurdly narrow definition of success to a sample nobody could properly stress-test. It also landed early, taking the temperature of chatbot-era pilots before the agentic wave had really begun. And the report’s own conclusion, conveniently forgotten in the panic, was that the failures came down to how companies were organised.
This year’s candidate is a behaviour with an absurd name: tokenmaxxing. The logic runs that if AI makes people more productive, then more AI use must mean more productivity, so you push employees (engineers above all) to run models and agents as hard as they can, sometimes with an internal leaderboard ranking who burned the most. Maximise the tokens and you maximise the output, or so the theory goes.
The (late) realisation that maybe things weren’t as simple is what poured fresh fuel into the sceptic machine.
Uber revealed it had burned through its entire 2026 token budget in the first four months of the year, largely on Claude Code.
Marc Benioff announced Salesforce's Anthropic bill would hit $300 million this year, and in the same breath wished aloud for a "smart router" to send easy queries to cheaper models - literally the first thing a 15,000-strong engineering team burning that much cash should have started with. An embarrassing admission in any other context.
Amazon scrapped its internal "Kirorank" leaderboard after workers started assigning agents to carry out needless tasks purely to climb the rankings. Meta took down the informal token leaderboard its employees had built. Microsoft cancelled Claude Code subscriptions for cost reasons across several key product divisions. One AI consultant told Axios that a client had accidentally run up HALF A BILLION DOLLARS on Claude in a single month, having never set a usage limit on its employees' licences. Gary Marcus, reliably, declared it breaking bad news for AI.
The narrative being assembled from these data points goes roughly like this: companies spent a fortune on tokens, got limited ROI, therefore AI doesn’t deliver on its value promise, therefore enterprise spend will contract, therefore the entire capex buildout is a write-off waiting to happen. Each step in that chain sounds plausible but sadly, each step is wrong.
Have we just discovered that incentives work?
Start with Uber. Engineers were given access to powerful AI coding tools, usage was tracked, and a "leaderboard" emerged. Why are we surprised that engineers optimised for the one legible, visible metric? It hardly helped that the most powerful man in AI hardware had been egging them on. Jensen Huang said he would be alarmed if a $500,000 engineer didn't consume $250,000 in tokens a year, and floated token budgets as a Silicon Valley recruiting tool. Token spend turned into a measure of ability. It might have been a serviceable proxy for adoption until it became the target on a leaderboard, which is the moment Goodhart's Law takes over: a measure stops being a good measure the instant you start optimising against it.
Uber's COO noted that high token spending didn't correlate with more consumer features delivered. The COO of a $140B company going on a podcast to explain his company torched its entire annual budget in a quarter with nothing to show for it would, with any other line item, be confessing his own mismanagement. The sort of thing no executive would dare say aloud about headcount or cloud. The AI context creates a strange exemption, where confessing to basic governance failures is somehow reframed as a principled critique of the technology, and as ever, the media obliges by running with the bubble angle rather than the obvious one.
That correlation was never going to hold, because engineers work in teams and rarely decide what gets built in the first place. Shipping a feature runs through the product managers who choose it, the customer research that justifies it, the designers who shape it, and the reviewers who sign it off. Coding is one link in a long chain. Run that single link at ten times the speed and the chain doesn't move ten times faster; the binding constraint simply shifts to whoever sits downstream. If anything, a wall of tokens yielding a trickle of features is precisely what you'd expect.
Hand a team a blank cheque they will cash it. The only surprise is that anyone is surprised.
Why are we treating input costs as an output measure?
The deeper error in the tokenmaxxing panic is treating token spend as a proxy for AI value creation. High token consumption never told you AI was working. A low yield per token does however tell you the work is badly organised. Tokens are an input cost, the electricity bill for computation, and judging AI ROI by tokens burned is about as useful as judging a factory’s productivity by its energy bill.
Amazon, to their credit, have figured this out. They have now replaced token consumption with "normalised deployments," a measure of engineers using AI to ship useful code. That is what a real output metric looks like. The companies still debating whether AI delivers ROI might want to start by measuring something that has ROI in it.
What the actual supply and demand picture says
Epoch AI put numbers on the macro picture the bubble narrative conveniently ignores: global inference capacity is more than tripling each year, while demand for tokens grows by roughly ten times a year. GPU rental prices are still up twofold on four months ago. That is the signature of a market under serious supply pressure.
The pricing correction now underway - companies switching to cheaper models for lower-complexity tasks, building routing logic, rationalising access - is a healthy market finding its equilibrium. A world in which the frontier labs subsidised unlimited consumption indefinitely would have been the actual bubble. Friction in the pricing signal is the system working.
The one thing worth worrying about
There is a genuine cost to the end of the free-experimentation era. Cheap, open-ended agent experimentation is how organisations and the non-technical people inside them figure out what agents are actually good for. The promise of agentic AI runs well past doing existing work faster, into discovering entirely new categories of work that have only just become tractable. That kind of discovery needs people free to play, without a CFO watching the token counter.
As pricing discipline tightens, that experimentation will concentrate in well-resourced organisations. The accelerating advantage is real: the firms that can afford broad discovery will find the highest-value applications first, and the gap between them and everyone else will widen. Organisations won’t stop using AI. The real risk is that only a handful can afford to discover what comes next and what works for them.
The story was never about AI
So, what have we learned?
That engineers burn tokens when you ask them to burn tokens.
That tracking consumption without tying it to outcomes produces consumption without outcomes.
That technology budgets need the same controls as every other budget.
None of this is new, and none of it is specific to AI.
It feels like a revelation only because the industry spent two years treating token spend as a proxy for AI seriousness, celebrating whoever burned the most. The reckoning, in the end, is with a metric that should never have been dignified.
Thanks for reading my substack! Subscribe for free to receive new posts and support my work.
About
I analyse AI progress beyond the headlines, focusing on enterprise execution, incentives, and real-world economic impact.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.