I seem to be having the same conversation everywhere at the moment… One person tells me the cost of AI is going to become enormous. Agents will consume extraordinary numbers of tokens, reasoning models will think for longer (or more deeply), and companies are dramatically underestimating the inference bill they’re creating. Someone else then quickly (and equally enthusiastically) tells me the opposite. Tokens are already getting cheaper, models are becoming more efficient, hardware is improving, and competition will drive prices down. In a few years, they argue, we’ll laugh that we ever worried about token costs at all.
Occasionally, I meet someone in the middle who thinks prices will simply settle. AI becomes another relatively predictable line in the technology budget, perhaps more like cloud infrastructure than anything genuinely novel.
The interesting thing isn’t that these people disagree. It’s how certain everyone seems to be. I don’t know which of them is right and I’m increasingly suspicious of anyone who claims that they do.
There is plenty of evidence for the optimists. The cost of achieving a given level of AI performance has already fallen extraordinarily quickly. Stanford’s AI Index found that the inference cost of achieving roughly GPT-3.5-level performance fell more than 280-fold between November 2022 and October 2024.
There are plenty of reasons to think this direction of travel might continue. Hardware gets better, models get smaller and more efficient, distillation improves, caching reduces unnecessary work, competition puts pressure on pricing, and open models create alternatives to proprietary ones. Organisations themselves should also get much better at using the right model for the right job rather than throwing the most capable, expensive model at everything.
We’ve seen versions of this story before. Storage used to be something we thought carefully about. Bandwidth was expensive. Computing power was scarce. All three became sufficiently cheap that entirely new products, behaviours, and industries became possible. Perhaps intelligence follows the same curve, with today’s expensive frontier capability becoming tomorrow’s commodity infrastructure. Things that currently require careful calculations about token consumption might eventually become cheap enough that nobody particularly cares.
It’s a compelling argument, but it contains an important assumption that when intelligence gets cheaper, we’ll consume roughly the same amount of it. History suggests otherwise.
The way we use AI is already changing. A relatively simple interaction with an LLM might involve sending it a prompt, giving it some context, and receiving an answer. An agentic process can be something quite different.
An agent might retrieve information from several systems, reason about it, call another tool, examine the result, decide it needs more information, search again, ask another model to check its work, discover a problem, retry part of the process, and finally produce an answer. What looks to the user like one action can represent a considerable amount of computation behind the scenes.
Reasoning models introduce another dimension. Rather than simply producing an answer, we’re increasingly prepared to spend more compute to get a better one. We can give the model more time to think, allow it to explore alternatives, and ask it to verify its own work. The economics are rather different from the simple chatbot interactions through which many organisations first encountered generative AI.
At the same time, we’re not just making individual interactions more sophisticated. We’re proposing vastly more of them: AI inside products, supporting employees, handling customer interactions, monitoring systems, producing content, analysing data, writing software, checking compliance, conducting research, and increasingly acting rather than simply answering.
This creates an important distinction. The cost of an individual token could fall dramatically while the number of tokens an organisation consumes increases even faster. Cost per token falls, tokens per task increase, and the number of AI-mediated tasks increases enormously.
The resulting bill could go almost anywhere.
This is the possibility I find most interesting. AI could become dramatically cheaper and dramatically more expensive at exactly the same time.
The cost of a particular level of intelligence could continue collapsing. Yesterday’s frontier model becomes today’s small model and tomorrow’s commodity. But businesses don’t necessarily keep buying yesterday’s intelligence.
If something that costs £1 today costs 10p tomorrow, that’s wonderful. But if tomorrow also offers something for £2 that can solve a problem the £1 model couldn’t touch, we may choose to spend the £2. The frontier moves, and our expectations move with it.
We could therefore end up with extraordinarily cheap commodity intelligence handling enormous volumes of routine work while organisations continue paying a significant premium for the best available reasoning. The average price of intelligence could fall while the economic value - and therefore the amount we’re willing to spend - rises.
There’s another effect too. Making something cheaper tends to create new demand for it. When computing became cheaper, we didn’t simply run the same software for less money. We built vastly more software. When storage became cheaper, we didn’t carefully preserve our existing storage requirements and pocket the saving. We started storing photographs, video, backups, telemetry, logs, and enormous datasets we would previously never have contemplated keeping.
Economists have a name for versions of this phenomenon: the Jevons paradox. Improvements in efficiency can sometimes increase overall consumption because they make entirely new uses economically viable.
Perhaps something similar happens with intelligence. A task that currently costs £10 of inference might cost £1 in a few years. That’s excellent news until the £1 price makes 100 other tasks worth automating too. The unit cost has collapsed, but the overall bill has increased tenfold.
This is why I think (and may well be wrong!) some conversations about token costs become slightly confused. There are at least three variables moving simultaneously:
What a token costs
How many tokens it takes to accomplish something useful
How many useful things we decide to do with them
All three could move dramatically, and they’re not independent. Cheaper tokens encourage more ambitious applications. More capable models make previously impossible applications viable. More autonomous systems create longer chains of reasoning. Better results encourage greater adoption, and greater adoption creates more use cases.
Trying to forecast total AI expenditure from the price of a million (or a billion) tokens is therefore a little like trying to predict the economic impact of the Internet from the price of a megabyte. It’s relevant, but it’s nowhere near sufficient.
And there is a further complication. “Token cost” itself is becoming a less useful shorthand for the economics of AI. Some workloads will be heavily cached. Some will run on small, specialised models. Some may run locally. Others will invoke expensive frontier models only occasionally. Agentic systems may involve model calls, search, tool use, external APIs, databases, infrastructure, and human review.
What we’re really trying to understand isn’t the future price of tokens. It’s the future cost of useful machine intelligence. That’s a much harder number to forecast.
This would all be an interesting intellectual exercise if we weren’t simultaneously being asked to make some very consequential decisions.
Companies are deciding today whether to hire people whose roles may look very different in three years. They’re redesigning workflows around agents, choosing architectures, negotiating model providers, setting prices, approving product roadmaps, and constructing investment cases. Finance teams are putting numbers into spreadsheets, technology teams are making architectural decisions, People teams are constructing workforce plans, and Boards are being asked to approve strategies extending several years into the future.
All of those activities demand a degree of certainty that the underlying technology stubbornly refuses to provide.
Take hiring. If machine intelligence becomes dramatically cheaper while capability continues improving, the economic case for automating increasing amounts of knowledge work becomes stronger. A sensible workforce plan might therefore assume that some future hiring never happens. But if the most valuable forms of machine intelligence remain relatively expensive, or require much more human oversight than anticipated, the economics could look quite different.
Architecture creates a similar problem. It might be tempting to optimise heavily around today’s model economics, but those economics could be largely irrelevant in two years. Equally, designing everything on the assumption that inference will inevitably become negligible could leave a company with a very unpleasant surprise if its increasingly agentic products consume vastly more compute than expected.
Pricing may be harder still.
Do you bundle AI into an existing subscription?
Charge separately?
Meter usage?
Absorb it as a cost of doing business?
If the cost of providing an AI-enabled feature falls by 90%, one model might look sensible. If consumption increases 100-fold, another might.
These aren’t decisions we can defer indefinitely while we wait for the answer. Some of them are being made now, and some will be difficult or expensive to reverse.
I don’t think the answer is to wait. Waiting for the economics of AI to settle before making decisions about AI would itself be a rather significant strategic decision, and we could be waiting for quite a while.
But I also don’t think the answer is to pick whichever forecast we find most persuasive and quietly convert it into fact.
The more useful question is whether the decisions we’re making work across several plausible futures. We should understand what happens to a product if inference becomes 90% cheaper, but also what happens if its agents consume 10 times as much of it. We should consider a world in which commodity intelligence becomes effectively free while the intelligence that creates genuine competitive advantage remains expensive. We should understand what happens if model capability continues improving extraordinarily quickly, and what happens if the rate of progress begins to slow.
There will still be bets. There have to be. Strategy without bets isn’t really strategy at all. But there’s a considerable difference between making a bet while understanding the assumptions underneath it and constructing a five-year plan that accidentally assumes one particular version of the future.
That distinction matters particularly when decisions are difficult to reverse. Hiring 100 people is different from experimenting with a new model provider. Rebuilding a product architecture is different from changing a prompt. Committing to unlimited AI usage within a fixed-price product is different from running a six-month pricing experiment.
The less certain the future, the more valuable optionality becomes.
We’re generally encouraged as leaders to have a point of view. Make the decision. Set the strategy. Build the roadmap. Commit to the outcome. I like all of those things, but good decision-making isn’t the same thing as certainty.
Right now, we’re being asked to make hiring decisions, architecture decisions, pricing decisions, product decisions, and organisational decisions against one of the fastest-moving cost and capability curves I’ve encountered. Pretending we know exactly where that curve ends doesn’t make those decisions better.
So I think we need to become a little more comfortable saying something that doesn’t always come naturally in leadership meetings: we don’t know. Not as an excuse for inaction, but as an input into better decision-making.
Have a view about where the economics of AI are going. Build a hypothesis and make the bet. But understand what needs to be true for that bet to work, know which decisions can be reversed cheaply and which cannot, and decide in advance what you’ll do if the world turns out differently.
Because right now, the people telling me that AI will become enormously expensive sound convincing. So do the people telling me it will become extraordinarily cheap. Which is probably the point - a little uncertainty might actually be the more responsible position. Especially when everyone else sounds so certain.
If the future economics are unknowable, that doesn’t mean today’s costs are unmanageable. In fact, one of the things I increasingly look for in AI products and AI-enabled organisations is evidence that somebody understands the economics of what they’re building. You don’t need to know what a token will cost in 2029 to make sensible decisions in 2026...

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.