RSS Amplifier

Offcuts · Jun 18, 2026

Traditional total cost of ownership is broken

0
Sign in to vote or save

Offcuts · Offcuts

Hello, and welcome to Offcuts: a podcast and newsletter about AI and commerce. Read and listened to by people at Stripe, Shopify, Dr. Martens, Loewe, H&M, and Mercedes-AMG F1. Brought to you by the team at ROIROI - engineering, design, performance marketing and AI for consumer brands.

So, what’s been going on?

Fable 5 launches. Genuinely the most capable commercial model anyone has released. The internet does what it always does, every business model declared dead before the day is out. Then three days later, gooooone for any foreign national, inside or outside the United States. US Commerce Department directive, Friday evening, no notice.

The ramifications of this directive are a debate for another edition. But what Fable’s disappearance actually surfaced is a question most brands still cannot answer cleanly: what does your AI programme cost, and do you actually know?

A team we work with at ROIROI built a Fable 5 workflow in an afternoon. A single automated process, nothing exotic, running against a frontier model API. It burned through their entire monthly token budget before anyone noticed. They are not an outlier. That pattern repeats consistently with teams moving from subscriptions to API-level build for the first time.

The tools feel familiar but the cost structure is completely different. In other words, your TCO model is now broken.

Let’s discuss.

Most organisations are currently running seats and building automated workflows, and treating them as the same category of spend. They are not.

A Claude Pro subscription or a Copilot licence covers the conversational layer - using AI as a sounding board, drafting emails, thinking through a problem. You know what it costs on the first of the month and you know what it costs on the last. The price does not change because someone on the team had a particularly productive Tuesday. These are predictable, bounded, easy to put in a budget line and forget about until renewal.

API access to a frontier model works entirely differently. You are not buying a seat. You are buying compute by the token, and the meter runs every time a workflow fires, every time a prompt gets longer, every time a new use case gets added to the endpoint. Gemini 2.5 Pro is $10 per million output tokens. Models at the Fable class run higher (much higher). A single workflow that passes large context windows to a reasoning model on every call can generate more spend in a quarter than an entire SaaS stack renewal. There is no invoice that prompts a conversation mid-month, unles one paradoxically builds a workflow for this (more on this later). The bill arrives after the consumption.

Interestingly, there is a behaviour pattern that emerges reliably at the API tier that does not exist with subscriptions: token maxxing. We spoke about this phenomenon a few weeks ago. But as a reminder - it is teams and developers, consciously or not, who push prompts to the limit of what a model can handle. Longer context, more examples, richer instructions. The outputs improve. The token count climbs. At an individual level it is good engineering practice. At an organisational level, with dozens of workflows doing the same thing simultaneously, it is an unmanaged cost driver that compounds fast and is invisible until someone pulls the billing report. Insert Uber reaching their annual token budget in a matter of months.

The practical consequence is that most AI budgets are built around the fixed cost tier and have no real model for the variable tier. Teams get approved for subscriptions, start building with APIs because the subscriptions are not powerful enough for what they want to do, and the two cost structures run in parallel with nobody owning the aggregate.

Getting this distinction clear is a good first step. Understanding what it means for how you model total cost of ownership is the harder one.

Myself and our Chief Solutions Officer Lucy Davis will be joined by Adam Yardley, Operations Lead at Hylo Athletics, for a live session on what using Claude well actually looks like inside an operations function - where most teams are today, what the next level looks like and what Adam has built at Hylo.
Free to attend, taking place on Thursday 25th June - 12pm UK / 1pm CEST.

If you lead operations, ecommerce or work inside a consumer brand and want to move beyond the chat window - this one’s for you. This event is 75% full so be quick, whilst spots last.

Save your seat →

For anyone who has run a commerce platform migration, Total Cost of Ownership (TCO) is a familiar exercise. You model platform licensing costs, ancillary technology, payment gateway fees, implementation, ongoing development resource, hosting (if you’re pursuasion is composable). Some of those inputs are fixed, some vary by usage tier or transaction volume, but the variables are bounded. You know the range. The output is a number you can put in front of a CFO with a reasonable confidence interval, defend in a board meeting, and track actuals against over a three-year period.

That predictability is what makes TCO useful as a planning tool. It works because the cost drivers are known in advance, the pricing is contractual, and the biggest variable, transaction volume, moves in a direction you can model from your own trading data.

AI breaks most of those assumptions.

The pricing floor is not fixed. The frontier labs are not yet charging what it costs them to run inference at scale. Current pricing reflects a land-grab for developer and enterprise adoption. Anthropic, Google, and OpenAI are all absorbing costs to drive usage.

That is not a permanent state. As models move from preview to general availability and labs move toward profitability, the pricing environment shifts. The baseline you build a business case around today is not the baseline you will be operating against in eighteen months.

The biggest cost driver is usage behaviour you cannot model in advance. With a payment gateway, you know roughly how many transactions you will process. With a frontier model embedded in an operational workflow, you do not know how often your team will use it, how long the prompts will get as people push what it can do, or how many adjacent use cases will emerge once the infrastructure is in place. Usage expands to fill available capability. That is a feature of good AI tooling. It is also a budgeting bug.

The natural response to all of this is to wait. Let the market settle, let the pricing stabilise, invest when there is more certainty about what the cost structure actually looks like. That response is understandable and it is wrong.

Current frontier model pricing does not reflect true inference costs. The compute costs underneath what labs are charging have not dropped fast enough to explain current rates. Labs are absorbing the difference to drive adoption, cross-subsidising frontier access with lower-tier revenue, or both. When that changes, and it will, the economics of building with AI shift materially. The brands that have already built infrastructure, workflows, and institutional knowledge of what works will operate those at the new price. The brands that waited will pay the new price to start from scratch.

What compounds is not the tooling. The tooling is available to everyone. What compounds is the organisational capability built around it: the prompts refined over hundreds of iterations, the workflows integrated into daily operations, the team’s understanding of where AI produces reliable output and where it needs supervision. A brand that has been running operational AI seriously for two years has a fundamentally different capability baseline than one writing its first business case. That gap cannot be closed by spending more later. It can only be built over time.

The Fable situation adds a further consideration. Political and regulatory exposure to frontier models is real, and it is escalating. The response is not to avoid frontier models. It is to build programmes that are model-agnostic by design, so that when a model disappears or becomes prohibitively expensive, the capability sits in the programme rather than in the dependency. That is an architectural decision you make at the start. It is very difficult to retrofit, and I’d argue the best decision one can make.

The window in which frontier AI is both highly capable and priced below cost is not permanent. Building now, with proper budget instrumentation, is not reckless optimism. It is a straightforward reading of the incentive structure before it changes.

ROIROI token guard example

Every AI budget conversation eventually arrives at the same hesitation: we are not sure the return justifies the spend.

Completely understandable given the thin veil of case study evidence and polarised media response to AI. However, it stops being useful when it becomes a reason to defer the decision indefinitely, which is what it usually becomes.

The cost of deploying AI at the wrong price point is recoverable. You renegotiate contracts, swap out expensive models for cheaper ones on lower-stakes tasks, retool workflows that turned out to be less valuable than expected. Every organisation that has been building seriously for more than a year has done all of those things. The iteration IS part of the process.

The cost of not building while your competitors do is not recoverable in the same way. The gap that opens is not a technology gap. The technology is available to everyone on the same API pricing. It is an operational gap, built from the accumulated decisions, refinements, and institutional knowledge that only come from running AI in production. A brand that can process customer returns, generate buying decisions, personalise at scale, and run its operational reporting through AI-native workflows is not just more efficient than one that cannot. It is operating with a fundamentally different cost structure. And that difference widens every quarter.

At ROIROI we have been working through this directly. We build token budgets into every AI project before anything goes near production, with hard caps and usage monitoring in place from the start. We think carefully about seat allocation - not every person in an organisation needs a frontier model subscription, but the people whose output changes materially with access to one should not be on the wrong tier. We map existing stacks before recommending new spend, asking whether a platform already in the client's environment has an enabling feature that removes the need to build. Tools like Peliqan sit across a data layer and can act as an MCP server connecting business systems to AI agents without custom infrastructure for every connection. And we stay deliberately nimble. While we were evaluating a custom data integration approach for a client, Xero launched an official MCP server. The right architectural decision the week before was different from the right decision the week after. Over-engineering a first version is as costly as not building at all.

The TCO model that worked for a platform migration does not work here. The pricing moves, the models change, and occasionally a government shuts the whole thing down on a Friday evening.

The folks that will be fine are the ones that built for that uncertainty rather than waited for it to resolve.

If you want to map where your business actually sits on this, our AI Readiness Workshop is built for exactly that conversation. A virtual or in-person session, a clear picture of what you have, what you are missing, and what to build first.

Book an intro call here.

“We brought ROIROI in to run an AI consulting project across the business. In a week they had spoken to every team, mapped our friction points, and produced a roadmap with real commercial impact. They did not try to sell us technology for the sake of it. They told us what we needed to hear, identified a significant saving, and gave us a foundation to build from.”

Gemma Stevenson, Chief Commercial Officer, Skin Rocks

Hi, I’m Tim. I’m the CEO of ROIROI, a technology and performance agency. We connect engineering, design, paid media and AI for consumer brands. I also host Offcuts, a weekly podcast and newsletter exploring the same themes. Outside of work, I’m usually cycling, swimming, running or in a sauna-cold-plunge trying to recover from all three.

No posts

Read the original on offcutsmedia.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.