For twenty years, the software business has been the best business in the world.
Once a SaaS company reaches scale, its unit economics become almost embarrassing. A new customer adds revenue that is nearly pure margin. The marginal cost of serving one more seat is close to zero. Gross margins of seventy to eighty-five per cent are not unusual. The “rule of 40,” the “magic number,” the stacking of cohort retention curves, all of it flows from one fact: at scale, software costs nothing to serve.
That fact is about to change. Arithmetically, not just philosophically.
The cost of serving an AI-powered workflow is not zero. It is not even close to zero. It is measurable, variable, and tied to volume. If the SaaS industry continues to price as if compute is free, a large number of companies are going to walk into a margin problem that their finance teams have never had to solve.
Call it the inference tax. And it is about to redraw the map of enterprise software.
The SaaS margin story has always been simple.
Hosting costs scale sub-linearly with usage. Once your infrastructure is provisioned, an additional user costs almost nothing. Customer support scales, but it scales with contracts, not with usage. The most expensive thing in your cost of goods is often the content delivery network bill and a handful of database licences. At a typical mid-market SaaS company, cost of revenue might run eighteen to twenty-five per cent, and gross margin clears seventy-five per cent cleanly.
That number is what makes the public markets love SaaS. It is what lets growth-at-any-cost make sense for a decade before anyone asks about profitability. It is what permits land-and-expand motions that look uneconomic for years and then print money. The entire playbook assumes that gross margin is a given and the game is distribution.
The inference tax breaks that assumption.
When an AI feature replaces a workflow, the cost structure changes in three places at once.
Per-request inference cost. Every meaningful action an agent or assistant performs costs real money; paid to the model provider, metered per token. Even after the significant price declines of the last two years, a non-trivial enterprise workflow, one that involves reading several documents, reasoning across them, calling tools, and producing structured output, still costs cents to dollars per invocation. At high volume, that is a significant line item not just a rounding error.
Retrieval and context cost. Serious enterprise AI is not just a model call. It is a retrieval pipeline, a vector store, a cache, an orchestration layer, an evaluation harness. Each of these has its own cost curve, and most of them scale with usage. Customers who “use the product more” now increase the bill in a way that a traditional SaaS customer never did.
Model improvement cost. To keep a vertical AI product competitive, the vendor has to invest continuously in evaluation, fine-tuning, safety scaffolding, and sometimes model training. This is research-and-development cost in the traditional sense, but it also leaks into cost of revenue, because the improvement only lands if it ships to the production inference path. The line between capex and opex blurs, and finance teams hate this.
The combined effect is that a product which replaces meaningful work, the kind buyers will actually pay a premium for, runs at cost of revenue that can reach forty to fifty-five per cent. That is gross margin in the forty-five to sixty per cent range. Still a good business, but not a software business in the way the public markets have priced the category.
The next two years will distribute the inference tax unevenly across the industry. Three patterns are already visible.
Pattern one: the incumbent who raises price but not capability. A legacy SaaS vendor bolts an AI assistant onto a thirty-year-old product, ships a fifteen per cent price uplift, and claims AI leadership. Behind the scenes, inference costs rise proportionally with usage. If the assistant is used lightly, margins hold but the feature fails to differentiate. If the assistant is used heavily, the price uplift does not cover the inference, and gross margin quietly compresses. The public narrative remains unchanged for two to four quarters, until the CFO gets asked an uncomfortable question on an earnings call.
Pattern two: the AI-native startup priced like SaaS. A new entrant builds a workflow-replacing product. The demo is extraordinary. The pricing is per-seat because that is what buyers are familiar with. Early customers love it, and use it heavily, which is the whole point. Gross margin prints at forty per cent on a revenue run rate that would not cover the CEO’s fundraising dinners. The investors notice first. The pricing model gets a hasty rework. Valuation multiples get a second rework.
Pattern three: the operator who prices per outcome. A smaller, more disciplined company looks at the economics of its workflow replacement honestly. It prices per successful task, or per ticket resolved, or per invoice processed. It absorbs the variable cost and marks it up. Gross margin is lower in headline terms, sixty per cent rather than eighty, but the business is coherent. Revenue grows with value delivered rather than with seats nominally licensed. These companies will be smaller at each stage than their SaaS predecessors. They will also be more defensible.
The interesting question is not which pattern is right in the abstract. It is which pattern a given workflow supports. Inference tax is not uniform. A product that replaces a rare, high-value action can absorb it easily. A product that replaces a high-volume, low-value action cannot.
The inference tax has a mirror image on the buyer side, and it is almost entirely un-thought.
Most enterprise software contracts today are priced per user per month. When a vendor adds an AI SKU at twenty to forty per cent uplift, the buyer is effectively paying a tax on every seat, regardless of how much the seat actually uses the AI feature. In many organisations, twenty per cent of users will drive eighty per cent of the inference-heavy usage. The buyer pays for a hundred.
The correct procurement response is not to resist the AI SKU. It is to change the unit of pricing. Ask the vendor for usage-based pricing on AI features, with a minimum and a cap. Ask for audit logs that prove usage. Ask what the vendor’s own cost of goods looks like, and refuse to subsidise their margin panic by paying per-seat for a per-call capability.
Very few procurement teams are structured to ask these questions, though they will learn. There is a real advisory opportunity here for anyone who speaks both languages.
Stand back and the shape of the next five years becomes clearer.
Horizontal SaaS incumbents with broad surface area, Microsoft, Salesforce, ServiceNow, Atlassian, can absorb the inference tax more comfortably than their margins suggest. They can cross-subsidise AI features from their broader platform margin, and they can negotiate preferential rates with model providers at a scale smaller competitors cannot match. They will look fine for longer than they should, and the gradual drag on blended gross margin will be explained away as “AI investment” for years.
AI-native vertical players will have a harder conversation with their investors earlier. The ones who survive will have rebuilt pricing around outcomes and will look nothing like classical SaaS on their P&L. They will also, increasingly, be the most interesting businesses to work inside, because the commercial model will enforce clarity about the value they deliver.
Old-school enterprise vendors with heavy services attach, the SAPs and Oracles of the world, are less disrupted on margin than the headline story suggests, because their pricing was never purely per-seat subscription. They have more commercial flexibility to price inference through than the SaaS pure-plays do.
And a new category will emerge: companies whose product is explicitly priced, measured, and sold on outcomes.
Not “number of agents deployed.” Not “seats of AI assistant.”
Tickets closed. Invoices processed. Claims adjudicated. Incidents triaged.
These companies will be mistaken for services firms for the first few years, they are not! They are the honest version of what AI software should look like when it is priced for what it does, not for what it displaces.
There is a reason most SaaS boards are not talking about this yet, and it is not because they do not see it.
Gross margin compression is the single most painful metric to explain to public markets. It rattles the narrative that drove valuations for a decade. No CFO wants to be the first to admit that their AI-powered product line is running at sixty per cent gross margin, when the rest of the business is at eighty. The incentive, for the next two to four quarters, is to muddle the reporting. Bundle AI usage into the base product, blur the variable cost and describe inference as “AI investment,” which sits below the line.
This will work, until it does not.
At some point in the next two to four earnings cycles, one large public company will be forced to break it out honestly, probably under pressure from an activist investor or a short report. When that happens, the market will reprice the whole category in a single week.
Software is not dying. It is just no longer free to serve. The industry has never had to think this way before, and most of it is not ready.
In the next piece in this series, I want to take on a related myth, and one that is even more widely believed. The idea that proprietary data is what protects enterprise software companies in the age of AI. I think the data moat was always a myth, and I want to say why.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.