RSS Amplifier

Damnang’s Substack · Jul 19, 2026

The Open-Source AI Era: Will Hyperscalers Keep Buying Chips?

0
Sign in to vote or save

Damnang · Damnang’s Substack

Semiconductor investors have had a run of unsettling news.

Hyperscalers are announcing far fewer new data centers, and chip stocks that rode the AI story have been sliding. On top of that, on July 16 Moonshot released Kimi K3, a 2.8-trillion-parameter model that placed level with the top US models in blind arena (LMArena) rankings, with its weight files scheduled for release on July 27. A wave of open models promising frontier-level performance at a fraction of the price has fed a simple worry,

Why would hyperscalers keep buying chips at this scale?

Has the market really moved past the semiconductor theme? This article works through the numbers in public view: how open models make money, what frontier developers’ margins actually look like, and what happens to hyperscaler profits when workloads move to cheaper models. The goal is to answer one question. Does the switch to open models make hyperscalers buy fewer chips, and where does compute and memory demand go from here?

Contents

  1. How open models make money, and how fast they are spreading

  2. The arithmetic of frontier margins

  3. How the switch to open models changes hyperscaler revenue

  4. Does the extra profit turn into compute investment?

  5. Will memory investment grow?

  6. Conclusion

Disclaimer

The figures cited here come from public sources: earnings releases, filings, public reports, and press coverage. All interpretation and outlook are the author’s personal analysis, and nothing here is a recommendation to buy or sell any security.

A token is the smallest unit of text an AI model processes, and model fees are charged per token. Frontier models are the top closed models, such as OpenAI’s GPT or Anthropic’s Claude, sold only through a paid API. Open models are models whose weight files are published, so anyone can download and serve them (they are also called open-weight models).

Open-model developers earn money differently from the frontier. Because the weights are public, a company that hosts the model itself owes the developer nothing. Even when the developer sells its own paid API, competing hosts selling the same model cap what it can charge.

The economics of that hosting business are real. Inference specialists report serving 70B-class models profitably at $0.90 to $2.00 per million tokens, and Together AI said its annualized bookings passed $1.15 billion. Serving margins exist in the open camp. What is zero, on the self-hosting path, is any mandatory premium for the model itself.

On capability, Epoch AI’s measurements put the best open models about four months behind the frontier. Kimi K3, released July 16, placed above the previous frontier generation in blind arena rankings right away, and level with the newest frontier models.

Usage data shows the shift too. OpenRouter is a routing platform that lets developers call hundreds of models through one API, so where its traffic goes is a widely used proxy for model market share. On this platform, the share of calls sent to models priced under $1 per million tokens rose from 18% in January 2026 to 41% in June. The share of tokens processed by Chinese open models such as DeepSeek, Qwen, and Kimi went from under 2% in late 2024 to about 61% in May and June 2026. Counting all open-weight models, the latest weekly snapshot is about 69%.

Figure 1
Figure 1. OpenRouter usage split two ways. Left: share of all calls sent to budget models (under $1 per million tokens), from 18% to 41% in five months. Right: share of all tokens processed by Chinese open models, from under 2% to about 61% in eighteen months (token-count basis, retrieved June 2026). Source: OpenRouter public data and press reports.

While volume moves at this speed, where has the money gone? Start with frontier developers’ margins.

Public reports put Anthropic’s gross margin (revenue minus the cost of serving) at -94% in 2024, improving to the mid-60s by 2026, with a 36% operating margin before model training costs (a research estimate). OpenAI’s inference cost reportedly quadrupled in 2025 and its adjusted gross margin slipped from 40% to 33%.

Add training costs and the picture changes. Reporting based on OpenAI’s 2024 investor projections put its 2026 loss at $14 billion. Anthropic told investors it expects its first quarterly operating profit in Q2 2026, $560 million at roughly a 5% margin, while stopping short of promising a profitable full year. The accounting bases are not identical, but for scale: Anthropic’s 36% before training (a research estimate) sits near AWS’s 35–38% operating margin (37.7% in Q1 2026), while the margin including training is around 5% even in the first profitable quarter, far below the hyperscalers’ 30s.

What matters more than the margin rate is that the premium is actually being paid. As of late May 2026, Anthropic processed about 11% of tokens on OpenRouter but took about 42% of estimated model spend on the platform. Benchmark gaps have narrowed to a few points, yet a 25x price gap ($25–30 per million tokens for Opus-class versus under $1 for budget open models) keeps getting paid. The reason is that companies do not pay for benchmark points. They pay for an agent that does not fall apart in the middle of a fifty-step job, for reliability where failure is expensive, and to avoid re-validating their whole workload on a new model.

Figure 2
Figure 2. Frontier share on OpenRouter, late May 2026: 11% of token volume, 42% of estimated spend on the platform. Volume is moving, but the money still sits with the frontier. Source: OpenRouter public data and press reports.

So the real risk that open models pose to frontier developers is not a falling margin rate. It is the loss of mid-tier volume. Open-model hosts now offer APIs that are drop-in compatible with OpenAI’s and Anthropic’s, so switching models can be a one-line change. Work with well-defined requirements, such as document classification, extraction, summarization, and simple code, is leaving first. What stays on the frontier is the hardest work, where the capability gap is worth money. Prices there are holding, but as the middle drains away, frontier revenue growth depends more and more on how fast that top tier expands.

Hyperscaler AI revenue comes through three accounts: renting compute to frontier developers, charging cloud customers for model usage, and selling their own products with models inside, like Microsoft Copilot. The open-model switch changes the customer account first.

Take a base case: an enterprise pays $25 for a job of one million tokens on a frontier API, and follow the money before and after the switch. Before, the payment reaches the hyperscaler indirectly. The $25 first becomes revenue for a developer such as OpenAI or Anthropic. At the reported 65% gross margin, the developer spends $8.75 of it on serving cost, that is, buying cloud compute, and keeps $16.25 to fund payroll and training. The hyperscaler books $8.75 of revenue. At a 35% infrastructure operating margin, that is $3.06 of profit and $5.69 of server operations.

For simplicity, this model treats the developer’s entire serving cost as hyperscaler compute purchases, assumes that after the switch the company deploys directly on the same hyperscaler’s cloud, and prices every tier on the same basis of one million output tokens.

When the same company moves to an open model, the pass-through disappears. The weights are free, so nothing is owed to a developer, and everything the customer pays to run the model in the cloud is infrastructure revenue for the hyperscaler. Structurally, this is better revenue: compute rental that used to ride on someone else’s model becomes a direct sale to the hyperscaler’s own customer. What comes with it is a price collapse. The same million tokens drops from $25 to under $1.

Figure 3
Figure 3. The base case, before and after. Before, only $8.75 of the $25 reaches the hyperscaler. After, the full payment is infrastructure revenue, but the payment equals $1 plus whatever portion of the saved $24 gets spent again.

That $1 is not a discounted version of the $25 model. It is the price of a small model whose active parameters, and therefore compute per token, are a tenth to a fiftieth of the frontier’s. Open serving prices form where competing hosts add a margin over cost (hence the profitable $0.90–2.00 price points). So on both paths, the compute the hyperscaler provides equals about 65% of the revenue it books.

Figure 4
Figure 4. Price and estimated compute for one million tokens, by model tier. Open models use the competitive serving cost ratio (price × 0.65); the frontier uses the reported 65% gross margin in reverse (price × 0.35). Prices are on an output-token basis. The budget model’s compute ($0.65) is a small fraction of the frontier’s ($8.75), while the top open model, Kimi K3 ($9.75), is frontier-class. Author’s estimates.

If the switch is to a top open model, there is little to weigh.

K3’s compute payment of $9.75 is larger than the frontier’s $8.75, so hyperscaler revenue and server operations rise the moment the switch happens, regardless of any reuse. The case where the outcome can go either way is the trade-down to a budget model, so the base case continues on the $1 path, the least favorable one for the hyperscaler.

Under the assumptions above, the size of the customer account after the switch is set by the reuse rate: how much of the saved $24 goes back into tokens. If none of it does, hyperscaler revenue is $1 and profit is $0.35, far below the $3.06 before. If 32% is reused, profit matches at $3.06. If all of it is reused, revenue is $25, profit is $8.75, and server operations are $16.25, which is 2.9 times the level before the switch.

Figure 5
Figure 5. How the same $25 splits at different reuse rates. The middle bar is the 32% break-even, where post-switch hyperscaler profit ($3.06) matches the pre-switch level; above it, the post-switch path earns more. Scenario analysis under the author’s assumptions (frontier gross margin 65%; infrastructure operating margin 35%).

If companies re-spend just a third of what they save, the hyperscaler earns as much as before. So where are companies actually sitting? Reports keep coming of annual AI budgets used up in a single quarter, and token usage is growing five to seven times a year. That is not re-spending savings; that is growing the budget itself. Reality sits to the right of the rightmost bar in the chart. In that zone, the $16.25 premium that used to go to the frontier developer disappears, and hyperscaler profit and server operations fill the space.

The other two accounts point the same way. In the own-product account, swapping in an open model removes the license fee from the cost line and improves margins. The developer-rental account carries a risk: if mid-tier volume loss squeezes frontier revenue, developers could slow their training compute purchases. But multi-year rental contracts with OpenAI and Anthropic sit in backlog, so this account is not one that breaks quickly.

So a hyperscaler whose profits grow with this switch faces the next question: raise, hold, or cut compute and memory investment?

The following sections take it up in detail.

Read the original on damnang2.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.