RSS Amplifier

The AI Company Builder Memo · Aug 13, 2026

Welcome to the Dot-AI Bubble

0
Sign in to vote or save

Karthik Ravi · The AI Company Builder Memo

Let’s get the obvious part out of the way. AI works.

The models are useful. People are paying for them. Enterprises are stuffing them into software development, customer support, sales, finance, research, healthcare and basically anything else involving a keyboard. The revenue growth is ridiculous. The infrastructure spending is even more ridiculous. You can watch a coding agent complete an afternoon of work before lunch and understand immediately why every company on Earth wants some version of this.

So if we’re in an AI bubble, it isn’t because nobody wants the product. That’s the old argument, and it’s increasingly silly.

The much more interesting question is how an industry can have overwhelming demand, enormous revenue growth and some of the most useful products ever created while still becoming a bubble. That sounds contradictory until you separate the AI economy into two factories.

The first factory turns silicon, memory, electricity and a truly stupid amount of capital into machine intelligence. The second factory has to turn that intelligence back into customer revenue, gross profit, free cash flow and an acceptable return for everyone who financed the first one.

We’ve gotten extremely good at building the first factory. The entire market is now leaning on the assumption that this means the second one will work too, and, uh, that part is considerably less obvious.

To keep the numbers concrete, we’re going to build a fake but physically plausible AI cluster and call it Atlas, mostly because repeatedly saying “our hypothetical billion-dollar B200 project” gets annoying. Atlas contains 8,192 Nvidia B200 accelerators and costs $1 billion after including servers, networking, storage, power equipment, cooling, the building, commissioning and financing carry. A cloud provider operates it. A frontier lab supplies much of the expected demand. Private credit finances 60 percent of the installed cost. None of this is a claim about one real project. It’s a whiteboard model assembled from the public specifications, filings, contracts and market evidence we’ll use as we go.

Atlas has three clocks running at the same time:

Here’s the weird part. Atlas doesn’t need AI to fail. Its lab customer can have millions of paying users. The models can keep improving. Every rack can eventually run at full power. The project can still disappoint its investors if construction finishes late, ordinary work moves to cheaper models, accelerator pricing falls or the hardware loses its premium economic life faster than the debt gets repaid.

That gives us the question at the center of this entire rabbit hole:

How can an industry have overwhelming demand and still be a bubble?

To answer it, we have to zoom all the way into one generated answer and then zoom all the way back out. We’ll start with memory traffic, batching and the cost of completing one useful task. Then we’ll follow the customer’s dollar through the lab, cloud provider, data center, semiconductor supply chain and lender. Finally, we’ll ask where the customer gets that dollar in the first place, because “AI spending” isn’t a source of money. It has to come from additional revenue, higher prices, lower costs or somebody’s payroll.

That last step is where the whole story eventually lands. Cost savings can finance AI adoption for a while. Labor savings can finance it for even longer because an eliminated salary is a recurring annual saving. But neither pool grows forever. A company can avoid the same hire only once, and it can’t reduce exposed payroll below zero. If laboratories, cloud providers, data centers, chipmakers, utilities and lenders all expect their AI revenue to keep compounding after the easiest costs have been removed, the broader market eventually has to put more dollars into customers’ bank accounts. Existing companies must sell more, new companies must form, prices must rise without destroying demand or entirely new markets must appear.

Labor savings can finance the bridge. Revenue expansion has to finance the destination.

The dot-com bubble was largely a bet that internet demand would arrive faster than it actually did. The dot-AI bubble may have almost the opposite problem. Demand is already here. The unanswered question is whether intelligence can be manufactured cheaply enough, sold at a high enough price and converted into enough customer value to repay the factory built around it.

The dot-com bubble overestimated how quickly demand would arrive. The dot-AI bubble may be overestimating how profitably that demand can be served.

The opening distinction becomes more useful if we make it concrete. Imagine two companies standing beside two enormous factories.

The first factory produces websites in 1999. It has already purchased servers, leased fiber and hired a small army of people who use the word “eyeballs” as though it’s a financial metric. The factory can make websites. That part works. Its problem is that relatively few people are online, even fewer are comfortable spending money there, and almost nobody knows which website will eventually become a real business.

The factory has production capacity waiting for demand. Now imagine a second factory producing AI intelligence. It’s 2026. Hundreds of millions of people already use the product. Developers leave coding agents running while they sleep. Enterprises feed customer-support tickets, financial documents, sales calls and internal codebases into models. The problem isn’t persuading anyone that generated intelligence might be useful.

The problem is that every unit of output requires the factory to turn on again. That gives us two very different diagrams:

The first bubble asked whether customers would arrive before companies ran out of money. The second asks whether serving customers will become profitable before the industry locks itself into an absurd amount of long-lived infrastructure. This sounds like a subtle distinction. It isn’t. It changes the variable that can break the system.

For a simplified dot-com business, value depended on demand catching up with capacity:

For a frontier-model provider, value depends on something more like:

The user number can be spectacular while the expression inside those parentheses remains disappointingly small.

Keep that visual in your head. Whenever somebody announces another billion tokens, another million users or another gigantic cloud commitment, mentally draw the parentheses. How much did the customer pay for the useful work? How much did the model provider spend producing it? And how much long-lived capital had to be committed before either number appeared?

The lazy version of dot-com history is that investors funded a collection of ridiculous websites, everyone discovered the internet was fake, and then Amazon somehow crawled out of the wreckage. That isn’t what happened.

The internet was already producing real productivity improvements. Between 1995 and 2000, US productivity growth averaged 2.8% annually, almost double the rate of the preceding 22 years. Technology-sector earnings also grew rapidly. A San Francisco Federal Reserve analysis found that earnings among a sample of publicly traded technology companies increased at an annual rate of 38% between 1998 and 2000.

The underlying technology was real. The productivity improvement was real. Some of the revenue was extremely real.

Investors took those true observations and extended them through a financial funhouse mirror.

If the internet was changing everything, every company associated with the internet could be assigned an extraordinary future. If traffic was growing, monetization would eventually appear. If monetization hadn’t appeared, the company simply needed to grow faster before some less visionary investor asked an irritating question about cash flow. The result was a positive-feedback loop:

Notice that nothing in this loop requires the technology to be fraudulent. It only requires capital to grow faster than durable returns.

That’s why “AI is real” isn’t a rebuttal to an AI bubble. It’s almost irrelevant. Railroads were real during railway manias. Telecommunications was real during the dot-com boom. Housing was real in 2007. A useful asset can be financed at a stupid price.

The Federal Reserve later described the late-1990s investment boom as plainly overdone. Companies overspent on technology, productive capacity and headcount to satisfy demand that proved unsustainable. The underlying causes included telecom deregulation, the web’s one-time arrival, the Y2K replacement cycle and equipment demand from dot-com companies that subsequently disappeared.

When the assumptions changed, the technology didn’t vanish. The financing did.

Bay Area nonfarm employment subsequently fell by about 350,000 jobs from its December 2000 peak through August 2003, a 9.5% decline. Roughly half the losses were in information technology.

The internet kept winning while a considerable number of internet companies, employees and investors got wrecked. That’s the historical analogy worth preserving.

The internet of March 2000 was real, useful and commercially tiny relative to the prices attached to it. Census estimated U.S. retail e-commerce at $5.8 billion in the first quarter of 2000, just 0.8 percent of retail sales under the agency’s then-current definition. By the fourth quarter it had reached $9.2 billion and 1.1 percent. That’s observed Census data, not a retrospective estimate. Demand was growing rapidly, but from a base small enough to hide behind the rounding error of total retail.

At the same time, telecommunications firms financed networks around forecasts that treated bandwidth demand like it had discovered nuclear fission. Global Crossing built a fiber network connecting more than 200 cities in 27 countries and entered Chapter 11 in January 2002 with $22.4 billion of assets and $12.4 billion of debt. The network worked. The company couldn’t generate enough cash to service the debt as bandwidth prices fell and expected customers failed to arrive. Equity holders were expected to be wiped out while service continued.

WorldCom filed for bankruptcy in July 2002 with more than $30 billion of debt after an accounting fraud accelerated the collapse. Fraud makes WorldCom an imperfect pure overinvestment example, but the financing lesson is still useful: a functioning network with 20 million customers continued operating while the capital structure failed around it.

The old boom therefore separated three outcomes that investors had bundled together:

The Dot-Com Boom Was Right About Demand and Very Wrong About Timing

Fiber had one enormous advantage over AI accelerators: glass in the ground didn’t become obsolete every time Nvidia announced a new architecture. Electronics on either end improved, but much of the route remained useful. A B200 cluster can be physically healthy and economically demoted by Rubin, custom silicon or a model requiring far less compute per task. That makes dot-AI’s duration problem potentially harsher:

The comparison shouldn’t be pushed too far. Today’s largest buyers are profitable hyperscalers rather than venture-backed websites with a Super Bowl commercial. AI labs already generate billions of dollars of revenue, and enterprise adoption is much further along than retail e-commerce was in 2000. The analogy isn’t “history repeats.” It’s that a transformational network can survive while the owners who financed the wrong capacity at the wrong price get erased.

People talk about “the AI bubble” as though it’s one wager with a yes-or-no answer. It’s really a bundle of connected assumptions.

Picture the AI economy as a tower of blocks:

The tower can survive if every layer is merely “pretty good.” But valuations near the top aren’t priced for pretty good. They depend on several layers expanding together for years.

Enterprise spending must keep rising. That spending must produce valuable work. Model providers must retain enough of the value as revenue. The revenue must outrun inference and training costs. Cloud providers must convert laboratory commitments into paid utilization. Hardware must remain useful long enough to recover its construction and financing cost.

The bubble thesis isn’t that every block falls to zero. It’s that the market has priced the top of the tower as though the blocks are largely independent. They aren’t.

If cheap models reduce frontier pricing, lab revenue weakens. If lab revenue weakens, enormous cloud commitments become less dependable. If cloud utilization disappoints, data-center returns fall. If data-center projects get delayed, chip, memory and power forecasts change.

One technical improvement can be socially wonderful while moving financial value downward through the entire stack.

That’s why AI contains such a strange contradiction:

The faster intelligence becomes abundant, the harder it may be for companies valued on its permanent scarcity to earn their expected returns.

Private labs don’t publish the audited segment disclosures we’d demand from a public company, surveys measure intentions rather than cash, and a management forecast isn’t revenue merely because somebody put it in a slide deck. So before we start throwing around enormous numbers, here’s the confidence ladder we’ll use. Higher rows deserve more weight than lower ones, and we’re not going to sneak a lower row into a higher category because the headline sounds better.

Not All AI Numbers Deserve the Same Amount of Trust

Public-company AI revenue is often undisclosed, so company-wide revenue can’t quietly become “AI revenue.” Private-company run rates can’t become recognized annual revenue. Announced compute commitments can’t become cash already spent. Every major figure below is identified in context by period, definition and evidence type. Where the available evidence can’t answer the question, the correct label is “unknown,” which is less exciting than a made-up number and far more useful.

Let’s also be fair to the other side. If this thesis can’t be proven wrong, it isn’t analysis. It’s just a personality trait.

The bearish thesis would weaken considerably if frontier labs demonstrate all of the following over several years:

  • Gross margins improve while agentic usage and context lengths expand.

  • Capital expenditures and compute commitments grow more slowly than recognized revenue.

  • Enterprise customers produce measurable revenue growth, not only internal efficiency.

  • Frontier pricing remains durable despite open and distilled competition.

  • Current hardware retains economically useful lives close to its financing schedule.

  • Labor productivity produces market expansion large enough to prevent broad headcount compression.

That world is possible. AI may make so many previously unaffordable services commercially viable that inference efficiency and customer revenue rise together. Personalized education, affordable legal assistance, custom software, drug discovery and autonomous research could create markets far larger than the clerical and software budgets AI initially consumes.

But notice how much has to go right.

The technology doesn’t merely need to improve. It needs to improve along a narrow financial path:

Actually, even that notation is confusing, which is exactly the problem. The labs want capability to rise rapidly, their own costs to fall rapidly, customer prices to fall slowly enough to protect margin, and total customer spending to rise despite those falling unit prices. A cleaner picture is four arrows:

If that combination holds, the industry can grow into the infrastructure.

If customer prices fall faster than production costs, margins get squeezed.

If production costs remain high, adoption becomes expensive.

If both fall rapidly, intelligence flourishes but infrastructure scarcity disappears.

So the whole game is figuring out how wide that path really is.

Atlas is enormous, but its economics are built one accepted task at a time. The enterprise doesn’t ultimately want accelerator-hours, HBM bandwidth or tokens. It wants a repository migrated, a support case resolved, a contract reviewed or an analysis it can actually use. So before pricing the factory, we need to understand what happens inside Atlas between a request entering the system and useful work leaving it.

The internet was expensive to build, but once the infrastructure existed, information became nearly free to distribute. Serving another webpage, delivering another search result or sending another email eventually became extremely cheap. AI doesn’t work quite like that.

Every time Claude writes code or ChatGPT reasons through a problem, a physical system has to perform the work. Model weights must be stored and moved through memory. The system maintains a KV cache containing information about the active conversation. GPUs, networking equipment, cooling systems and data centers consume electricity while generating every token.

The internet made information nearly free to distribute. AI makes intelligence possible to manufacture, but it hasn’t made intelligence free to manufacture.

A traditional SaaS company can build its product once and serve millions of additional customers at a very low marginal cost. An AI company can train a model once, but it still has to manufacture every response individually.

Frontier AI is therefore closer to an industrial process hiding behind a chat box than ordinary software.

It helps to stop thinking about “a GPU” as one magical black box.

Imagine walking through an intelligence factory with four connected rooms.

The storage room determines whether the model and active conversations fit.

The loading dock determines how quickly weight and cache data reach the arithmetic units.

The machine floor determines how much mathematical work can be performed each second.

The shipping desk determines whether many customer requests can be combined efficiently without making everybody wait forever.

AI hardware marketing tends to point at the machine floor because the numbers are enormous. A modern accelerator can advertise multiple petaflops of low-precision matrix performance. That sounds like the whole story until the machine floor runs out of material because the loading dock can’t deliver weights quickly enough.

This gives us four constraints that people constantly mix together:

The Bottleneck Moves Depending on the Workload

A bigger memory doesn’t necessarily move data faster. Faster memory doesn’t help if the model doesn’t fit. More FLOPs don’t help if the arithmetic units are waiting. Larger batches improve utilization but consume additional KV cache and can hurt latency.

There is no single “make AI cheaper” knob.

There is a panel of knobs connected to one another with string, and turning one can make a different constraint worse.

Suppose a library contains one million books. That tells us its capacity.

Now suppose the librarian can deliver only one book per hour. The library contains plenty of knowledge, but the reader will spend most of the day staring at an empty desk. Capacity answers:

Bandwidth answers:

The distinction is visible in current hardware. AMD’s MI350X contains 288 GB of HBM3E and advertises up to 8 TB/s of peak memory bandwidth. An eight-accelerator platform therefore contains about 2.3 TB of HBM, although each accelerator retains its own local 8 TB/s pathway and the devices still require a high-speed interconnect to cooperate.

Micron’s HBM3E provides more than 1.2 TB/s per memory stack. Its HBM4 generation increases that to more than 2.8 TB/s per stack, using a bus twice as wide and claiming more than a 20% improvement in power efficiency over HBM3E.

Those improvements are extraordinary. They’re also evidence that memory movement remains important enough to justify redesigning and stacking some of the most complicated silicon on Earth.

If raw arithmetic were the only thing that mattered, the industry wouldn’t be spending this much money teaching memory to run faster.

Now we can build a deliberately simplified model.

Imagine a dense 100-billion-parameter model stored at four bits per parameter. Its weights occupy roughly 50 GB:

For a low-batch autoregressive decode, generating each new token may require streaming a large portion of those weights through memory. In that regime, memory bandwidth frequently constrains throughput. It isn’t a universal law. Increase the batch, lengthen the context, distribute sparse experts, add speculative decoding or push enough matrix work through the accelerator and the bottleneck can migrate into compute, KV-cache movement, interconnect, scheduling, communication or tail latency. You aren’t managing one permanent bottleneck. You’re managing a bottleneck that moves every time the workload changes.

If the system sustained 2 TB/s of usable bandwidth and had to move approximately 50 GB per decoding step, the crude low-batch upper bound would be:

That isn’t a performance guarantee. Real systems have imperfect bandwidth utilization, communication, cache behavior, attention work, synchronization and software overhead. MoE changes the amount of weight data activated per token. Batching allows one weight read to contribute to several sequences.

But the toy equation gives us the right intuition:

Now quantize the same model from four bits to two. Weight storage falls from about 50 GB to 25 GB. If quality survives and every other assumption remains unchanged, the bandwidth-bound ceiling can approximately double.

That’s why quantization can improve both economics and speed. It doesn’t merely squeeze a model onto cheaper hardware. It reduces the amount of material crossing the loading dock.

But “if quality survives” is carrying a refrigerator on its back. Aggressive quantization can damage rare knowledge, reasoning stability, attention behavior or output quality. A compressed model that finishes the wrong task twice as quickly hasn’t halved the cost of useful work.

Now place four customers in the factory. Without batching, the system reads the model weights separately for each customer:

With batching, one coordinated weight pass can help advance several sequences:

The weights are being reused across more paid work. Arithmetic intensity rises, memory bandwidth is amortized and total throughput improves.

This is the core economic magic of shared inference infrastructure.

It’s also why accelerator utilization matters as much as the accelerator’s sticker price. An expensive GPU serving a steady, batchable workload can produce cheaper tokens than a cheaper machine sitting mostly idle. We can write the intuition as:

The denominator is where inference businesses live or die. But batching has enemies.

Requests arrive at different times. Prompts have different lengths. Some customers want an immediate answer while others tolerate delay. One sequence may finish while another generates 20,000 tokens. Long contexts occupy more cache. Agentic workloads pause for tools and then return. Safety checks and routing add more stages.

The scheduler can wait to assemble a more efficient batch, but the customer experiences additional latency. It can serve requests immediately, but the hardware does less useful work per weight read.

That’s not a temporary engineering embarrassment. It’s a queueing problem built into a shared service with unpredictable demand.

When a customer submits a large prompt, the system can process many prompt tokens in parallel. This is prefill. It tends to create large matrix operations that use the GPU’s arithmetic machinery efficiently.

Once the model begins responding, it generates tokens autoregressively. Token 501 depends on token 500. Token 500 depended on token 499. This sequential phase is decode.

NVIDIA describes prefill as highly parallel and decode as autoregressive. Its own optimization research says LLM decode is typically memory-bandwidth-bound, with long-context decode spending substantial time moving KV-cache data rather than performing arithmetic. This creates two customer-facing latency metrics:

  • Time to first token: How long the prompt processing and queueing take before the response begins.

  • Inter-token latency: How quickly subsequent tokens appear once generation starts.

A system can be good at one and bad at the other. A huge prompt may delay the first token while the eventual response streams smoothly. A small prompt can begin immediately but crawl through a bandwidth-limited decode.

That distinction becomes financially important for agents. Human chat tolerates a visible stream of 30 or 50 tokens per second. A background coding agent may care less about pretty streaming and more about total time to complete ten tool calls, three searches and a code-editing loop.

The product metric changes from “does this feel fast?” to “how much useful work does this hardware finish before the billing hour ends?”

We usually discuss AI chips in terms of computation. Nvidia announces more FLOPs, models train faster, Jensen signs another leather jacket, and everyone gets excited.

But during token-by-token generation, the GPU is often limited less by arithmetic than by how quickly it can move model weights and cached information through memory. Inference has two broad stages:

  • Prefill: The model processes the prompt. This tends to be compute-heavy.

  • Decode: The model generates tokens sequentially. This is frequently constrained by memory bandwidth.

During decode, the model repeatedly reads enormous amounts of weight data to produce one token after another. Its arithmetic units can be ready to work while waiting for memory to deliver the next pile of numbers.

Adding theoretical compute therefore doesn’t automatically generate tokens faster. You can keep installing larger engines, but if the freeway feeding them is jammed, congratulations, you’ve built a very expensive parking lot. Model size creates the first memory problem:

Before including KV cache and runtime overhead, the rough storage requirements are:

How Quantization Shrinks the Model-Weight Memory Bill

Quantization helps tremendously, but it doesn’t make the physical requirements disappear. Neither does mixture of experts. An MoE model activates only part of its network for each token, reducing arithmetic, but the expert collection still has to be stored somewhere accessible. “Somewhere” doesn’t mean one accelerator’s HBM. Experts can be distributed across accelerators, nodes and memory domains through expert parallelism. The saved arithmetic is partly exchanged for routing, communication, load-balancing and availability problems.

Imagine twenty checkout lanes, only four of which open for any customer. The store doesn’t pay four lanes’ worth of rent. It maintains all twenty, directs each shopper correctly and copes when everybody suddenly chooses the same specialist lane. MoE has the same shape:

Expert parallelism can spread capacity across the cluster, but selected activations then cross an interconnect. A hot expert can create a queue while other experts sit idle. MoE therefore separates total representational capacity from active computation without making storage or the network disappear.

Then there’s KV cache. It prevents the model from recomputing the entire conversation whenever it generates another token, but it grows with context length, concurrent requests and batch size. The relationship is deeply annoying:

The model becomes more useful while the infrastructure becomes less efficient.

Model weights are mostly a fixed cost for a loaded model. Whether one customer or thirty customers are using the server, the weights have to be resident somewhere.

KV cache behaves differently. Each active sequence brings its own pile of notes.

Return to the library analogy. The model weights are the books on the shelves. The KV cache is the notebook the reader builds while working through a particular problem. A longer conversation creates a thicker notebook. A second customer needs a second notebook. Thirty-two customers need thirty-two notebooks.

The exact cache size depends on the architecture, precision, number of layers, attention heads and tokens retained. But the directional relationship is simple:

This means context length and concurrency compete for the same memory.

Suppose the weights and runtime consume 70% of available HBM. The remaining 30% is the space available for active conversations. Doubling the average KV footprint doesn’t merely add a modest expense. It can approximately halve the number of conversations that fit at once.

The feature the sales team advertises as “more context” can therefore appear in the infrastructure layer as “fewer simultaneous customers per accelerator.”

NVIDIA’s own work on four-bit KV-cache quantization illustrates the importance of this constraint. Moving from an FP8 cache to NVFP4 can reduce KV-cache memory by roughly 50%, enabling larger contexts, larger batches or more concurrent users. The same optimization also reduces bandwidth pressure during decode.

That’s a major improvement. But notice what the improvement gets spent on. The industry rarely banks all of it as lower cost. It often reinvests efficiency into longer context, more reasoning and more concurrent agent state.

This is AI’s version of Jevons paradox. Make inference cheaper, and products discover new ways to consume it.

A chatbot conversation is relatively simple. The user asks a question, the model responds, and eventually the session ends. An agent accumulates state.

It may need the original instruction, repository map, tool definitions, retrieved documents, prior edits, command outputs, failed attempts, test results and a plan for what remains. Some systems summarize or evict older context, but that creates another tradeoff: compression saves memory while risking the loss of something the agent later needs. Picture a contractor working in a room.

At first, the desk contains one blueprint. After an hour it contains building codes, invoices, photographs, revised drawings and notes about three mistakes already fixed. Clearing the desk makes the contractor faster until somebody throws away the one page explaining where the gas line runs.

Agent memory management is the same problem in digital form:

This is why cost per million tokens can become actively misleading. A cheap token used to repeat work forgotten during context compression isn’t actually cheap. A costly token that prevents an agent from corrupting a production database may be a bargain. The useful unit remains the completed task.

Now we can see the pattern that keeps showing up everywhere else.

Suppose a new generation of hardware doubles effective memory capacity and bandwidth. A normal software comparison might hold the workload constant and celebrate a large cost reduction. The AI industry often does this instead:

The customer receives a more capable system. But the lab’s cost per completed task may not fall by anything close to the hardware improvement because the definition of a task expands to consume the new headroom.

This doesn’t mean efficiency work is pointless. Without it, the new capabilities might be economically impossible.

It means investors shouldn’t automatically translate a two-times hardware improvement into a two-times gross-margin improvement.

Some portion becomes lower cost. Some becomes higher quality. Some becomes longer context. Some becomes additional usage. And some disappears into the operational complexity of serving irregular agent workloads.

The allocation between those buckets is one of the most important unknowns in frontier-model economics.

There’s a second memory problem hiding underneath the hardware problem.

An LLM’s parameters are themselves a kind of compressed representational memory. The model’s knowledge of language, programming, history, science, human behavior and millions of obscure patterns is distributed across billions or trillions of numerical weights.

Those parameters don’t store facts like neat rows in a database. But the basic constraint remains:

A smaller model can be outstanding within a narrower distribution. It can also be trained far more intelligently than earlier models of the same size. Parameter count isn’t a clean intelligence dial. Architecture, data quality, training compute, post-training and the allocation of capacity all matter. A well-trained smaller model can humiliate a wasteful larger one.

The stronger claim is that broad, long-tail capability requires representational capacity somewhere in the complete system. Some of it may live in parameters. Some can live in retrieval, persistent memory, tools, search, sparse experts, verifiers, specialized modules or interaction with an external environment. Current frontier systems don’t prove that all of it must permanently reside in dense parameters. They do show that removing capacity without replacing its function tends to lose rare knowledge, subtle behavior or reliability somewhere in the distribution.

That distinction matters financially. If capability migrates from one enormous resident model into cheap specialists plus retrieval and tools, the demand for premium centralized inference can fall even while the complete system becomes more capable.

That’s why the leading labs continue pushing the frontier even while releasing smaller tiers. Their business, branding and infrastructure plans are still organized around building the smartest broadly capable model, not merely the smallest sufficient model.

The industry loves talking about falling prices per million tokens. But price per token is becoming a less useful measure of economic value.

An older model might receive one prompt and generate 500 tokens. A modern agent might search the web, inspect files, call tools, maintain a huge context, generate internal reasoning, retry failed actions and ask other agents to verify the result.

The token price can fall while the number of tokens and model calls required to complete one useful task explodes.

Suppose token prices fall by 90%, but an agent uses 25 times more tokens and model calls:

The nominal price fell by 90%. The total model cost increased by 150%. The metric that matters isn’t:

The metric we actually care about is:

Now give the denominator teeth:

That’s what enterprises ultimately care about, and it’s what determines whether agents produce gross margin or simply convert payroll into an enormous cloud bill. Nobody buys tokens for spiritual fulfillment. Customers buy accepted work. A lab can lower its advertised token price while retries, supervision and rejected outputs make the completed task more expensive.

This is the first number Atlas has to beat. More accelerator-hours, tokens or benchmark points matter only insofar as they reduce the cost or increase the value of accepted work. With that unit established, we can zoom back out and assemble the physical system required to produce it.

When analysts describe the AI buildout, they usually begin with capital expenditures. That makes sense because capex is visible in cash-flow statements and company guidance.

Epoch AI estimates that combined capital spending by Alphabet, Amazon, Meta, Microsoft and Oracle grew at an average annual rate of 72% from the second quarter of 2023, shortly after GPT-4’s release. It approached half a trillion dollars during 2025. The calculation combines cash purchases of property and equipment with newly obtained finance-lease assets.

Meta’s filings provide a clean example. It spent $69.69 billion on property and equipment during 2025 and initially expected approximately $115 billion to $135 billion of capital expenditures during 2026 to support AI and its core business. Its first-quarter filing subsequently placed the expected range at $125 billion to $145 billion.

But capex is only one way to acquire productive capacity.

A company can buy a data center. It can lease one. It can sign a capacity agreement with a third party. It can commit to buying chips or electricity in the future. It can finance equipment through a special-purpose vehicle. It can ask a developer to borrow the money, build the facility and recover the cost through a long-term lease.

Economically, these arrangements can point to the same building full of accelerators.

Accounting makes them appear at different times and in different places.

Reuters calculated that Microsoft, Meta, Oracle, Amazon and Alphabet had accumulated about $1.09 trillion in future lease commitments, much of it related to AI data centers. That compared with roughly $285 billion of lease liabilities already recognized on their balance sheets. The gap exists partly because accounting recognition can wait until a facility is operational and available for use.

This doesn’t mean somebody hid a trillion-dollar bill under the couch. The commitments are disclosed, may be conditional and stretch across many years.

It means a cash capex chart can understate how much future infrastructure the industry has already promised to support.

Now draw two timelines on top of each other.

The thing generating demand changes every few months. The thing financed to serve it may remain under contract for fifteen or twenty years.

Long-lived infrastructure isn’t automatically reckless. Railroads, power plants and semiconductor fabs also require long commitments. The danger appears when the asset’s economic value depends on one fast-changing customer, architecture or hardware generation.

Oracle demonstrates the concentration problem. Reuters reported approximately $260 billion of pending data-center lease commitments, nearly seven times its recognized lease liabilities, with terms extending roughly 15 to 19 years. Oracle also carried $129.5 billion of debt and faced scrutiny over customer concentration around OpenAI.

We can express the risk as two useful lives:

Financing works comfortably when:

It becomes painful when:

A building can remain physically functional while becoming economically stranded. The power still works. The cooling still works. The racks still blink. But the chips may consume too much power, support the wrong numerical formats, lack sufficient memory or produce tokens at a cost customers no longer accept.

That’s the AI version of owning a perfectly functional DVD factory in 2012.

Cloud backlog sounds reassuring because it represents contracted future business. It’s considerably better than a founder pointing at a total-addressable-market slide and making spaceship noises. But backlog still contains assumptions.

A long-term agreement may include conditions, ramp schedules, cancellation rights, minimums, construction dependencies and customer-credit risk. Revenue arrives only when capacity becomes available and the customer consumes or pays for it under the contract. Think of three increasingly solid layers:

AI coverage frequently jumps from the first or second layer to the fourth.

That’s how a spectacular headline about future compute demand becomes treated as though an enterprise customer has already paid for profitable inference.

The distinction matters because infrastructure construction begins before the final layer is known. Developers borrow against expected rent. Utilities build generation and transmission around expected load. Memory suppliers reserve production. Cloud providers order hardware to meet expected utilization.

The financial system acts on the promise before the economics complete the journey.

A gigawatt is one billion watts. A one-gigawatt data-center campus running continuously would consume 8.76 terawatt-hours in a year before adjusting for downtime:

That mental model matters because a gigawatt isn’t “a large electric bill.” It’s power-station territory. The campus needs generation or contracts, transmission, substations, transformers, switchgear, backup systems and cooling that can remove essentially the same energy as heat. Nearly every electrical watt entering computing equipment ends up as heat somewhere. The intelligence factory is also a very organized space heater.

Berkeley Lab’s June 2026 bottom-up update estimated that data centers could consume 11.8 percent of U.S. electricity in 2030, with a scenario range of 9.5 to 15.3 percent. Those are modeled scenarios based partly on planned equipment shipments, device energy and cooling performance, not promises about realized construction. The lab’s earlier estimate put 2023 data-center consumption at about 176 TWh, or 4.4 percent of U.S. electricity.

The arithmetic from electricity to useful AI task has several multipliers:

PUE, or power usage effectiveness, is total facility energy divided by IT-equipment energy. A theoretical PUE of 1 means every watt reaches computing. If PUE is 1.2, facilities consume 20 percent on top of IT energy. Now add retries. If an agent requires 1.5 attempts on average and succeeds 80 percent of the time, the energy per successful task is multiplied by (1.5/0.8=1.875) before considering the PUE. A faster chip can be swallowed by longer reasoning, more attempts or worse utilization.

Construction creates a second clock. Chips may improve every year. High-voltage transmission, gas pipelines and turbines don’t appear because a product manager changed the roadmap. Projects wait for studies, permits, interconnection, equipment and skilled trades. EIA’s April 2026 long-run scenarios said data-center load had become a dominant driver of U.S. electricity growth and projected installed generating capacity rising 50 to 90 percent by 2050 across cases, with natural gas, solar and wind accounting for most additions. EIA explicitly describes these as alternative scenarios, not predictions.

During delay, the revenue clock stops but the finance clock doesn’t:

Suppose a $5 billion project is half funded during a one-year delay at a 7 percent cost of capital. Financing carry alone is roughly $175 million before labor escalation, storage, redesign or lost revenue. If accelerators are delivered early, their competitive lives can begin decaying while they’re waiting for power. That is the most Silicon Valley form of tragedy imaginable: obsolete equipment still in the original packaging.

Grid queues also contain speculative or duplicate requests, so announced gigawatts can’t be treated as certain demand. In August 2026, Texas paused approvals for major new data-center grid connections pending an audit. Reuters reported that roughly 90 percent of 474 GW of proposed demand under review was associated with data centers, more than five times the state’s peak load. The number is a queue, not a forecast of facilities that will all be built. Its absurd size is precisely why grid operators require deposits and feasibility tests.

If gas turbines are scarce, turbine manufacturers gain pricing power. If transformers have multi-year lead times, electrical-equipment suppliers can raise prices and fill backlogs. If HBM is scarce, Micron and SK hynix enjoy better mix and margins. Investors looking only at those suppliers can correctly observe a boom. The lab sees the mirror image:

The scarce component initially validates the AI story because orders and margins soar. Yet every extra dollar paid for memory, power equipment or construction raises the amount of model revenue required to earn the project’s target return. Shortage beneficiaries can flourish while the end customer’s unit economics deteriorate. Gold-rush shovel sellers don’t need every miner to find gold. They need miners to keep financing the search.

HBM is an especially clean example. More bandwidth can raise throughput and lower task cost, but advanced stacks require specialized fabrication and packaging. Micron says its HBM3E delivers more than 1.2 TB per second per stack and its HBM4 more than 2.8 TB per second, with over 20 percent better power efficiency in the company’s comparison. Those are vendor technical claims, not independent workload benchmarks.

Cooling and water complete the loop. Berkeley Lab notes that nearly all data-center electricity becomes heat and estimates U.S. data-center water consumption could reach 0.14 to 0.28 billion cubic meters by 2028 in the cited scenario work. Water use varies dramatically with climate, cooling technology and whether measurement counts onsite use or electricity-generation water.

None of this proves a permanent physical ceiling. Higher-voltage power delivery, direct-to-chip liquid cooling, onsite generation, batteries, geographic load shifting and more efficient models can all help. The bear case is about price and time, not impossibility. The buildout can arrive late and over budget, then enter service into a market whose task prices changed while concrete was curing.

Picture a brand-new cluster after commissioning. The servers pass diagnostics, the fabric is stable, the cooling loops don’t leak and a nervous operations engineer has finally stopped sleeping beside the pager. Nothing is broken. Now suppose the cluster can’t earn enough to repay its financing and replace the hardware before customers migrate to something better. That’s economic stranding: the machine still computes, but the cash flow no longer supports what investors paid for it.

We need a real physical cluster for the example, not fuzzy “accelerator-equivalents.” So let’s build one from public specifications and rental prices. These are hypothetical assumptions, not a forecast for Nvidia, CoreWeave or any specific project.

Start with 1,024 eight-GPU DGX B200-class systems, or 8,192 B200 GPUs. Nvidia specifies 1,440 GB of aggregate HBM3e, 64 TB/s of aggregate memory bandwidth, 10 rack units and approximately 14.3 kW of maximum power per DGX B200. CoreWeave listed an eight-GPU HGX B200 instance at $68.80 an hour in North America in July 2026, equivalent to $8.60 per GPU-hour before discounts. Those are observed vendor specifications and a posted on-demand price, not the cost or realized rate of this hypothetical project.

Assume each complete eight-GPU server costs the project $450,000 including CPUs, host memory, local storage, support and system integration. That’s a hypothetical procurement assumption. Reuters reported in March 2024 that Nvidia expected individual B200 accelerators to cost roughly $30,000 to $40,000, but a working server costs much more than eight loose GPUs because customers also buy CPUs, memory, NVSwitch, power supplies, storage, chassis and vendor margin.

What It Costs to Build Atlas

The compute and fabric account for $575.8 million. Storage and software add $84.2 million. The long-lived building, power and cooling plant account for $250 million, and contingency completes the billion. Different projects will allocate these costs differently. The useful point is that we can now see which assets become obsolete quickly and which can be reused.

Finance 60 percent of the project with seven-year amortizing debt at 7 percent and 40 percent with equity. Annual debt service is approximately $111.3 million. That payment is calculated from the standard annuity formula, not guessed:

The cluster contains 8,192 GPUs and each calendar year contains 8,760 hours:

“Available” doesn’t mean billable. Atlas enters its steady-state year with contracts and expected on-demand traffic covering 78 percent of gross capacity. Planned and unplanned downtime remove 3 percent of gross hours, customer credits and failed billable attempts remove another 1 percent, and ramp plus workload fragmentation remove 4 percent. The bridge is:

These percentages are hypothetical steady-state underwriting assumptions, not observed CoreWeave or hyperscaler utilization. A real contract would also specify reservation deposits, start dates, service credits, minimum consumption, termination rights and whether unused take-or-pay capacity can be resold. At 70 percent billable utilization:

Assume the project realizes $6.50 per GPU-hour after mixing reserved contracts, on-demand bursts and discounts. That’s below CoreWeave’s posted July 2026 B200 on-demand rate of $8.60 and remains a hypothetical realized price. Annual revenue becomes:

Now power it. The 1,024 systems draw a documented maximum of roughly 14.64 MW before the scale-out network and storage. Assume average IT load of 14 MW across the whole compute and network plant, then multiply by a base-case PUE of 1.20. Facility load averages 16.8 MW:

That’s 147.2 GWh per year. At a blended energy rate of seven cents per kWh, raw energy costs $10.3 million. Add $4.7 million for demand charges, backup testing, water and other utility costs, producing a $15 million annual power-and-water bill. The relatively modest number is a useful correction to loose commentary: for a high-priced B200 rental business, hardware depreciation and utilization can matter far more than electricity. Power becomes existential when it delays the project, restricts deployment or combines with collapsing rental prices.

Atlas in the Base Case

The depreciation schedule assumes four years for the $575.8 million of servers and fabric, five years for $45 million of storage, three years for $39.2 million of software and spares, and ten years for the $250 million building, electrical and cooling plant. That produces about $191.0 million of annual economic depreciation. The remaining $90 million is contingency and construction-financing carry, which is part of invested capital but not itself a separately depreciating productive asset. This is an economic-depreciation estimate, not a claim about any company’s GAAP policy.

The base case covers debt and reports a small operating profit, but “cash after debt service” still isn’t the equity return. Atlas needs maintenance capital, may owe cash tax and eventually has to replace the equipment that produces the revenue. A proper waterfall keeps those checks separate.

Assume $12 million of annual maintenance capex beyond the service contracts already included in operating cost. In the first year, the $111.3 million debt payment consists of about $42 million of cash interest and $69.3 million of principal. Assume $5 million of cash taxes after available deductions. Atlas then reaches $88.2 million of distributable cash before funding future replacement:

Now comes the part that cheerful project decks tend to leave in the appendix. The servers, fabric, storage, software and spares consume about $166 million of annual economic life. If Atlas expects replacement assets to maintain the same 60/40 debt-to-equity financing, equity must reserve roughly 40 percent of that amount, or $66.4 million a year, while continuing to preserve debt capacity. That leaves approximately $21.8 million of normalized distributable equity cash, a 5.5 percent annual cash yield on the original $400 million equity contribution before considering growth, terminal shell value or changes in replacement cost. It is not a complete equity return. Principal amortization increases the owner’s residual claim, economic depreciation reduces it, and whatever the shell and accelerators are worth at exit can move the final internal rate of return in either direction. Calling 5.5 percent “the return” would quietly treat a wasting hardware asset as though it were a bond that hands the original principal back untouched.

This reserve is an underwriting convention, not GAAP. If lenders refuse to finance replacement equipment, the project needs the full $166 million reserve and distributable cash turns negative. If replacement systems cost less per useful task, the required reserve falls. If the reusable shell, substation and cooling plant retain substantial terminal value, equity gets protection Atlas’s accelerators don’t provide. The central point is that investors supplied $400 million and can’t treat the entire $105.2 million of cash after debt service as profit while the productive core wears out underneath them.

At 55 percent utilization, billable hours fall to 39.47 million and revenue falls to $256.6 million. Assume variable power and support lower cash operating cost from $110 million to $95 million. Site EBITDA becomes $161.6 million, cash after debt service falls to $50.3 million and operating profit after economic depreciation becomes negative $29.4 million.

Utilization is the denominator across which the fixed factory spreads. A 21 percent decline in billable utilization, from 70 to 55 percent, erases more than half the cash remaining after debt service.

Return utilization to 70 percent and cut the realized rate from $6.50 to $5.50. Revenue becomes $276.3 million, cash operating profit becomes $166.3 million, cash after debt service becomes $55.0 million and operating profit after depreciation becomes negative $24.7 million. The cluster is busy, liquid-cooled and economically underwater.

This is the central commoditization risk. Price doesn’t need to fall to zero. It only needs to fall below the rate assumed when the asset and debt were sized.

At ten cents per kWh instead of seven cents, raw energy costs rise by about $4.4 million. If PUE also worsens from 1.20 to 1.35 because the site uses less efficient cooling, annual facility energy rises from 147.2 to roughly 165.6 GWh. At ten cents, the combined raw-energy difference from the base case is about $6.3 million. That hurts, but it doesn’t destroy the base case by itself.

This result is important because it keeps the argument honest. Electricity is essential, and delays can be catastrophic, but the depreciation of expensive computing equipment dominates this particular cluster’s annual cost structure. A model claiming otherwise needs to show its wattage, PUE and power price.

Suppose one tenant represents 40 percent of billable hours and receives a 20 percent discount at renewal. The weighted realized rate falls 8 percent:

Revenue falls by $26.1 million. Cash after debt service falls from $105.2 million to $79.1 million, and operating profit turns slightly negative. The tenant asks for relief precisely when excess capacity weakens the owner’s alternatives.

If compute and fabric fall from a four-year economic life to three years, annual depreciation on the $575.8 million layer rises from $144.0 million to $191.9 million, a $48.0 million increase. At a two-year life it becomes $287.9 million, a $143.9 million increase from the base schedule.

The old GPUs still run. The problem is that a new accelerator completes the same useful task so cheaply that customers won’t pay the old rental rate. Obsolescence attacks both variables:

Assume the project has drawn an average $600 million of capital during a twelve-month delay. At a 7 percent financing cost, carry adds roughly $42 million before change orders, storage, labor escalation or lost revenue. If GPU systems arrive before the site has power, part of their competitive life decays inside crates. The cluster then enters service one product cycle closer to replacement.

If cheaper models move 30 percent of this cluster’s workload elsewhere, utilization falls from 70 to 49 percent. At $6.50 per hour, revenue falls to about $228.6 million. Assuming $90 million of cash operating cost, cash after debt service falls to $27.3 million and the operating loss after depreciation reaches roughly $52.4 million.

Total AI usage might still be exploding. The displaced tasks could be running on older GPUs, custom silicon or local hardware. The project loses because demand for this generation in this location underperforms, not because society stopped using AI.

The bear case shouldn’t assume zero residual value. Suppose the compute and fabric retain 15 percent of their installed value after four years because they can serve batch inference, embeddings, scientific workloads or distilled models. Economic depreciation on that layer falls by about $21.6 million per year:

That lifts base-case operating profit from $25.5 million to about $47.1 million. A secondary market materially protects the owner. It doesn’t solve a simultaneous utilization and pricing collapse, but it’s why “old GPU” and “worthless GPU” can’t be used interchangeably.

What Happens When One Atlas Assumption Breaks

These are our own sensitivities, not analyst estimates. Taxes, reservation deposits, curtailment, financing covenants, failure rates and workload-specific performance would make a real underwriting messier. The geometry is the point. Small misses in price, utilization or useful life consume the narrow layer between cash generated today and capital that must be replaced tomorrow.

Now, Amazon and Microsoft aren’t underwriting one isolated project in a vacuum. They can pool workloads, move jobs between regions, reuse buildings, build custom silicon, negotiate power and fund construction with diversified operating cash flow. That’s a real advantage, and it makes a hyperscaler much safer than a leveraged single-tenant developer. The stranded-asset thesis weakens if we see high utilization across several hardware generations, stable realized pricing per useful task, durable secondary values and returns on invested capital recovering as AI revenue scales.

We can now stop admiring Atlas as a pile of equipment and ask what the pile has to earn. In the base case, 8,192 B200 accelerators produce 50.23 million billable GPU-hours at 70 percent utilization and a realized price of $6.50 per hour. That creates $326.5 million of annual revenue. After $110 million of cash operating cost, Atlas produces $216.5 million of site EBITDA. After $111.3 million of debt service, $12 million of maintenance capex and $5 million of cash taxes, it has $88.2 million available before replacing the equipment that actually earns the money.

That last phrase matters. If Atlas reserves only the equity-funded portion of its estimated equipment replacement need, normalized distributable cash falls to roughly $21.8 million on the original $400 million equity contribution, or about a 5.5 percent annual cash distribution. The owner may also build equity as debt principal amortizes, lose equity as the equipment economically depreciates and recover value from the shell or secondary hardware market at exit. A proper investment return needs all four. Even so, the annual cash available after a credible replacement reserve is thin enough that a modest utilization miss, price decline or shorter hardware life can make the project unattractive without turning off a single rack.

This is the equation the rest of the article will keep changing. The numerator is not zero. AI creates obvious value. The bearish question is whether enough of that value reaches Atlas after the application, laboratory, cloud provider, customer and financing structure take their pieces. The bullish answer is equally concrete: utilization can rise, task cost can fall, old hardware can find secondary work and induced demand can fill every rack. Great. Now we need to find the customer dollar that makes those things happen.

Now let’s give Atlas a customer. Suppose an enterprise pays $100 for an accepted AI task. The application keeps its workflow margin, the lab charges for inference, the cloud provider recovers capacity cost, the data-center owner collects rent, suppliers recover hardware and electricity cost, and the lender receives interest and principal. This isn’t literal invoice accounting for every workload. It’s a map of how many economic claims are stacked on the same $100 of customer value.

If the customer receives less than $100 of value, somebody must accept a lower margin, subsidize the task or finance the gap. “AI stocks are expensive” is too mushy to identify who. The application company selling an agent, the laboratory training the model, the cloud provider hosting it, the chip designer, the memory manufacturer, the data-center landlord and the electric utility don’t own the same economics. A capability breakthrough that crushes one layer can enrich another. We need to draw the tower.

We also need to distinguish a dollar moving upward through the tower from a new dollar entering it. A cloud provider investing in a lab, a lab signing a compute commitment and an enterprise redirecting payroll into model spending can all create legitimate revenue for somebody in the chain. But they are transfers among existing pools of capital and expense. The tower becomes durably larger only when end customers spend more, new customers enter or lower prices create enough additional volume to expand total revenue. Every layer can report growth for a while before that distinction becomes visible. Eventually, it becomes the only distinction that matters.

Each layer is making a different promise to its capital providers.

Every Layer of the AI Capital Tower Is Making a Different Bet

Start at the top. An AI application can have wonderful economics if it owns a workflow, proprietary data, distribution or a regulated relationship. If a legal assistant saves a firm $100,000, its model bill might be $5,000 and its subscription $30,000. Falling inference cost widens the application’s margin. But if twenty competitors buy the same model and offer the same feature, competition passes the savings to customers. “Powered by AI” isn’t a moat when everyone has the same outlet.

The labs have the inverse exposure. They supply the intelligence and carry frontier research cost. Their pricing power depends on being better enough, for long enough, on tasks important enough that customers won’t route away. The hyperscalers are safer in one sense because they already own diversified cash engines. Microsoft can monetize AI through Azure, Microsoft 365, GitHub, security, advertising and its developer ecosystem. Amazon has AWS and commerce. Alphabet has Cloud, Search and advertising. A lab can lose while its cloud provider still fills capacity with another model or ordinary computing. Yet diversification doesn’t repeal arithmetic. In Microsoft’s quarter ended March 31, 2026, servers, network equipment and software at cost reached $190.9 billion, up from $132.8 billion at June 30, 2025. Nine-month depreciation expense rose to $24.0 billion from $15.7 billion a year earlier, and Microsoft said cloud gross-margin percentage declined to 67 percent partly because of AI infrastructure investment and usage. Those are observed filing figures, not an estimate of AI-only assets.

That filing gives us a simple lesson. Capital spending doesn’t disappear when the ribbon is cut. It walks into the income statement as depreciation:

If a $10 billion asset has a five-year life and no residual value, annual depreciation is $2 billion. Change the useful life to three years and it becomes $3.33 billion. Cash went out earlier, but reported margins feel the asset for years. Extending an accounting life improves current profit. Economic obsolescence doesn’t ask the accountant’s permission.

Alphabet’s June 30, 2025 10-Q showed the same conveyor belt at an earlier stage. Six-month property-and-equipment depreciation rose to $9.5 billion from $7.1 billion, and the company disclosed $23.9 billion of not-yet-commenced leases, primarily data centers, scheduled to begin from 2025 through 2031 with noncancelable terms ranging from one to 25 years. That is observed lease disclosure, not an AI-only commitment, though Alphabet explicitly said its technical-infrastructure investment supported AI products and services.

Chip designers sit in a fascinating position. Nvidia can win when labs compete, because every contestant needs shovels. It can also win when efficiency matters, because a new accelerator that halves cost per task gives buyers a reason to replace the old one. But the same improvement shortens the economic life of installed clusters. What’s a product moat for Nvidia can be an impairment problem for Nvidia’s customer. AMD benefits from buyers wanting a second source and from enormous demand exceeding one vendor’s supply, but software compatibility, ecosystem maturity and performance on actual workloads matter more than brochure FLOPs.

Memory suppliers such as Micron and SK hynix have stronger near-term scarcity economics and harsher long-term cyclicality. HBM isn’t ordinary commodity DRAM. It stacks dies, uses advanced packaging and delivers extreme bandwidth close to the accelerator. When capacity is tight, supplier pricing and margins can rise even as labs’ inference margins worsen. Then capital arrives, yields improve, capacity catches up and buyers discover they ordered through the shortage twice. Semiconductor history contains enough inventory corrections to stock a museum.

The bottom layers have longer asset lives and more local monopolies. A transmission line can serve factories, homes and future data centers. A gas turbine can sell power to the grid. A generic powered shell may host non-AI computing. Those are reusable assets. A building designed around one tenant’s liquid-cooled rack density, proprietary interconnect and fifteen-year power contract is less fungible than the word “data center” suggests. Reusability is a gradient:

This is why there won’t be one universal AI crash trade. A cheap-model breakthrough could hurt premium labs and accelerator utilization while helping applications and enterprises. Persistent giant-model demand could enrich Nvidia, HBM and power suppliers while preventing the labs from earning software margins. The transformation can succeed while value migrates down, up or sideways through the stack.

Public-market analysis gets silly when analysts assign every dollar of cloud growth to AI and every dollar of capex to a chatbot. Most companies don’t disclose an audited AI segment. Search, advertising, databases, ordinary cloud workloads and AI often share the same infrastructure. The honest scoreboard therefore uses company-wide and disclosed segment figures, then labels management’s AI attribution instead of inventing our own.

The operating snapshot below uses the latest results publicly available through August 7, 2026. Quarterly flows, trailing-twelve-month flows and fiscal-year figures are identified rather than blended. The valuation snapshot uses July 23, 2026 closing market capitalizations so eight moving securities aren’t silently measured on eight convenient dates. Market capitalization is a market observation calculated from price and shares, not an audited financial-statement item. Enterprise value is omitted where a same-date reconciliation of debt, cash, investments and minority claims isn’t supportable from the cited records. That limitation is better than decorating a table with false precision.

What Public-Market AI Valuations Already Assume

These equity values are approximate third-party market records, useful as a common-date price lens rather than primary accounting evidence. The price panel uses the same market-cap methodology and measurement date for Microsoft, Amazon, Alphabet, Oracle, Nvidia, AMD, Broadcom and Micron. These are pricing references, not primary accounting evidence.

This still isn’t a complete intrinsic-value model. A market cap doesn’t reveal the expected cash-flow path by itself, and a high multiple can be rational when growth and reinvestment returns are exceptional. It does force the correct investment distinction:

Amazon crossed $3 trillion of equity value on August 3, 2026, after this measurement date, illustrating why the price panel needs a fixed clock. The durable question is whether AWS cash generation and the rest of Amazon can justify its $220 billion 2026 capital-spending plan.

The Operating Evidence Behind the Valuations

The entries are observed company results unless described as management guidance. Microsoft figures come from its fiscal Q4 report. Amazon and Oracle figures come from company releases, while Alphabet figures come from its Q2 release and call. Nvidia’s inventory and margin figures come from its Form 10-Q, and AMD’s Q2 data were reported by Reuters. Broadcom’s AI revenue is a company-provided result. Micron and SK hynix figures come from company releases. None proves AI-only return on invested capital because the required segment disclosures don’t exist.

The scoreboard also reveals why net income can mislead during a circular boom. Amazon’s Q2 2026 net income included $53.4 billion of non-operating pre-tax income, primarily from revaluing its Anthropic investment. Alphabet recorded $98 billion of other income, primarily unrealized gains on equity securities. Those marks are valid accounting events under the relevant rules, but they aren’t cash generated by renting compute. A lab’s rising private valuation can make its cloud investor report enormous earnings while both parties continue consuming cash to build the commercial relationship.

Nvidia’s moat isn’t just a fast matrix engine. It’s CUDA, libraries, networking, rack-scale design, developer familiarity and the ability to ship a supported system. That’s why a theoretically cheaper chip doesn’t automatically win. Customers compare completed-task cost, deployment time and software risk.

But scarcity rent creates its own enemy. If a hyperscaler spends $20 billion a year on merchant accelerators and believes a custom chip can lower equivalent workload cost 30 percent, the theoretical savings pool is $6 billion a year. That pays for an absurd amount of silicon design.

Google has TPUs. Amazon offers Trainium and Inferentia. Microsoft has Maia. Meta develops internal accelerators. Broadcom and Marvell sell the design, networking and intellectual property required to turn those ambitions into hardware. Broadcom said fiscal Q2 2026 AI semiconductor revenue reached $10.8 billion, up 143 percent year over year, driven by custom accelerators and AI networking. Management also identified six core custom-chip customers and warned that the mix toward rapidly growing AI semiconductors diluted total-company gross margin relative to software. That’s company-provided segment commentary, not evidence that custom chips have already displaced Nvidia broadly.

Custom silicon wins where workload volume is enormous and relatively predictable. A TPU or Trainium deployment can remove general-purpose features, optimize memory and interconnect around internal software, and avoid an external supplier’s margin. Merchant GPUs win where flexibility, time to market, frontier performance and ecosystem compatibility matter. The market can support both:

This segmentation changes Nvidia’s risk. The threat isn’t that every customer leaves. It’s that hyperscalers reserve Nvidia for the highest-performance workloads while moving the stable high-volume middle onto internal silicon. Nvidia keeps the glamorous frontier and loses some of the workload that amortizes the customer’s factory.

Semiconductor filings report customers invoiced directly. Economic concentration can be higher. A server manufacturer, distributor and cloud provider can all appear as separate customers while the same frontier lab or cloud buildout drives their purchases.

Counting invoices at the bottom doesn’t create independent demand at the top. The right stress test asks how much supplier revenue ultimately depends on the same hyperscaler capex budgets, anchor labs and token-growth assumptions.

HBM uses vertically stacked DRAM dies connected through through-silicon vias and placed close to the accelerator through advanced packaging. The finished supply chain is a relay race:

A perfect GPU die without qualified HBM is inventory. Perfect HBM without packaging capacity is also inventory. Nvidia’s Blackwell generation shifted advanced-packaging demand toward TSMC’s CoWoS-L process while Hopper continued using CoWoS-S, illustrating how a bottleneck can change form rather than disappear.

HBM also consumes disproportionate wafer capacity. Micron said HBM3E required roughly three times as much wafer supply as DDR5 to produce the same number of bits on the same node, with later HBM generations expected to raise the trade ratio. That means an HBM boom tightens ordinary DRAM too. It also means every capacity decision has opportunity cost.

The cycle looks harmless while everyone is sold out:

SK hynix approved $38.3 billion of additional Korean manufacturing investment through 2031 in August 2026, with the relevant cleanrooms arriving years after the board decision. That’s observed investment approval, not proof of future oversupply. It demonstrates the lag. The memory supplier must decide today what model demand, accelerator architecture and competitive capacity will look like when the cleanroom opens.

The bull case is structural scarcity, higher HBM content per accelerator and long-term agreements that discipline supply. The bear case is familiar memory physics: high fixed costs, slow capacity response and brutal pricing when supply finally exceeds revised demand. The scarce supplier can earn extraordinary margins today while helping make the lab’s task economics worse. Both statements can remain true until the cycle turns.

Atlas is the point where these public-market bets become one physical object. Its accelerator invoice becomes Nvidia revenue, its memory content supports HBM pricing, its network becomes supplier backlog, its electricity contract supports utility investment and its financing becomes a private-credit asset. One construction decision can appear as growth at six public and private layers before Atlas has completed a single customer task. That is why counting revenue down the stack doesn’t tell us whether six independent sources of demand exist. They may be six claims generated by the same anchor tenant.

This is where the labs become the whole game. A frontier lab can promise to consume compute years before every future enterprise task has produced collected cash. That makes OpenAI and Anthropic more than glamorous software companies. Their revenue assumptions help support cloud backlog, infrastructure construction and supplier orders all the way down the stack.

So the bear case can’t be that OpenAI and Anthropic have no demand. They obviously do.

The real question is why some of the fastest revenue growth in corporate history hasn’t yet liberated frontier AI from enormous capital consumption.

OpenAI represents one side of the problem. It has staggering usage and revenue growth, but it also has staggering compute requirements. The optimistic interpretation is that the company is investing far ahead of demand and will eventually produce extraordinary margins. The pessimistic interpretation is that capital consumption isn’t merely a temporary growth expense. It’s an intrinsic property of manufacturing increasingly capable intelligence.

Reuters reported that OpenAI generated approximately $13 billion of revenue during 2025 while spending roughly $8 billion. It expected about $600 billion of cumulative compute spending through 2030 and more than $280 billion of revenue in 2030. Reported inference expenses quadrupled during 2025, contributing to a decline in adjusted gross margin from 40% to 33%.

Be careful with those numbers because they don’t all describe the same period or accounting concept. Cumulative compute spending through 2030 can’t be divided casually by revenue in one future year. Spending can support future customers, training and owned capacity. Revenue projections can compound sharply near the end of the forecast.

Still, they let us draw the central question.

Traditional gross margin asks how much revenue remains after delivering the current product.

Frontier AI complicates the boundary because the next product is partly necessary to defend the current one. If OpenAI stopped training more capable models, competitors could erase its premium. Some research spending therefore behaves like optional growth investment, and some behaves like the cost of remaining in business. The distinction is crucial:

If training expense eventually stabilizes while a durable model serves expanding demand, margins can improve dramatically.

If each generation is quickly matched, distilled or commoditized, the laboratory must spend billions again merely to recreate a temporary lead.

OpenAI then resembles a pharmaceutical company whose patent expires every year, except the next drug also requires a new power plant.

Anthropic appears closer to the counterexample because its enterprise-heavy business can generate much better unit economics. But its model creates a different vulnerability.

Anthropic’s reported growth has been extraordinary. Reuters reported in October 2025 that the company was targeting a $9 billion annualized revenue run rate by year-end and projected between $20 billion and $26 billion for 2026. Claude Code alone had approached a $1 billion run rate.

Annualized run rate is useful for a rapidly growing business, but it isn’t annual recognized revenue.

If a company produces $750 million during December, multiplying that month by twelve produces a $9 billion run rate. The calculation tells us the speed at the finish line. It doesn’t mean the company collected $9 billion during the race.

Neither number is dishonest. They answer different questions. Run rate asks, “How fast are we moving right now?” Recognized revenue asks, “How far did we actually travel?” Gross profit asks, “How much remained after delivering the service?” Free cash flow asks, “After the entire machine consumed cash, what was left?”

Every serious analysis of Anthropic needs to keep all four numbers separate.

Anthropic isn’t simply raising its published price per token every generation. Recent frontier generations have broadly maintained their price bands. That’s the important point. While underlying compute should become cheaper over time, Anthropic is attempting to preserve the market price of frontier intelligence while encouraging longer sessions, more reasoning, more agent calls, premium speed tiers and deeper integration into enterprise workflows.

OpenAI’s consumer subscription, enterprise/API business and frontier research machine should be mentally separated even when private-company reporting doesn’t provide audited segment accounts. The consumer business converts a fraction of a vast free audience into monthly subscriptions. The enterprise business sells seats, API consumption and workflow infrastructure. The research machine trains the models and creates the capability both commercial businesses sell.

Each engine has a different test. Consumer economics depend on paid conversion, retention and serving cost per user. Enterprise economics depend on task value, reliability, security, integration and price competition. Research economics depend on how long a capability lead lasts and whether the resulting models generate enough gross profit before the next training cycle.

Reuters reported that OpenAI’s recognized 2025 revenue was $13 billion, above a $10 billion internal projection cited by the source. That is a full-year flow, unlike the company’s annualized run-rate milestones during the year. Reuters also reported that OpenAI expected cumulative compute spending of roughly $600 billion through 2030 and 2030 revenue above $280 billion, divided approximately evenly between consumer and enterprise units. Those are management plans relayed by sources, not audited forecasts.

The gap between gross margin and cash burn tells us which question each number answers. Gross margin asks whether revenue covers direct serving cost. In Q1 2026, documents reported by The Information put revenue at $5.7 billion, cost of revenue at $3.5 billion and gross margin at 39 percent. Reuters separately reported $3.7 billion of quarterly cash burn and said it couldn’t independently verify the underlying report. If the gross-margin figures are comparable:

That’s genuine economic progress at the delivery layer. The Information’s reported figures, also summarized by Reuters, put Q1 research and development expense at $8.6 billion, including model training, and operating loss at $9.3 billion including $2.3 billion of share-based compensation. Gross profit can be positive while the frontier treadmill consumes far more.

Stock compensation isn’t current cash, but calling it free would be cute accounting. It transfers value through dilution. Training expense may create an asset-like future benefit, but accounting usually expenses research because success and useful life are uncertain. Cash burn adds another distinction because customer prepayments, compute credits, financing and working capital can make cash timing differ from operating loss. For a private company with complicated partner agreements, every number needs its label attached like luggage at an airport.

Advertising could become the subsidy that closes the consumer equation. Reuters reported in April 2026 that internal projections contemplated roughly $2.5 billion of advertising revenue in 2026 and $100 billion by 2030. These were reported projections, not recognized results. The strategic logic is obvious: hundreds of millions of nonpaying users create attention, and advertising monetizes users who won’t buy subscriptions.

But ads don’t repeal inference cost. They change who pays. Suppose a free user costs $1.20 per month to serve and generates $0.80 of ad gross revenue. The user is still contribution-negative by $0.40 before R&D and overhead. If better targeting raises ad revenue to $1.80, the free tier contributes $0.60. Now increase agent usage threefold and serving cost to $3.60. The subsidy breaks again. The relevant ratio is:

OpenAI’s paid-conversion assumptions therefore matter enormously. If 900 million weekly users convert at 5 percent to an average $25 monthly net subscription, that simple hypothetical produces $13.5 billion a year. At 10 percent, it produces $27 billion. The arithmetic is illustrative, not an estimate, and weekly users aren’t the same as unique monthly billable accounts. It shows why one percentage point of conversion can be worth billions and why consumer enthusiasm alone doesn’t reveal profitability.

The business becomes self-financing only when operating gross profit covers research, sales, administration, interest, working capital and the capital expenditures or leases needed to sustain growth. A useful milestone ladder is:

What OpenAI Has to Prove Before the Model Self-Funds

OpenAI doesn’t need SaaS-like gross margins tomorrow. A fast-growing infrastructure-heavy business can rationally consume cash. The thesis fails if margins improve, capital intensity falls per useful task, conversion deepens and each compute cohort earns attractive returns before obsolescence. It strengthens if revenue keeps exploding while the capital required before self-financing explodes faster.

OpenAI’s reported 2030 revenue target doesn’t tell us what gross margin, recurring research cost or cash infrastructure burden accompanies it. We can still put the missing variables on the whiteboard without pretending somebody left OpenAI’s spreadsheet in our inbox. The following are our own scenarios, not management forecasts or analyst estimates. They show what has to be true for a reported $280 billion revenue year to become an economic success.

Three Ways OpenAI’s 2030 Math Could Land

The bull scenario requires something close to software-like delivery economics despite agentic workloads, while research and capital spending rise much more slowly than revenue. The middle scenario produces extraordinary revenue and still consumes cash because the laboratory remains on the capability treadmill. The bear scenario isn’t “nobody uses ChatGPT.” It’s a $120 billion business whose price, serving cost and competitive investment never line up.

Advertising changes the revenue column, not the laws of arithmetic. Paid conversion changes consumer revenue. Custom silicon and inference optimization change gross margin. A slower frontier race changes R&D. Owned infrastructure and leases change capital timing. The model works when those improvements arrive together.

An investor can update the table quarterly with five observations:

If revenue and gross margin rise while R&D growth, burn and incremental commitments decelerate, the bull path is becoming real. If revenue grows while the other four demands grow just as quickly, scale hasn’t yet produced self-financing economics.

For Atlas, these scenarios aren’t an abstract debate about a private-company valuation. Imagine the project’s cloud operator has reserved 45 percent of Atlas for an OpenAI-like anchor tenant. In the bull scenario, the tenant’s gross profit increasingly funds the reservation, its credit quality improves and Atlas can refinance against a customer that no longer depends on repeated equity infusions. In the middle scenario, the tenant may keep paying while outside capital remains available, but Atlas’s lender is underwriting both AI demand and the willingness of future investors to bridge the laboratory’s cash deficit. In the bear scenario, even a $120 billion laboratory can be a dangerous anchor tenant if serving costs, research and infrastructure commitments leave it structurally cash hungry. Atlas doesn’t need the tenant to disappear. A request for lower pricing, shorter contract duration or less reserved capacity is enough to change the project’s equity math.

Anthropic’s enterprise weighting makes it the strongest natural objection to the claim that frontier economics are broken. Reuters reported in October 2025 that more than 300,000 business customers generated about 80 percent of Anthropic’s revenue, with an annualized run rate approaching $7 billion at the time. Claude Code had reached nearly a $1 billion annualized run rate. Again, run rate annualizes the current pace. It isn’t cumulative recognized revenue, gross profit or free cash flow.

Enterprise concentration is both a strength and a risk. Businesses can spend far more than consumers because the value ceiling is a labor or revenue budget, not an entertainment subscription. They also negotiate, route workloads, demand service credits and notice when one workflow suddenly eats a small intern’s salary in tokens. Concentration among large customers makes sales efficient but gives buyers leverage at renewal.

Claude Code makes the consumption mechanism visible. A coding agent doesn’t merely answer a question. It reads repositories, searches, edits, tests, retries and sometimes delegates. Cost per active developer is therefore:

Monthly cost depends on active days and the distribution’s fat tail. T

Read the original on cuecloud.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.