Happy Sunday and welcome to Investing in AI. Be sure to check out the AI in NYC podcast for highlights on what is happening the New York - the applied AI capital of the world.
Also, my book, “Investing in AI: Six Mental Models To Tell Moats From Traps” will be published September 15th. I will be doing lots of events and fireside chats about the topics in the book (contact me if you want me to do one for a group you are affiliated with) and my first one is with 10 East, a private markets investment group. Readers of this newsletter are welcome to attend. Here are the details:
I am doing a webinar with 10 East next Wednesday, August 5th at 11:30am. For those interested in attending, you will need to first sign up to join 10 East. 10 East is a private markets investment platform that provides qualified individuals and family offices with access to curated private equity, private credit, venture capital, and real estate opportunities. Members invest alongside the firm’s principals on a deal-by-deal basis, with no upfront commitments or membership fees. Learn more and apply here.
Now on to this week’s commentary about Neoclouds and Hyperscalers and the AI trade.
The AI infrastructure trade of the last three years was a training trade. Capital chased whoever could stand up the largest coherent GPU cluster fastest, and the winners were the companies that solved for peak FLOPs. That era is closing. Inference — running trained models against live traffic, billions of times a day — is becoming the dominant workload, projected to account for roughly 80% of the neocloud market by 2030.
That shift matters because training and inference are different businesses wearing the same hardware. Training is a capital event: a handful of buyers, long commitments, tolerance for scarcity pricing. Inference is an operating expense that recurs on every query, with buyers who optimize relentlessly on cost and latency.
Two architectures are competing to serve it. On one side, the legacy hyperscalers — AWS, Azure, GCP — general-purpose, virtualized, CPU-first platforms retrofitted for accelerated compute. On the other, the neoclouds — CoreWeave, Lambda, Crusoe, Together — purpose-built, bare-metal, GPU-native. My thesis: neoclouds are winning today on scarcity and speed, but scarcity is a temporary asset. Long-term survival against hyperscaler balance sheets requires moving up the AI software stack, and most of them haven’t started.
Speed to deployment. Hyperscalers carry the tax of serving everyone. Their architectures are multi-tenant, multi-purpose, and layered with virtualization and compliance machinery accumulated over fifteen years. Standing up a large accelerated cluster inside that environment is a procurement and integration project. Neoclouds, unburdened by legacy, have deployed clusters approaching 80,000 GPUs in a matter of weeks. When a foundation model lab’s roadmap slips by a quarter because compute arrived late, that delta is worth more than any price concession.
Cost efficiency. Bare-metal economics let neoclouds price GPU capacity at a fraction of hyperscaler list — reports of up to 85% cheaper are common in the market. Part of this is real architectural advantage: no hypervisor overhead, no à la carte networking and storage line items. Part of it is that neoclouds are pricing for share while hyperscalers price for margin. Either way, for a startup burning venture dollars on tokens, it is irresistible.
The co-opetition paradox. The most interesting structural fact in this market is that hyperscalers are themselves anchor tenants of neoclouds. Microsoft accounted for roughly 62% of CoreWeave’s revenue in 2024, effectively white-labeling neocloud capacity to avoid losing customers to its own bottlenecks. That arrangement bootstrapped the category — and it is also the category’s central risk. When your largest customer is also your largest competitor, your revenue quality is a function of their capacity planning, not your sales execution.
Bargaining power of suppliers: extremely high. NVIDIA sets the clock. It controls allocation, generational cadence, and the software layer that makes the hardware usable. AMD provides a second source at the margin, not a counterweight. For a pure-play neocloud, this produces a brutal margin structure: the input price is dictated by a monopolist and the output price is set by an increasingly commoditized market. Gross margin is the spread between a depreciation schedule and a rental rate, and neither end is under management’s control.
The second supplier is energy. Power grid interconnects, substation capacity, and advanced liquid cooling have become as binding a constraint as silicon. Multi-year interconnect queues in key markets mean that data center real estate with committed power is now the genuinely scarce asset. This is why the smartest capital in the space is buying megawatts, not GPUs.
Bargaining power of buyers: increasing fast. Inference is fungible in a way training never was. A training run is a multi-week commitment to a specific cluster with specific interconnect. An inference request is a millisecond-level decision that can be routed anywhere the weights are loaded. As orchestration and routing layers mature, buyers dynamically shift workloads to whichever provider is cheapest and fastest at that instant. Switching costs collapse toward zero. Vendor loyalty in inference is a rounding error, and every layer of abstraction added above the compute strips more pricing power out of the compute itself.
Competitive rivalry: fierce and asymmetric. Hyperscalers are projected to spend $600–700 billion in CapEx in 2026. They are not conceding the workload; they are absorbing it. And they compete with a weapon neoclouds do not have: bundling. A CIO negotiating an eight-figure enterprise agreement can have AI compute folded into an existing contract alongside storage, networking, security, and committed-spend discounts. The neocloud shows up with a better per-hour price and loses to a worse per-hour price embedded in a relationship. Meanwhile neoclouds fight each other in a straight price war over undifferentiated infrastructure.
Threat of substitutes: moderate to high. Three substitutions are live. First, the edge: local inference on AI PCs, smartphones, and autonomous systems is projected to represent the majority of global inference volume — some forecasts put it above 70% by 2026 — and volume that runs on a device never touches a cloud invoice. Second, custom silicon: Google TPUs, AWS Inferentia and Trainium, and the next wave of inference ASICs are explicitly designed to move workloads off NVIDIA and off merchant capacity. Third, and most underrated, model efficiency itself. Distillation, quantization, and task-specific small models substitute engineering for compute. Every efficiency gain is demand destruction for someone renting GPUs by the hour.
Threat of new entrants: moderate. Capital intensity is a real barrier, but it is a narrower one than it looks. GPU-backed debt and private credit have made it possible to enter without equity-funding the full asset base. And a new class of entrant is arriving with a non-economic mandate: state-backed sovereign AI clouds across Europe, the Gulf, and APAC, built to keep data and inference inside national borders. They compete on jurisdiction rather than price, and they don’t have to clear a venture hurdle rate.
The asymmetry is worth sitting with. The hyperscaler threats are regulatory and competitive — survivable, expensive, slow-moving. The neocloud threat is existential and arrives on a timeline set by someone else’s supply chain.
Pure compute brokerage is a stopgap business. It exists because demand outran supply, and it earns excess returns exactly as long as that gap persists. Every quarter of capacity additions narrows it.
Neoclouds have two credible paths. The first is moving up the stack — becoming AI-native platforms rather than landlords. Managed inference, training orchestration, fine-tuning pipelines, observability, routing: these are software margins layered on infrastructure, and they create the switching costs that raw capacity never will. The second is defensible niches: sovereign compute mandates, ultra-low-latency edge deployments, and regulated verticals where compliance is the moat.
The hyperscaler endgame is more straightforward. Use custom silicon to structurally lower inference cost, use full-stack integration and bundling to absorb mainstream enterprise AI demand, and buy distressed neocloud capacity when the consolidation wave hits. Expect that wave. A market with fixed-cost assets, commodity pricing, and leveraged balance sheets consolidates on schedule.
For investors, the diligence question has changed. It is no longer “how many GPUs do they have.” It is what percentage of revenue comes from software and managed services, how concentrated the customer base is, and whether contracted power outlasts the depreciation schedule on the hardware.
The infrastructure that built the cloud is not the infrastructure that runs AI efficiently, and that gap is the entire neocloud opportunity. But gaps close. The companies that treated the GPU shortage as a window to build a platform will look like the next generation of infrastructure titans. The ones that treated it as a chance to rent hardware at a markup will be acquired, refinanced, or forgotten. We will know which is which within about eight quarters.
Thanks for reading.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.