RSS Amplifier

Macrocosmos · Jul 14, 2026

The Economics of Liquid Training

0
Sign in to vote or save

Macrocosmos · Macrocosmos

There has rarely been as much interest in physical infrastructure build-out as there is today [1]. Data centres, and the compute they contain, have become central to the conversations in corporate boardrooms and government chambers alike.

A colleague once characterised the new entrants to the data centre market as “real estate opportunists.” It is a crude description, but it captures the mood. The race to build sheds and pack them with GPUs is a modern day gold rush.

Nearly fifteen years in the construction industry has taught me that accelerating an infrastructure build-out at this scale isn’t easy, and that Western nations (rich in NIMBYism) will struggle with pulling it off.

The real north star of this effort is not the sheds themselves. It is the compute within and the golden age that compute heralds. Whether you are an international conglomerate, a nation state, or a start-up, the critical resource of the coming generation will be compute.

Supply gluts and squeezes will produce different market effects, but the unprecedented level of capital deployed today will consistently lead to a single core question - how to maximise effective returns on this investment?[1]

The next economic question will not be how much you can build, but how much useful work can be extracted from the capacity already installed. That question emerges at every scale, from the global market to the nation state, the conglomerate, the single data centre, all the way down to the discrete GPU.

In this piece, we lay out the shape of the market today, how different compute workloads create value, how the capital model for GPUs works, and how liquid training on iota will drive us toward better economic outcomes.

The current phase of AI infrastructure is almost crude in its articulation: as much compute as possible, as quickly as possible, at almost any price.

That model emerged because of a variety of primary factors. The scale of demand is enormous, the urgency to be first (and the fear of being second), and capital markets have offered deep liquidity to fund the companies chasing this opportunity. Bond markets are straining with new forms of private debt, all in the pursuit of a maximal build-out. [1]

The earliest stages of the bet have paid off. Token utilisation has followed a hockey-stick that venture capitalists could not have dreamed of, model capabilities continue to scale at a frightening rate, and demand for AI compute keeps being underforecast.[1] [2]

The first infrastructure phase was built around peak performance. That worked, it unlocked today’s scaling laws, and it still matters for the most demanding workloads, but the market has moved beyond a single dominant workload pattern.

As the market matures, workload-specific optimisation becomes more and more important, within both hardware and software. Numerous providers are coming to market with Nvidia alternatives - specialised inference chips such as Cerebras, or OpenAI’s Jalapeno, but similarly, software evolution to drive MFU and fleet performance up is accelerating. [3]

Different AI workloads now place different demands on the same scarce GPU base, bottlenecked in different ways: by fragility, latency, memory and network bandwidth, or by raw compute.

As frontier pretraining grows, the constraint is no longer only the performance of a single cluster. It is whether enough suitable capacity can be assembled, kept available, and coordinated for the duration of the run.

That is why distributed training has become commercially important. If the required compute cannot be found in one place, training has to work across multiple, spread out pools of compute. Google has already trained Gemini across multiple sites [4], and Anthropic’s Fable release treated distributed training as valuable closed knowledge. [13]

iota enters as the product expression of the same shift: training designed for a market where compute capacity is fragmented, priced differently, and constantly reallocated.

By making training inherently flexible, we can achieve both greater utilisation per installed GPU-hour, and greater aggregate training capacity from a given installed base.

“Allocation is not the same as utilisation.”

Those words were repeated to me many times in conversations with data centre owners and AI labs, and they sit at the heart of compute economics. Allocation tells you whether capacity has been sold, rented, reserved, or contracted. Utilisation tells you whether the GPU is actually doing useful work. [5]

Put simply, commercial output and productive work are not the same thing.

The incentive to close that utilisation gap differs by actor, be they renters, suppliers, or internal fleet operators. The market incentive, however, is clear - higher productive utilisation of an asset will drive greater capital returns, and greater market competitiveness.

We should be honest about this: underutilisation is not universally bad for every actor. What matters for the industry as a whole is productive GPU-hours, not merely booked ones.

Utilisation means ‘how much is the compute unit used’. At a single GPU (or training run level), this is expressed as MFU (or model flop utilisation) - how much is the GPU itself actually working. At data centre level it refers to the asset being hot and serving demand (e.g. having a workload deployed against it). Even optimised data centres achieve values at circa 60-70%, with the best in the industry citing figures in the 80% and 90% region.

Utilisation is also rarely predictable. Time zones create imbalances, inference is inherently peaky [2] [6], while regions are oversupplied during local off-hours and constrained during other markets’ working hours. These patterns leave capacity that is genuinely difficult for today’s most valuable workloads to use.

This is where the opportunity for liquid training emerges. What if training became the workload that could fill the gaps in the compute market — absorbing spare, flexible, fragmented, heterogeneous, or under-monetised capacity and turning it into productive work?

iota is designed for exactly that future: safely occupying compute that conventional training cannot reliably use, and converting otherwise stranded capacity into new economic output.

The market already tells us that not all GPU-hours are equal, because it already prices the difference, as the charts below showcase the differences between reserved and spot compute across AWS. [7]

Certainty, immediacy, location, hardware generation, reliability, and contract structure all move the price of a GPU-hour. Guaranteed capacity commands a premium.

Fragmented, interruptible, or time-limited capacity is cheaper, precisely because most valuable workloads cannot reliably consume it. [7] [8]

The core point is not that cheap compute is abundant. It is that the market discounts imperfect compute [7] [8], and conventional training cannot use enough of it. The economic opportunity is to make the discounted part of the market usable: to take capacity that is cheaper for a good reason, and still turn it into training work that is reliable enough to be valuable.

Frontier and large-scale training represent a large enough share of the market that, if they became liquid, all major participants could raise their utilisation. In a supply squeeze, iota harvests the long tail of compute (heterogeneous, off-peak, fractional), as well as in a glut, it helps to become a floor bid for overavailable supply. [9]

iota is designed to make more of this discounted or imperfect supply usable for training.

Training has properties that make it incredibly attractive for this class of demand (putting technical complexity to one side for a moment), because it is valuable, persistent, and less latency-sensitive than inference. It is also strategically necessary because it creates the models that later become products through inference, applications, agents, and APIs.

It is also persistent: there is always more pre-training, post-training, fine-tuning, evaluation, and research to do, and the markets reflect this with training compute expected to grow at a double-digit rate into the 2030s. [1]

The persistence of pretraining matters in particular. The industry has moved toward training base models for longer [11] than earlier compute-optimal assumptions suggested [10], because stronger base models can make downstream post-training more efficient [10] [11] across many later tasks and use cases. If a model is pretrained once and adapted many times, there is a rational case for spending more compute upfront to reduce the cost and difficulty of repeated downstream work.

While other workloads like batch inference and RL are already interruption tolerant, we believe they are already factored in, and even as agentic use scales, mechanisms to improve workload liquidity are the key to driving greater utilisation.

Compared with inference, training has a different time-sensitivity profile, measured in days and months rather than seconds and minutes. That profile is exactly what makes it a candidate for flexible demand.

The obstacle is that conventional training is rigid.

It generally wants stable availability, homogeneous hardware, strong networking, and a low tolerance for disruption [9] [12]: when nodes drop, slow down, or need to be reallocated, the cost shows up as delay, wasted work, and operational complexity.

In that paradigm, interruptible compute is only cheap on paper, because the discount is consumed by delay, wasted work, and operational complexity.

Liquid training changes the shape of the workload so it can occupy the gaps conventional training cannot use.

For that to work, a handful of technical properties matter: tolerance to interruption, tolerance to heterogeneous hardware, tolerance to constrained bandwidth, and fast reallocation so capacity can leave for higher-value work and return cleanly.

Our recent Orion 100B work gives us an early indication of the shape of that trade-off: if flexible training can preserve a meaningful share of co-located training throughput while accessing compute at a material discount, the economic case can become attractive quickly.

Spot and flexible capacity can trade at discounts large enough to make that question worth testing [7] [8], provided the discount is not consumed by slower training, operational overhead, or model-quality degradation.

We treat these in depth in the next piece in this series. Here, it is enough to say that iota turns cheap-on-paper capacity into genuinely usable training supply.

The anchor of our thesis is simple: driving useful work across installed capacity carries tremendous value.

The market has installed a great deal of compute that is not uniformly useful to every workload. The missing piece is a workload that can turn fragmented, discounted, heterogeneous, or time-limited capacity into useful training. iota is designed to aggregate that low-order capital and convert it into a higher-order commodity: frontier-scale training flops.

This expands effective supply without manufacturing a single new GPU. It improves supplier economics, because owners can monetise capacity across more time steps, more geographies, and more conditions. It improves buyer economics, because model builders gain another route to useful training capacity, provided quality, reliability, and timeframe stay within acceptable bounds.

Fractional, hour by hour, minute by minute settlement mechanisms for compute will help to unlock this [9] [13] - as the compute market grows, different vehicles for aggregating and acquiring compute on demand in a liquid fashion will be transformative. In the same way that plastics and different fuel classes transformed the oil sector, we see the cracking processes emerging for compute.

This is not a direction unique to iota. Google’s work on Kubernetes-native AI infrastructure and scheduling [3] [13], and platforms such as Shadeform that aggregate GPU supply across providers [12], both point to the same underlying market shift: compute is becoming more liquid, more workload-aware, and more actively routed.

For sovereign AI participants [1], iota also unlocks the ability to aggregate compute in the way hyperscalers aggregate campuses, changing the fundamental infrastructure parameters of training their own models.

We see a world where the overall market for training increases - more start-ups, labs, enterprises, sovereign entities and beyond will have growing compute budgets, growing token budgets, and a need to balance training and inference costs to achieve financial performance. Harnessing disaggregated compute to achieve these effects can help reduce the barrier to entry, as well as allowing for new mechanisms to collaboratively train models.

We are still on the hedonistic, upward swing of the AI investment cycle. We will keep building at tremendous speed and scale as the industry races to harness the opportunity in front of it. That build-out remains necessary.

But economic efficiency will become the decisive secondary question. Those who win will be more profitable, more able to reinvest, and ultimately able to deliver superior value to those who cannot.

The mechanism we believe drives maximal global utilisation is the advent of liquid, flexible, heterogeneous, interruptible training: a way to turn training from rigid demand into flexible demand, and to expand effective training supply without waiting for every new substation, data centre, and cluster to arrive.

The next phase of the AI infrastructure race will not only be built. It will be refined.

iota is the refinery for that future.

Will Squires, Macrocosmos

If these ideas resonate, get in touch: hello@macrocosmos.ai. Or sign up to our newsletter to receive the rest of the series direct to your inbox.

[1] International Energy Agency — Energy and AI
https://www.iea.org/reports/energy-and-ai

[2] Aubakirova et al. — State of AI: An Empirical 100 Trillion Token Study with OpenRouter
https://arxiv.org/abs/2601.10088

[3] Malleni et al. — Evaluating Kubernetes Performance for GenAI Inference
https://arxiv.org/abs/2602.04900

[4] Google DeepMind — Gemini 1.5 Technical Report
https://storage.googleapis.com/deepmind-media/gemini/gemini_v1_5_report.pdf

[5] Microsoft — Singularity: Planet-Scale, Preemptive and Elastic Scheduling of AI Workloads
https://arxiv.org/abs/2202.07848

[6] AWS — Amazon EC2 Spot Instances
https://aws.amazon.com/ec2/spot/

[7] Google Cloud — Spot VMs
https://cloud.google.com/spot-vms

[8] Hoffmann et al. — Training Compute-Optimal Large Language Models
https://arxiv.org/abs/2203.15556

[9] Gadre et al. — Language models scale reliably with over-training and on downstream tasks
https://arxiv.org/abs/2403.08540

[10] Khatua and Mukherjee — Application-centric Resource Provisioning for Amazon EC2 Spot Instances
https://arxiv.org/abs/1211.1279

[11] Shadeform — GPU Cloud Marketplace

https://www.shadeform.ai/

[12] Shadeform — Documentation

https://docs.shadeform.ai/

[13] System Card: Claude Fable 5 & Claude Mythos 5

https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf

Read the original on macrocosmosai.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.