By Farhad Billimoria and Conleigh Byers
This post reports on a new academic working paper released by the both of us on “Bits-to-Watts: Connecting Markets and Prices for Compute and Power”. The paper is available here as well as an informational website www.valueofcomputeload.info
The story so far:
Huge demand for interconnection from AI and data centers. Data point: ERCOT’s large load interconnection queue grows to 226GW, up 270% from December 2024.
Power remains the critical constraint in scaling up load.
Catalysed by Tyler Norris and the Duke ‘Rethinking Load Growth’ study, the concept of flexible load provides a pathway for interconnecting these volumes within the existing grid.
Much of the discussion seems to focus on flexibility through administrative means (operator-controlled load curtailment etc), but this belies the notion that with the right market design, aspects of flexible load can be directly incorporated into the spot market.
Indeed the fact that build of data center and AI load is purely commercially driven means that we should be able to integrate flexibility through the revealed value of load.
Back to the future: The original Schweppian market design vision
Let demand bid its flexibility and be dispatched through the spot market (SCUC, SCED etc) either directly or through intermediaries1. This was the original market design vision envisioned by Schweppe as far back as 1978, and is evident in more recent initiatives, such as LLP and work by Arushi Sharma Frank and others.
Indeed with the right market design, new demand is not inherently harmful for reliability or affordability, and may indeed provide resilience for the system. We show this in a recent working paper by the two of us (Conleigh Byers and Farhad Billimoria). For a more digestible version, see our Harvard Belfer Policy Brief on ‘Participatory Demand and New Large Loads in Electricity Markets’.
New demand (even if that new demand has a high value of lost load and is inflexible) is not inherently harmful for reliability or average prices in the long run, provided supply is elastic and permitting/interconnection procedures allow new generation to be built and connected in commercial time-frames. (Byers & Billimoria, 2026).
Much work is being dedicated to flexible load control in data center and AI applications. Recent demonstrations of viability for many applications (see EPRI’s DC Flex workstream) suggest this is increasingly technically feasible (Colangelo et al (2025). This work will likely continue to demonstrate newer ways of shifting and deferring load. Hence the question becomes less a technical and more of an economic one.
Ultimately the ability and willingness to be flexible depends upon the the economic value of compute. While anecdotal discussions suggest that the Bit-Watt spread “has never been higher”, there is a relative dearth of analysis that translates the economic value of data to their power market equivalent. Like many things, the answer may well be hidden in plain sight.
Cloud computing is a massive, albeit concentrated market (>$700bn), with transparent pricing across the main competitors including Amazon Web Services (AWS), Google Cloud Platform (GCP) and Microsoft Azure.
There are 1000s of cloud instances available for purchase across the major cloud providers (AWS, GCP and Azure), varying by specification (e.g. GPU/CPU architecture, memory, storage etc). The choice of instance depends upon a range of factors including the type of application (enterprise, HPC, ML/AI/GPT training, inference etc), user willingness and ability to pay; and the Quality of Service (QoS).
The Quality of Service (QoS) for cloud compute is governed by the Service Level Agreement (SLA), and has typically two forms of service.
High priority service - On-demand computing: guarantees high uptime (typically 99.99% or greater) with service crediting for failures to achieve these levels. This can be purchased on hourly basis (or longer terms)2. See as an example Amazon EC2’s region-level SLA below.
Low Priority or Interruptible Service - Spot computing: This offers computing at much lower rates but is interruptible by the provider. See an excerpt from Amazon EC2’s spot service interuption policy:
The On-demand service can be seen as a physically and financially hedged version of the Spot service.
This product differentiation allows the consumers elasticity to be reflected in service selection — on-demand for essential / critical uses; and spot instances for non-essential / interruptible / elastic applications.
What does the transparent data on the price of cloud compute reveal about the value of load and the potential impact on price formation? We can gain insights on this by translating the price of compute to its energy equivalent.
We move away from the term VoLL — as lost load implies involuntary curtailment. We call it the Value of Compute Load (VoCL) or more generally the Value of Load (VoL).
We define the Value of Compute Load (VoCL) as a metric linking the economic exposure associated with compute services to the incremental electrical load required to provide those services. Formally the VoCL is given by:
\(VoCL = \frac{MCE}{MPD} \)
MCE is Marginal Compute Exposure, measured in $/hour - the opportunity cost for the cloud provider of providing the compute. This includes the revenue from selling the cloud compute capability, as well as any liabilities associated with reliability guarantees.
MPD denotes Marginal Power Draw of the compute workload, measured in MW.3
The spot case is the simplest to understand as its freely interruptible by the cloud provider with the MCE purely attributable to cloud computing cost for the instance in the immediate time period. For on-demand pricing, the calculus is more complex, given there is a real contingent liability associated with interrupting load based on the SLAs (and non-financial costs e.g. for reputational risk).
Important to note that there are limits; this is a snapshot of instances available as infrastructure-as-a-service (IaaS) via the AWS cloud platform. Of course, compute may be self-hosted, on-edge premises as well as offered directly to AI service providers or other hyperscalars. So this represents a part of the market rather than all of it, or even most of it. Nevertheless it still provides a relevant datapoint to understand the energy value of load.
We computed VoCLs for 75+ Amazon AWS cloud computing configurations with GPU capabilities (most applicable for AI applications). We show in the figure below the calculated VoCLs for on-demand cloud service for the instances assuming utilization based on the median of inference benchmark data.
The key message is that there is no one-size fits all number. The results revealed a big spread of outcomes, varying by the type and vintage of the GPU chip - newer, frontier chips have a higher VoCL driven by factors incl. consumer demand and scarce supply for faster compute.
There was also variation based on the reliability of the service - between the high-priority on-demand service and the interruptible (spot service). We show this for AWS’ p6-b300.48xlarge instance, which uses the frontier Blackwell B300 GPU chipset.
The blue line is the Interruptible (Spot) service and the black line is the On-demand over time. The difference between the two is the value of flexibility for the cloud customer (essentially what you can save by being interruptible). Interesting to see that the value varies over time, initially little value but has gotten much more valuable in recent times.
There’s also variation by location and by the power utilization of the computing instance. Much more of this variation shown in the paper.
The VoCL has important implications for data center policy, for market design and for data center flexibility. By considering the raw-economics of AI-based cloud compute we are able to get a bottom up view of the implicit willingness-to-pay for compute (at least at this snapshot in time) and the natural flexibility that may be embedded in large load. This suggests a few potential opportunities to thread the needle on enabling the economic opportunity of AI-driven large load, while mitigating the impacts on energy costs and broader distributional impacts on consumers in the grid. Not easy… but important and worth it.
More to come… including market design implications.
This can be either through direct wholesale participation, retail, or demand response contracts. Note I have specifically said ‘demand response contracts’ here because the way in which DR has been implemented in spot markets to date has incentive and operational considerations. See Hogans work here and here.
Savings plans can also be purchased to reserve capacity for forward periods at a discount to the base on-demand rate.
The marginal power draw MPD takes into account the instantaneous energy use of the system including the combined draw from GPUs, CPUs, storage, memory etc. This is the more complex component to decipher. Data from MLCommons provides benchmarks for a set of common AI models across a range of different system configurations. From this we can extract some typical bounds on utilization. It is however critical to note that this is a highly heteroenous value, and can vary signficantly by application, resources and time. We also include the power draw required for ancillary cooling of the facility.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.