Fears about exhausting the supply of training data for AI frontier models have spiked over the last year. No matter how powerful (or expensive) your array of Nvidia clusters, you need a growing supply of high quality data to eke out further scaling gains for your AI model. That has prompted discourse from across the opinion spectrum, from AI doomers and boomers alike, including Elon Musk, who claimed that we have reached the point where “the cumulative sum of human knowledge has been exhausted in AI training.” The more optimistically-inclined have suggested that we’ll be able to replace “natural” human-created data with synthetic data created by AI itself.
Perhaps. But I’m not sure we need to hope for innovation around reliable synthetic training data to forestall a data exhaustion cliff. In this regard, peddlers of data exhaustion alarmism can learn an important lesson from the failed predictions of resource exhaustion alarmists in the last century.
In the mid-20th century, environmentalists operated from a different set of fears than they do in the 21st century. Rather than worrying about carbon emissions and global warming as today — worries I share, it so happens — it was fears of overpopulation and resource exhaustion. Many people, Very Serious People™, dedicated their careers to warning of skyrocketing energy prices as oil reserves emptied out, of an exploding human population consuming all the usable minerals, and of overtaxed agricultural land with declining yields. The consequent wars, droughts, and famines would place humanity in an inescapable Malthusian bind. Surely, only a fool (or an industry plant) would dare disagree with the weight of this expert consensus!
One of the most (in)famous advocates of this point of view, Paul Ehrlich — a butterfly expert and a proponent of mass, coercive sterilization — passed away just this month. Ehrlich’s warnings about imminent waves of mass starvation killing hundreds of millions of people made newspaper headlines and sold thousands of copies of his bestselling book The Population Bomb. But in his less prominent work he focused on resource exhaustion. As he wrote in 1971:
“No matter how you slice it, the resources of the planet are finite, and many of them are non-renewable. Each giant molecule of petroleum is lost forever when we tear it asunder by burning to release the energy of sunlight stored in it millions of years ago. Concentrations of mineral wealth are being dispersed beyond recall, senselessly scattered far and wide to where we cannot afford the energy to reconcentrate them.”
In 1980, Ehrlich’s claims about resource exhaustion were challenged by an economist, Julian Simon, who made a $10,000 wager that the supply of common minerals would increase (and thus prices fall) over the next decade. Ehrlich took the bet, and lost the bet. One key moral of that story is that while Ehrlich may have been a superlative biologist, he didn’t understand basic economics.
First, there is a difference between absolute and nominal scarcity. Think of a resource existing in a series of concentric circles. The largest circle is an unknown, which is all of that resource which exists somewhere on or in this planet. The next, much smaller circle is our estimate of how much of it there is that is theoretically accessible (with a bunch of caveats about current extraction technology). The next, even smaller circle is the proven reserve of that resource, where geologists and other researchers have gone to test for its presence, like the identification of a new, vast oil field off of Guyana in 2015. The final, smallest circle is the percentage of those reserves that are currently feasible to extract and sell.
Second, Ehrlich failed to appreciate that as resource supply falls or demand grows, the rising price of that resource would propel the discovery of new sources of that resource (or alternative materials that can provide the same function, although that’s a conversation for another time). Think of the rising price as moving the line of the smaller circles outward. And just like when you move out the line of a circle, even a small movement can greatly increase the area within the bounds. It’s like how increasing the size of a pizza you’ve ordered from a small (10”) to a large (14”) actually increases the volume of pizza that you’ll receive by two and a half times.
Ultimately, the fatal flaw of the resource exhaustion alarmists was that they failed to appreciate the power of incentives, in which even small movements in price can generate massive capital investment to unlock remarkable quantities of untapped resource reserves.
So it is with data exhaustion. There is a massive, unknown circle of all the information that 8.3 billion humans could theoretically produce if they dedicated all their time and energy to that task. There is then a much smaller circle of all the information that they actually do produce, most of which is lost into the ether because it is unmarked and unrecorded. And then within that circle is the even smaller circle of all the information that is recorded but is currently inaccessible to LLMs, eg, non-digitized books, paper archives, etc. And then there is the smallest circle of all, the information that is created AND recorded AND digitized AND currently enjoys low enough transaction costs to be mined by LLMs for training data.
Data exhaustion worriers make the same mistake as the resource exhaustion crowd once did a generation ago: assuming that the supply of information is static and finite rather than expandable and functionally infinite under the correct set of incentives.
Let me provide a concrete example. I wrote a fairly well-received book about the careers of several now mostly-obscure, mid-20th century radio broadcasters, including one named Carl McIntire. Now, consider the amount of “usable” data about McIntire that could be provided to an LLM trained on that book through some agreement with the publisher, Oxford University Press. That gives us 320 pages of data. Throw in a few other published books that touch on McIntire and we might have a few thousand total pages.
Not shabby, and one could learn a lot about the man and his times from that amount, but consider how much wider the circle is if we include everything that Carl McIntire himself wrote but which has not yet been digitized. That includes multiple books as well as thousands of issues of a weekly paper, The Christian Beacon, which mostly exists only on microfilm today. We’re now up to more than 10x the amount of data, 10,000s of pages.
But wait, there is also all of the unpublished material in Carl McIntire’s archives at the Princeton Theological Seminary Library including correspondence and notes. That’s some 650 linear feet of material, each of which contains between 1,500 - 3,000 pages. We now have another 10x or perhaps 100x more material, comprising 100,000s or even millions of pages of untapped data.
But even then, that’s merely the “proven reserves” of McIntire-related data. Beyond those reserves is a larger theoretical reserve: think of all the collections of papers from his associates, family, and friends sitting in attics and church basements, functionally inaccessible. And there are also completely unproduced sources of data, like the potential oral histories that could be gleaned from his associates, family, and friends.
Yet only the tiniest slice of that data is available as training data because there is insufficient incentive to collect from that vast theoretical reserve. Now, theoretically even small amounts of financial incentive could be sufficient to expand the circle of usable data. A grant in the $10,000s would be enough to digitize his published works. $100,000s would be enough for the full archive. $1,000,000s would lead to oral history projects and lead to discoveries of currently untapped collections. (Do note that the value of the marginal dollar invested is particularly high at the lowest end of that spectrum.)
Now, think about the fact that I’m describing the vast, untapped potential of a single, relatively obscure historical figure. But that same iceberg-shaped potential — where the visible portion is a small fraction of the size — also exists for a functionally inexhaustible quantity of people and topics and sites of untapped data.
Of course, potential data isn’t training data. But data exhaustion is a problem that contains its own solution. Companies doing frontier model training are spending literal trillions on chips and electricity in order to squeeze out additional scaling gains and thus gain a competitive advantage in performance. And they’ve already demonstrated their willingness to spend hundreds of millions on slurping up sources of already digitized but gated data, like their deals for access to newspapers’ back catalogs. Given that trajectory, it makes sense that once they’ve exhausted those sources they’ll start spending money excavating other, novel sources of usable data, including the vast proven reserves of archival data.
As an aside, this is why I’ve never been more excited about the future of the historical profession. (Relatively) cheap digitization made it possible for researchers to access so many archives that were previously inaccessible because of the cost of physical travel and the difficulty of collecting the information. But as any archivist will tell you, grant-funding for digitizing archives is the bottleneck. And research historians make poor paying customers. But I anticipate that data exhaustion will raise the value of onboarding those archives, leading to a (relative) flood of investment in the medium term.
But thinking more broadly, you can imagine the ways that AI will create a large, priced, addressable market for novel information that goes beyond mere frontier model training. To go back to my example of Carl McIntire, imagine the near future version of me, an ambitious grad student interested in right-wing radio preachers but who is studying in the 2030s instead of the 2010s.
I would take my little $5,000 research grant, but instead of spending it to travel to physical archives and spend my time pouring over undigitized records, I would task an AI agent with contacting the AI agent for the university library where those records sat. My agent would have a budget to pay out bounties in exchange for the university AI agent arranging to hire a local researcher to digitize a set of those papers and transmit them to me.
In fact, someday the humanity / roboticity of that local researcher will be negotiable. After all, just as warehouses for consumer products are increasingly automated and run by a series of robotic systems, so too will be future historical archives. Just instead of picking a gee-gaw off the shelf to pack away in an Amazon box, it will be pulling down a box of historical documents off of a shelf to scan.
And that capital expenditure will create positive societal spillover. My little $5,000 investment will, yes, allow me to write my scholarly paper, but it will also reduce informational deadweight loss by digitizing those records, ultimately pushing out the circle of accessible data further than it was before. And that increases the value of that archive for the holder, in this case the university library, which subsequently makes it more likely to be contacted by other researchers’ AI agents. Every novel piece of information added increases the value of every other piece of information previously gathered. It’s pure, positive sum, informational network effects.
Ben Thompson has a nice illustration of the incentive structure that will guide this kind of informational commons in an agentic AI future. These agents will become the primary audience for digitized content, information that they will then use to incept the intent of their human users. It’s the shift from an attention economy to an intention economy, to use Shuwei Fang’s terminology.
While I’ve focused on the beneficial effects for one, small slice of data — that generated by and about an obscure historical figure — the value of all novel information will increase markedly. To return to the concentric circle idea, I’ve been talking about expanding out the addressable market to include offline data. But the returns to true creativity will increase as well, which means stuff that currently exists in the outermost circle that contains information that could be produced but never would have been otherwise.
This is very good news for folks who are very used to very dire discourse. After a period of painful disruption, artists, musicians, philosophers, and creatives of all kinds are ultimately going to become more valuable (and potentially better-compensated) than ever before.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.