RSS Amplifier

Light Drafts · Jul 2, 2026

Do We Want Machines to Learn the Way We Do?

0
Sign in to vote or save

Charles Eldering, Underground Producers Alliance · Light Drafts

AI music is yet another jaw-dropping consequence of AI, already capable of producing catchy, seemingly original-sounding, chart-topping songs. Yet it raises, in an exceptionally clear way, a fundamental issue with these master machines. They have ingested a tremendous amount of copyrighted material in "learning" their craft of music generation. Is this simply fair use, or is it a more nefarious indication of what is to come: machines that can replace musicians, artists, animators, and authors after having been trained on their own works?

In the next few months, my colleague Raz Mesinai and I will break down AI music generation, looking at it from various aspects, including the underlying technology, the fundamental copyright issues AI music raises, and the societal issues it raises with respect to our art and how we want it treated in this next technological era.

This week, we'll look at a fundamental issue at the heart of AI music generation: the use of massive amounts of copyrighted material to train these models. The litigation is already in full swing. On one side are the major-label suits, UMG v. Suno in the District of Massachusetts and UMG v. Udio in the Southern District of New York, both backed by the RIAA. On the other are the independent-artist class actions, Justice v. Suno, also in the District of Massachusetts, and Justice v. Udio, also in the Southern District of New York. A third front, the Woulard cases in the Northern District of Illinois, pushes past copyright into biometric-privacy and right-of-publicity claims over artists' voiceprints and identities, a theory we'll take up in a future installment.

Although there are claims directed at both the “input” as well as the “output” of these systems, we want to first focus on the input side and the unauthorized-copy claim: that copying songs into a training set, and the further copies made as the model learns, is reproduction of a copyrighted work without permission. And with the training of AI systems, it’s a massive unauthorized copying problem

And this isn't just a music problem; it's a machine-learning problem that happens to have reached music. The clearest illustration is The New York Times v. OpenAI and Microsoft, filed in late 2023. The first cause of action is exactly the claim the record labels make against Suno and Udio: that the defendants made millions of unauthorized copies of protected works, in this case, Times journalism, to train their models, without a license.

And here is the discomfort at the center of it. The problem is that using copyrighted material to learn is almost inherently part of all our learning. As humans, there's no question that our brains are formed through the use of tons of copyrighted material. A teenager learns guitar by playing along to copyrighted records until she can reproduce a solo note-for-note. That's literal copying, done thousands of times, as the price of getting good.

Human copying is often messy and seldom gives credit where it's due. Rock was a genuine innovation, but it borrowed heavily from the blues, with that borrowing at times crossing a line. Led Zeppelin's "The Lemon Song" (1969) was a direct reworking of "Killing Floor," the blues classic Howlin' Wolf recorded in 1964; after a 1972 infringement suit, Zeppelin settled, and Howlin' Wolf was added as a co-writer. And Hip-hop went further still, building itself directly out of fragments of older recordings and using tools like turntables as instruments. Yet from that act of copying it forged something wholly its own in the form of a genre that today dominates the U.S. streaming market.

Which brings us to the question I think this whole dispute is really about: are we prepared to let machines learn the way we let people learn, or do we want to hold them to a different standard? Do we need to make a distinction between kids in the South Bronx using whatever samples and tools they can lay their hands on to make music and a machine that is capable of generating a nearly infinite number of variations and combinations based on what it has been fed, which in theory could approach all of the recorded music and sounds known to man and made available to it?

It's tempting to say, of course, that learning for machines and learning for humans is the same; copying to learn is copying to learn. But the case for treating machines differently doesn't rest on what they do; it rests on the scale, fidelity, and purpose with which they do it. When a person internalizes a record, the copy is lossy, partial, and trapped in one skull. The machine makes a perfect digital copy of millions of works at once, holds them in a form that can be reproduced, and does it to build a commercial product that competes in the very market the originals occupy. Copyright law has always cared about exactly those things, namely how much was taken, and whether the use harms the market for the original, so a different outcome for machines need not be hypocrisy. It can simply be the same principles applied to a difference in degree so large it becomes a difference in kind.

Suno and Udio argue strongly that ingesting all of that copyrighted material is simply fair use. The doctrine, codified at 17 U.S.C. § 107, that lets you use a copyrighted work without permission for purposes the law treats as socially valuable. Courts weigh four factors: the purpose and character of the use (above all, whether it's commercial, and whether it "transforms" the original into something with a new meaning or function); the nature of the original work; how much of it was taken; and the effect on the market for the original. No single factor decides it, but the first and the fourth, transformation and market harm, tend to do most of the work, and they're where these cases will be won or lost.

Similarly, OpenAI and Microsoft argue that training their foundation models was also fair use, and their version is essentially identical: that a model reading the world's writing to learn the statistical patterns of language is a transformative use and no more an act of publication than a person reading the same material. They argue that the model's parameters are not a stored library of articles but something genuinely new. The only real difference from Suno's position is the medium: words instead of music. Which is exactly why the two fights are really one fight, and why a ruling in either could echo through both.

It's worth being clear about what rides on that one word, transformative. Fair use isn't a side issue in these cases; it's the whole defense. If a court rejects it, the copying that happens during training stops being a hard question and becomes plain infringement. The case turns from a debate into an accounting exercise, and three things follow. The companies are liable, and a court can order them to stop and potentially to destroy not just the copied files but the models trained on them. And, most consequentially, a precedent is set that training on copyrighted work requires a license. That last one is the real prize. A loss wouldn't just cost Suno and Udio money; it would turn training data from something you scrape into something you have to buy, for every AI company, in every medium.

Those of you in the music licensing business might scratch your head: isn’t this the classic need for a license? You duplicated my recording; you owe a licensing fee. This is how it worked before, at least when other businesses made copies without permission.

But in this brave new world of AI, how would the artists actually get paid? Two ways, and they reach very different people. The first is damages. Copyright runs per work infringed, and the owner can elect statutory damages, ranging from $750 to $30,000 per work, rising to $150,000 for willful infringement. The labels’ catalogs are registered, so they qualify for those numbers. But here’s the part that should give us pause: in the label cases, the money goes to the labels, who own the master recordings. How much of it ever reaches the musicians who made those recordings depends entirely on their contracts. That is exactly why independent artists, and even the musicians’ union, have filed their own suits. They watched the settlements roll in and asked who, exactly, was being paid.

The second way is the one actually happening: settlement and some new-fashioned music licensing. Rather than gamble on a verdict, the major labels have started cutting deals, Universal with Udio and then Warner, Warner with Suno, others following, that pair a lump sum with an ongoing royalty on every track the machines generate, sometimes plus a stake in the business. This is the future taking shape in real time: not a courtroom bloodbath, but the labels licensing the very catalogs that were copied, and collecting a toll going forward.

And how much is at stake? Suno's case began with a list of around 560 registered recordings. At the willful maximum, that's about $84 million. Then the labels fingerprinted Suno's training data and moved to expand the list to more than 61,000 works, which, at $150,000 apiece, is north of $9 billion. Suno is fighting the expansion, and for good reason: the company raised $400 million at a reported $5.4 billion valuation, which means a judgment at that scale isn't a fine, it's a death sentence. And that, more than any principle, is why both sides keep drifting toward the licensing table.

But even if the major labels all settle, where does that leave the independent artist? The settlements are the labels' to make as they own the masters, and they're cutting deals for catalogs they control. The musician recording on her own is not at that table. And when she tries to collect on her own, she hits a quieter obstacle: the law rewards paperwork. The copyright is hers the instant she records the song; she doesn't have to register anything to own it. But to actually sue, and above all to claim the eye-watering statutory damages that make these cases worth bringing, those $750-to-$150,000-per-work figures, the recording has to be registered with the Copyright Office, and registered before the infringement, not after. An artist who never filed can still go to court, but only after registering, and even then she's left with "actual damages": whatever she can prove the harm was really worth, which for one independent track is usually a number too small to interest a lawyer. That gap is the whole game. The registered major-label catalog collects in the billions; the unregistered independent collects, in practice, almost nothing.

And this is no small corner of the music world. Artists and labels operating outside the three major groups now account for close to half of the global recorded-music market. These are the people who seed new genres before the majors notice them, the soil our music actually grows from. If the law lets the catalog owners settle and leaves everyone else holding a copyright they can't afford to enforce, we'll have built a system that pays for music's past and quietly disinherits its future.

Suno and Udio have, in barely two years, marched the music world to a crossroads and handed us the choice. Down one road, we take copyright seriously: we set the fair-use arguments aside, decide that feeding the whole recorded history of music into a machine is a use that has to be paid for, and build the licensing, the registries, and the accounting to make that real. Down the other, we look the other way, seduced by the promise of magic machines that conjure more music than we could ever have imagined, and we let them train on the entire corpus of human sound for free, because stopping them feels like standing in the way of the future

Neither road is clean. Actually paying artists will be a mess with millions of works, tangled rights, independents who never registered, royalties measured in fractions of a cent. It will be slower, more expensive, and less magical than the alternative, and it may hold the technology back, at least at first. But it preserves something the other road quietly removes: the principle that human achievement is worth paying for, even when a machine can imitate it for nothing.

The machines are clever; that was never in question. Whether the people who taught them to sing are owed anything, that is the only question, and we are answering it right now.

No posts

Read the original on charleseldering.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.