A quiet revolution is sweeping through the realm of popular music. It begins not in a studio filled with instruments, nor in the imagination of a big-name producer, but in the rarefied field of artificial intelligence. There, digital pulses are believed to be ready —or almost ready— to unleash the next Hot 100 songs. This transformation may not be obvious, but has already sparked a vigorous online debate, blending anticipation with unease. Renowned jazz critic and music historian Ted Gioia's recent video interview with producer Rick Beato underscores the intensity of this discourse. Amassing over a million views in less than 24 hours, the clip features a subtle critique of AI-generated music on both economic and cultural grounds. Ultimately, though, even if Gioia’s assessment cannot be easily dismissed as yet another case of “senior citizen yelling at cloud”, his cautious stance misses the mark in locating the technology's potential to foster human creativity.
The following article draws on the sense of wonder I experienced during a short period of intense use of the technology: namely, three weeks of testing two generative AI music models found online, Suno and Udio. As the phenomenon is still very new, any conclusions found herein should be treated with a healthy dose of incredulity. Yet I still believe it is of vital importance —even at this early stage of development— to engage in the discussion from a user perspective.
[Listen to my published playlists on Udio.com: Anachronisms, Barbarella Jazz]
Gioia’s critique of AI music hinges on a contentious practice: Spotify's alleged surreptitious inclusion of AI-generated tracks into its official playlists as a cost-cutting strategy. The critic argues that this practice not only threatens musicians' livelihoods, but also degrades the listening experience. His concerns highlight two primary risks: cultural homogenization and the destabilization of the (already fragile) musical ecosystem.
A third, minor aspect of this critique focuses on the deceptive behavior of some of the highest profile companies involved with AI, as shown in Spotify’s allegedly routine non-disclosure of specific instances of AI use, or in Sports Illustrated’ well-documented attempt to deceive its readers by publishing AI-generated articles and then obscuring the fact by posting fake bios of the authors. “If AI is really so wonderful, so exciting, if it’s so tremendous, why did they have to keep it a secret?”, Gioia asks.
Yet criticism of AI music extends well beyond Gioia's sensible arguments, often fixating on the music’s perceived inferiority to traditional human efforts. Critics mock the obvious flaws in many current AI compositions, namely: unnatural melodic shifts, structural inconsistencies, low resolution and occasional glitches. While these criticisms are certainly fair, they often seem to stem from little actual user experience. More importantly, the perception of some of these flaws is likely to depend heavily on implied normative ideas about music. In this framework, everything that does not conform to the rule (be it a certain understanding of pitch relations, such as tonality, or other generally accepted features of musical syntax) is set to the side. This position has obvious risks: chief among them is a crippling conservatism.
Furthermore, these same critics overlook two important aspects of AI-generated cultural products. First: the rapid evolution of AI technology, which promises significant improvements in quality and reliability. Second: differences in quality between generative AI models which are apparent even today. Some AI services are indeed more advanced than others, with Udio being currently the most impressive one by far.
Founded by a team of former researchers from Google DeepMind and launched just about a month ago (on April 10, 2024), Udio will probably surprise even hard-line detractors of AI-generated music with its uncanny ability to mimic the human voice and write idiomatic lyrics in many musical genres. I was delighted by the punchiness of its vocal delivery in my own hip hop experiments, or by the smoky aroma of its saxophone solos. But it is in the field of contemporary instrumental composition that Udio surprised me the most. The following ideas draw on this experience.
In any case, one should take into account that the very imperfections mocked by today’s critics will probably endure in some form for the rest of the decade. I argue that this “artificial cognitive detritus” will present challenges for users, who will have to deal with unwanted “uncanny valley” effects. That is, users will either dismiss these effects as tolerable errors, or incorporate glitches and other non-intentional sounds into a kind of intentional musical discourse.
It is quite possible that these effects will become endearing in their own right in some form of future nostalgia, or serve as a new frontier for aesthetic exploration. Consider, for example, how trip hop, glitch and vaporwave artists have historically embraced technological imperfections, turning phenomenoms such as the surface noise of a vinyl recording into expressive features as deliberate as the thick, gestural brush strokes of an impressionist painter. Similarly, AI-generated music’s quirks are bound to inspire new aesthetic practices that celebrate its distinct characteristics.
This will in turn foster a critical evaluation of the specificity of generative AI as a medium. Indeed, it can already be argued that its ability to mimic and blend artistic styles is what defines it. This realization will probably guide AI-enhanced creativity for the next decade; we should therefore prepare for what might be called postmodernism 2.
However, generative AI technology’s capacity to engage in the production of speculative sound or imagery, by which I mean its capacity to generate output with singular expressive qualities not directly traceable to its concrete data set, should not be dismissed at this stage.
A critical distinction must be made between AI as a creator and AI as an enhancer of human creativity. While initial resistance to AI-generated music is understandable, this skepticism is likely to wane as the role of human ingenuity in AI-generated music becomes more apparent. In my view, the promise of this technology lies not in its ability to make AI music indistinguishable from traditional human-made music, but in the integration of AI capabilities into services designed to amplify human creativity.
We are thus confronted with two models of musical consumption:
Model A: Passive Consumption. Currently, AI-generated music predominantly caters to passive consumers. Platforms like Spotify employ algorithms to tailor playlists, embedding AI subtly into our daily listening habits. According to Ted Gioia, these measures are purely driven by commercial motives to cut costs by avoiding royalties. In this way, the “Passive Consumption” model undermines both the value of musicians’ work and that of the listeners’ attention.
Model B: Active Creation. In contrast, services like Suno and Udio envision a future where AI fosters active engagement, transforming listeners into creators. While these claims should be confronted critically, it is certainly true that generative AI music solutions of the sort afford a very different, active experience, by allowing users to obtain musical results based on simple text prompts. The companies behind these models assure us that AI will democratize music creation by enabling even those without formal training to produce and download tracks. At any rate, by empowering users to create, this approach is already redefining the musical experience for Zoomers as a dynamic, participatory venture.
Professional musicians often fear AI might render their skills obsolete. However, this perspective overlooks AI's potential as an empowering tool. Models like Udio and Suno are conceived with a simple goal in mind: that of assisting musically untrained users in compositional tasks. But traditional musical training and literacy still offer a competitive edge in this new context. Users now act as curators, producers or even players which clear the way toward desired outcomes (more on this later). Accordingly, a certain musical culture is expected of them, as well as strategic thinking and musical taste. Being musically trained is a way to get there. But the practice of steering generative AI models in “the right direction” depends to an even larger extent on the ability of users to craft effective prompts and engage creatively with AI systems. Musicians who wish to experiment with the technology cannot escape this. They can, however, easily learn how to master the medium, while applying political pressure to make sure that the goal of the technology is kept in check: AI must be used for enhancement, not for replacement.
Creating music with AI can be likened to the strategic role of a curling player who carefully prepares the path to a desired outcome. Users guide AI to avoid pitfalls such as structural inconsistency, abrupt endings, unpleasing melodic or harmonic shifts, etc. Similarly, the metaphor of a chess player is apt, as composing music with AI involves both short-term tactical decisions and long-term strategic planning.
To fully understand what a service like Udio can offer, direct usage is essential. However, here is an outline of its functionality. Currently, in its free version, Udio generates two 30-second audio clips each time you use it (in the paying versions it’s three or four clips, depending on the subscription tier). This setup provides material for the user to select and refine. A few aspects to consider:
Lyrics: Udio can write lyrics based on a prompt, but it can also compose music for lyrics fed by users. At this level it is easy to see how user creativity can be enhanced by the model.
Stylistic Labels: By default, Udio assigns stylistic labels to user prompts (e.g., "expressionistic," "piano solo") before generating sound. The model does it allegedly to optimize outcomes, as not all users understand the technology well enough to write effective prompts. Nevertheless, and crucially, it also features a “manual mode” aimed at experienced users. This allows them to talk to the machine directly, so to speak.
Prompt Sensitivity: Results vary significantly based on how prompts are written: for example, prompts like "in the style of" and "in a hybrid style between X and Y" carry special power, mainly when used in combination with manual mode.
Complementarity of generative AI models: It is of course possible to combine the use of several generative AI technologies to produce music. For example, one can use ChatGPT to write lyrics or to produce prompts that will be read by Udio. Complex situations can arise from these practices. For example, using ChatGPT to write prompts for Udio is akin to asking the model to reverse engineer another AI model in order to maximize what the first model interprets as the user’s desired outcomes.
Experimental prompt writing: Some experiments with prompt writing have already been shown to render interesting results with Udio. For example, using a fragment of a poem as a prompt. This can be used to set an atmosphere, in combination with commands such as “in the style of” or “in a hybrid style between X and Y”.
Once the first batch of material has been created (always two 30-second audio clips at once, as a minimum), there are two main operations that can be performed:
Remix: The material is rearranged with different instrumentation, voicings or style. On the one hand, this suggests Udio is not only capable of statistical aggregation, but that it understand (at least in a primitive way) instrumentation, voicing, harmonic styles, etc. Interestingly, the amount of variability can be manually adjusted: this means that a generally pleasing outcome can be modified just enough to maintain whatever characteristic makes it work, while replacing unwanted features.
Extend: Udio produces two new 30-second clips that build on the previous one. This means that the model is able to generate structural continuity between the clips.
Dismiss Section: Users can now erase a section of the original clip before launching the extension.
Memory Adjustment: Users can now calibrate how much of the clip Udio takes into consideration (or, conversely, how much it forgets) when generating an extension. This provides a measure of control over large-scale structural consistency vs. local variability.
Despite the simplicity of these features, creating a short three-minute piece might require dozens of iterations of material and many choices as well. Similarly, let’s note that pieces by different users that originated from the same material can be vastly different in terms of length, structure or style.
Currently Udio already excels at:
Imitating Composers: In fact the model can imitate some composers better than others, probably on account of its training materials (which remain undisclosed). For example, Udio imitates tonal music better than atonal music. It can do a lovely Prokofiev.
Mixing Styles: Combining styles, preferably of compatible composers. Some composers, like Bach, can be useful for general purpose (i.e. providing a canonic texture).
With this in mind, I will posit that the most important insight for creating music with generative AI models at this juncture is that complex operations involving these capabilities can be conducted in multiple steps. For example, users might try this set of prompts:
“An instrumental composition for piano in the style of Bach.”
Rewrite the prompt as: "An instrumental composition for piano trio in a hybrid style between Brahms and Schoenberg,” using the Remix function.
Rewrite the prompt as: “An instrumental composition for violin and piano in the style of Prokofiev,” using the Extension function.
This hyper remixing capacity is set to significantly impact culture in the foreseeable future (i.e. TV series, meme production, etc.). It is also potentially important for aesthetic discussions on generative AI.
There is a fundamental tension between two approaches for creating music with AI: 1) The art of prompting, akin to conceptual art; and 2) The more traditional path of selecting material (deemed interesting in terms of harmony, melody, rhythm, texture, etc.) and developing it.
I prefer the former for poetic reasons, but the latter has been more useful in my own (admittedly conservative) musical practice. The second approach requires at least two or three consecutive steps to yield interesting results. The key function here is Remix, as used in combination with the practice of rewriting the prompt at each step. For example, the user might wish to start making music with a textural idea: a canonical passage reminiscent of Bach; and then remix the clip to introduce a modern harmonic palette. A third step might be to transform this piece into a composition for jazz quartet, which will eventually evolve into a cello concerto. Crucially, the original 30-second clip can be used to generate several sequences which can serve as thematically linked movements of a larger work.
Obviously, at this point the discussion about human creativity in the age of generative AI has shifted completely. It’s no longer about AI producing automatic results with minimal user input. Instead, it involves a distinct form of human creativity, akin to both the activity of playing a game, and to traditional artistic endeavors. Indeed, users can be seen as chess players moving not on chessboards, but in the space of possibilities of the generative AI model. This is only a metaphor, of course; but I believe that with enough experience, users can actually begin to understand this space; that is, to recognize its singularities: small regions where an unusual convergence of factors conveys especially interesting or weird results.
The user might know, for example, that the model imitates Russian composers well, that it “understands” tonal music better, that it can “write” canons, and that it yields better results when mixing generally compatible styles, instead of highly contrasting combinations. I posit that, eventually, users will start to consciously use those singular points rather than follow ideas external to the medium. This will pave the way for a radical aesthetics of generative AI.
None of this will be automatic. These cultural changes will rely on the musical experience of users, their taste and their rational analysis of the model’s capabilities. As I said, this is where the aesthetics of generative AI will begin to flourish. They will focus on the specificity of the machine: that is, its combinatorial nature and the weird behaviors that can be unleashed when the model is led to produce results in certain specific conditions.
I posit that three features will be needed to unlock generative AI’s full musical potential:
Bridge Clips: Seamlessly connect two clips.
Add File: Allow users to input their own audio files.
Export as Sheet Music: Convert AI-generated music into traditional notation, bridging the gap between AI content and live performance.
These features will likely materialize in the near future, despite the obvious legal challenges they pose. The most important one is Add File. Indeed, as users feed personal audio and video libraries into AI systems, the production of culture will transform drastically.
One aspect of such evolution will be the radical customization of entertainment. For example, porn will be made using specific user libraries of videos as training materials. This will ensure that the user’s sensibility and taste “colors” the outcome. In the same way, users of AI music services will be able to create music by combining recordings of their favorite bands, not by pairs but in combinations of hundreds of specifically chosen recordings at a time. Users will also be able to download other users' setups, or to acquire “stylistic add-ons”: libraries of files that purport to represent any given musical style.
Clearly, AI is set to reshape the entertainment industry. But what does this entail exactly? On the one hand, mainstream musical consumption will change in order to accommodate more active ways of engaging with music enhanced by AI models. But the adoption of AI generated imagery and sound by vast populations of users will also nudge our collective understanding of musical syntax and other formal aspects of music in ways compatible with the development of this technology. In other words, human musical creativity and AI generative technologies will co-evolve.
The most pressing question now is how the large players in the industry envision their new role, and if they can be pressured to comply with desired future outcomes we might have as a society. Will these companies focus their efforts on the conservative paradigm of automation? Will they embrace the idea of treating customers as creators instead? The opening of new opportunities for professional and amateur musicians alike depends on these decisions, as well as the creation of a more interactive and inclusive musical culture.
As we stand on the brink of this new era, the music that awaits us promises to be disruptive rather than sedative; wild rather than predictable; and engaging in weird new ways. The quiet revolution of AI in music is not about machines making music: it is about the rediscovery of our own culture through the lens of generative technologies. It is also about collaboration at a massive scale: between users scattered around the globe, between different generations of musicians, and even between species (depending on how you see AI). This cultural transition will prove difficult for many reasons, not least of all legal. But it is bound to happen: it will make fortunes and redefine the very essence of music in the 21st century.
Note: This article was made with the assistance of ChatGPT. I first manually outlined the main ideas; then reorganized the material with the help of the AI model; then asked the model to write an article based on this material, “in the style of The New Yorker”; and then, finally, to rewrite it “in the style of The Paris Review”. The result was not outstading by any stretch of the imagination, but it was enough for a first draft. The last, most important step was to make all kinds of corrections and additions manually. The point I am getting at is this: formally trained musicians will be able to engage in the same dynamic in the near future with mainstream generative AI music models. That is, to combine material generated with AI with their own precisely notated input. This should be exciting news.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.