Zendoric
← Back to the day · September 2, 2026

AI that composes music with diffusion models puts authorship and human creativity in check

🕒 Published on Zendoric: September 2, 2026 · 08:27

✨ AI-generated · how it's made

An MIT Technology Review article explores how diffusion models — the same kind of technology behind the AI image and video boom — are now arriving in music, a field the author considers especially vulnerable because music is deeply tied to emotions, memories and…

An article in MIT Technology Review explores how diffusion models —the same type of technology behind the boom in AI-generated images and video— are now reaching music, a field the author considers especially vulnerable because music is deeply tied to people's emotions, memories and social lives. The piece opens with a historical recap: in 1956, at the famous Dartmouth conference where the term "artificial intelligence" was coined, the organizers already included among the field's challenges the building of machines capable of creativity and originality, and they proposed a striking formula: the difference between creative thinking and merely competent thinking would lie in injecting some randomness, guided by intuition. Almost seventy years later, diffusion models follow that recipe almost to the letter.

The article explains in accessible terms how these models work, using the example of an image of an elephant: they are trained by progressively adding random noise to real photos until they become pure static, while the model statistically learns to reverse that process, that is, to "denoise" the image step by step. Once trained, generating new content means starting from pure noise and, guided by an input text (the prompt), removing noise until a coherent image emerges, without the model really "knowing" what it is depicting. In music, the mechanism is analogous but applied to waveforms or spectrograms, that is, visual representations of sound signals: the model learns from millions of song fragments labeled with descriptions, and then generates new waveforms —and therefore new complete songs, with all their instruments and vocals at once, not instrument by instrument— from random noise shaped by the user's prompt.

The report names the players leading this race: Udio, based in New York and co-founded by David Ding (previously a senior engineer on image and video diffusion models at Google DeepMind), with Andrew Sanchez as chief operating officer, and Suno, based in Cambridge, Massachusetts. According to the figures cited, Suno says it has more than 12 million users and raised a $125 million funding round in May 2024, in addition to establishing a collaboration with producer Timbaland. Udio, for its part, raised a $10 million seed round in April 2024 with investors such as Andreessen Horowitz and musicians including Will.i.am and Common. Both companies want people with no musical training to be able to generate complete songs, and Suno already has "artist" pages that in fact belong to users skilled at writing prompts, some with notable followings, even accompanied by "artist" images that are also AI-generated.

The article notes that both companies face lawsuits from the major record labels —including Universal and Sony— filed in June 2024 and still ongoing. The labels allege that the models were trained on copyrighted music "on an almost unimaginable scale" and that they generate songs imitating qualities of real human recordings; the suit against Suno specifically mentions an ABBA-style song called "Prancing Queen" as an example. Suno did not respond to questions for the report, but its chief executive, Mikey Shulman, said on the company's blog in August that they train on music available on the open internet, which "does in fact contain copyrighted materials," while arguing that "learning is not infringing." Udio declined to comment on the ongoing litigation, although it stated at the time that its model includes filters to avoid reproducing protected works or the voices of specific artists. The text adds that in January the U.S. Copyright Office published guidance stating that AI-generated works can be registered if they include substantial human input, and that a month later a New York artist obtained what may be the first copyright for a visual work created with AI assistance, with music flagged as the likely next frontier. It also notes that YouTube has reportedly negotiated music licenses for AI training with major labels, and that Meta expanded its agreements with Universal Music Group, suggesting that such licenses could become widespread regardless of how the lawsuits end.

Beyond the legal plane, the article digs into what makes AI-generated music good —or not— pointing to three factors: the training data, the architecture of the diffusion model itself and the quality of the prompt. Neither company has revealed what music makes up its training set, although that information will foreseeably come to light in the litigation. Ding explains that song labeling —from basic genre descriptions to nuances such as whether a piece is "melancholic" or technical details like a chord progression— is an active area of research at Udio, which combines annotators with formal musical training and enthusiasts with more informal vocabulary in order to serve a broad user base. The text also stresses that these models need to be continuously fed new human-made music so as not to become "frozen" in time, though it mentions that, following a trend already explored in other areas of AI, in the future they could end up being trained on their own outputs.

One curious aspect the article highlights is these systems' deliberate randomness: because they start from a noise sample, giving the same prompt to the same model produces a different song every time, and companies such as Udio inject additional randomness by slightly distorting the waveform at each step to achieve imperfections that come across as more "interesting" or "real." Sanchez comments that this non-deterministic nature surprises even the artists who work with Udio, who are used to software always giving the same answer to the same input; according to him, when they ask why the system behaves this way, the company's honest answer is that "we don't really know," which illustrates that the generative era requires accepting messier, more inscrutable programs.

The article devotes much of its analysis to contrasting these processes with what cognitive science knows about human creativity. It revisits the classic definition formulated in 1953 by psychologist Morris Stein —a creative work must be novel and useful (or "satisfying"), and some add that it must also be surprising— and how, since the 1990s, techniques such as functional magnetic resonance imaging have made it possible to study the neural mechanisms of creativity. It quotes Roger Beaty, of Penn State's Cognitive Neuroscience of Creativity Laboratory, who describes the "associative theory" of creativity: more creative people would connect semantically very distant concepts, a phenomenon research has linked to semantic memory and to attention that is more "permeable" to information that is in principle irrelevant. Beaty and other researchers, among them Dean Keith Simonton, insist that creativity does not reside in a single brain area —"nothing in the brain produces creativity the way a gland secretes a hormone"— but in dispersed networks of activity: some for generating ideas by association, others for identifying the most promising ones and others for evaluating and modifying them. The article also mentions a February study led by researchers at Harvard Medical School suggesting that creativity may even involve suppressing certain brain networks, such as those associated with self-censorship.

Against this, the report gives voice to defenders of AI-generated music who maintain that human creativity also rests on a kind of "statistics" accumulated through experience: composer Anthony Brandt, a professor of music at Rice University, argues in a recent study that both humans and large language models use past experiences to evaluate future scenarios and make better decisions, and he recalls that much human art, especially in music, is itself "borrowed," which often ends up in lawsuits over copying or unauthorized samples. Even so, Brandt himself draws a clear line where, in his view, human creativity clearly surpasses the artificial: what he calls "amplifying the anomaly." As he explains, AI models operate through statistical sampling, that is, reducing errors and seeking probable patterns, whereas humans are drawn precisely to oddities, which in human creation do not remain isolated events but come to permeate the entire work. In the available excerpt, the article begins to illustrate that idea with an example of a compositional decision by Beethoven, but the text breaks off there, so it is not possible to reconstruct precisely what that specific example was.

Overall, the report does not settle whether what these models produce is genuine creation or merely a sophisticated replica of their training data, and it leaves open the question of whether that same ambiguity might not also apply, to some extent, to human creativity. What it does make clear is that, with or without a court ruling on the lawsuits against Suno and Udio, AI-generated music is already seeping into playlists, soundtracks and streaming platforms, often without the listener ever knowing whether there is a person or a machine behind it, something the author himself describes as unsettling after hearing an AI-generated song he found genuinely good.

🔗 Related on Zendoric

Sources & references