Who Owns the Culture AI Was Trained On? Training Is Legal, Pirating Is Not — the Line Deciding the Trial of the Century
Anthropic will pay $1.5 billion, the largest copyright settlement in US history — but not for training: for pirating. While the New York Times and OpenAI trade accusations over hidden evidence and the White House points at China's Moonshot for 'distilling' Anthropic's model, courts on three continents are drawing the same line. We map every key lawsuit, run the impossible math of licensing everything, and tackle the deeper question: if a human can read the whole library and create, why can't a machine?
🎬 Our Short
THE THESIS. The 'trial of the century' over AI data is not one trial: it is dozens of lawsuits across three continents. But a surprisingly clear line is emerging from the noise. Courts in the US, the UK and Germany are accepting that training an AI on copyrighted works can be legal. What they punish is something else: how the data was obtained (piracy is proving ruinously expensive) and what the model spits back out (regurgitating song lyrics or full articles is not covered). Our thesis: print-era copyright will not stop AI training, but training will not be free either. The outcome will be a hybrid system — legal exceptions, a licensing market and transparency obligations — being written right now, one ruling at a time. Who gets paid, and how much, will define the social contract between AI and the culture that feeds it.
THE LITIGATION MAP. The cases must be kept apart, because they say different things. The most expensive one is closed: Bartz v. Anthropic. In July 2026 a federal judge gave final approval to the $1.5 billion settlement — the largest in US copyright history — under which Anthropic will pay roughly $3,000 per book to the authors of some 482,000 works downloaded from pirate libraries such as LibGen. Mind the nuance, because it is the key to everything: Judge William Alsup had already ruled in 2025 that training on legally purchased books is fair use — the US doctrine that allows 'transformative' uses of others' works without permission — and that the illegal part was pirating them. Anthropic did not pay for training; it paid for how it got the books. And the settlement sets no binding precedent: it is a deal, not a ruling.
The New York Times case against OpenAI and Microsoft, by contrast, is alive and has turned into trench warfare. It sits in discovery (the pre-trial exchange of evidence) with no trial date. In November 2025 a judge ordered OpenAI to hand the Times a sample of 20 million ChatGPT conversations, a decision affirmed in January 2026. In July 2026 the publishers requested sanctions, accusing OpenAI of withholding evidence; the judge has not ruled, and OpenAI rejects the accusations and maintains its use is lawful. This is the case that could settle whether training on paywalled journalism is fair use. Getty v. Stability AI in the UK ended in November 2025 with a different lesson: Getty lost almost everything — not because training is legal there, but on territoriality grounds, since the training happened outside the UK and the court could not judge it. Only a very limited trademark claim survived (Getty watermarks showing up in generated images). Translation: labs can choose where they train, which turns jurisdiction into strategy. In music, the script was different again: the labels sued Suno and Udio… and ended up signing. Universal settled with Udio in October 2025 and Warner with Suno in November; only Sony remains in court. Forbes summed it up as 'launch, train, settle': infringe first, license later — and it paid off.
THE EMERGING DOCTRINE. In the US, three rulings draw the map. Alsup (Anthropic): training is transformative; pirating is not. Chhabria (Meta, June 2025): Meta won, but the judge warned in his own opinion that the authors might have prevailed with a different argument — 'market dilution', the idea that an AI trained on your books can flood the market with substitutes and devalue your work even without copying a single sentence. That is the door plaintiffs will now try to push open. And Thomson Reuters v. Ross (February 2025) marks the boundary from the other side: when the AI competes directly with the work it copied (a legal search tool trained on Westlaw's case summaries), there is no fair use; its appeal, argued in June 2026, is the first time a US appeals court examines fair use and AI. Europe plays by different rules. There is no fair use: there is a 'text and data mining' (TDM) exception in the 2019 copyright directive, which allows extracting information from lawfully accessible works unless the rightsholder opts out through machine-readable means. The problem: no recognized standard exists for that opt-out, so in practice it is a veto many cannot exercise. Two German rulings in 2025 confirmed the picture with nuance: LAION won (data collection was covered), while OpenAI lost to GEMA, the music rights society — the Munich court accepted that the TDM exception covers training in general, but sanctioned ChatGPT for memorizing and reproducing song lyrics. Same line as in the US: the problem is not learning, it is regurgitating. And the EU AI Act adds the missing piece: it requires publishing a 'sufficiently detailed summary' of training content. Without transparency, neither opt-outs nor licenses can work.
THE MATH THAT DOESN'T ADD UP. What if everything were licensed? Here is the number that dismantles the easy answer. The deals are real and move real money: News Corp–OpenAI, over $250 million across five years, the largest disclosed; Reddit–Google, about $60 million a year; Amazon–New York Times, $20–25 million a year; Axel Springer–OpenAI, about $13 million a year; the Financial Times, $5–10 million a year, per Quartz's tally. Reddit disclosed $203 million in data contracts in its IPO prospectus. But the arithmetic is merciless: Llama 3 was trained on 15 trillion tokens (the text fragments a model learns from), and it would take the New York Times about 316,000 years to produce that volume, by ProMarket's calculation. The marginal contribution of any single work to a model is nearly zero, so negotiating work by work is unworkable on transaction costs alone. The uncomfortable but honest conclusion: licensing everything is impossible. The deals compensate large, concentrated rightsholders — publishers, labels, platforms — and leave the individual creator out. Which is why the realistic path is neither 'license everything' nor 'except everything', but exception plus collective licensing plus mandatory transparency.
WHEN EVERYONE ACCUSES EVERYONE. The final twist: those accused of drinking from other people's culture now claim others are drinking from theirs. On July 22, 2026, Michael Kratsios, the White House's technology policy director, accused China's Moonshot AI of building its Kimi K3 model through covert, industrial-scale 'distillation' of Fable, Anthropic's model — distillation means training one model to imitate the outputs of a more powerful one — and of accessing banned Nvidia chips through servers in Thailand; the Treasury threatened sanctions. These are accusations, not proven facts: independent researchers doubt a model released days after Fable's launch could have been distilled that fast. Anthropic, meanwhile, is itself being sued by Reddit — a mass-scraping case, with over 100,000 unauthorized accesses according to Reddit, sent back to California state court in April 2026 — for taking without permission exactly what Reddit charges Google $60 million a year for. The irony is complete: the labs that invoke fair use to train are locking down their own outputs against others' distillation. Data is no longer an input: it is a strategic, even geopolitical, asset.
OUR READ. The philosophical question remains: if a human can read the entire library and create from it, why can't a machine? Our answer: because scale changes the species of the act. A human reads hundreds of books in a lifetime and competes with other humans; a model ingests millions and can replace, at a stroke, the very market it learned from. The 'diligent reader' analogy is useful but insufficient — which is why Chhabria's 'market dilution' strikes us as the most important legal concept of this cycle: it does not ask whether the machine copied, it asks whether it impoverishes the ecosystem it feeds on. That said, we reject both extremes. 'AI is theft' ignores that courts are validating machine learning as transformative use — and that without that path, only incumbents (or those who ignore the law, usually outside the West) could ever train. 'Information wants to be free' ignores that without revenue there will be no journalism or books left to train the next generation of models: an AI that exhausts its own source poisons itself. In the short term, the harm is real and unequal: Reddit and News Corp get paid, the essayist and the illustrator do not — which demands collective compensation mechanisms like those that already exist in music. In the long term, we keep our underlying optimism: an AI legitimately trained on human knowledge is the best tool we have ever had to cure, teach and generate abundance. The goal is not to stop it — it is to make the culture that made it possible a partner, not a quarry. Anthropic's $1.5 billion is not the end of the trial of the century; it is the first invoice of a social contract we are still negotiating.
Sources & references
- Fortune — Anthropic to pay authors $1.5 billion over pirated books used to train Claude
- Authors Guild — What Authors Need to Know About the $1.5 Billion Anthropic Settlement
- Variety — NYT and Other News Outlets Accuse OpenAI of Lying in Discovery, Demand Sanctions
- DLA Piper — Getty Images v Stability AI: the UK High Court decision
- Skadden — Fair Use and AI Training: Kadrey v. Meta and Bartz v. Anthropic
- Davis Wright Tremaine — Thomson Reuters v. Ross Intelligence: Copyright, Fair Use, and AI
- Taylor Wessing — GEMA, Getty and beyond: AI and copyright litigation in the EU and the UK
- Quartz — The price of AI training data, from $5M to $250M


