Zendoric
← Back to the day · July 23, 2026

Anthropic will pay 1.5 billion for pirating books, but AI training is shielded as fair use

🕒 Published on Zendoric: July 23, 2026 · 00:24

A federal judge in San Francisco approved the largest copyright class-action settlement in history: 1.5 billion dollars for authors whose works Anthropic downloaded from pirate libraries to train Claude. The legal key is not the training itself —already declared fair use— but how the books were obtained.

By Zendoric · July 23, 2026.

Judge Araceli Martínez-Olguín, of the United States District Court in San Francisco, signed final approval on July 20 of a $1.5 billion settlement between Anthropic and a class of authors and publishers, the largest sum ever paid in a class-action copyright lawsuit, according to the court order itself as cited by Mashable. The lawsuit, led by Andrea Bartz and Kirk Wallace Johnson among others, accused the company of downloading books from the pirate libraries LibGen and PiLiMi to build Claude's training corpus.

The detail that really matters is what exactly is being paid for. The dispute did not revolve around whether training AI models on protected works is legal —that point had already been resolved by the same court in a prior ruling, describing it as "fair use" (the U.S. doctrine that allows the use of protected material without permission under certain conditions)— but about the method of acquisition: pirating copies instead of buying or licensing them. That distinction, illegal acquisition versus legitimate training, is what has been established as precedent.

The distribution figures are concrete. Authors and publishers whose works appear on Anthropic's so-called "Works List" can claim about $3,000 per book, roughly four times the usual minimum in copyright infringement cases in the U.S., according to the settlement itself. More than 91% of eligible works, over 440,000 books, had already been claimed even before final approval. Anthropic, moreover, will have to delete the pirated files it downloaded.

There are two nuances that limit the scope of the settlement and that are worth keeping in mind. First, the judge was explicit: the deal does not release Anthropic from future lawsuits over what the chatbot generates, nor from claims "based on the output of the AI models." In other words, the chapter of how the training library was built is closed, but the potentially more complex front of whether Claude's responses reproduce or imitate protected works remains open. Second, the court rejected the 54 objections filed by class members and third parties, including requests to expand the list of covered works, require source attribution in the model's responses or even delete the models trained on that material. The court held that those requests went beyond what this specific lawsuit could resolve.

Our reading is that this settlement works as a kind of regularization fee for the foundational stage of generative AI, when several companies built their training corpora in a hurry and with little scrutiny of the data's provenance. $1.5 billion is a considerable figure in absolute terms, but it is manageable for a company that in 2026 is negotiating valuations well above $100 billion; it is not a settlement that threatens Anthropic's viability, and that raises the question of whether the financial penalty is enough as a real deterrent for the rest of the sector, or whether it simply becomes the cost of doing business during the data-accumulation phase.

What does change, and this is what matters for the rest of the industry, is that from now on the provenance of training data goes from being a gray area to a quantified legal risk with a market price: around $3,000 per pirated work that ends up in a commercial model. That pushes hard toward licensing agreements with publishers and authors, something various companies in the sector are already signing, and undercuts the argument that "it'll all get sorted out in court anyway." In parallel, the front of the outputs, what the model does with what it has learned, remains alive and will probably be the setting for the next big lawsuits against generative AI.

This connects with something we have been pointing out at Zendoric: the problem of AI is not a single legal battle, but a succession of them, each more nuanced than the last. First it was debated whether training was legal (now resolved, in general, in the sector's favor). Now the debate is how the data was obtained (here Anthropic has paid dearly for taking the pirate shortcut). The next chapter, predictably, will be whether what the model generates infringes on others' rights, a far more slippery terrain because there is no list of books to audit, but rather millions of responses generated in real time. The abundance that AI promises in the long term does not exempt anyone from building, along the way, a fair compensation framework for those who created the knowledge on which it is trained; this settlement is a first imperfect but real attempt to put a price on that past harm.

🔗 Related on Zendoric

Sources & references