A federal judge approved Anthropic's $1.5 billion settlement with authors over pirated books used to train its Claude AI models.
A federal judge approved Anthropic's $1.5 billion settlement with authors over pirated books used to train its Claude AI models.

A federal judge in San Francisco approved Anthropic's $1.5 billion settlement with authors who accused the company of using pirated copies of more than 480,000 books to train its Claude chatbot, closing the first major US copyright case against an AI developer to end in a payout.
"This is the first of its kind in the AI era and sets a precedent requiring AI companies to pay copyright owners," said Justin Nelson, class counsel for the authors.
The settlement covers more than 480,000 works at roughly $3,000 per book. Claims came in for more than 90 percent of eligible works — over 440,000 — a rate far above the roughly 10 percent typical of class actions, class counsel told the court. Anthropic will pay in four installments through September 2027, with an initial $300 million already in an interest-bearing escrow account.
The payout grew out of a split ruling by now-retired Judge William Alsup, who found that training Claude on books was "exceedingly transformative" and therefore fair use — but that downloading and storing more than seven million pirated books from shadow libraries Library Genesis and Pirate Library Mirror was not. Because US copyright law allows damages of up to $150,000 per work for willful infringement, a library that size left Anthropic facing theoretical liability in the hundreds of billions of dollars. The company settled rather than test that figure at a trial set for December 2025.
A $3,000-Per-Book Price Tag Reshapes AI Data Economics
The settlement values each work at roughly $3,000, though objectors argued the total is too small to deter a company of Anthropic's size. Anthropic, backed by Amazon and Alphabet, admitted no wrongdoing. The company must also destroy the original files it pulled from the two shadow libraries, along with any copies made from them. Class counsel trimmed its fee request to about 12.5 percent of the fund, below the roughly 30 percent common in such cases, but objectors still argued the lawyers take too large a cut and that the deal wrongly shuts out works not registered in the United States. A handful of authors and publishers opted out to pursue their own claims, which remain live.
While a settlement binds only the parties and sets no legal precedent that future courts must follow, it establishes a market price for unlicensed training data that will reverberate across the industry. Defending one of these cases to trial can run into the millions, and statutory damages make the downside catastrophic. For developers building large language models, the operational takeaway is clear: those who license or lawfully buy training material face less exposure than those who reach for free, infringing copies. The distinction also rewards companies that scrub questionable data provenance from their training pipelines.
What the Settlement Means for OpenAI, Meta and Google
Anthropic's is the first of dozens of AI copyright suits — including cases against OpenAI, Microsoft and Meta — to settle at this scale. For the rest of the industry, the narrow release covers only how Anthropic acquired the books in the past. It grants no license to keep training on them, and it releases nothing about what Claude actually generates. Authors keep the right to sue over model outputs and over any book left off the settlement list.
The unresolved question of whether scraping the open web constitutes fair use will determine the next phase of the copyright fight over AI training. For now, the settlement creates a template for how AI companies might resolve similar claims: pay for past data use, destroy infringing copies, and negotiate licenses going forward.
For AI developers, the settlement raises the cost of training data procurement. Anthropic's $1.5 billion payout — equivalent to roughly 15 percent of the $10 billion the company has raised from investors including Amazon and Alphabet — represents a significant capital outflow that could pressure startup valuations across the sector. OpenAI, which faces similar lawsuits from authors and news outlets including the New York Times, may need to set aside reserves for potential settlements or licensing agreements. Meta, which has also been sued over its use of copyrighted books to train its Llama models, faces comparable exposure. The market is now pricing in the likelihood that training data will carry a material cost — a shift from the era when web-scale data scraping was treated as free.
This article is for informational purposes only and does not constitute investment advice.