📰 Key Highlights

Anthropic received final approval from a federal judge this Monday to pay $1.5 billion to settle a copyright class action lawsuit — the largest settlement in U.S. copyright law history. Judge William Alsup of the U.S. District Court for the Northern District of California had preliminarily approved the settlement last year, finding that Anthropic had illegally downloaded and stored millions of copyrighted books; after Alsup retired, Judge Araceli Martinez-Olguin officially signed the approval this Monday. Under the settlement terms, roughly 500,000 works are eligible for $3,000 in compensation each, shared among the rights-holding authors and publishers. However, many authors and creators don’t view this as a victory, because the ruling on the case’s core legal question actually favored Anthropic: Alsup determined that training AI models on copyrighted text qualifies as fair use, a decision seen as a pivotal turning point for the AI industry. But the ruling didn’t absolve Anthropic of responsibility for how it sourced its books — its training data came from two channels: books it purchased and scanned (legal), and books downloaded from piracy sites like Library Genesis and Pirate Library Mirror. Alsup ruled the latter itself constitutes illegal behavior, and said the piracy claims could proceed to jury trial. Anthropic then agreed to settle to avoid going to court and facing potentially massive damages. Because the settlement prevents the case from being appealed to form binding precedent, broader industry-level legal questions remain unresolved. Google, Meta, Midjourney, OpenAI and others still face similar copyright lawsuits. Just last week, a group of publishers and authors including Hachette, Cengage, Elsevier, and writer Scott Turow sued Google, alleging it used copyrighted works to train Gemini.


💬 JudyAI Lab Perspective

JudyAI Lab today is watching a ruling with major implications for the AI industry: Anthropic settled a copyright lawsuit for $1.5 billion — the largest settlement amount in U.S. copyright history.

The key point of this case isn’t the settlement amount — it’s Judge Alsup’s earlier finding that “training AI models on copyrighted text qualifies as fair use.” That’s a significant signal for the entire AI industry. But the settlement also exposes a separate layer of the problem: part of Anthropic’s training data came from legally purchased and scanned books, while another part came from piracy sites like Library Genesis and Pirate Library Mirror. The latter was deemed illegal on its own, and that’s the real pressure that drove the settlement. This highlights something the AI builder community often overlooks: “whether your training data counts as fair use” and “whether the means of obtaining that data is legal” are two independent legal questions — a favorable ruling on the first one doesn’t give you free rein on the second. And since this settlement can’t be appealed to establish precedent, Google, Meta, OpenAI and others still face similar lawsuits — the industry-wide legal disputes aren’t truly over yet.

Something worth thinking about for readers: if your project or product uses any AI training or RAG data sources, it’s worth taking the time to confirm whether the acquisition channel for every piece of material is legal — not just whether the “use case” is reasonable.


📅 Source Info


🔗 Further Reading