📰 Key Takeaways

Taiwan’s AI community is once again focused on the legal controversy surrounding AI companies’ use of copyrighted books to train large language models. The article points out that the AI models behind mainstream chatbots like ChatGPT, Gemini, and Claude have training databases covering hundreds of millions of published books, web articles, academic papers, and more — and the vast majority of authors had their work used to train these tools, which could threaten their livelihoods, without being notified or asked for consent. Last year, Judge William Alsup issued a landmark ruling ordering Anthropic to pay a copyright settlement of up to $1.5 billion, but notably, the judge actually found Anthropic’s AI training itself to be legal — what was actually being penalized was Anthropic’s acquisition of books through illegal online “shadow libraries” of pirated content. The judge compared an LLM’s process of learning from trillions of words to a writer studying literary works, emphasizing that AI isn’t meant to copy or replace the original work, but rather to “make a dramatic departure and create something different.” IP lawyer Cathy Gellis argued that the ruling is actually more favorable for AI companies, since the judge analogized AI training to “reading” rather than “copying” — and copyright law is fundamentally about copying, not merely using or reading a work. Compared to Anthropic’s projected annual revenue of roughly $200 billion by 2028, the $1.5 billion fine has limited impact. The article also notes that current copyright law hasn’t been updated since 1976, forcing judges to rule on brand-new legal questions shaping the AI industry’s future using half-century-old statutes. The controversy largely centers on the “fair use” doctrine — whether using copyrighted content is sufficiently “transformative.” See the original article for full details.


💬 JudyAI Lab’s Take

Taiwan’s AI industry has once again been buzzing about the copyright controversy over LLM training data, and this lawsuit is worth watching for every AI builder.

Judge William Alsup’s ruling sends a subtle signal: AI training itself was found to be legal, and what actually got penalized was Anthropic acquiring pirated books through illegal “shadow libraries.” The judge compared an LLM’s process of learning from text to a writer studying literary works, with the key question being whether the AI “makes a dramatic departure and creates something different” rather than simply copying. IP lawyer Cathy Gellis’s take also highlights something important: copyright law governs “copying,” not “reading” — which actually works in AI companies’ favor. That said, current copyright law hasn’t been updated since 1976, and using half-century-old statutes to handle new AI-era disputes means we’ll almost certainly see more rulings like this down the line.

For AI builders, the takeaway here is: the legality of how you source your training data may be a bigger legal risk point than the training process itself.


📅 Original Article Info


🔗 Further Reading