πŸ“° Key Takeaways

Amazon has reportedly been buying up rare books in bulk, cracking their spines, and scanning the pages to use as AI model training data. The investigation comes from 404 Media, whose reporter planted a tracking device inside a rare book and traced it to an Amazon facility in Las Vegas. The facility is codenamed VGT3, marked with a logo showing a dinosaur claw gripping a book. In response, Amazon told 404 Media that the company “purchases books through commercial channels to improve products and services used by customers.” Companies like Amazon need an enormous volume of text to train large language models (LLMs), and the text available online has basically been wrung dry (Anthropic has even been accused of illegally pirating books). Rare books β€” especially out-of-print editions or ones that simply don’t exist online β€” have become a hot new source of training data. What makes this text especially valuable is that anything published before 2022 couldn’t possibly be AI-generated. When LLMs train on AI-generated text, it can trigger “model collapse” β€” where output quality gradually degrades as the model absorbs too much of its own AI-generated content.


πŸ’¬ JudyAI Lab Commentary

Based on the original summary, Amazon’s VGT3 facility dismantling rare books to scan for AI training data points to a fast-shrinking supply of quality text data.

What stands out to us is that this case highlights a structural shift in where LLM training data comes from. With the text available online already chewed over repeatedly by countless models, rare books published before 2022 β€” never digitized, never online β€” have become scarce assets, precisely because they’re guaranteed to be free of AI-generated text and avoid the “model collapse” problem that comes from training on AI-contaminated data. It also shows data quality is steadily overtaking data quantity as the key bottleneck in the next phase of the model race β€” even a giant like Amazon has to go back to physical books to fill the gap.

For AI builders, this is a reminder: when evaluating any model or dataset, it’s worth asking about the source and timing of the training data β€” that often tells you more about quality than parameter count ever will.


πŸ“… Original Article Info


πŸ”— Further Reading