Online retail giant Amazon is allegedly destroying thousands of rare books at a Las Vegas warehouse to scan their pages and train artificial intelligence language models, according to an investigation. The discovery followed an independent technology outlet placing a tracking device inside a suspected rare book purchase and tracking its cross-country journey to a Nevada facility.
Amazon Warehouse Operations and the Search for Clean AI Data
Once shipments of obscure volumes arrive at the VGT3 warehouse, workers reportedly remove the bindings so the loose pages can be processed through high-speed scanners. The physical copies are destroyed during digitization, leaving behind discarded paper and digital files that investigators claim serve as training material for Amazon systems. Employees at the location stated that their daily tasks involve receiving commercial shipments of printed books and removing spines for scanning. The internal logo for the team operating inside the warehouse features a dinosaur holding a book. Amazon has not publicly confirmed that books processed at the warehouse are destroyed specifically for AI training.
The Value of Pre-LLM Literature and Industrial Book Destruction
Technology companies investing in artificial intelligence require enormous text quantities to train large language models, facing challenges over data availability and legality. Rare books unavailable through digital archives have become a potential text source, with older literature considered valuable training material.
AI developers have raised concerns about training models on computer-generated material, a phenomenon known as model collapse.
Researchers warn that repeatedly training systems on AI-generated content can reduce output quality over time, leading to less accurate or less diverse results. Services like ISBNdb facilitate bulk orders ranging from 1,000 to one million books per order, promising to keep buyers anonymous. Print books from the pre-LLM era are sought after because they are structurally guaranteed to be free of AI-generated contamination.
Copyright Scrutiny, First-Sale Doctrine, and Legal Precedents
Technology companies face widespread scrutiny over how they obtain written material for AI development. Anthropic faced legal challenges over its use of copyrighted material in training its Claude language model, resulting in a $1.5 billion copyright settlement with authors over allegations involving pirated books. Anthropic denied wrongdoing while maintaining that its use of training material was connected to developing AI systems.

In the Anthropic case, a judge ruled that using a hydraulic-powered cutting machine to remove pages from purchased books and scanning them with industrial equipment took advantage of the first-sale doctrine. This legal concept allows a buyer to do what they want with a purchase without the original copyright holder’s permission. Because the texts were turned into digital files rather than redistributed as new copies, the practice was found to be transformative and protected by fair use.
Industry Response and the Sourcing of Training Materials
Despite growing legal disputes over copyright and artificial intelligence, companies have provided limited public detail about how they source training materials. Regarding its warehouse operations, Amazon stated that it purchases books through commercial channels to improve the products and services that customers use daily.
Booksellers and archivists continue to question whether increased demand for rare and out-of-print books is linked to AI companies seeking new training data sources.

