SATURDAY, OCTOBER 10, 2026|No. 18216
Technology · Ethics

Amazon Reportedly Disassembles Rare Books for AI Training Data

An investigation reveals Amazon may be acquiring and destroying rare books to use their content for training artificial intelligence models, raising concerns about data sourcing ethics.

A rare book is shown, representing the potential data source for AI training.
A rare book is shown, representing the potential data source for AI training.
1 sources
Pipeline ingest
3 reads
Positive / Neutral / Negative
1 countries
Related coverage

Amazon is buying tons of rare books, cutting off their spines, and scanning them for AI training, according to 404 Media, which placed a tracking device in a rare book that ultimately arrived at an Amazon facility in Las Vegas.

The facility, known as VGT3, identifies itself with a symbol of a dinosaur holding a book in its claws. Amazon told 404 Media in a statement that it “purchases books through commercial channels to improve the products and services customers use.”

Companies like Amazon need unfathomably large amounts of text to train their LLMs, which have already ingested what they can from the internet (and, in Anthropic’s case, illegally pirated books). Rare books, especially ones that are out of print or impossible to find on the internet, offer a new source of coveted training data.

These texts are especially valuable since there’s no chance that anything published before 2022 was written by an LLM. When LLMs train on AI-generated text, they risk “ model collapse,” which can occur when the quality of an LLM’s outputs degrade after ingesting too much AI-generated text.

PAN's pipeline reviewed approximately 1 open sources for this article. No human editor reviewed this article before publication.

Related Reads

Show on timeline →

Earlier on PAN

More in Technology →