AI Firms Reportedly Used Book-Buying Middlemen to Hide Mass Destruction
A 404 Media report linked Anthropic and other AI providers to bulk purchases of physical books that were scanned, stripped, and discarded for training data.

Anthropic and other AI providers reportedly turned to third-party book-buying services to obtain training material without putting their own names on orders, according to a 404 Media investigation. The books were sought for their human-written text, then scanned and discarded after the pages had been processed.
The demand comes from an uncomfortable problem for generative AI companies. As AI-generated material spreads across the internet, older physical books offer large amounts of text created before the current wave of chatbot writing. Rare and out-of-print editions may also contain material that does not already appear in competing training collections.
Anthropic’s earlier book-collection effort became public through court proceedings. The company acquired millions of printed books, removed their bindings, scanned the pages, and disposed of the physical copies while building training data for Claude. A 2025 court ruling described that process as legally defensible under the fair-use analysis in Section 107 of the Copyright Act.
That legal finding did not remove the public-relations problem. Internal documents made public during the copyright case showed that Anthropic understood people might react badly to the destruction of millions of books. The company later agreed to a $1.5 billion settlement in the author-led lawsuit.
ISBNdb marketed bulk book sourcing for AI training
One service named in the reporting was ISBNdb. The company recently removed a page that advertised printed-book sourcing for large language model training. An archived copy of the page said ISBNdb could obtain books from libraries, used bookstores, and out-of-print catalogs, with orders reaching as many as 1 million books.
A separate archived ISBNdb blog post promoted confidentiality for each engagement. Its former wording said clients’ identities, plans, and acquisition targets would remain undisclosed. The post also acknowledged that a public story about an AI company destroying millions of books would be damaging.
ISBNdb later said in a news update that it removed the service page because it had been part of testing demand and the company decided to move in another direction.
Booksellers saw a sharp rise in unusual orders
The reported buying activity was not limited to one supplier. A bookseller who focuses on rare and low-circulation titles said dealers had seen an unprecedented increase in sales since April. The seller was filling orders for hundreds of books each week, compared with an earlier goal of selling about 20.
The books did not share an obvious subject or genre. Their common feature was that they had ISBNs, which may indicate that buyers selected them through a database rather than by browsing for specific topics. Book dealers in the Netherlands reported a similar pattern earlier in 2026.
For sellers, the orders can clear old inventory and bring in money. The same bookseller said the work was financially helpful, especially for stock sourced overseas and in foreign languages, but objected to uncommon editions being pulped after their contents were captured.
The reporting does not identify every AI company using these services, and it does not establish that every bulk purchase ends with a book being destroyed. It does show how companies can place distance between themselves and a practice that may be lawful in some circumstances but remains difficult to defend publicly.
Share your thoughts on AI training data and the treatment of physical books in the comments, and follow us on X, Bluesky, YouTube, and Instagram.





