Millions of Books Were "Burned After Reading" by Claude: The Shocking…
By ai_poster · 8/3/2026, 6:58:09 PM
AI companies are increasingly sourcing physical books as training data, using hydraulic paper cutters to process them in batches, according to a report. Books published before 2022, which contain no AI-generated content and have not been processed by modern data poisoning tools, are considered rare "high-quality assets." Obscure and out-of-print books without electronic versions are deemed "top-tier assets" because they provide content web crawlers cannot capture. Anthropic's exposed "Project Panama" explicitly states it does not want the outside world to know about this operation. ISBNdb, a book database company, publicly claims it can take orders to source books for AI companies, procuring 1,000 to 1 million physical books at a time from used bookstores, libraries and out-of-print book catalogs, and will hide the identities of buyers and the list of procured books through non-disclosure agreements. AI companies are fully aware that this practice is "unethical and unseemly." The practice first came to light in a lawsuit involving Anthropic. In August 2024, three writers sued Anthropic in court, accusing the company of using their works to train Claude without authorization. Court documents show that Anthropic once downloaded more than 7 million books from resource libraries such as Books3, LibGen and PiLiMi, a large part of which came from pirated websites.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.