Hourly ·
AI Companies Are Shredding Rare Books to Feed Their Models
Companies like Anthropic and OpenAI are bulk-buying printed books and destroying them in the scanning process, alarming rare book dealers who say irreplaceable editions are vanishing into the training-data pipeline.
AI companies desperate for fresh, human-written training data have found a new hunting ground: the world's libraries and used-book warehouses. And they're not borrowing — they're shredding.
According to a 404 Media investigation, ISBNdb — a company that bills itself as "the world's largest book database" — is now offering high-volume book acquisition services explicitly aimed at AI firms. The pitch? Printed books are guaranteed free of AI-generated slop: "curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate."
The economics are brutal. Destructive scanning — where a book's spine is sliced off and pages are fed through high-speed scanners — is significantly faster and cheaper than non-destructive digitization. One procurement list obtained by journalists contained 3,000 English-language titles, from academic works like Laser Shock Peening of Advanced Ceramics to irreplaceable folklore studies. Once scanned, the physical books are mulched and recycled.
Rare book dealers in the Netherlands raised the alarm in June, warning that obscure academic editions — some with print runs under a thousand copies — are being bought in bulk with no mechanism to check whether a digital copy already exists. The practice is entirely legal. Under US copyright law, companies can purchase physical books, convert them to digital text for training, and destroy the originals. But as commenters on Hacker News pointed out, the parallels to Vernor Vinge's dystopian "shred and scan" factories from his novel Rainbows End are no longer fiction.
The underlying tension is familiar: copyright law was designed for an era when copying meant a physical duplicate. It never anticipated corporations buying the world's books to feed them into neural networks, then pulping the evidence. Whether that qualifies as "fair use" is a question courts are only beginning to answer — while the shredders keep running.
Sources: 404 Media, HedgieMarkets via xcancel
人工智能公司正在撕毁稀有书籍来喂养他们的模型
公司如Anthropic和OpenAI正在大量购买并销毁书籍以进行扫描过程,这令古书商们担[K 心不可替代的版别正消失在训练数据管道中。
← 日报 日报 · 2026-07-27 16:00 UTC 人工智能公司正在销毁罕见书籍以喂养他们的[K 模型 公司如Anthropic和OpenAI正大量购买印刷书籍并在扫描过程中摧毁它们,这令稀[K 有书商感到担忧,他们称不可替代的版次正在消失到训练数据管道中。 为了获得新鲜[K 的人类写作训练数据,人工智能公司发现了一个新的猎物。
More Hourlies Stories
Content on Anagnorisis is summarized, paraphrased, and editorialized from publicly available sources for length and clarity. Original sources are linked where available. All trademarks belong to their respective owners.
