AI labs are buying rare books by the pallet, scanning them, and shredding the originals. A federal judge says it's legal.
AI companies are bulk-purchasing obscure and rare physical books, running them through high-speed spine-cutting scanners, and destroying the originals to build training datasets free of AI-generated text. A service called ISBNdb facilitates anonymous orders of up to a million books. After Judge Alsup ruled the practice qualifies as fair use because destruction means only one copy exists at a time, European booksellers say the buying has accelerated.

How Destroying Books Became AI's Fair Use Strategy
Judge William Alsup ruled that scanning a book and destroying the physical copy qualifies as fair use. Now AI labs are buying obscure texts by the pallet, and European booksellers are raising the alarm.
Anthropic called it "Project Panama." An internal planning document unsealed in January 2026 court filings described the effort as a push to "destructively scan all the books in the world" 1. Anthropic spent tens of millions buying books in bulk, cutting off their spines, feeding pages through industrial scanners, and discarding the originals
2.
When Judge Alsup ruled in Bartz v. Anthropic that this pipeline qualified as fair use, the decision gave every AI lab a legal blueprint 2. Destruction is not a side effect of the scanning process. It is the mechanism that makes the scanning defensible.
Under the first-sale doctrine, a buyer can do what they want with a physical copy they own. Scanning a book creates a second copy, which introduces a copyright problem. Destroying the original after scanning means only one copy exists at any given time: the digital version replaced the physical one rather than duplicating it. Judge Alsup found this made the use transformative and therefore protected 2
3.
Under this framework, preservation is a liability. Keeping the physical book alongside the digital file creates two copies, which weakens the transformative-use argument. The ruling provides a template that other AI companies now explicitly cite 2.
Standard pallets hold 800 to 1,200 books, with buyers scaling from pilot orders to batches exceeding 10,000 volumes. Destructive scanners process 80 to 120 pages per minute once spines are removed 2. A service called ISBNdb, which describes itself as the world's largest book database, facilitates anonymous orders ranging from 1,000 to one million books per transaction
3. ISBNdb keeps buyer identities confidential and acknowledged on its own website that "AI company destroys two million books is not a headline that generates sympathy"
3.
The demand targets books published before 2022, when large language models began saturating the internet with synthetic text. ISBNdb describes pre-LLM printed books as "structurally clean of modern poisoning" 3. Researchers have warned that training models on AI-generated output degrades quality across successive generations, a problem called model collapse
2.
In Europe, the buying has spread. Pieter de Vries, who runs the antiquarian bookstore De Vries De Vries in Haarlem, received an email from a Singapore-based company called 2077AI containing a list of 3,000 English-language titles organized by ISBN 4. Similar bulk-purchase requests have surfaced in Switzerland, Spain, and Germany. A German bookseller reported a surge of orders arriving between 3 and 5 a.m. every night from a Canadian company called Zoom Books, targeting obscure, specialized titles that would be unprofitable to resell
4. Zoom Books told Swiss broadcaster SRF the purchases were part of a "regular recycling and trading model"
4.
Booksellers can typically only speculate about whether AI labs are behind the orders, because services like ISBNdb keep buyers anonymous 3. One small bookseller told 404 Media that his weekly sales jumped from no more than 20 books to hundreds, with selections appearing random except that every title carried an ISBN
3. Many of the targeted titles are out of print, foreign-language, or low-circulation editions that exist in small numbers
2. Once scanned and pulped, the text survives only inside a proprietary dataset.
Anthropic separately settled copyright claims for a reported $1.5 billion covering roughly 500,000 works 3
2. That settlement covered content Anthropic had already ingested through other means. Project Panama is the forward-looking strategy: acquire new training data through a buy-scan-destroy pipeline backed by a ruling that says destruction is part of what makes it legal.
Harvard University, partnering with Google and Microsoft, released nearly one million public-domain digitized books in 254 languages without destroying a single physical copy 2. Microsoft's Burton Davis argued that public-domain data is a prudent starting point because libraries contain cultural and historical records absent from online sources
2. But the supply of public-domain books is finite. The bulk of published writing remains under copyright, and for that material the cheapest route to clean training data runs through a spine-cutting scanner and a shredder. A federal judge has ruled that the shredder is part of what makes it legal.
References
Cite this story
ProvenBrief (2026). "AI labs are buying rare books by the pallet, scanning them, and shredding the originals. A federal judge says it's legal.." ProvenBrief. https://provenbrief.com/story/ai-labs-are-buying-rare-books-by-the-pallet-scanning-them-and-shredding-the-orig
Free to quote and link with attribution. Republishing in full or AI-training use requires a license.
Get the next brief in your inbox
One weekly email. Every claim verified against primary sources before we hit send.
This story
WordsProduced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our editorial standards.