Print books beat the open web for one simple reason. Machines cannot rewrite them after the fact.
A Company Selling Books to Feed AI
A firm called ISBNdb now sources physical books in bulk for AI labs to scan into training data. It calls itself the world’s largest book database, and its site claims the world’s best AI training data is sitting on a shelf. The company used to help libraries and bookstores manage catalogs. Now it fills orders that can run past a million volumes at a time.
Why Pre-2022 Books Carry a Premium
Date matters more than content here. ISBNdb argues that books printed before 2022 predate the large language model era, so they cannot contain AI-generated text. That distinction protects against a real problem called model collapse. Feed a model on the output of earlier models, and each new generation gets a little worse than the last. A pre-2022 print run carries no risk of that. It is a fixed, human-authored record, and nobody can quietly rewrite it after publication.
How the Books Actually Get Processed
The process is physical and irreversible. Companies use high-speed machines to cut the spines off books and scan the pages before disposing of the originals. Anthropic ran a large version of this effort under an internal program called Project Panama. Anthropic used a hydraulic-powered cutting machine to remove pages from books bought through resellers such as Better World Books, then scanned the pages with industrial equipment. A vendor proposal tied to the effort described converting between 500,000 and two million books over six months.
The Legal Question That Settled It
A federal judge in San Francisco addressed the core legal question directly. The judge ruled that scanning legally purchased physical books into digital copies, even after destroying the originals, counted as transformative fair use. Judges in separate cases involving OpenAI and Meta reached similar conclusions. That ruling did not clear every legal risk for Anthropic. A different case in the same court ended with a $1.5 billion settlement, after Anthropic trained Claude on pirated copies of copyrighted books rather than purchased ones. Boing Boing
Anonymity Has Become Part of the Business Model
Buyers want distance from the story. AI providers now use third-party middleman services to source physical books anonymously, and these middlemen strip and destroy the books at industrial scale on the providers’ behalf. ISBNdb tells its own clients the concern is real. The company acknowledges to clients that the optics problem is real. Rare and low-circulation titles have started selling at a faster pace, driven by demand from this new buyer class.
Why It Matters
Training data quality decides how good the next generation of AI models will be. Human-authored books, especially the ones printed before generative AI existed, offer a clean signal that the open web can no longer guarantee. That is pushing AI companies toward paper, even as the same industry keeps facing lawsuits over how it built its earlier training sets.