A Training-Data Pipeline That Ends in a Shredder
Reports from used-book dealers, court documents, and an investigation by 404 Media describe a practice that sounds closer to fiction than AI supply chain logistics: companies training large language models are bulk-buying used and out-of-print books, scanning every page with high-speed machines that cut the spines to feed the pages through, and disposing of the physical originals once the text has been digitized into training data.
This isn't a fringe rumor. Brokers like ISBNdb, which normally maintains a database of book metadata for the publishing industry, are now reportedly facilitating anonymous bulk orders on behalf of AI clients, in batches said to reach as high as one million volumes. One Dutch bookseller told 404 Media that a request for 3,000 copies of a single title initially looked like spam or phishing, before turning out to be a real order tied to AI training data sourcing.
Why Pre-2022 Editions Are the Target
The sourcing pattern isn't random. Pre-2022 editions are reportedly preferred because they predate the flood of AI-generated text now circulating online and in print, making them "cleaner" training material, less likely to contain synthetic content that could degrade a model trained on it. That preference is exactly why some of the rarest and least-reprinted editions, books with only a handful of surviving physical copies, are the ones dealers say are disappearing from circulation permanently.
The Legal Backdrop: Bartz v. Anthropic
None of this is happening in a legal vacuum. In Bartz v. Anthropic, Judge William Alsup of the U.S. District Court for the Northern District of California ruled in June 2025 that digitizing legally purchased print books to train a large language model qualifies as fair use. The court's reasoning: a digital copy simply replaces the physical one the company already owns, without adding new copies into circulation or redistributing them to third parties.
That ruling is doing a lot of work here. It's the legal foundation that makes buying a book, scanning it, and destroying the original defensible in court, provided the copy was legitimately purchased in the first place. Anthropic separately settled a related claim for $1.5 billion over books sourced from pirated digital libraries rather than legally purchased copies, a reminder that the fair-use protection has real limits tied to how the material was acquired.
Fair Use Covers the Copying. It Says Nothing About the Destruction.
The Bartz v. Anthropic ruling addresses copyright: is scanning a legally purchased book to train a model an infringing act? The court said no. What the ruling never addresses is the physical fate of scarce cultural objects once the scanning is complete. A rare, out-of-print edition with a handful of surviving copies has a claim to preservation that has nothing to do with copyright law, and everything to do with what happens to be true if that copy is shredded rather than resold, donated, or archived.
Why This Matters Beyond the Headline
Set aside the "AI destroys books" framing for a moment, because the more durable story here is about how training data sourcing is evolving as an industry practice.
- Training data sourcing is quietly becoming a supply chain problem. Used bookstores, rare book dealers, and metadata brokers like ISBNdb are now upstream vendors in AI training pipelines, a role none of them were built for and few disclose publicly.
- Legal and irreversible are not the same thing. A court can validate the copying step while the outcome, permanent loss of a scarce physical object, remains something no ruling was ever asked to weigh in on.
- Provenance and annotation are lost with the object. A digital scan captures the text. It does not capture marginalia, printing variants, binding history, or the physical evidence that book historians and archivists rely on.
- Libraries build preservation policy around scarcity over decades. A single high-volume scanning operation can undo years of that work in an afternoon of processing, with no equivalent institutional review.
For context: the used and rare book trade has historically operated on the assumption that even out-of-print titles remain findable somewhere, in a private collection, a secondhand shop, or a library sale. Bulk acquisition for destruction breaks that assumption at a scale the trade has no established mechanism to track or resist.
What This Means If You're Building With AI
Whatever your view of the ethics here, there are practical implications for teams that build products on top of foundation models or that make their own decisions about training data:
- Data provenance is becoming a real due-diligence question. If you're evaluating a model or a vendor's training practices, "was this fair use" and "was this sourced responsibly" are now two separate questions worth asking.
- Court precedent moves faster than public expectations. Bartz v. Anthropic settled a legal question over a year ago; most people are only now learning what companies have been doing under its protection. Assume the legal floor for what's permitted is lower than what your users assume is happening.
- Reputational risk doesn't require illegality. A fully legal practice can still generate the kind of backlash that damages trust in a brand or a model, independent of any court ruling.
- Custodianship matters for anyone sourcing data at scale. Digitize-then-donate or digitize-then-archive pipelines exist and cost little more than digitize-then-destroy. The choice not to build one is a choice, not a neutral default.
Frequently Asked Questions
My Take
Courts settled the copyright question. Nobody settled the preservation question, and those are genuinely different problems wearing the same headline. A ruling built on the logic that "we're just replacing the print copy" quietly assumes the print copy was disposable to begin with, and for a rare edition with a handful of surviving copies, it plainly wasn't.
Data sourcing decisions get treated as a legal or technical checkbox almost everywhere I look. This story is a reminder that they're also cultural decisions, made at a speed and scale that leaves no room to reconsider once the shredder has already run. If training pipelines can legally erase originals as a byproduct of building a product, "it's legal" is a low bar for an industry with this much influence over what gets preserved and what quietly disappears.
Related Articles: