AI Labs Are Buying Rare Books, Scanning Them, Then Shredding Them

Reports from booksellers, court filings, and 404 Media describe a training-data pipeline that ends in a shredder: AI companies bulk-buying used and out-of-print books, scanning every page, then destroying the physical originals. Here's what's actually legal, what isn't settled, and why it matters for anyone building with AI.

A Training-Data Pipeline That Ends in a Shredder

Reports from used-book dealers, court documents, and an investigation by 404 Media describe a practice that sounds closer to fiction than AI supply chain logistics: companies training large language models are bulk-buying used and out-of-print books, scanning every page with high-speed machines that cut the spines to feed the pages through, and disposing of the physical originals once the text has been digitized into training data.

This isn't a fringe rumor. Brokers like ISBNdb, which normally maintains a database of book metadata for the publishing industry, are now reportedly facilitating anonymous bulk orders on behalf of AI clients, in batches said to reach as high as one million volumes. One Dutch bookseller told 404 Media that a request for 3,000 copies of a single title initially looked like spam or phishing, before turning out to be a real order tied to AI training data sourcing.

1M Volumes reportedly orderable in a single anonymous bulk batch via ISBNdb
3,000 Copies of one title requested in the order a Dutch bookseller mistook for spam
$1.5B Anthropic's settlement over separately sourced pirated digital books
2022 Cutoff year for preferred editions, chosen to avoid AI-generated content

Why Pre-2022 Editions Are the Target

The sourcing pattern isn't random. Pre-2022 editions are reportedly preferred because they predate the flood of AI-generated text now circulating online and in print, making them "cleaner" training material, less likely to contain synthetic content that could degrade a model trained on it. That preference is exactly why some of the rarest and least-reprinted editions, books with only a handful of surviving physical copies, are the ones dealers say are disappearing from circulation permanently.

Worth noting: Snopes, which fact-checked the claim, found the full scale of the practice hard to verify independently given the anonymized nature of the bulk orders. But the underlying sourcing and destruction reports, sourced from booksellers and industry databases, have not been denied by the companies named in them.

The Legal Backdrop: Bartz v. Anthropic

None of this is happening in a legal vacuum. In Bartz v. Anthropic, Judge William Alsup of the U.S. District Court for the Northern District of California ruled in June 2025 that digitizing legally purchased print books to train a large language model qualifies as fair use. The court's reasoning: a digital copy simply replaces the physical one the company already owns, without adding new copies into circulation or redistributing them to third parties.

That ruling is doing a lot of work here. It's the legal foundation that makes buying a book, scanning it, and destroying the original defensible in court, provided the copy was legitimately purchased in the first place. Anthropic separately settled a related claim for $1.5 billion over books sourced from pirated digital libraries rather than legally purchased copies, a reminder that the fair-use protection has real limits tied to how the material was acquired.

01
The Core Gap

Fair Use Covers the Copying. It Says Nothing About the Destruction.

The Bartz v. Anthropic ruling addresses copyright: is scanning a legally purchased book to train a model an infringing act? The court said no. What the ruling never addresses is the physical fate of scarce cultural objects once the scanning is complete. A rare, out-of-print edition with a handful of surviving copies has a claim to preservation that has nothing to do with copyright law, and everything to do with what happens to be true if that copy is shredded rather than resold, donated, or archived.

Why This Matters Beyond the Headline

Set aside the "AI destroys books" framing for a moment, because the more durable story here is about how training data sourcing is evolving as an industry practice.

For context: the used and rare book trade has historically operated on the assumption that even out-of-print titles remain findable somewhere, in a private collection, a secondhand shop, or a library sale. Bulk acquisition for destruction breaks that assumption at a scale the trade has no established mechanism to track or resist.

What This Means If You're Building With AI

Whatever your view of the ethics here, there are practical implications for teams that build products on top of foundation models or that make their own decisions about training data:

Frequently Asked Questions

Are AI companies really destroying rare books?
Multiple booksellers and industry reports, including an investigation by 404 Media, describe AI companies bulk-buying used and out-of-print books, scanning them at high speed, and destroying the physical originals. Snopes notes the full scale is difficult to verify independently given the anonymized nature of the bulk orders, but the companies named have not denied the underlying sourcing and destruction reports.
Is scanning books to train an AI model legal?
In Bartz v. Anthropic, a federal judge ruled in June 2025 that digitizing legally purchased print books to train a large language model qualifies as fair use, on the reasoning that the digital copy simply replaces the print copy the company already owns. The ruling does not address what happens to the physical book afterward, only the legality of the copying itself.
What is ISBNdb's role in this?
ISBNdb, a company that normally maintains a large database of book metadata for the publishing industry, is reported to now broker high-volume, anonymous book acquisitions on behalf of AI companies sourcing training data, in batches reportedly reaching up to one million volumes.
Did Anthropic pay a settlement related to book training data?
Yes. Separately from the Bartz v. Anthropic fair-use ruling on legally purchased books, Anthropic settled a related claim for $1.5 billion over books it sourced from pirated digital libraries rather than legitimately purchased copies.

My Take

Courts settled the copyright question. Nobody settled the preservation question, and those are genuinely different problems wearing the same headline. A ruling built on the logic that "we're just replacing the print copy" quietly assumes the print copy was disposable to begin with, and for a rare edition with a handful of surviving copies, it plainly wasn't.

Data sourcing decisions get treated as a legal or technical checkbox almost everywhere I look. This story is a reminder that they're also cultural decisions, made at a speed and scale that leaves no room to reconsider once the shredder has already run. If training pipelines can legally erase originals as a byproduct of building a product, "it's legal" is a low bar for an industry with this much influence over what gets preserved and what quietly disappears.

Related Articles:

Kodjo Apedoh

Kodjo Apedoh

Network Engineer & AI Entrepreneur

Founder of TechVernia & SankaraShield. Certified Network Security Engineer with 4+ years of experience specializing in network automation (Python), AI tools research, and advanced security implementations. Also builds iOS and Android applications. Holds certifications from Palo Alto Networks, Fortinet, and Cisco. Based in Arlington, Virginia.

Connect on LinkedIn →