AI firms are shredding physical books because copyright law is quietly rewarding them

To legally train AI, tech companies are buying millions of physical books—just to slice off their bindings and shred them.

Editorial collage showing a humanoid AI figure with exposed machinery beside a courthouse, library shelves, and an open book releasing a stream of fragmented legal documents.
Image by CryptoSlate
3 min read

Quick Take

  1. ISBNdb markets physical-book orders of up to one million titles to AI developers.
  2. A 2025 ruling treated Anthropic’s destructive scanning of purchased books as fair use.
  3. No title-level evidence shows a rare book was destroyed, but preservation incentives remain.

ISBNdb is marketing physical-book orders of up to one million titles to AI developers, including material it describes as non-digitized, rare, or out of print.

The offer, reported by 404 Media, puts a 2025 copyright ruling in an uncomfortable new light: discarding a purchased book can support the legal premise that its internal scan replaced the original while keeping the library's copy count unchanged.

ISBNdb's current service and Anthropic's historical scanning program follow separate source chains. Physical disposal is the shared incentive created when books become scalable training data.

The one-copy incentive

In February 2024, Anthropic hired Tom Turvey, the former head of partnerships for Google's book-scanning project, to develop a lawful route to a much larger research library.

The company later spent many millions of dollars buying millions of print books. Service providers removed the bindings, cut the pages to size, scanned them, and discarded the paper originals, according to the federal court record.

The Washington Post later reported on the project using material that had become public. The operation predates ISBNdb's current marketing and stands on its own documented procurement chain.

The June 23, 2025 order reached two distinct fair-use conclusions. It held that Anthropic's use of copies to train specific large language models was transformative fair use on the record before it.

Separately, it held that converting lawfully purchased print books into non-distributed digital library copies was fair use because the PDFs replaced the purchased books without increasing the library's copy count.

The same order treated Anthropic's pirated central-library copies differently. The court denied Anthropic summary judgment on those copies and left the claims for trial. Its ruling stopped short of a general license to acquire books by any means or to treat every destructive scan as lawful.

Authors Guild claims OpenAI used pirated eBooks to train ChatGPT on copyrighted material
Related Reading

Authors Guild claims OpenAI used pirated eBooks to train ChatGPT on copyrighted material

George R.R. Martin angry AI can generate alternate ending to Game of Thrones? Or an existential threat to authors?
Sep 21, 2023 · Liam 'Akiba' Wright

The order's one-for-one reasoning implies that destruction did practical work. Discarding the paper original preserved the premise that one owned copy had been exchanged for another format. The court never declared destruction mandatory.

Keeping both the book and its scan would present a different copy-count fact pattern, making preservation legally inconvenient even when the physical object carries value beyond its text.

ISBNdb brings that preservation question into the present through a separate commercial offer. Its pages advertise book data and physical acquisition filtered by ISBN, subject, publication year, language, and edition, with orders of up to one million titles. The catalog can include non-digitized, rare, and out-of-print material.

A separate ISBNdb sourcing article promotes a legally binding nondisclosure agreement for each engagement and describes destructive scanning followed by verifiable destruction or recycling.

It also acknowledges the reputational problem created by headlines about AI companies destroying books. These are vendor marketing and compliance claims. Public material identifies no completed engagement, buyer, or disposal record for a specific title.

ISBNdb says pre-2022 print books are less exposed to AI-generated text and modern data-poisoning techniques than newer online material. That positioning makes the physical publishing record attractive as a source of human-produced text. Transaction data and comparable pricing needed to demonstrate a measurable premium have not surfaced.

AI training dataset used by tech giants allegedly created by scraping YouTube videos in violation of terms
Related Reading

AI training dataset used by tech giants allegedly created by scraping YouTube videos in violation of terms

Non-profit AI research group EleutherAI created the dataset called "the Pile."
Jul 16, 2024 · Mike Dalton

Anthropic's program and ISBNdb's offer remain separate source chains. Neither ISBNdb's marketing nor 404 Media's public report identifies an AI buyer behind a completed ISBNdb order or shows that such an order ended in destructive scanning.

CryptoSlate Daily Brief

Daily signals, zero noise.

Market-moving headlines and context delivered every morning in one tight read.

5-minute digest 100k+ readers

Free. No spam. Unsubscribe any time.

You’re subscribed. Welcome aboard.

What a scan cannot preserve

ISBNdb uses terms such as rare and out of print. Those labels say nothing about bibliographic scarcity on their own. A title can be hard to buy without being unique, and a particular copy can be replaceable even when its edition is uncommon.

Any claims of cultural loss which are now spreading on social media need title-level evidence identifying the book, the copy destroyed, and the number of comparable copies that survive.

The public record contains no named rare, unique, nearly extinct, or last-surviving book or edition destroyed by Anthropic or an ISBNdb customer.

The July 27, 2026 HedgieMarkets post that drove wider attention fused Anthropic's historical project, the 2025 ruling, and ISBNdb's 2026 marketing into one account. Its claims about rare-book shredding and a price premium remain unsubstantiated by the underlying sources.

Title-level proof is absent. The preservation incentive remains. Copyright analysis asks whether protected expression was copied and how the copy was used. Conservation asks about a binding, an annotated page, a particular printing, or an object's provenance.

The court's copy-count logic and ISBNdb's acquisition pitch assign value to different things: one to control over reproductions and the other to text that can be extracted at scale. The physical artifact can fall between them.

We've seen something similar in crypto before. In a 2021 episode involving Banksy's Morons, we saw the limit of digital representation from another direction. Injective Protocol, not Banksy, bought and burned the physical print before selling a token representing it, according to the BBC.

The blockchain recorded ownership and provenance without preserving the artwork. Destructive book scanning has a different purpose. A scan likewise retains information while leaving the object's survival to a separate decision.

Banksy’s renowned Wharf Rat will be auctioned as an NFT
Related Reading

Banksy’s renowned Wharf Rat will be auctioned as an NFT

The NFT drop will commence from the 17th to the 19th of December.
Dec 16, 2021 · Ricardo Rivas

The documented facts reveal a systemic preservation tension. A court's case-specific one-copy reasoning now sits beside a vendor's industrial-scale pitch for pre-AI books. Cultural loss has not been demonstrated, while the incentive to treat preservation as expendable is already visible.