← all musings

AI Companies Are Eating the Library

The labs didn't just burn books. They burned the optionality to ever have a different conversation about who owns the training data that powers AGI.

AI companies are physically destroying books to train their models. A federal court has already ruled it’s legal. That combination should end the debate about whether this will keep happening — it won’t. It will accelerate.

The mechanics are documented now, not just alleged. ISBNdb, a book-metadata company, is openly brokering bulk print-book purchases for AI labs — orders from 1,000 units up to a million, with a strict NDA on every engagement. Their pitch: pre-2022 books are structurally guaranteed to be free of AI-generated text. Their marketing copy is unsettlingly self-aware: “The optics problem is real,” and “‘AI company destroys two million books’ is not a headline that generates sympathy.” They know exactly what they’re selling.

The legal foundation comes from Judge Alsup’s June 2025 summary judgment in Bartz v. Anthropic. Anthropic purchased millions of copyrighted books, removed the bindings, scanned them, and discarded the paper originals. Alsup found this qualified as fair use — the books were legally purchased, each physical copy was destroyed after scanning, and the files stayed internal. He likened it to conserving space through format conversion. Anthropic also hired Tom Turvey, former head of partnerships for Google Books, and tasked him with obtaining “all the books in the world.” That’s not hyperbole. That’s a job description from a court filing.

Two things the ruling leaves out that change the picture: Alsup also ruled against Anthropic on pirated-library ingestion, and the company settled for $1.5 billion over those books. And this is one district court — a different court could reopen the fair-use question on the destructive scanning side.

One piece of the circulating story that’s real but unverified: the rare-books angle. A bookseller told 404 Media that his weekly volume jumped from roughly 20 books to hundreds in April, that the selections looked random and ISBN-keyed, and that his stock includes rare and out-of-print titles — so an AI company could be destroying some of the last findable copies. No one has confirmed a specific rare title was shredded, or which labs are placing the current ISBNdb orders. Anthropic’s documented destruction is from 2024 court filings, not evidence they’re buying through ISBNdb now. The emotional core of the story — irreplaceable manuscripts being pulped — is plausible and unverified.

Everyone says this is a story about AI ethics. The opposite is closer to true. If you understand the data economics, the behavior is completely legible. The easy training data ran out. Common Crawl is degraded — it’s increasingly full of AI-generated text, which is training poison that causes model collapse. Web licensing got expensive after the NYT sued OpenAI and others followed. Physical archives offered something increasingly scarce: high-quality, pre-contamination text with no lawyers and no digital footprint. The labs found the next frontier of extraction, and it happened to be the physical patrimony of human knowledge.

The steelman for the labs is real and worth taking seriously: these books may have otherwise decayed unscanned, undiscovered, unread by anyone. A medieval marginalia joke that three scholars have ever read arguably does more cultural work embedded in a model that a billion people use than it does sitting in an acid-free box in a climate-controlled basement. There’s a version of this story where the labs are unglamorous preservationists.

I don’t buy it. Not because the preservation argument is wrong in principle, but because it’s not what’s driving the behavior. If preservation were the goal, the digital copies would be made publicly available and the physical originals would be returned intact. That’s what legitimate digitization projects — the Internet Archive, Google Books, university initiatives — actually do. What’s documented here is extraction in service of competitive advantage, with preservation as a retroactive alibi.

The deeper issue is the data moat. Compute commoditizes. Data doesn’t. Nvidia’s margins will eventually compress, custom silicon is proliferating, inference costs are dropping. You can build more GPUs. You cannot recreate a destroyed manuscript. If the labs that moved fastest on physical archive acquisition built a durable training-data advantage, they may have locked in model quality improvements that compound for years. The destroyed books are a one-time cost. The capability gain is permanent.

That’s the stakes. Not the loss of any single book — tragic as that is — but the structural fact that private companies may be converting shared human heritage into proprietary competitive advantage, permanently and irreversibly, with no compensation to the public and no accountability mechanism in place. Copyright law wasn’t designed for this. Libraries weren’t funded for this. And the regulatory frameworks that might address it are years behind the problem.

The labs didn’t just burn books. They burned the optionality to ever have a different conversation about who owns the training data that powers artificial general intelligence.