news / 2026 / ai-training-data + used-books A used-book warehouse sorting table and densely packed shelving show the physical procurement layer behind large training-data acquisitions.

news

AI Turned the Used-Book Market Into a Data Supply Chain

The AI data race has reached secondhand warehouses, where ownership, secrecy, and destructive scanning can convert a physical copy into an internal dataset.

A Galway bookseller received an order for roughly 5,000 books. The list jumped from local history to health manuals to Internet Explorer for Dummies. The buyer came through an intermediary, did not identify itself, and showed none of the taste or bargaining behavior that usually explains a large collection purchase.

Kennys suspected AI training. Booksellers in Australia and Germany have reported similarly strange, price-insensitive orders for boxes of unrelated old titles. Nobody has publicly proved that these particular shipments entered a model-training pipeline. That uncertainty is part of the machinery. The buyer can remain hidden while marketplaces, brokers, recyclers, warehouses, scanning contractors, and model labs each handle one leg of the transaction.

The used-book market has become an unmarked data-acquisition layer.

The book is the license-shaped object

The mechanism became public in the 2025 Bartz v. Anthropic order. The court found that Anthropic spent millions buying millions of print books, often used. Contractors stripped the bindings, cut the pages, scanned them into searchable PDFs, and discarded the paper. Anthropic retained the digital library and drew training sets from it.

Judge William Alsup held that converting a purchased print copy into an internal digital replacement qualified as fair use on the record before him. The ruling treated the operation as a format change: one lawfully purchased object went in, one non-redistributed digital copy came out, and the paper original ceased to exist. The same order refused to bless Anthropic’s separate library of pirated books.

That distinction supplies the economic incentive. Ebook licensing is fragmented, encumbered by DRM, restricted by contracts, and largely excluded from first-sale rights. A used paperback is brutally simple. Buy it. Own it. Process it. Keep the resulting file internal.

The cutting machine sits downstream of copyright law.

Scarcity is invisible to a bulk order

A warehouse can contain 500,000 books and still lack a reliable answer to a basic question: how replaceable is this exact copy?

ISBN counts editions, not every surviving physical instance. Marketplace listings reveal what is for sale today, not how many copies remain in attics, institutional stores, private collections, or landfill queues. Local histories, technical manuals, small-press poetry, cookbooks, and community records often look worthless to a pricing algorithm until the last accessible copy disappears.

The current reporting needs discipline. Anthropic says its acquisition programs do not buy or destroy rare or antiquarian books. Zoom Books, named by Australian sellers as a bulk buyer, says it resells books intact and does not digitize them. The company declined to identify customers because of confidentiality agreements. The Atlantic found no direct evidence connecting ISBNdb to the recent orders described by booksellers. Headlines claiming a proven campaign to shred rare books outrun the evidence.

The proven facts are already ugly enough. Industrial destructive scanning exists. Millions of purchased books were processed that way. Companies now market physical-book procurement for AI. Booksellers across several countries are seeing anonymous orders with the statistical shape of corpus assembly. The custody chain does not expose destination, preservation status, scan method, or whether another accessible copy survives.

Two digitization systems

The Internet Archive offers the useful counterexample because it treats the physical volume as evidence. Its Scribe workflow turns pages by hand, captures foldouts, checks image quality, and returns books to their libraries or preserves a physical copy. The Archive links digital records back to physical locations so the source can be inspected again.

That last part matters. OCR can fail. Pages can be skipped. Marginalia, binding structure, paper, plates, inserts, and edition-specific mistakes can carry information absent from plain text. A digital file is an access object. The physical book remains an authority object.

Industrial AI ingestion optimizes a different target: maximum clean text per dollar and per hour. Destructive scanning flattens bindings quickly and produces uniform page feeds. It treats the book as disposable packaging around tokens.

Non-destructive scanning costs more. That excuse gets weaker when the customer is an AI company spending billions on compute and competing for marginal benchmark gains. The expense is a choice about where the system assigns value. GPU time is capital. Human handling is friction. The book becomes feedstock because the accounting says its future existence is somebody else’s problem.

Booksellers are becoming the control point

The acquisition stack has no regulator checking cultural scarcity at checkout. Booksellers are improvising the policy themselves.

Tomás Kenny said his shop plans to refuse orders that appear destined for AI training. Rare-book dealers already screen buyers when they fear a volume will be broken apart for individual illustrations. Their judgment is imperfect and commercially painful, but they occupy the only point in the pipeline where physical custody and domain knowledge still meet.

A sane procurement standard would make buyers disclose the beneficial customer, intended scan method, destruction policy, and preservation plan. It would require non-destructive handling for scarce editions and deposit a verified preservation copy with a library or public archive before any destructive conversion. It would record source edition and page-level provenance instead of reducing a book to anonymous text in a data mix.

This would slow acquisition. Good. Speed is the source of the damage.

The weirdest part of the story is the loop. Models contaminate the open web with cheap synthetic text. Labs then prize older books because their publication dates certify human origin. The industry reaches backward into physical culture to escape the pollution it created, while using a procurement system capable of removing the clean source material from circulation.

A corpus can preserve words while destroying the object that proves where those words came from. That is extraction wearing the clothes of digitization.