AI book scanning and the so-called destruction of rare books has dominated headlines for weeks, but one word in that framing deserves a hard look: rare. The books being purchased, de-spined, digitised and discarded by AI companies are overwhelmingly post-1970 ISBN-catalogued volumes, and the distance between that reality and the image of irreplaceable monastic manuscripts being fed into a shredder is rather large.
What AI Book Scanning Reveals About ‘Rare’
The operation at the centre of this story has a name. According to The Washington Post, an internal planning document described the effort as follows: ‘Project Panama is our effort to destructively scan all the books in the world.’ The Dallas Express reported that Project Panama began in early 2024. The ambition declared in that internal document is breathtaking in its scope, and there are plenty of legitimate concerns about it. But ‘rare books’ may not be the sharpest of those concerns.
Here is the detail that cuts through the mythology: the books are being ordered by ISBN number. The International Standard Book Number system was only introduced in 1970. If a volume is being located and purchased by its ISBN, it is not a medieval illuminated manuscript. It is not a hand-copied codex. It is a mass-produced commercial publication from the last five decades, available through ordinary bookselling channels. The word ‘rare’ has been doing a great deal of heavy lifting in the coverage, and it does not quite bear that weight.
The aim appears to be training data that is guaranteed pre-2002, and therefore free of AI-generated text. That is a sensible technical goal from a model-quality perspective. The content of concern is post-1970 commercial publishing, not the contents of a rare-book room.
The Copyright Strategy and Where It Stands in Court
The destruction of the physical copies is not random vandalism. There are two practical reasons for it. First, as anyone who has ever tried to digitise a bound volume will know, it is considerably easier to feed loose pages through a scanner than to wrestle with a spine. Removing the binding first is simply an efficiency decision. Second, and more legally interesting, if a company buys a book, digitises it, and destroys the physical copy, it can argue that only one copy of the work exists at any moment, and hopes, by that route, to sidestep copyright claims from publishers.
Whether that strategy holds legally is, as of the time of writing, a live question with at least one significant ruling already on the books. Campus Technology reported that Judge William Alsup of the U.S. District Court for the Northern District of California ruled that Anthropic’s use of copyrighted books to train its large language models constituted fair use under copyright law. The authors who brought that lawsuit were Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson, according to TechPolicy.Press, which also noted that the ruling left ample room for future defeats on related questions. The legal picture, in other words, is not settled.
The copyright angle is the genuinely consequential argument here. The destruction of the physical copies matters primarily because of what it signals about the companies’ legal reasoning, not because the books themselves are irreplaceable artefacts.
Books Are Pulped in Their Millions Every Year
There is a broader context that rarely appears in the outrage cycle. The publishing industry has been a mass-production operation for centuries, and the world has always been awash with old books. Every year, vast quantities of unsold new books are pulped by publishers, and enormous numbers of older volumes are recycled by the paper industry. This is not a secret. Anyone who has worked in or near publishing is familiar with the economics of print runs and returns.
The books most likely to be targeted by AI scanning operations are, almost by definition, the ones nobody has particularly valued up to now. Any post-1970 volume that is both genuinely scarce and genuinely prized will have found its way into a collection, a library, or an archive long before a bulk buyer with an ISBN list comes along. The books that are physically uncommon but commercially next-to-valueless are in that position precisely because the market decided their survival was not worth organising around.
The symbolism of book destruction is powerful and historically loaded, and it is entirely understandable that it triggers strong reactions. The 1933 photograph taken in Bebelplatz in Berlin is not a neutral image. But the symbolic charge of book burning does not automatically transfer to a commercial digitisation operation working through post-1970 ISBN-catalogued paperbacks. The two things are not the same, however uncomfortable it may feel to say so.
There are serious questions to put to AI companies about Project Panama: about consent, about compensation for authors, about the legal durability of the fair-use argument that Judge Alsup’s ruling has, for now, accepted. Those are the questions worth pressing. The ‘rare books’ framing, however satisfying as shorthand, risks directing attention toward the wrong target entirely.

