Get Off My Back: AI Destructive Book Scanning, Training Data and the Provenance We Are Losing
AI companies destructively scan books to train models. Source founder Angelina Giovani-Agha on what's lost when the text survives but the book doesn't.
Every time a book is destroyed, a provenance disappears with it, and although that provenance may not appear significant to the person holding the book at that particular moment, its disappearance is irreversible because historical value has an inconvenient tendency to reveal itself only after the evidence that establishes it has been recognised, connected and understood.
I have spent most of my professional life working with provenance, which means learning never to dismiss the apparently insignificant mark, inscription, label, number or annotation simply because its meaning is not immediately obvious. In art history we accept that the reverse of a painting can sometimes tell us as much as the front, because a faded dealer label might reconstruct a transaction, a handwritten number might reconnect an artwork with a dispersed collection, and a stamp nobody noticed fifty years ago might become decisive once an archive makes it legible.
Books deserve the same intellectual seriousness, because a book is not simply the text contained between its covers but a physical source with a history of its own, one that may include an ownership inscription, a library stamp, marginalia, a dedication, a bookseller's ticket, an unusual binding, inserted correspondence, annotations, corrections, signs of censorship, evidence of use, or any number of other traces left by the people and institutions through whose hands it has passed.
This is why the practice of destructive book scanning for artificial intelligence troubles me far beyond the already substantial debates surrounding copyright, authorship and AI training data.
Court documents concerning Anthropic revealed the existence of Project Panama, described internally as an effort to “destructively scan all the books in the world.” The actual operation involved the acquisition of huge quantities of physical books, after which bindings were stripped away, pages were cut to dimensions suitable for scanning, digital copies were created, and the original print copies were discarded. The court itself described a process in which the digital version effectively replaced the physical copy in Anthropic's internal research library.
The story has since widened beyond a single company. Booksellers in Britain, Ireland, Australia and the United States have reported unusual bulk purchases of obscure, older and sometimes difficult to replace titles, leading to growing concern about the supply chains through which books may be entering artificial intelligence data acquisition programmes, although attribution must be handled carefully and companies including Anthropic have stated that their programmes do not buy and destroy rare or antiquarian books.
What interests me most, however, is the philosophy concealed inside the logistics, because before anyone cuts the spine from a book, somebody has already made a decision about what part of that book matters.
A book contains more than its text
If we understand a book purely as a container for language, then destructive scanning has an impeccable internal logic. The words are valuable, the paper carries them, the binding keeps the pages together, and once those words have been captured accurately enough to become machine readable data, the physical structure that delivered them has completed its function.
There is only one problem with this reasoning, which is that books have never been merely delivery systems for words.
Every physical book contains at least two histories. There is the history written in it, whether that is a novel, an exhibition catalogue, a scientific argument, a memoir, a local history or a work of scholarship, and there is the history that has happened to that particular copy while it has existed in the world.
The first can frequently be digitised extremely well, whereas the second is far more difficult to preserve because provenance is contextual, material and often invisible until somebody knows what they are looking for.
The handwritten name on a flyleaf may mean nothing today and become enormously important when another archive is catalogued tomorrow. An apparently insignificant annotation might acquire meaning once the identity of the reader is established. A cheap exhibition catalogue might contain the only surviving evidence connecting an artwork with a particular collector, dealer or exhibition. A mass produced book can therefore become a singular historical object because uniqueness does not depend exclusively upon how something was manufactured; it can also be created by everything that happens afterwards.
This is the problem with deciding, at industrial speed, that a physical source has no further value once its textual content has been extracted.
Historical significance rarely announces itself before we destroy the evidence.
When preservation becomes extraction
For much of the history of library digitisation, copying has been associated with preservation and access. Fragile manuscripts were photographed so that researchers could consult them without repeatedly handling the originals, deteriorating newspapers were microfilmed before their paper failed, and digital collections allowed a researcher in London to examine material held thousands of miles away without threatening the survival of the object itself.
The underlying relationship was straightforward: reproduction allowed the information to travel while preservation allowed the source to remain.
Destructive scanning reverses that relationship, because the copy is no longer created to protect the original; the original is physically dismantled in order to create the copy.
This is a much more important reversal than it initially appears, because it represents a shift from preservation to extraction, in which the purpose is no longer to make a source accessible while safeguarding it, but to obtain the component of the source that has been identified as useful.
The book is gradually translated into another vocabulary. Its pages become images, the images become text, the text becomes tokens, the tokens become training data, and the training data contributes to a model, while everything that cannot participate in that conversion risks being treated as residue.
This is the thinking behind our Source campaign, Get Off My Back, which refers literally to the backs of books being cut away but also to the broader idea that materiality itself has begun to be treated as an inconvenience standing between technology and the information it wishes to extract.
When the library becomes a dataset
Martin Heidegger's philosophy of technology feels unexpectedly relevant here because his concern was never simply that machines might become more powerful, but that technological thinking might change the categories through which we understand the world.
He described a process in which things increasingly reveal themselves as resources awaiting ordering and use, so that a forest is encountered as timber, a river as energy and land as mineral potential.
The contemporary equivalent is remarkably easy to imagine: a library becomes data.
Once the book has been conceptually transformed into latent training material, the physical features that previously constituted the book begin to look inefficient. The binding obstructs the scanner, the spine slows the process, page turning consumes time, and the object's integrity becomes friction within a system designed around extraction.
At that point, cutting the spine can stop looking like destruction and begin looking like optimisation, which is precisely why the language we use matters so much.
“Destructive scanning” has a clinical neutrality that obscures an uncomplicated physical reality: when the process is complete, that particular book no longer exists.
The fire and the scanner
The history of book destruction makes this transformation especially uncomfortable, although the comparison requires care because the motives behind historical library destruction and artificial intelligence data acquisition are profoundly different.
The libraries of Warsaw, Jaffna, Sarajevo and Mosul were destroyed in circumstances involving war, ethnic violence, political domination and ideological persecution, and it would be both historically and morally wrong to suggest that a corporate scanning operation belongs in the same category.
The comparison becomes interesting precisely because the intentions are so different.
The historical biblioclast often destroyed books because the ideas, identities or memories contained within them were considered dangerous and therefore had to disappear. The artificial intelligence company wants the contents of books for almost the opposite reason, because professionally produced human writing is extraordinarily valuable material from which language models can learn. Anthropic's own search for books reflected their value as complex, organised, long form human text.
The traditional biblioclast destroys the book because the knowledge must not survive, whereas what I would call the extractive biblioclast destroys the book after ensuring that the knowledge has been captured.
One seeks erasure while the other seeks ingestion, but both require a decision about which aspect of the book is allowed to continue into the future.
The historical destroyer concludes that neither the book nor its knowledge should survive. The extractive system decides that the informational content should survive while the material source may not.
The fire and the scanner therefore operate in opposite directions, yet they lead us towards the same philosophical question about who has the authority to determine the form in which human memory deserves to endure.
Reproduction without preservation
Walter Benjamin famously examined what happens to cultural objects when technologies of reproduction separate them from their original material circumstances, although the printed book complicates his idea of the unique original because books have always existed as multiples.
What feels particularly strange about destructive AI scanning is that it produces something more radical: reproduction that depends upon the destruction of the object being reproduced.
For centuries, one of the great advantages of reproduction was redundancy. Copies protected knowledge from the vulnerability of a single object, which is precisely why texts could survive fires, wars, accidents and the deterioration of individual volumes.
We have now arrived at a situation in which the creation of the digital copy can itself become the reason that the physical copy is destroyed.
The technology that should make preservation easier has, under a different economic logic, made destruction efficient.
From circulation to enclosure
There is another loss that is harder to quantify because books do not simply exist; they circulate, and through circulation they acquire biographies.
Someone buys a book, reads it, writes inside it and eventually sells it. Another person finds it twenty years later, carries it somewhere else, places something between its pages and eventually gives it to a library. Its price changes, its associations accumulate and its historical meaning develops through encounters nobody could have predicted when it left the printer.
A book acquires provenance because it remains available for history to happen to it.
A copy purchased for destructive scanning enters a terminal transaction because there is no next reader of that physical object, no later bookseller, no future donation and no researcher discovering it unexpectedly on a shelf fifty years from now.
Its language may continue to generate enormous value inside a proprietary computational system while the biography of the object from which that language was extracted has permanently ended.
This is why the concept of enclosure is useful, because knowledge that once circulated through a physical object between human beings is acquired, extracted and incorporated into private computational infrastructure, after which its informational value can continue to expand while the source itself has disappeared from circulation.
AI needs books because books are human
Perhaps the greatest irony is that artificial intelligence wants books precisely because they represent such concentrated expressions of human intelligence.
Books contain the accumulated labour of writers who spent years constructing ideas, historians who searched archives, scholars who verified evidence, editors who refined arguments, translators who moved thought between languages, publishers who exercised judgement, librarians who catalogued collections and readers who preserved particular copies long enough for them to reach us.
That human density is precisely what makes books so attractive as AI training data.
We have spoken casually for years about “feeding” information into artificial intelligence, but destructive scanning gives the metaphor an uncomfortable physical reality because the machine is being taught to write by consuming the objects through which generations of human beings taught one another how to think, while the process of consumption can involve dismantling the very artefacts from which that knowledge was obtained.
This does not mean that every book should be preserved forever, because libraries must deaccession material, publishers pulp stock, collections change and countless genuinely interchangeable copies exist.
The argument is neither sentimental nor anti technology.
It is about intellectual humility.
We should be cautious about assuming that because we have extracted the information we currently recognise as valuable, we have exhausted everything a physical source might ever be able to tell us.
Why this matters to Source
Source exists because much of the world's evidence remains physical, non digitised and inaccessible through ordinary search engines, which is why our researchers physically enter archives, libraries, registries and private collections to retrieve primary sources that databases cannot reach.
That work continually reminds me that information and evidence are not synonymous.
Information can be copied, summarised, indexed and generated, whereas evidence has context, materiality, location and provenance, and those qualities frequently become important only when a researcher arrives with a question that nobody previously thought to ask.
There are archives that have not yet been catalogued, relationships that have not yet been reconstructed, inscriptions whose names remain unidentified, and future researchers who will approach familiar objects with methodologies that do not yet exist.
When we destroy a physical source, we therefore lose more than the information we failed to capture at the time. We also lose the possibility of asking that source a different question in the future.
If you want the wider picture of why so much evidence like this still isn't digitised, our complete guide to archive research covers the same access problem across genealogy, due diligence and academic research, and our piece on artwork provenance research goes into how this plays out specifically for art.
This is ultimately what Get Off My Back is about, because the danger of destructive book scanning is not simply that paper is being discarded, but that we are becoming increasingly comfortable with a philosophy in which the informational value of an object is assumed to exhaust its cultural value.
Perhaps the defining image of book destruction in the twenty first century will therefore not be the spectacle of the bonfire, but something much quieter: a warehouse filled with books, a blade removing their spines, a scanner capturing page after page, and a system recording everything it has been designed to recognise while discarding the object that might one day have told us something it was never designed to see.
For centuries, we have worried about those who wanted to destroy books because they feared the knowledge contained within them.
The challenge before us now is more subtle, because we must ask what happens when we value the knowledge inside a book so highly that we convince ourselves the book itself has nothing left to say.
If there's a record, an inscription, or a detail you need checked in a book that only exists in someone else's hands, you can post a request and get matched with a researcher who already has access to it.
Post a request