The Facility with a Dinosaur Eating a Book
Somewhere in Las Vegas, there is a warehouse that captures everything happening at the collision of AI and the physical archive: books go in, and data comes out. Not scanned-and-shelved, the way a library might digitise a fragile pamphlet. The books are dismantled.
Industrial equipment slices the spines clean off so the pages lie flat, and then the pages feed through high-speed scanners at a rate no intact book could survive. The operation belongs to Amazon, and the team running it is called VGT3.
Here is the strange part. The VGT3 facility has a logo. It is a dinosaur holding a book in its teeth.
Whoever designed that logo was either deeply self-aware or accidentally honest in a way that should give everyone pause. The dinosaur is not reading the book. It is gripping it, the way a predator grips something it has caught.
That image, small and cartoonish on a warehouse door in Nevada, turns out to be the most candid piece of corporate branding in the current AI moment.
The story took a while to surface, and it started not in Las Vegas but in the stock rooms of independent booksellers. Sellers across the country began reporting something strange: anonymous entities were placing unusual bulk orders for rare titles, foreign-language volumes, anything printed before the digital noise of recent years.
The orders were large. The buyers were opaque. Services like ISBNdb, which brokers book sales at scale, were apparently part of the pipeline.
Booksellers had no idea where the books were going. Now they do.
The logic is brutal in its simplicity. A physical book printed before 2022 is, to an AI developer, something genuinely precious: human-written text that no language model has touched. The internet is increasingly full of synthetic output, a kind of digital pollution that contaminates training data. Paper, it turns out, is clean.
Why the Internet Ran Out of Clean Words
Here is the strange part. The internet, which once seemed an inexhaustible reservoir of human expression, has begun to poison itself. Every time a language model generates text that ends up indexed online, it contaminates the pool that the next model will drink from.
Researchers call this model collapse: a degradation loop where each generation of AI trains on the errors and hallucinations of its predecessors, compounding them like a game of telephone played at civilizational scale. The signal degrades. The noise wins.
This is why 2022 has become a magic number in AI procurement circles. Printed books published before that year carry a particular guarantee: no generative AI existed yet to have written a single sentence in them. Every paragraph is provably human.
That guarantee, once something only a rare-book librarian would care about, now has a market price attached to it.
Now hold that thought, because the supply chain here is deliberately obscured. Services like ISBNdb operate as brokers, allowing AI companies to make anonymous bulk purchases of thousands to millions of volumes without revealing who is buying or why. Independent booksellers report unusual surges in orders for rare and foreign-language titles from entities they cannot identify. The buyer hides; the books move; the bindings come off.
What is happening is a quiet repricing of the physical archive. A volume that a university library might have valued for its scholarly annotations or its scarcity is now being assessed by a different metric entirely: how many clean, pre-synthetic sentences does it contain? The collectible value and the training value used to be unrelated. They are not anymore.
The Law That Says You Can Destroy What You Own
Here is the strange part. The legal principle protecting all of this is one designed to protect you reselling a used paperback. The first-sale doctrine holds that once you purchase a physical object, it is yours, entirely.
You may resell it, annotate it, or feed it through an industrial spine-slicing machine. The law does not distinguish between the paperback you sell at a garage sale and a pallet of rare editions processed in Las Vegas.
A judicial ruling has gone further, classifying the destructive scanning of legally purchased books for AI training as fair use. That is a consequential sentence. Fair use is the provision that lets a critic quote a novel, a teacher photocopy a poem, a scholar reproduce a diagram.
It was built for commentary, education, and transformation at human scale. Now hold that thought - because the scale here is millions of volumes, processed at industrial speed, the physical originals discarded once the pages are through the scanner.
The ruling does not cover everything. Legally purchased does not mean legally obtained in every case, and fair use is determined case by case. But for the practical reality of what is happening in facilities like VGT3, the framework holds.
The gap between legality and consequence is exactly what makes this uncomfortable. A law designed so that a reader owns what they buy now provides the architecture for something the legislators of that doctrine could not have imagined: a corporation owning the only remaining high-fidelity copy of a book, in digital form, after the physical original is gone.
A law designed so that a reader owns what they buy now provides the architecture for something the legislators of that doctrine could not have imagined: a corporation owning the only remaining high-fidelity copy of a book, in digital form, after the physical original is gone.
The Physical Archive That Is Swallowing Two Million Files
Picture a single archivist, sitting down on a Monday morning, opening a folder. The folder contains two million records. She has one working lifetime.
She does not stand a chance.
This is not a hypothetical. The National Archives and Records Administration in the United States faces something archivists call the "dark archive" problem: vast collections that have been digitized but remain invisible because nobody has written the metadata tags that make a record findable. A document without a tag is a book shelved with its spine facing the wall.
NARA's answer is Azure OpenAI, currently working through approximately two million digital records and generating those tags automatically, at a pace no human team could match in a generation.
Here is the strange part. The same cloud infrastructure that powers commercial AI is now doing the unglamorous, essential labor of keeping public memory legible.
NARA also runs Amazon Kendra inside its internal "NARA@WORK" system, so that an archivist can type a plain question - "find all records relating to the 1962 missile correspondence" - rather than memorizing the rigid logic of a legacy index.
Ask a question. Get an answer. The index works for you, not the other way around.
Across the Atlantic, the UK National Archives has made this direction official. Its 2025-2030 strategy formally designates the institution a "living digital archive," with AI written into routine appraisal and sensitivity review, not bolted on as an experiment but built into the spine of the operation.
The contrast with Las Vegas is stark and deliberate. These institutions are not consuming physical heritage. They are deploying AI to rescue heritage that would otherwise drown in its own unprocessed abundance.
The tool is the same. The direction is entirely different.
Teaching Machines to Read the Handwriting of the Dead
Most of the written past is, to a computer, pure noise. Standard optical character recognition - the software that reads a typed page in seconds - collapses entirely when it meets 19th-century Cyrillic script, cramped ecclesiastical hands, or ledgers whose ink has faded to the color of old tea. These documents are not lost exactly.
They sit in climate-controlled rooms, perfectly physical, perfectly illegible to any machine trained on clean modern fonts. The gap between "exists" and "can be read" is, for an enormous portion of the historical record, still wide open.
That gap is what the ArchXAI Project, running from 2025 to 2028, is trying to close. A multi-national collaboration, it is building Handwritten Text Recognition models - systems trained specifically to parse 19th- and 20th-century Cyrillic scripts. Here is the contrast worth holding in mind: while one set of institutions feeds physical books into spine-cutting machines to generate training data, another set is doing the harder, slower work of teaching AI to rescue scripts that no one thought to train it on in the first place.
The preservation work runs deeper than transcription. The Natural History Museum in the UK now watches its physical collections with a network of more than 50 environmental sensors, feeding live readings into AI climate-control models that respond to humidity and temperature at the molecular scale.
AI-driven vertical storage systems, meanwhile, can shrink an archive's physical footprint by up to 90 percent - which sounds like a triumph until you ask what disappears when the stacks do, when the serendipitous wander through a physical row of shelves becomes impossible. Efficiency and discovery have never been perfectly aligned, and they still are not.
One Institution Preserves. Another Incinerates. Now What?
The FRAIM project, involving the British Library and partner institutions, is currently drafting the ethical guidelines that should have existed before the first spine hit the cutting blade. That sequencing matters. The rules are arriving after the destruction, not before it.
The divergence could not be sharper. National archives are deploying AI as a tool of stewardship, tagging millions of records so they survive and become findable. AI corporations are deploying physical archives as disposable raw material, feeding them into scanners and discarding the husks.
Both sides call what they are doing "preservation of knowledge." Only one of them is preserving the objects.
Here is the strange part. We do not know how many rare or unique books have been destroyed at VGT3. We do not know whether the physical destruction is contractually required or simply the fastest way to run the machinery. Those answers may never surface.
The question is not whether AI belongs in the physical archive. It already does, irreversibly, and in many places it is doing genuine good. The question is who decides which books survive the encounter with it, and whether anyone will answer for the ones that did not.