By John Wayne on Thursday, 30 July 2026
Category: Race, Culture, Nation

The Great Book Burning of 2026

 There is a particular kind of horror in watching something irreplaceable destroyed for something disposable. The large language models these companies are training will be obsolete within a few years, replaced by the next architecture, the next funding round, the next pivot, yet the books they are shredding were meant to last centuries. Judge William Alsup's June 2025 ruling in the Anthropic case is one of those decisions future historians may read with the same disbelief we bring to certain earlier judicial missteps. The logic is striking: turning a lawfully purchased physical book into a digital copy was treated as transformative fair use in part because the print version was replaced and the format change served the company's central library needs without multiplying copies or redistributing them. Destruction of the physical object, in this framing, became intertwined with the legal justification. The practical incentive that follows is clear. If retaining the original book complicates the analysis or creates storage burdens, the rational course is to scan and discard. Alsup's decision did not invent the buy-scan-dispose pipeline, but it placed a judicial blessing on a process already underway at industrial scale.

The urgency behind the rush for older printed books has a technical name: model collapse. When successive generations of models train primarily on text produced by earlier models, the output grows more formulaic, more repetitive, and more detached from the variety and texture of human writing. By the mid-2020s the public web had become so saturated with synthetic content: AI-generated articles, product descriptions, summaries, and commentary, that it had grown toxic for high-quality training. Pre-2022 printed books remained one of the last relatively clean reservoirs. They were written, edited, and published by humans before the flood of machine-generated text. They offer structured, often carefully curated knowledge that web crawls cannot easily replicate. The irony is thick. The same industry that accelerated the pollution of the open internet has rendered large portions of that internet less useful for its own purposes and turned instead to physical books as the remaining uncontaminated source, only to destroy many of them in the process of extraction.

The pipeline itself is clinical and large. Bulk buyers place orders measured in the hundreds of thousands or millions through intermediaries that emphasize nondisclosure. The selections are often subject-agnostic: history and botany sit beside regional law treatises and foreign-language textbooks in the same carts. Booksellers have described the orders as strange not only for their volume but for the indifference to content or condition. Once acquired, the books move to industrial processing. Hydraulic cutters remove spines; high-speed scanners process dozens of pages per minute; the loose sheets are digitised and the physical remains are recycled into fiber. The text continues as training data locked inside corporate systems, inaccessible to the public, to scholars, or to anyone outside the lab. Anthropic's Project Panama, documented in court filings, spent tens of millions on precisely this approach, contracting with scanning vendors and operating under internal instructions that treated the effort as something best kept quiet. The project was overseen by figures with prior experience in large-scale digitisation, including work connected to Google Books, an earlier effort that, whatever its controversies, at least aimed to make texts searchable without systematically destroying the physical objects.

What disappears most readily are the volumes with the least commercial value and, frequently, the greatest vulnerability. Rare and out-of-print titles can vanish when a bulk purchase claims the last few surviving copies; the text may persist as a file, but if the server is retired, the company fails, or the format becomes unreadable, the record is gone. Foreign-language works with small print runs: Dutch monographs, specialised German legal texts, collections of Persian poetry, are particularly exposed. AI developers value them for linguistic diversity, yet once the physical copies are pulped the communities that produced them may have no practical way to recover the originals. Low-circulation academic works, obscure dissertations, conference proceedings, and specialised reference volumes that never entered the major digital libraries exist only as physical objects. When the final copy is shredded, the knowledge does not pass into the public domain; it passes into a proprietary model that will not return the source text for ordinary reading or citation. Intermediaries themselves have acknowledged the public-relations difficulty: headlines about companies destroying millions of books do not generate sympathy. The preferred solution has often been opacity rather than preservation.

Non-destructive alternatives already exist. Harvard, working with Google and Microsoft, has released large collections of public-domain digitised books across hundreds of languages without requiring the destruction of originals. Libraries and cultural institutions have long demonstrated that careful, high-quality digitisation at scale is possible. The AI companies are not forced into destructive scanning by technical necessity. They choose it because it is faster, cheaper, and, after Alsup's ruling on lawfully acquired print copies, legally advantageous. The choice adopts immediate training data over the continued existence of the physical record.

Every civilisation that allowed its books to disappear thought it was being practical at the time. The Library of Alexandria did not vanish in a single dramatic fire; it eroded through accumulated decisions that valued short-term utility over long-term stewardship. Scrolls were scraped and reused; texts were lost not always from hatred but from neglect and the calculation that shelf space or labour was better spent elsewhere. Today's AI companies are not motivated by malice. They are motivated by convenience, by legal incentives that reward the conversion of books into training tokens, and by the judgment that the physical objects are worth more as data than as durable cultural artifacts. The question is no longer solely whether the practice is lawful: Alsup's decision settled much of that for purchased copies. The deeper question is whether a society that systematically converts its printed textual heritage into proprietary model weights, discarding the originals in the process, is exercising the kind of care that earlier generations expected of custodians of knowledge. The books being shredded: the last copy of an obscure botanical treatise, a rare medical text, a dissertation that represented years of human labour, do not belong exclusively to any corporation. They form part of the shared record of human thought. Allowing them to be reduced to recycled fibre for a training run that will itself be superseded in a few years is a choice, not an inevitability.

https://www.naturalnews.com/2026-07-28-ai-firms-destroying-rare-books-training-data.html