AI’s Silent Bonfire: How Tech Giants Are Shredding Rare Books to Fuel Tomorrow’s Models

Books have always carried risk. Floods ruin them. Fires consume them. Wars scatter their pages. Yet few imagined the latest threat would come from sleek offices in San Francisco, where engineers chase the next leap in artificial intelligence.

Millions of physical volumes now vanish into industrial scanners. Spines sliced by hydraulic cutters. Pages fed at high speed. Remains shredded or recycled. The process bears a clinical name: destructive scanning. Its purpose? Feed clean, human-written text into large language models before the internet drowns in synthetic content.

The Scale of the Operation

Anthropic stands at the center of this effort. Court records unsealed in the copyright lawsuit Bartz v. Anthropic PBC laid bare Project Panama. Launched in early 2024, the initiative aimed to “destructively scan all the books in the world,” according to an internal planning document. The same memo warned employees to keep quiet. “We use a ‘soft codename’ for it because we don’t want it to be known that we are working on this,” it read. “This document is visible to all Anthropic employees, but you should avoid talking about it in public areas, and the fact that we are working on this should not be shared with anyone outside Anthropic.”

Vendor proposals spoke of handling 500,000 to two million books in six months. A hydraulic-powered cutting machine removed spines. Industrial scanners followed. The physical copies? Disposed of or recycled. Anthropic even hired Tom Turvey, former head of Google Books partnerships, to secure “all the books in the world.” Purchases flowed through used-book wholesalers such as Better World Books and World of Books.

Other firms joined the hunt. Services like ISBNdb broker orders from 1,000 to one million titles. They target older, rare and specialist volumes. Pre-2022 books carry special value. They remain “structurally guaranteed” free of AI-generated text. As one broker’s site put it, per reporting in the Dallas Express, “The world’s best AI training data is sitting on a shelf.” The same site acknowledged the optics. “The optics problem is real. ‘AI company destroys two million books’ is not a headline that generates sympathy.” It offered non-disclosure agreements and suggested clients describe the work as “digital preservation.”

Booksellers noticed immediately. One told Decrypt via Yahoo Finance that weekly sales jumped from about 20 books to several hundred. The inventory included hard-to-find foreign language works and out-of-print titles. “It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell,” the seller said. “On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped.” Similar bulk orders hit dealers in the Netherlands, Switzerland, Spain and Germany. One Singapore-based buyer requested 3,000 English titles spanning folklore to technical manuals. No collector would want such a random mix.

And. These aren’t mass-market paperbacks alone. Rare editions with few surviving copies enter the pipeline. Books that endured centuries, wars, fires. Once pulped, they disappear. The digital scan becomes the sole record. Subject to whatever interpretation or alteration the model applies later.

A federal judge in San Francisco gave the practice legal cover. In the Bartz case, U.S. District Judge William Alsup ruled that buying print books, digitizing them and destroying the originals qualified as transformative fair use. Only one copy existed at any time. No new copies proliferated. Similar findings emerged in cases against OpenAI and Meta. Yet a separate matter produced a $1.5 billion settlement. Anthropic agreed to pay thousands of authors roughly $3,000 per book for using pirated digital copies to train its Claude model. The distinction matters. Legal purchase plus destruction passes muster. Piracy does not.

But the destruction continues. Reports surfaced this week across platforms. The Herald framed the episode as the Library of Alexandria burning once more. Author Derek McArthur highlighted the irreversible cultural cost. Books that survived history’s scourges now face shredders in service of chatbots. Public pressure or fresh legal shifts could change course. Without them, the losses mount.

Even Elon Musk weighed in. On X he directed his xAI team “to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning.” The remark drew mixed reactions. Some praised the stance. Others noted his own companies’ data practices and past controversies around generated content.

The hunger for training data grows desperate. The open web has been scraped clean. Pirated libraries triggered lawsuits. Now physical shelves supply the gap. Pre-AI text offers purity that synthetic slop cannot match. Models trained on AI output risk collapse into repetitive loops. Human books from earlier eras provide the antidote. At least until those run dry.

Small wonder the industry prefers silence. NDAs. Soft codenames. Anonymous brokers. The phrase “AI company destroys two million books” lands poorly with authors, librarians and readers. Yet the economics prove compelling. Destructive scanning runs faster and cheaper than careful, non-destructive methods. One copy in, one digital file out. Courts largely agree it passes fair use.

Still, questions linger. What happens when the last physical copy of an obscure 19th-century text vanishes? Does the scan preserve every nuance of layout, annotation or paper quality? Can future scholars trust a version born from a machine that may later edit or summarize at will? These concerns echo beyond technology. They touch memory itself.

Booksellers face their own dilemma. The surge clears warehouses and boosts revenue. Many welcome the business even as they lament the destination. One dealer cleared foreign-language stock that had sat for years. Profit clashed with principle. The pattern repeats across Europe and North America.

Recent coverage adds urgency. Stories published in the past week describe the same frenzy. Yahoo Finance compared the scene to “Fahrenheit 451,” where firemen burn books rather than save them. The parallel feels uncomfortable. Here the burners wear hoodies and optimize neural nets.

Defenders argue progress demands sacrifice. Better models improve medicine, science and daily tools. Training data scarcity threatens that path. Yet the counterargument grows louder. Cultural heritage isn’t raw material. Rare volumes aren’t mere data points. Their physical survival carries meaning separate from any digital twin.

So far courts favor the scanners. One-for-one replacement avoids the sin of extra copies. Destroy the original and fair use holds. The $1.5 billion settlement addressed piracy, not purchased destruction. That boundary may face new tests as more authors and publishers object.

Alternatives exist. Non-destructive scanning. Partnerships with libraries that retain originals. Dedicated preservation projects. Musk’s instruction points one direction. Scan carefully. Keep the books. Build libraries rather than bonfires. Whether others follow remains uncertain.

The Library of Alexandria legend tells of irreplaceable knowledge lost to flame. Modern versions arrive not with torches but purchase orders and NDAs. The fire burns slower this time. Pallet by pallet. Title by title. The ashes still accumulate. And the machines keep learning.


Discover more from Web and IT News

Subscribe to get the latest posts sent to your email.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top

Discover more from Web and IT News

Subscribe now to keep reading and get access to the full archive.

Continue reading