GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Amazon Is Buying Rare Books in Bulk, Stripping Their Spines, and Pulping the Rest

Amazon has been purchasing rare and out-of-print physical books in bulk, running them through a scanning pipeline at a warehouse in Nevada, and sending the physical copies for pulping and recycling. Workers at the facility describe stripping book spines to allow faster flatbed scanning — a process that destroys the physical object in exchange for a high-resolution digital file.

The practice, reported by 404 Media in July and confirmed by multiple sources since, puts Amazon alongside Anthropic and Google as labs engaged in systematic bulk purchasing of physical books for AI training data. The scale and operational infrastructure involved — a dedicated facility with workers specifically tasked with spine removal — indicates an organised, ongoing programme rather than opportunistic acquisition.

What the Archive Sees

Anna’s Archive, the largest shadow library aggregator, published a direct response this week calling for a coordinated effort to scan rare books before AI labs buy and destroy them. The post frames the situation as a race between preservation and consumption: once a physical book is pulped, any digital copy that existed before the purchase becomes the only surviving record.

Rare book dealers have noted unusual buying patterns — bulk orders from anonymous purchasers across used bookstores in multiple countries — since at least late 2025. The suspicion that AI companies were behind the purchasing predated confirmation by months.

Training Data at Scale

Physical books offer something digital text rarely does: pre-copyright-reform material, out-of-print academic works, and regional publications that never made it to the web. For frontier model training, the value is in coverage — text that did not appear in Common Crawl, Wikipedia, or public archives.

Anthropic’s training data practices became public via litigation; a $1.5 billion settlement with authors was approved in July 2026, with the court finding that digitising 7 million books stored on servers was not fair use, though training itself was. Google’s book scanning program predates the AI era and was itself litigated for over a decade.

Amazon entering this space at scale, with dedicated warehouse infrastructure, suggests the company views physical book acquisition as a competitive necessity rather than a supplementary option.

The Preservation Gap

No mechanism currently exists to require AI labs to deposit digital scans with national libraries or archives before pulping physical copies. The books being purchased are often unique — single surviving copies of regional histories, specialist technical manuals, and literary works printed in small runs decades ago.

Anna’s Archive’s call to action is essentially a crowdsourced counter-programme: identify rare holdings before they move through AI lab supply chains, and digitise them into public-access repositories first.