Over a dozen organizations urge FTC to investigate AI companies for destroying books
More than a dozen civil society organizations have written to the U.S. Federal Trade Commission (FTC), requesting a review of the practices of large AI companies that purchase, scan, and destroy physical books to obtain training data. They describe this as a destructive method of data acquisition that may constitute unfair competition.
These organizations include Demand Progress Education Fund, Consumer Federation of America, and Institute for Local Self-Reliance. The open letter points out that Anthropic was previously revealed through court documents to have spent tens of millions of dollars on bulk book purchases, removing spines to scan pages for training Claude, and subsequently destroying the physical books for recycling. An investigation by 404 Media also tracked Amazon's VGT3 team at the LAS8 warehouse in Las Vegas receiving large quantities of books, cutting bindings, scanning, and then destroying the originals.
The open letter emphasizes that some rare or out-of-print books may permanently disappear, with digital companies holding the last physical copies. This hoarding and destruction practice raises the costs for competitors to acquire similar original materials, cutting off key non-renewable resources that startup AI companies rely on for training models, thereby expanding competitive barriers for leading firms.
The organizations urge the FTC to use its regulatory powers, including Section 5, to determine whether these actions constitute unfair methods and to ascertain the scale of destruction and the number of remaining copies. However, they explicitly do not seek to limit AI model training itself, focusing only on the destruction of existing works, hoping for intervention before large companies establish systemic moats.
Google, Microsoft, and OpenAI have also faced lawsuits over similar copyright issues, while internal documents from Anthropic refer to the related initiative as Project Panama, which aims to process 500,000 to 2 million books within six months.
From a market mechanism perspective, AI giants are acting as buyers, acquiring physical books in bulk from second-hand booksellers and online platforms and permanently removing them from market supply. This situation is driven by the quality of training data and copyright risks, with funds flowing from tech companies to booksellers and scanning service providers, ultimately directed towards model training computing power. Leading firms benefit from exclusive digital datasets and higher entry barriers, while startups and the public face pressure due to reduced availability of original texts.
Source: Public information
ABAB AI Insight
Project Panama, initiated by Anthropic in 2024 and led by former Google Books executives, explicitly aims for "destructive scanning of all books" and requires confidentiality. Previously, the company acquired over 7 million pirated e-books through channels like LibGen, later paying $1.5 billion in settlements due to copyright lawsuits. Amazon has established a dedicated cutting and scanning team at its Las Vegas facility, where workers scan ISBNs before dismantling books, continuing the large tech companies' path of large-scale physical control over high-value texts.
From a capital perspective, these companies invest tens of millions to over a hundred million dollars to bulk purchase second-hand and out-of-print books from sources like Better World Books and World of Books, then pay for scanning and logistics costs, transforming limited non-renewable texts into proprietary digital assets. Their motivation lies in avoiding online low-quality and AI-generated content pollution while utilizing the first sale doctrine to reduce copyright risks, with resources flowing from open-source markets to closed training corpora.
Similar cases can be seen with Google Books, which used non-destructive scanning and returned library books, ultimately standing firm in copyright challenges. Currently, Anthropic and Amazon are transitioning from data expansion to controlling input resources, creating scarcity through physical destruction, making it impossible for later entrants to replicate high-quality text foundations at the same cost.
Structurally, this represents capital concentration and industrial chain reconstruction: once limited physical book supplies are permanently removed by leading firms, training data becomes a non-renewable barrier. The mechanism involves buying out and then destroying to cut off competitors' access, thereby transferring pricing power of text resources from the open market to a few giants that possess digital copies.
ABAB News · Cognitive Law
- Destroying supply creates barriers.
- Once scarce resources are privatized, competition shifts from price to thresholds.
- The ultimate form of a data moat is to leave competitors without resources.