ISBNdb Exposed for Bulk Purchasing Used Books for Anonymous AI Clients, Anthropic's 'Panama Plan' Faces Public Backlash Again
404 Media reveals that the book database service provider ISBNdb is providing bulk physical book purchasing services for AI companies, with single orders ranging from 1,000 to 1 million copies, while concealing buyer identities through strict confidentiality agreements. Its website and marketing materials emphasize that "used books are the best training data," assisting clients in systematically acquiring, scanning, and avoiding duplicate purchases by ISBN.
The report does not name specific AI companies, but Anthropic's previously exposed "Panama Plan" has resurfaced in public discourse: the company reportedly spent millions of dollars purchasing millions of physical books, cutting off the spines with hydraulic cutters, scanning the entire content, and then destroying the originals. Internal documents explicitly state that "we do not want the outside world to know about this plan." A federal judge in California, William Alsup, ruled that if books are legally purchased and the physical copies are destroyed after scanning, without distributing digital versions, it can be considered fair use.
The controversy centers on the possibility that rare and out-of-print books could also be caught up in this process: while ordinary commercially available books can be reprinted, once extremely rare editions are destroyed, their physical existence in the world is lost. Although the court has determined that "legal purchase + destructive scanning + destruction of originals" can meet fair use criteria, the public and creators remain strongly opposed to this systematic "book scanning and destruction" for the purpose of training data.
Source: Public Information
ABAB AI Insight
Historically, Anthropic's "Panama Plan" continues its aggressive approach to data acquisition: the company has massively scraped texts from piracy databases like Books3 and LibGen, and upon realizing the legal risks of piracy, quickly switched to a scheme of legally purchasing physical books, destructive scanning, and destroying the originals. Court documents indicate that its internal planning documents describe the Panama Plan as "an attempt to scan all the world's books in a destructive manner" and explicitly state that they do not want external knowledge of it, indicating that their core goal is not to preserve books but to maximize the acquisition of structured high-quality text training data.
In terms of capital pathways, ISBNdb acts as a "data asset procurement intermediary" alongside similar suppliers: it utilizes a vast bibliographic database to help AI companies select books published before 2022 that have not been contaminated by AI content for bulk purchases, providing a full-service process from selection, bulk buying, logistics, book disassembly, and scanning. AI companies use funds to convert physical books into high-value training data assets, while ISBNdb sells not just books but also "clean, high-quality corpus streams," with its revenue coming from each bulk order and additional services. This pathway effectively reconstructs traditional books from "cultural consumer goods" to "model training raw materials."
Compared to other historical cases, Google Books scanned library collections without disassembling books and defended its actions in years of copyright litigation based on public interest and transformative use. Anthropic's model treats books as consumable training materials rather than objects to be preserved—by purchasing, disassembling, scanning, and recycling, it reduces copyright negotiation and licensing costs, and also circumvents the need to pay systematic compensation to authors through fair use precedents. This practice is seen as "exploiting legal gray areas for large-scale data harvesting," which directly conflicts with the interests of traditional publishing and authors.
Structurally, this controversy essentially reflects the manifestation of "technological replacement + transfer of pricing power" at the text and copyright level: when AI companies can obtain high-quality training data through bulk book purchases and destructive scanning without paying ongoing copyright compensation for each work, the economic value of books shifts from "ongoing sales and copyright revenue" to "one-time absorption by AI as model parameters." Whoever controls large-scale bibliographic procurement and scanning capabilities can establish competitive barriers on high-quality text data while circumventing the pricing power of traditional publishing and copyright systems—the existing legal framework for fair use reveals structural vulnerabilities in the era of generative AI: as long as the conditions of "purchase + destruction of physical copies + no distribution of digital versions" are formally met, capital can legally transform the entire set of physical cultural assets into training materials for closed-source models, while creators and the public can only respond to "cultural loss" and "double standards" in public discourse.