Flash News

Amazon's Book Disassembly Scanning, Suspected for AI Training

404 Media tracked an order of about 1,000 Biblio books, which were ultimately sent to Amazon's LAS8 warehouse in North Las Vegas. Investigators collaborated with second-hand and rare book dealers who received the order, embedding an AirTag in one of the books; tracking showed that the book arrived at LAS8 after being transferred through multiple states, where a team named VGT3 operates. According to employee posts cited by the investigation, VGT3 cuts off the spines of large batches of physical books to expedite page-by-page scanning; this process permanently damages the original physical books. Amazon confirmed that it purchases books through commercial channels, stating the purpose is to "help develop and improve products and services used by customers"; the company did not confirm which AI models, products, or training processes the scanned texts specifically enter. Amazon did not disclose the total number of books procured for the project, the project's start time, selection criteria, or whether rare, out-of-print, or scarce editions were excluded before disassembly. Existing public reports also do not prove that every book sent to the warehouse has been scanned or destroyed. Market mechanisms indicate that the demand for high-quality, long texts with low AI contamination for AI training drives bulk acquisitions of second-hand and out-of-print books, providing dealers with one-time selling opportunities, while buyers of scarce editions and libraries face pressure from declining circulating physical stock. By procuring physical copies and digitizing them in-house, Amazon internalizes the content acquisition, processing, and model training chain, potentially benefiting its AI products and cloud model service system. Source: Public information

ABAB AI Insight

Amazon's starting point is book retail: the company launched as an online bookstore in 1995 and subsequently expanded into integrated online retail, AWS, and digital content through book catalogs, inventory fulfillment, and reader search and recommendation systems. The tracked procurement of physical books indicates that it is repurposing the book supply chain built in its early years for data acquisition in the model era; the difference is that books were previously traded commodities, while the reported process transforms them into machine-processable text inputs. The path of data capital does not stop at buying books. Amazon's centralized reception, disassembly, and scanning at LAS8 allow book procurement, text digitization, and subsequent computing resources to connect with its AWS and proprietary AI product systems. For a large number of old, out-of-print, or long-tail publications that cannot be obtained through licensing, purchasing physical copies is faster than negotiating with copyright holders one by one, but whether this constitutes a licensed source for model training still depends on copyright law and specific usage; Amazon has not disclosed the final use of the scanned data. This stands in historical contrast to Google Books' large-scale book digitization: Google collaborated with libraries to scan millions of volumes, leading to long-term copyright litigation and settlement disputes. The difference is that the core output of Google Books is searchable bibliographies and snippet displays; the current competitive focus of generative AI is to convert long texts into training corpora to enhance model answering, reasoning, and content generation capabilities. Companies like Anthropic and Meta continue to face lawsuits and public scrutiny regarding the sources of training data and copyright, as book data has shifted from search index assets to model capability assets. Essentially, this represents a restructuring of the supply chain: model companies are re-evaluating publications from cultural commodities to training materials. The mechanism is that high-quality human writing features structural integrity, high knowledge density, and relatively low error rates, while publicly available internet texts are increasingly mixed with AI-generated content; as quality corpora become scarce, platforms with procurement budgets, logistics networks, scanning capabilities, and computing power access can centralize and process dispersed physical knowledge assets, gaining benefits at the model level, while publishers, authors, and physical book preservation systems face new rights and stock pressures. ABAB News · Cognitive Laws 1. Old assets, once computable, will be repriced 2. The scarcer the free data, the more expensive the physical assets 3. Content belongs to the author, value belongs to distribution and computing power.

Source

·ABAB News
·
4 min read
·10 hrs ago
分享: