Back to news

Microsoft Executive Claims OpenAI's Scraping is the Largest Labor Theft in History

The New York Times reports that in newly unsealed summaries from the copyright lawsuit between OpenAI and Microsoft, Brent Hecht, Microsoft's Director of Applied Science, referred to the large-scale scraping of news for model training as an "unprecedented astonishing theft" in an internal memo from January 2023, and stated it could be the "largest labor theft in human history."

The unsealed materials indicate that OpenAI's mid-training dataset contains at least 91,692 copies of works from The New York Times, Daily News, and the Center for Investigative Reporting; a dataset sourced from Common Crawl includes over 2 million documents from nytimes.com alone. Both parties are also accused of scraping plaintiff content through Bing indexing, with OpenAI providing the entire GPT-3 training data to Microsoft for commercial evaluation, while Microsoft supplied training data to OpenAI through projects like Project Taxi and Project Mango.

Nick Turley, head of ChatGPT, wrote in an internal memo in June 2023 that AI poses a "survival threat" to publishers; in February 2024, he noted that related products are "generally substitutable" and will become increasingly so as they improve. Internally, ChatGPT has been referred to as a "modern newsstand." A Microsoft document from 2023 warned that large models "sucking up" others' labor could be perceived by the public as astonishing theft, potentially creating a "doomsday loop" that harms both model performance and the entire web, as end products threaten the economic foundation of their content supply chain.

Employee Nick Ryder informed President Greg Brockman about discovering a "hack to bypass The New York Times paywall," to which Brockman replied, "ah nice." The publisher claims the company also used covert methods to scrape content behind paywalls and removed copyright management information from the training data. Microsoft CEO Satya Nadella testified that materials behind paywalls "should be authorized by the user"; had he known at the time that OpenAI was training on paid content, he would have exercised contractual rights to require retraining of the model. He also acknowledged that the chatbot provides information directly to platforms, eliminating the need for users to visit the original sources.

Hecht further stated that extensive news scraping renders the concept of "fair use" a complete mockery, claiming that large models are "products that destroy their own supply chains." OpenAI engineers noted in internal discussions in 2023 that no matter how prominently links are displayed, users will not click on them. In 2020, then-policy head Jack Clark warned Brockman and Sam Altman in a memo that the system would increasingly replace the labor of those defining society's "culture." Most original documents remain sealed, with only excerpts from the lawsuit publicly available.

From a market mechanism perspective, this is a battle for copyright pricing power driven by events: publishers seek to reclaim diverted subscriptions, advertising, and licensing fees; model providers turn unpaid text into substitutable reading products, keeping clicks at the dialogue layer rather than returning to the original sites. Beneficiaries are model companies and their cloud partners that control training power and distribution access; those under pressure are newspapers and investigative agencies that rely on original articles to maintain editorial budgets. The funding path is "first swallow data, then sell capability"—externalizing training costs while internalizing licensing costs. If the court determines that bypassing paywalls and large-scale copying exceed fair use, the licensing market will be forcibly opened; if fair use is upheld, scraping will continue to depress content supply prices and shift traffic from sites to model interfaces.

On the supplementary data level, the lawsuit also points out that self-built corpora like WebText and WebText2 are overly reliant on news sites, with Microsoft and OpenAI exchanging indexes and corpora, creating a closed loop where the same content is copied into different generations of models multiple times.

Source: Public Information

ABAB AI Insight

Microsoft's strategy regarding OpenAI has been to transform the latter from a non-profit lab in 2015 into a model supplier that can be embedded in Azure and Office: first supplying computing power and distribution, then tying training corpora and product evaluations to the same cloud. Brent Hecht is not the decision-maker but plays a role as a "asymmetric perspective" internally, with his memo directly labeling scraping as theft, indicating that while the company publicly insists on fair use, it has internally calculated the repercussions of news organizations' bankruptcy on the quality of future corpora using supply chain language. This aligns with Bing's dual use of web indexing as both a search asset and training material.

The capital mobilization method is "data for valuation, valuation for computing power." Microsoft supplies training data and supercomputing to OpenAI, which in turn returns the entire GPT-3 data package to Microsoft for product evaluation, effectively removing the copying costs of news texts from the licensing market and embedding them into model weights and subscription revenues. Nadella later stated that paywall content should be authorized, or retraining could be requested, acknowledging that the contract included a correction clause, but the correction occurs after the product has already replaced clicks. The motivation is not merely to reprint individual articles but to make "source-less reading" the default interface, turning authorization negotiations into after-the-fact payments.

A similar structure has occurred before: Google News snippets keep clicks on the search page, and publishers retaliate with antitrust and neighboring rights legislation; the music industry litigated against Napster and YouTube's early unauthorized libraries, ultimately transforming copying into revenue sharing. Cases like Anthropic have taken a different path: some courts have accepted that training on published materials can constitute fair use. The news industry is currently caught between the late stage of expansion and the early stage of control—model companies are still expanding corpora and users, while publishers attempt to use class action lawsuits to shift the industry from "default to scrape" to "authorize before training."

Structurally, this represents a transfer of pricing power combined with a reconstruction of the industry chain. The mechanism is: the marginal cost of original texts is high, while the marginal cost of copying is nearly zero; once the dialogue interface can fulfill the function of "reading this news article," the pricing power of site advertising and subscriptions shifts from producers to aggregators. The doomsday loop speaks not of morality but of the supply curve: not paying will reduce the availability of high-quality new articles, and the next round of models can only consume increasingly scarce, older, and more homogeneous texts, meaning that today's saved licensing fees will turn into tomorrow's data depreciation. The court must decide whether this layer of depreciation should be accounted for by model companies or continue to be recorded on the newspapers' books.

ABAB News · Cognitive Laws

  1. The end of free training is to incorporate suppliers into one's cost function.
  2. Substituting clicks is more lethal than copying sentences.
  3. Once fair use is scaled, it becomes a relocation of pricing power.

Source

·ABAB News
·
9 min read
·11 hrs ago
分享: