Unsealed NYT Filings: Microsoft Director Called AI Training Scraping the Largest Theft of Labor in Human History
New unredacted material in the New York Times’ three-year copyright lawsuit against OpenAI and Microsoft, unsealed this week, contains internal admissions that the companies’ AI training and deployment practices constituted theft and posed an existential threat to publishers.
The most explosive document: a January 2023 internal memo from Brent Hecht, Microsoft’s director of Applied Science, in which he described the industry-wide practice of scraping the open web to train large language models as “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.”
Internal Documents, External Consequences
The unsealed filing — from the Times’ own brief rather than from underlying exhibits, which remain sealed — lays out a pattern of internal acknowledgment that directly undercuts the fair use defense both companies have advanced in court.
Hecht’s January 2024 internal presentation reveals data on what Microsoft Copilot has done to web traffic: click-through rates to the New York Times’ domain dropped as much as 93% compared to traditional Bing search. The document describes this as a “doom loop” that would “hurt the performance of our models and the entire web at the same time.”
“It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” the document reads, as quoted in the filing.
Scale of the Copying
The filing quantifies the scope of what was taken. OpenAI’s mid-training datasets contain more than 91,692 copies of works published by the New York Times, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone.
Two internal data assembly initiatives are named. “OpenAI delivered the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI’s models within its own commercial products,” the filing reads. “Microsoft similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango.” The Project Mango dataset contains copies of at least 160,903 unique works from the news publishers.
The companies also allegedly bypassed paywalls undetected, assembled training datasets via mass scraping, and deliberately stripped copyright notices from training data.
What the Executives Said
The unsealed material contains statements from multiple executives that cut against the substitution prong of the fair use test — the requirement that use not substitute for or harm the market for the original work.
Nick Turley, OpenAI’s head of ChatGPT, wrote internally that publishers face an “existential threat” from products like ChatGPT, which are “largely substitutive” and “will get more and more substitutive as they get better.”
Greg Brockman described the models as “excellent at news.”
Satya Nadella testified in a deposition that “anything that is paywalled should be licensed by anyone who wants to use it for grounding or training,” and said that if he “had been made aware that OpenAI had scraped and trained on information that was behind a paywall,” he would have “invoked Microsoft’s right to require OpenAI to retrain its models.”
Nadella also agreed under oath that conversing with chatbots “has substituted giving you the information right there on the website on the AI platform versus needing to go to the underlying source.”
A Microsoft document states that there is a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.”
The Legal Backdrop
Fair use cases hinge on four factors, with the substitution test particularly important. Multiple internal statements in the unsealed filing directly address this factor — and in each case, the companies’ own executives concluded the products were substitutive.
The unsealing comes as judges have been broadly favorable to AI companies’ fair use arguments. Earlier this month, the Trump administration filed a brief in defense of OpenAI’s unlicensed use of copyrighted material. The new disclosures have complicated that picture by demonstrating that the companies themselves, privately, did not believe what they were arguing publicly.