Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal

New unredacted court filings reveal that a Microsoft executive described AI scraping as “the largest theft of labor in human history,” an…

By Vane September 17, 2026 4 min read
Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal

New unredacted court filings reveal that a Microsoft executive described AI scraping as “the largest theft of labor in human history,” an admission made in the copyright lawsuit filed by The New York Times against OpenAI and Microsoft three years ago.

Internal admissions

The unsealed documents show that a top Microsoft leader privately called the companies’ AI training practices “theft.” OpenAI leadership stated its models posed an “existential threat” to the publishers and journalists whose work trained them.

The material also details how the firms allegedly obtained content by bypassing paywalls undetected, building training datasets via mass scraping, and deliberately stripping copyright notices from the data.

Most of this new information comes from The Times’ own brief, while the underlying exhibits remain sealed. The quotes below are presented without their original context.

The filing is the latest escalation in the three-year-old lawsuit, in which The New York Times initially alleged the firms violated copyright law by training generative AI models on its content.

Whether AI firms can legally use copyrighted material to train AI has no clear answer, but judges have largely sided with AI companies arguing that training constitutes “fair use.” This rule allows use of copyrighted work without permission in specific cases, such as parody, news reporting, or criticism. Earlier this month, the Trump administration contributed a brief defending OpenAI’s unlicensed use of copyrighted material to train its large language models.

Several new admissions run counter to OpenAI’s fair use defense, particularly the requirement that use does not substitute for or harm the market for the original work.

Microsoft data shows its Copilot “answer engine” caused click-through rates for The New York Times’ domain to drop as much as 93% compared to traditional Bing search. An internal Microsoft presentation written by Brent Hecht, the Director of Applied Science, in January 2024 describes the decline as a “doom loop” that would “hurt the performance of our models and the entire web at the same time.”

“It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,'” reads the Microsoft document, as quoted in the filing.

Microsoft CEO Satya Nadella testified in a deposition earlier this year that “anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training.” He made clear that, if he “had been made aware that OpenAI had scraped and trained on information that was behind a paywall,” he would have “invoked [Microsoft’s right to] require OpenAI to retrain its models.”

Other admissions cut against different pillars of the fair-use test. OpenAI’s Head of ChatGPT, Nick Turley, wrote in internal communication that publishers face an “existential threat” from products like the chatbot, which are “largely substitutive” and “will get more and more substitutive as they get better.”

OpenAI President Greg Brockman described the models as “excellent at news.” Nadella agreed under oath earlier this year that conversing with chatbots “has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source.”

That language speaks to how the technology could directly compete with, rather than transform, the original work.

A Microsoft document states there is a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.”

The scale of copying

The documents reveal for the first time that OpenAI’s mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone.

In a January 2023 internal memo, Hecht called it “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.”

The filing lays out in new detail how OpenAI and Microsoft acquired the plaintiffs’ content, including scraping it from the Bing Index.

“OpenAI delivered the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI’s models within its own commercial products,” the filing reads. “Microsoft similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango.”

The companies allegedly assembled the Project Mango data into a training dataset that contains copies of at least 160,903 unique works from the news publishers.

To get the most out of their scraping, OpenAI employees allegedly devised a plan to circumvent paywalls without detection. The filings show that when OpenAI researcher Nick Ryder told Brockman about a “hack to get around nytimes paywall,” Brockman replied: “ah nice.”

OpenAI employees also allegedly built training datasets like WebText and WebText2 that disproportionately relied on scraped news content. They also allegedly pulled millions of articles from Common Crawl, a free, open repository of web crawl data. The findings also describe deliberate efforts to strip copyright notices from training data before it reached the model, since researchers “wouldn’t want model outputting” “copyright notices” to users.

OpenAI and Microsoft did not return requests for comment.

Scroll to Top