AI training built on fair use looks shaky when the companies’ own people call it “astonishing theft”

The New York Times and a group of other media companies are asking a US court for billions of dollars in damages…

By Vane September 18, 2026 6 min read
AI training built on fair use looks shaky when the companies’ own people call it “astonishing theft”

The New York Times and a group of other media companies are asking a US court for billions of dollars in damages from OpenAI and Microsoft. They claim these tech giants stole their content to train artificial intelligence models. The plaintiffs say the companies’ own internal messages and sworn testimony prove that the defence of “fair use” is collapsing.

Executives’ internal statements undercut the fair use defense

A summary judgment brief filed in New York District Court lists The New York Times, the Daily News group (which includes the Chicago Tribune and Denver Post), and Ziff Davis (owner of CNET, IGN, and PCMag). The Center for Investigative Reporting, which runs Mother Jones, and The Intercept are also part of the suit. This 92-page document is part of a consolidated multidistrict litigation that began with the Times’ case in December 2023.

The filing relies on emails, Slack messages, and testimony that directly contradict the companies’ legal position. Microsoft’s director of applied science, Brent Hecht, described the practice as “an astonishing theft of unprecedented proportions” and possibly the “largest theft of labor in human history.” He wrote that a successful fair use defence would arguably “make a complete mockery of the idea of ‘fair use.'” Microsoft told the Financial Times these comments “reflect one employee’s individual perspective” and “are not a legal analysis.”

The US Copyright Office also concluded in May 2025 that fair use cannot apply broadly given the sheer scale at which AI companies copy data. The official who oversaw the report was later fired by the Trump administration.

Nick Turley, OpenAI’s head of ChatGPT, wrote that publishers face an “existential threat,” according to the brief. He said the products “are largely substitutive, period” and would increasingly replace publishers’ offerings “as they get better.” An OpenAI engineer added that “no matter how prominently we show the links, users won’t click.”

Microsoft CEO Satya Nadella confirmed under oath that chatbot conversations had replaced visits to original sources. In a statement to the FT, Microsoft said his testimony concerned “broad principles and changes under way in how people find and consume information.” It was not a “conclusion about copyright questions” at the center of the case, the company said.

An internal Microsoft document describes a self-reinforcing cycle it calls a “doom loop.”

Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.

Microsoft’s own data shows that click-through rates on Copilot were much lower than on traditional Bing search. Rates were down 87 to 93 percent for the New York Times, 83 to 91 percent for the Daily News group, and 51 to 94 percent for Ziff Davis.

OpenAI internally described local news, a core business for the Daily News group, as a “pretty common quer[y]” in ChatGPT. Other studies have also found that chatbot answers sharply reduce traffic to the open web.

OpenAI co-founder Greg Brockman discussed the models’ ability to handle news in an internal message.

we are excellent at news btw. every time i do any generative stuff on NYT it seems to predict the next sentence pretty well.

Paywall workarounds and license restrictions complicate OpenAI’s defense

The plaintiffs describe how OpenAI systematically bypassed paywalls and ignored terms of service. When an employee told Brockman about “a hack to get around nytimes paywall,” he replied “ah nice.” Around 2017, Brockman also wrote that he was “deeply motivated by the gazillions” he hoped to earn by commercializing OpenAI’s technology.

OpenAI’s corporate representative testified that he knew of no method for detecting paywalled content in the training data and no effort to remove it. The company’s standard crawling process “did not include reviewing websites[‘] … Terms of Use or Service.”

Nadella testified under oath that “anything that is paywalled should be licensed by anyone who wants to use it.” He said he would have forced OpenAI to retrain its models had he known the company had scraped paywalled content and used it for training.

OpenAI also acquired the “New York Times Annotated Corpus,” a collection of 1.8 million articles, through a third party. Its license restricted use to “non-commercial linguistic education, research and technology development.” OpenAI employees knew using the corpus to train models “would not be appropriate,” but did so anyway.

Plaintiffs accuse OpenAI of suppressing evidence and challenge all four fair use factors

Immediately after the lawsuits were filed, OpenAI built a filter to suppress output most likely drawn from the plaintiffs’ publications, according to the brief. Content from companies that hadn’t sued remained unaffected. The plaintiffs argue that the filter was designed not to protect copyrights but to prevent them from gathering evidence.

Microsoft’s Hecht had raised concerns that OpenAI might deploy such a filter, calling it an “accidental cover up.” He warned it would result in “people who have a right over the content having less visibility into what was used for training.” Even so, the brief includes numerous examples of problematic outputs, with ChatGPT reproducing exact copies and summaries of NYT articles on request, including articles behind a paywall.

The plaintiffs also argue that the fair use defence fails on all four statutory factors. They say the use is substitutive and commercial rather than transformative, while the articles are expressive works at the core of copyright protection. The defendants also copied entire works, even though their own experts admitted that no single work was necessary.

On market harm, the plaintiffs point to existing licensing markets and those that could realistically develop. OpenAI and Microsoft have signed licensing deals with other publishers, as have Amazon, Google, Meta, and Perplexity. But rather than pay for the plaintiffs’ content, the defendants took it for free and traded it among themselves, the brief argues.

A functioning licensing market makes the fair use argument harder to sustain. The defendants can’t credibly claim licensing wasn’t possible when they’re already paying for comparable content from other publishers.

AI can generate a million “pink slime” articles for ,800

The plaintiffs also challenge the argument that AI models don’t compete with publishers, further undercutting the fair use defence. They argue that OpenAI’s products let users flood the market with “pink slime,” meaning low-quality, often plagiarized pseudo-news. At OpenAI’s API prices, generating one million news-style articles of 500 words each costs about $6,800, without involving a single journalist. The brief cites Prism News, which used just four employees to run 200 AI-generated publications posing as local newsrooms and hobby sites.

OpenAI announced a tool called “Media Manager” in 2024 that was supposed to let publishers opt out of scraping. The project has since been shelved, according to the brief.

The plaintiffs are asking the court to find copyright infringement and reject the fair use defence. They also want each article treated as a separate work when calculating any statutory damages.

Previous court rulings have leaned toward fair use, and the US Department of Justice, under the Trump administration, has also sided with the AI labs. OpenAI and Microsoft will likely present their side soon, focusing on fair use and the transformative nature of their work. The internal documents now in the public record won’t make that argument any easier.

Scroll to Top