US Department of Justice backs fair use for AI training in landmark copyright case

The US Department of Justice has sided with AI developers in the consolidated copyright lawsuit brought by The New York Times, stating…

By Vane September 2, 2026 3 min read
US Department of Justice backs fair use for AI training in landmark copyright case

The US Department of Justice has sided with AI developers in the consolidated copyright lawsuit brought by The New York Times, stating that training artificial intelligence models on protected text qualifies as fair use.

The core dispute

The New York Times filed suit in late 2023 against OpenAI and Microsoft in Manhattan federal court. The newspaper alleged that millions of its articles were used without permission to train models such as GPT-4 and to build competing information products. The Times claimed damages in the billions and demanded the destruction of the language models trained on its content.

That case has escalated considerably and is widely considered a bellwether for how courts will handle copyright and AI training.

The DOJ now argues that the copyrighted text used for large language model training in this case does not amount to infringement. The distinction between the training process and the model output is what matters. During training, entire works are copied but never made publicly available. The outputs often if not always lack substantial similarity to the originals.

The department contends that a blanket theory of market harm which conflates the two is legally wrong. Others disagree.

A literary analogy

The filing invokes Joan Didion as an analogy. As a teenager, she copied Hemingway’s stories to understand how his sentences worked.

The DOJ argues that, under the logic of the Kadrey ruling, Didion could have faced liability whenever she published because her learning process and subsequent writing would have been treated as a single use. Citing an earlier ruling, the department contends that it would be unthinkable to require people to pay whenever they later draw on a book to write something new in a new way.

The department also argues that large language models have creative and public value. “Human beings create original works using LLMs,” the DOJ writes. Liability for AI training would stifle the creativity copyright law is supposed to protect. Even New York Times writers reportedly use LLMs to draft and edit articles, though the linked source ironically supports mostly the opposite conclusion.

And all these arguments gloss over scale. A single author copying text to learn is one thing. A multibillion-dollar company turning that content into competing mass-market products is another, a point the US Copyright Office has made explicitly.

Administrative friction

That is the point the US Copyright Office made in a report that rejected blanket fair use for AI training. The Copyright Office argued that AI works with perfect copies and generates content at a speed and scale far beyond human creation. Commercial applications that compete with original works in existing markets exceed what fair use allows.

In its filing, the DOJ goes after that report directly. Former Copyright Register Shira Perlmutter’s assessment carries no binding legal authority, the department argues, and the report ignored case law on case-by-case analysis and the types of market harm that actually count under the statute.

“The fair-use inquiry hinges on the specific facts and uses at issue in each case. But it would be problematic—and legally incorrect—to impose broad copyright liability that would generally render training of AI models impermissible without licensing,” the DOJ writes.

Perlmutter was fired by the Trump administration shortly after her report came out. Democrat Joe Morelle said she was let go because she refused to legitimize AI training on copyrighted works, a position favored by Trump ally Elon Musk. In a footnote, the DOJ notes that Perlmutter is currently challenging her dismissal.

What it means

For writers and developers, this ruling suggests that the legal risk of using existing books and articles to train an AI system is significantly lower than previously feared. However, the distinction between training data and public output remains critical. If a model reproduces specific passages too closely, it may still face liability. The decision protects the act of learning from data but does not necessarily shield the commercial sale of content that mimics the original work too closely.

Scroll to Top