In this article
Delhi High Court rejects Indian news agency’s copyright injunction against OpenAI
The Delhi High Court has dismissed a preliminary injunction sought by Asian News International (ANI) against OpenAI.
ANI, one of India’s largest news organisations, had sued OpenAI alleging copyright infringement over how the company uses material for training its models and generating outputs in ChatGPT. Judge Amit Bansal denied the requested relief on both counts.
The ruling touches on memorisation, Retrieval Augmented Generation (RAG), and the legal standing of AI training data. Andres Guadamuz, an AI copyright law expert, described the decision as a significant early victory for OpenAI.
ANI’s evidence undermined its copyright claims
ANI submitted several ChatGPT outputs to the court, claiming they were substantial copies of its articles. The strategy backfired because OpenAI demonstrated that the models in question, GPT-4 and GPT-4o, were trained on data collected up to April 2022 and April 2024.
The articles ANI cited as evidence were mostly published in August and September 2024. Consequently, they could not have been part of the training data used to build the models.
The judge’s preliminary view was that the similarities stemmed from RAG, a feature that allows a language model to retrieve online information in real time, functioning much like a search engine. ANI had not addressed RAG in its filing, so the court could not make a final ruling on that specific issue. The judge stated that RAG-based outputs might qualify as “communication to the public,” a question reserved for the main proceedings.
ANI’s case weakened further because the agency had used adversarial prompts, explicitly instructing the model to reproduce articles “exactly.” Despite this, ANI could not produce a single verbatim copy. The judge found that facts within news articles are generally not copyrightable and that reproducing topics and headlines did not amount to direct competition with ANI in this instance.
Evidence also failed to support ANI’s claim that OpenAI permanently stores training data within its models and can reproduce the agency’s work verbatim on demand. The court will revisit that question during the main proceedings.
Court tentatively treats AI training as fair use
ANI also failed to demonstrate that copying its work for AI training constituted copyright infringement. Both parties agreed OpenAI used ANI content during the training phase. OpenAI argued the material represented a tiny fraction of the overall dataset and that the model extracted only non-expressive elements such as grammar, syntax, and language patterns.
The judge examined exceptions under Indian copyright law and relied on a clause covering “private or personal use, including research.” He interpreted “research” broadly enough to encompass AI training.
For this exception to apply, the judge set conditions. Training copies must originate from lawful sources, excluding shadow libraries or paywalled sites accessed without permission. OpenAI also never made the training copies public, processing them solely internally. Guadamuz noted this marks the first time a court has explicitly found that AI training falls under a private use exception.
The court applied a three-part fairness test and sided with OpenAI on all three counts. OpenAI’s use of ANI’s works was limited to training, as no memorisation or reproduction was proven. ANI could not show economic harm because the two companies operate in different sectors. Even when users ask ChatGPT about ANI headlines, the model returns topics and, at most, a few article titles.
The judge cited U.S. cases including Bartz v. Anthropic and Kadrey v. Meta, where language model outputs were deemed transformative. He also referenced the earlier Google Books ruling.
The judge also found that trained language models improve access to information, support education, advance scientific research, assist with software development, enable translation, and create tools for people with disabilities.
International AI copyright cases paint a mixed picture
The Delhi ruling joins a growing list of court decisions worldwide that have reached conflicting conclusions. In the U.S., a judge dismissed the lawsuit filed by Raw Story and AlterNet against OpenAI because the plaintiffs could not show sufficient harm and the odds of exact copies were low. That court also held that facts are not copyrightable.
The GitHub Copilot case failed similarly, with plaintiffs unable to present a single example of identical code. The Intercept, however, won a partial victory through a DMCA complaint over copyrighted material that had been stripped out before training.
In Ross Intelligence v. Thomson Reuters, a court denied fair use because the AI research tool directly competed with Thomson Reuters’ legal database Westlaw, making the use non-transformative. The court stressed that this ruling applied only to this non-generative use case and could not be extended to large language models.
In the Anthropic case, a federal court in San Francisco called AI training with copyrighted works “spectacularly” transformative, a strong signal favouring fair use. But the court drew a “Napster comparison” because Anthropic had used pirated books from shadow libraries as training data. Fair use does not cover unlawfully obtained material. Anthropic later paid $1.5 billion to book authors for using those pirated copies.
The U.S. Copyright Office rejected the AI industry’s argument that training on “vast troves of copyrighted works” broadly qualifies as fair use. The official who wrote the report was fired by the Trump administration shortly after it was published.
In Europe, two courts reached opposite conclusions almost simultaneously. The Munich Regional Court ruled in the GEMA case that song lyrics were reproducible in the model weights, making it a copyright-relevant reproduction. The High Court in London dismissed the Getty Images v. Stability AI lawsuit, ruling that an AI model is not an “infringing copy.” New research showing that language models can memorise copyrighted books could ramp up the memorisation debate further.
Across all these cases, the same core questions remain unresolved. Do AI models permanently store training data? Can training qualify as fair use? Where is the line between lawfully and unlawfully obtained data? And do copies generated through adversarial prompts reflect normal use?
AI and media face a problem beyond copyright
Beyond copyright, another question looms for the media industry. Even if courts rule that AI training is lawful, AI-powered search products could still erode the news market. A recent Pew Research Center study shows that click-through rates to external websites drop to just 8 percent with Google’s AI Overviews, compared to 15 percent without an AI summary. Users tend to stop searching right after getting the AI response. They do not check other sources.
For news agencies like ANI, this means AI systems that summarise news and make clicking the original source unnecessary could erode the industry’s business model over time, even without direct copyright infringement.
The Munich I Regional Court also recently ruled that Google is directly liable for false claims in its AI summaries, since these count as independent content rather than search results. The limited liability that traditionally shielded search engine operators does not extend to AI-generated summaries.
That ruling could become relevant for ChatGPT’s RAG-based responses too. When AI systems summarise news and make independent claims, their operators effectively become media providers, with all the liability that comes with it. That shift could also force courts to reconsider fair use. One key factor in the fairness test is whether the new product competes with the works it was trained on. If AI summaries replace the need to visit news sites, judges may have a harder time ruling that the use is non-competitive.




