Microsoft says virtually nobody was grabbing NYT articles through its chatbot

Microsoft claims its Copilot chatbot almost never reproduces full sentences or substantive chunks from news articles, a point central to its defence…

By Vane September 4, 2026 1 min read
Microsoft says virtually nobody was grabbing NYT articles through its chatbot

Microsoft claims its Copilot chatbot almost never reproduces full sentences or substantive chunks from news articles, a point central to its defence against copyright claims from publishers including The New York Times and book authors. As part of the lawsuit’s discovery process, the company provided 8.2 million chat logs to an expert hired by the publishers. Microsoft states these logs were specifically chosen because they hit keywords implicating use of the plaintiffs’ websites, making them the most likely to contain the disputed works. The resulting analysis showed that out of 8.2 million interactions, only 59,545 instances contained text from the news publishers. In those specific cases, the company argues the output was merely a brief phrase or a single word rather than a meaningful excerpt that could substitute for the original article.

This data challenges the core assertion that the training process or the chat interface directly appropriates large volumes of copyrighted text for commercial use. If the model rarely outputs protected material during active conversation, the legal argument regarding direct infringement shifts from output to the underlying training methodology. The case now hinges on whether the mere ingestion of data for training constitutes a violation even if the final user interaction does not visibly replicate the source material.
* 8.2 million logs were reviewed for specific keyword matches
* Only 59,545 instances contained any text from the publishers
* Most outputs were limited to brief phrases or single words

Scroll to Top