3 hrs ago
Unredacted Filings Expose AI Training Dispute Over Scraped News
The New York Times is suing OpenAI and Microsoft over how their AI systems used news articles.
New court papers describe very large collections of news stories used for training.
The papers also allege that some researchers accessed articles protected by a paywall.
Microsoft and OpenAI reportedly worried that AI answers could reduce visits to news websites.
One Microsoft analysis found much lower click-through rates from Copilot than from regular Bing search.
The Times says this could hurt the business that pays for journalism.
OpenAI and Microsoft have not necessarily agreed with all of the claims in the filing.
A court will decide how much these facts matter and whether the use of the articles was legal fair use.
Newly unredacted filings in The New York Times’ lawsuit describe extensive copying of news articles by OpenAI and Microsoft for AI training.
The Times says datasets included more than 91,692 copies of works from several publishers and over two million documents originating from nytimes.com.
The filings allege that researchers found ways to access paywalled New York Times content and that some preparation removed copyright notices.
Microsoft analysis reportedly found that Copilot generated up to 93% fewer click-throughs to The Times than conventional Bing search.
The companies have not necessarily endorsed The Times’ characterizations, and courts must still decide whether the alleged uses qualify as fair use.
- Who
- The New York Times, OpenAI, Microsoft, and other publishers named in the case.
- What
- A copyright lawsuit examining alleged copying and use of news material to train and develop AI systems.
- Where
- The dispute is before a United States court; the datasets included material from nytimes.com and other web sources.
- When
- The disclosures were made three years into the dispute; the filing cites events and testimony from 2023, 2024, and this year.
- Why
- The Times argues that the alleged copying was unauthorized and that AI-generated answers may substitute for publisher websites and weaken their revenue.
The New York Times’ Claims
OpenAI and Microsoft’s Position
Copyrighted training material
The New York Times’ Claims
The Times says the companies made extensive copies of protected news material, including works in large training datasets and material allegedly obtained from paywalled sources.
OpenAI and Microsoft’s Position
The companies have not necessarily endorsed the characterizations in The Times’ filing, and the legal significance of the alleged copying remains for the court to determine.
Effect on publishers
The New York Times’ Claims
The Times argues that AI-generated answers can substitute for visits to original publications, potentially damaging publishers’ traffic, audiences, and revenue.
OpenAI and Microsoft’s Position
The cited materials include internal concerns that AI could create a damaging cycle for publishers and AI systems, but those concerns do not by themselves establish liability.
Fair use
The New York Times’ Claims
The Times says the newly disclosed evidence undermines important parts of the companies’ fair-use defense, especially evidence of market substitution.
OpenAI and Microsoft’s Position
Whether AI training is fair use remains unsettled, with courts reaching different outcomes depending on the facts and circumstances of individual cases.
Key facts
- Case
- The New York Times’ copyright dispute with OpenAI and Microsoft
- Mid-training datasets
- The Times says they contained more than 91,692 copies of works from The New York Times, The Daily News, and the Center for Investigative Reporting
- Common Crawl dataset
- The Times alleges it contained more than two million documents originating from nytimes.com
- Project Mango
- The Times says it included at least 160,903 distinct works belonging to publishers involved in the case
- Copilot traffic
- Microsoft analysis reportedly found click-through rates to The Times were as much as 93% lower than with conventional Bing search
- Paywall allegation
- The filing says an OpenAI researcher described a method for bypassing The New York Times’ paywall
- Legal question
- The court must assess whether the alleged uses of copyrighted news material qualify as fair use




