1 day ago
Unsealed Filings Expose AI Copyright Dispute Over News Content
The New York Times says OpenAI and Microsoft used news stories to help train AI tools without permission.
New court documents show that some people inside the companies worried about this practice.
One Microsoft researcher described the copying as extremely large and unfair.
OpenAI and Microsoft also discussed how chatbots might give people answers without sending them to news websites.
The filings say Bing’s AI sent far fewer visitors to The New York Times than regular Bing Search.
Microsoft gave OpenAI data from its Bing database and helped gather additional web content.
The companies say using copyrighted material for AI training may be allowed as fair use.
A court will decide whether the companies broke copyright law.
The documents provide evidence for the publishers’ arguments but are not a final legal ruling.
Newly unsealed filings detail concerns inside Microsoft and OpenAI about copying news content for AI training.
A Microsoft researcher called large-scale copying an “astonishing theft” and possibly the “largest theft of labor in human history.”
OpenAI and Microsoft acknowledged that improving AI products could substitute for visits to publishers’ websites.
Microsoft supplied Bing data to OpenAI and later used Project Mango to collect more than 160,000 publisher works, according to the filings.
The documents may support The New York Times’ case, but do not establish that Microsoft or OpenAI infringed copyright.
- Who
- The New York Times is suing OpenAI and Microsoft; the filings also cite statements by Microsoft and OpenAI employees and executives.
- What
- Newly unsealed court documents reveal internal discussions about copying news content for AI training and the possible effect of AI products on publishers.
- Where
- The dispute is being considered in a United States court under U.S. copyright law.
- When
- The lawsuit was filed in December 2023; the documents concern events and discussions including 2019–2022 and January 2023.
- Why
- The New York Times alleges that copyrighted articles were used without permission to train AI systems and that AI products could harm publishers’ markets and website traffic.
Publishers’ position
Microsoft and OpenAI’s position
Copyright and AI training
Publishers’ position
The New York Times argues that its copyrighted articles were used without permission to train AI models and that the resulting products can reproduce or substitute for its journalism.
Microsoft and OpenAI’s position
Microsoft and OpenAI argue that using copyrighted material for AI training can qualify as fair use, depending on the circumstances and the four legal factors.
Effect on publisher markets
Publishers’ position
Publishers can point to lower click-through rates and internal company discussions as evidence that AI products substitute for publisher websites and harm the value of original journalism.
Microsoft and OpenAI’s position
Microsoft says Copilot is not a substitute for publishers’ journalism and maintains that the reported traffic figures do not by themselves establish copyright infringement.
Internal criticism and data collection
Publishers’ position
The filings highlight Microsoft employee concerns about extensive copying and allege that Microsoft helped OpenAI obtain copyrighted and potentially paywalled material.
Microsoft and OpenAI’s position
Microsoft says Brent Hecht’s comments were personal views from a researcher whose role included presenting divergent perspectives, not the company’s official position. Microsoft CEO Satya Nadella said paywalled material should be licensed and that he would have required retraining if he had known it was scraped.
Key facts
- Plaintiff
- The New York Times
- Defendants
- OpenAI and Microsoft
- Lawsuit filed
- December 2023
- Relevant legal issue
- Whether using copyrighted works to train AI qualifies as fair use
- Reported traffic decline
- Click-through rates from Bing’s AI product to New York Times websites were 87% to 93% lower than from conventional Bing Search
- Project Mango dataset
- The filing says it contained copies of at least 160,903 unique works belonging to publishers involved in the litigation
- Data transfer period
- Microsoft supplied Bing Index data to OpenAI over three years between 2019 and 2022
Quotes
Microsoft spokesman
Spokesperson representing Microsoft’s position on the copyright litigation
“Microsoft’s position is set out in its court filings, which explain why these transformative uses are consistent with copyright law and why Copilot is not a substitute for publishers’ journalism”
indianexpress.com
“anything that is paywalled should be licensed by anyone who wants to use it.”
indianexpress.com








