14 hrs ago
New York Times Alleges AI News Training Was Theft
The New York Times says OpenAI used millions of news stories to teach its artificial-intelligence models without permission.
The newspaper says more than 10 million articles were scraped.
It says nearly one-third came from The New York Times.
A Microsoft science director reportedly called this an enormous theft of workers’ efforts.
Microsoft says those comments were only the employee’s personal opinion.
Microsoft and OpenAI say using the stories to train models is a legal form of fair use.
Other publishers are also part of the lawsuit.
The publishers want money for each article they say was used.
A judge could decide the case without a trial, but a decision is not expected until 2027.
The New York Times alleges OpenAI scraped more than 10 million news articles to train its artificial-intelligence models.
The filing says nearly one-third of the articles came from The New York Times.
A Microsoft director allegedly called the conduct an unprecedented theft, while Microsoft said his views were personal.
Microsoft and OpenAI argue that using news content for model training is transformative and protected by fair-use law.
The publishers seek damages and summary judgment; any ruling is not expected until 2027.
- Who
- The New York Times and other publishers are suing OpenAI and Microsoft; Microsoft director Brent Hect is quoted in the court filing.
- What
- The publishers allege that OpenAI used copyrighted news articles to train its models without authorization.
- Where
- The case is in New York federal court; OpenAI is based in San Francisco.
- When
- The relevant court document was unsealed on Thursday; the lawsuit was filed three years earlier, and a ruling is not expected until 2027.
- Why
- The publishers seek damages for articles they say were copied and used by OpenAI’s models.
Publishers’ allegations
Microsoft and OpenAI’s position
Use of news articles
Publishers’ allegations
The New York Times alleges that OpenAI committed an unprecedented theft by using millions of copyrighted articles to train its models.
Microsoft and OpenAI’s position
Microsoft and OpenAI argue that using news content to train artificial-intelligence models is transformative and falls under fair-use law.
Microsoft employee’s statements
Publishers’ allegations
The court filing attributes statements to Brent Hect describing the alleged conduct as an unprecedented theft and suggesting there may have been an accidental cover-up.
Microsoft and OpenAI’s position
Microsoft says Hect’s statements reflected one employee’s individual perspective and do not represent the company’s views.
Impact and public interest
Publishers’ allegations
The publishers seek damages and say generative AI may reduce traffic to internet news sites.
Microsoft and OpenAI’s position
The United States Department of Justice filed a brief supporting OpenAI and Microsoft, citing scientific progress, economic growth, and national security.
Key facts
- Articles allegedly scraped
- More than 10 million
- Share allegedly from The New York Times
- Nearly one-third
- Lawsuit plaintiffs
- The New York Times and several other news publishers
- Defendants
- OpenAI and Microsoft
- Defendants’ legal position
- The content use was transformative and protected by fair use
- Requested legal action
- The plaintiffs seek summary judgment
- Expected ruling
- Not expected until 2027
Quotes
An OpenAI engineer
An unnamed engineer at OpenAI quoted in the court document
“No matter how prominently we show the links, users won't click.”
NDTV
deccanchronicle.com




