New York Times Reveals OpenAI Withheld Evidence in ChatGPT Copyright Dispute
The New York Times and The Daily News have accused OpenAI of falsely representing its ability to access customer chat logs and training datasets linked to their copyrighted content. This development comes amid a two-year legal dispute with the AI company, which faces allegations of violating copyright laws by using the Times’ material to train its generative AI models and reproducing such journalism in user-generated outputs.
Throughout the legal proceedings, OpenAI argued that it could not search its own training corpus. It also asserted that retrieving or analyzing its vast collection of ChatGPT conversations would be technically complex and could compromise user privacy, as it would require accessing, processing, and anonymizing the logs. The media outlets aimed to acquire this data to determine whether their copyrighted journalism was present in OpenAI’s training datasets and, if it was, how often ChatGPT utilized or reproduced their content.
In a court-ordered deposition from April, OpenAI data privacy engineer Vinnie Monaco reportedly revealed that the company had performed internal searches and evaluations of its training corpus to locate instances of copyrighted journalism.
Monaco’s testimony also suggested that prior to the NYT filing its lawsuit, OpenAI had created a database comprised of around 78 million anonymized ChatGPT conversations for internal assessment of potential copyright infringements. Furthermore, OpenAI allegedly employed a “Bloom” filter as part of a toolkit named “Project Giraffe” to detect and log instances of content reproduction shortly after the lawsuit was filed.
These two revelations hold significant importance. Initially, the plaintiffs asked OpenAI for a sample of 120 million chat logs, but later agreed to reduce that number to 20 million. OpenAI submitted this sample to the court last December, but it was reportedly so heavily redacted that the court deemed it “unusable.” The plaintiffs further claimed that OpenAI deleted billions of ChatGPT outputs after the lawsuit commenced, thus violating the court’s preservation directive, and that the company swapped out millions of logs in the requested sample.
Essentially, they argue that OpenAI intentionally made it difficult to retrieve information that the company had already collected.
“If OpenAI genuinely believed that reproducing our clients’ journalism was fair and legal, it would not have hidden the truth about having done so,” stated Ian B. Crosby, the lead counsel for the plaintiffs.
Currently, the NYT and The Daily News are urging the judge to impose sanctions on OpenAI for allegedly concealing evidence and disrupting the discovery process. They seek to prevent OpenAI from using the 20 million chat log sample as evidence, claiming it lacks reliability; to accept as fact that ChatGPT logs would have shown significant reproduction of the plaintiffs’ content; to bar OpenAI from arguing that its submitted chat logs do not indicate substantial reproduction; and to compel OpenAI to cover the legal expenses incurred in pursuing this evidence.
In reply, OpenAI spokesperson Drew Pusateri refuted the allegations, claiming the Times is trying to invade user privacy as its case weakens.
“As the Times’ case falters and they are forced to retract claims against us, they continue to attempt to breach the privacy of individuals not involved in this case, including through these blatantly false accusations,” Pusateri remarked. “We will persist in defending our users’ privacy and upholding the established principles of fair use.”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.


