The New York Times and the New York Daily News have further escalated the charges against OpenAI in their copyright suits. Two media sources stated that OpenAI had repeatedly stated in court that it was difficult to retrieve training materials and large-scale ChatGPT chat records, but that there was already a search and assessment capability within the company.

The proceedings have been ongoing for about two years. According to the plaintiff, OpenAI used its copyright-protected news content in training the Generated AI model and reproduced the reports in some user outputs. One of the main points of the dispute between the parties was whether OpenAI was able to verify whether training data and chat records included plaintiffs and whether there were large-scale repetitions of model outputs.

Internal testimony leads to new charges

According to the plaintiff, during a court-requested deposition procedure in April this year, the OpenAI Data Privacy Engineer Vinnie Monaco disclosed that the company had previously conducted an internal search and assessment of the training language for copyright-protected news works.

The plaintiff also claimed that Monaco ' s testimony showed that before the New York Times filed the lawsuit, OpenAI had accumulated approximately 78 million de-marked ChatGPT conversations, which were used internally to assess the extent to which it had violated the work of others. Shortly after the proceedings were instituted, OpenAI also allegedly launched a set of tools called Project Giraffe and identified and recorded duplicates in the output through the so-called Bloom Pilot.

20 million samples focused

The plaintiff initially requested 120 million samples of chat records from OpenAI, which was reduced to 20 million after consultation. OpenAI submitted the sample to the court last December, but the plaintiff claimed that it was too redacted and the court described the sample as “unusable”.

The plaintiff further alleged that after the commencement of the action, OpenAI had removed billions of ChatGPT exports, violated the court ' s request for preservation of evidence and replaced millions of records in the sample to be submitted. According to the plaintiff, this meant that OpenAI, on the one hand, claimed to have difficulty in obtaining evidence, and on the other hand, had already acquired and used relevant data, complicating the process of external evidence.

The plaintiff demanded court sanctions.

On the basis of the above allegations, The New York Times and The New York Daily News have requested judges to impose sanctions on OpenAI, including, inter alia, not allowing it to continue to use the 20 million samples as evidence and limiting it to refuting the large-scale repetition of allegations on the grounds that the available samples are insufficient.

  • Request the court to exclude 20 million sample evidence
  • Request to make sure that chat records show a large number of repetitions
  • OpenAI is required to bear the relevant legal costs

OpenAI denied this. The company spokesman, Drew Pusateri, stated that, as the claim of the New York Times party diminished, the plaintiff was trying to obtain a private user dialogue unrelated to the case and making false allegations. OpenAI states that it will continue to maintain the privacy of users and maintain its position on reasonable use.