Several publishers and authors have filed a class action against Google in the United States Federal District Court for the Southern District of New York, alleging unauthorized use of copyright-protected books and works for training in the Gemini model. The case once again brought the copyright issue of AI training data to the fore and raised the tension between technology companies and the content industry.

The plaintiff pointed to the contents of the books and shops.

The plaintiffs at this time included Hachette, Cengage, Elsevier and writers Scott Turow and S.C.R.I.B.E. According to the prosecution, Google used not only copies of books originally provided for the Google Books search service, but also book content training at Google Play, which were not authorized.

The application stated that Google Books cooperated on the premise that books were searchable and limited bibliographic information and footage were shown to users, rather than opening the entire book for model training. According to the plaintiff, Google transferred these elements to AI training, which went beyond its original mandate.

The petition stated that copyright information had been altered.

The plaintiff also alleged that Google had intended to remove or modify copyright information in the work in order to cover up the Gemini model's use of stolen materials for training. The complaint also cites a document allegedly from within Google, which mentions that the use of copyright-protected book training AI may be “highly problematic” for Google and may entail a potential risk of a fine of tens of billions of dollars.

By the time the report was released, Google had not responded immediately to the request for comments.

AI training copyright litigation is still spreading.

This is not the first time that AI has faced copyright litigation for training data sources. In the past, Google, Meta, OpenAI and Anthropic have been sued by publishers, authors or other copyright holders.

Although two earlier courts in California have ruled in favour of AI and concluded that the use of copyrighted works for training AI could be considered “reasonable use”, there is still no uniform conclusion in the case. The case of Google was referred to the Federal Court in New York, meaning that another judge would judge on a similar dispute.

Additional information:TechCrunch mentioned that Anthony had previously been awarded $1.5 billion in compensation for the theft of training materials, and that approximately 500,000 authors were eligible for at least $300,000, but some chose to withdraw from settlement and continue to seek follow-up action.