Hachette Book Group, Cengage Learning, Elsevier, author Scott Turow, and SCRIBE, Inc. have filed a proposed federal class action lawsuit against Google.
They accuse the company of copying copyrighted books and magazine articles without permission to train their employees. Gemini AI models. The lawsuit was filed on July 10, 2026. The case, Hachette Book Group Inc. et al. v. Google LLC, is filed under case number 1:26-cv-05870 in the United States District Court for the Southern District of New York.
The complaint describes Google’s alleged actions as “one of the most prolific infringements of copyrighted materials in history.” However, these accusations have not been proven in court and no ruling has been made regarding Google’s liability.
Legal claims and agreements with publishers at center of Gemini lawsuit
The plaintiffs allege four legal claims: direct copyright infringement, contributory copyright infringement, deletion or alteration of copyright management information, and violations of the Digital Millennium Copyright Act.
They argue that Google obtained protected material through its book services, Internet scraping and other sources before using the content to train Gemini.
Publishers provided books and journal articles to Google Books, Google Play Books, and Google Scholar under agreements that covered specific uses, such as displaying searchable excerpts, distributing e-books, and helping users discover academic publications.
The plaintiffs claim that these agreements did not authorize Google to copy entire works in data sets used for commercial development of generative AI, and that Google reused material received through established publishing relationships for uses not covered by the contracts.
The complaint also claims that Google collected books and other protected material through extensive Internet crawling, including from sources such as pirate websites and publications behind subscriptions or paywalls.
How this Gemini case is different from Google Books and what the internal emails say
Google previously won a copyright case related to digitizing books for search and snippet previews. In 2015, a federal appeals court ruled that these uses were transformative and protected by fair use.
The plaintiffs argue that this earlier ruling covered databases of searchable books and limited excerpts, not the use of entire books to train generative AI models.
They claim that AI training serves a different business purpose and results in systems that can compete with the original works involved in the training. The key difference, they say, lies between search-focused digitization and generative model training, which is critical to their case.
According to the complaint, a Google employee warned during an internal discussion that using books submitted through Google Play publishing agreements for the development of artificial intelligence could create serious legal risks. The employee reportedly estimated that such practices could result in fines of between ten and one hundred billion dollars.
Other internal documents described in the complaint suggest that Google sought professionally written books to improve the performance of its Gemini AI system.
The plaintiffs claim that tests showed that models trained only on public domain books performed worse than those trained on collections that included copyrighted works. The complaint claims that Google intended to include works with selected facts, organized analyses, fictional narratives, and professionally edited writings.
Works cited in the complaint and what Gemini is accused of generating
The complaint highlights several works as examples. Hachette lists Peter Brown’s The Wild Robot, NK Jemisin’s The Fifth Season, Becky Lomax’s Moon Glacier National Park, and Who Could That Be at This Hour? by Lemony Snicket.
He also mentions Turow’s Innocent. Cengage includes Cognitive Psychology, Principles of Economics, Milady Standard Barbering, Nutrition: Concepts and Controversies, and Calculus: Early Transcendentals. Elsevier notes copyrighted journal articles in its section of the complaint.
Turow and SCRIBE refer to Presumption Innocent, Innocent and Testimony. The plaintiffs claim that these examples are just a sample of the books and articles allegedly copied in connection with Gemini.
The presentation presents examples to show that Gemini can produce material related to specific protected books. It claims that the system generated content based on The Fifth Season and produced answers involving characters, events and details from Who Could It Be at This Hour? The plaintiffs also argue that Gemini can generate low-cost substitutes for professionally published works.
They estimate the service could produce a 100-page murder mystery set in a sleepy coastal town in about 20 minutes for 39 cents. The complaint states: “No publisher or author can compete with that.” These time and cost figures are allegations of the plaintiffs and have not been verified by the court.
Who could join the group and what solutions the plaintiffs are seeking
The proposed class would include registered copyright owners in books that have international standard book numbers, as well as journal articles identified by digital object identifiers or international standard serial numbers.
To be part of the class, members would have to prove that Google copied their work from one of its services, obtained it through web scraping, or used it in connection with Gemini training. The court has not yet certified this proposed class.
The plaintiffs seek statutory damages or compensation based on Google’s claimed losses and alleged profits, along with legal fees and other remedies available under federal copyright law.
They ask the court to stop Google from continuing with what they claim to be unauthorized copying through an injunction that would restrict the use of protected books and articles in Gemini training and related AI development.
The complaint also requests an accounting of the work and methods Google used to train Gemini. This includes details about the materials obtained, where they came from, and how they were incorporated into Google systems.
The plaintiffs seek court-supervised destruction of infringing copies and data sets derived from their works. It is unclear whether such an order would be legally feasible or appropriate in this case.
Publishers, authors and academic rights holders who might be part of the proposed class should take some practical steps as the case develops. They should verify which of their titles are registered with the US Copyright Office, as the proposed class relies on registered works marked with ISBN, DOI, or ISSN.
It’s also a good idea to review existing agreements with Google Books, Google Play Books, and Google Scholar to understand what uses those agreements allow.
Additionally, they must keep records of publication dates, distribution channels, and any communications with Google regarding the use of their work in AI training data sets. It may also be important to track the progress of class certification, as inclusion in any potential class will depend on the court’s decision.
Where Hachette and Cengage’s lawsuit against Google is now
Hachette and Cengage had previously attempted to join a separate copyright lawsuit filed against Google in California in 2023, involving authors and visual artists. Google objected to their involvement, and the publishers later withdrew that effort before filing the current case in New York.
Google has not yet issued a detailed public response to the specific allegations in the New York complaint. The company will have the opportunity to respond as the case progresses through motions, potential discovery and class certification processes. No trial date has been set and the court has not ruled on any of the four claims.






