Publishers and Authors Sue Google Over Gemini Training Data
Hachette Book Group, Cengage Learning, Elsevier, author Scott Turow and S.C.R.I.B.E., Inc. have sued Google in federal court, accusing the company of copying copyrighted books, journal articles and other written works to develop and train its Gemini artificial-intelligence models without permission or compensation.
The proposed class action was filed July 10, 2026, in the U.S. District Court for the Southern District of New York. The docket lists the case as Hachette Book Group, Inc. et al. v. Google LLC, case number 1:26-cv-05870. The complaint demands a jury trial.
The filing is an allegation, not a court finding. The court has not ruled on the merits, Google has not been found liable, and the proposed class has not been certified.
What the complaint alleges
The plaintiffs allege that Google obtained copyrighted works through Google Books, Google Play Books, web scraping and other Google services, then reproduced those works repeatedly during Gemini’s development and training.
According to the complaint, the alleged copying included placing works into computer memory, converting them into machine-readable formats and incorporating them into training datasets. The plaintiffs also allege that Google copied works again as it developed successive versions of its models.
The complaint says the alleged conduct threatens established markets for books, textbooks, scholarly journals and licensing. It further claims that Gemini can produce summaries, substitute passages, near-verbatim material and new works that imitate the expressive choices of particular authors. Those claims remain disputed allegations to be tested through the litigation.
What the plaintiffs want
The filing seeks certification of a proposed class of authors and publishers. It also requests damages, an injunction, disclosure of Google’s training materials and methods, and destruction of allegedly infringing copies.
None of those remedies has been ordered. The court could reject some or all of the claims, limit the case, allow it to proceed, or oversee a settlement or licensing-related resolution.
Why Google Books matters
The plaintiffs’ argument centers partly on the distinction between using books for limited Google Books functions and allegedly reusing those works for a separate commercial AI-training purpose.
That distinction matters because earlier legal disputes over Google Books involved search and snippet displays. The prior history does not automatically decide whether the alleged use of books and journals to train Gemini is lawful. The outcome could turn on the specific works, copying practices, licenses, model-development steps and defenses presented in this case.
A broader fight over AI training
The lawsuit is part of a wider publishing dispute over whether commercial AI developers must license copyrighted text used to build models. Publishers and authors have also sued Meta over allegations involving books and the training of its Llama models, but that is a separate case and does not determine the claims against Google.
The U.S. Copyright Office continues to examine copyright and artificial intelligence, including questions involving training data and AI-generated material. That federal policy work does not resolve the claims in this lawsuit.
What readers and the industry should watch
For authors and publishers, the case could affect whether books and scholarly works become a formal licensing input for commercial AI systems. For educators, libraries and students, future licensing rules could influence digital-book access, research databases and AI-assisted learning tools.
The filing does not immediately change Gemini, the availability of books or existing copyright permissions. The next important steps are likely to include Google’s first response, possible motions to dismiss, disputes over discovery of training data and any proceedings on class certification.
Sources
- Federal complaint in Hachette Book Group v. Google
- Southern District of New York docket listing
- Publishers Weekly report
- U.S. Copyright Office AI initiative
Look for updates to this story
Discover more from Interactive News
Subscribe to get the latest posts sent to your email.