The Meta Book-Piracy Lawsuit: What the Hachette-Macmillan-McGraw Hill Case Means for Authors
On May 5, 2026, five publishing houses — Hachette, Macmillan, McGraw Hill, Elsevier, and Cengage — filed a class-action complaint against Meta and Mark Zuckerberg in the U.S. District Court for the Southern District of New York. Bestselling author Scott Turow joined as a named plaintiff.
The publishers describe it as the first AI action brought by major publishing houses, spanning academic, educational, and trade sectors at once. That framing is accurate, but it undersells what makes the case legally interesting. This complaint is built on the one theory that has already extracted money from an AI developer — and it deliberately avoids the theory that has repeatedly failed.
The Distinction That Organizes Everything
Two 2025 decisions split AI copyright liability into separate questions, and understanding the split is the difference between following this case usefully and following it as noise.
Training on lawfully obtained text has fared well for AI developers. In Kadrey v. Meta, Judge Vince Chhabria of the Northern District of California granted summary judgment for Meta on June 25, 2025, holding that training a language model on copyrighted books — without reproducing recognizable expression in output — was fair use. He was explicit that the plaintiffs’ failure to prove concrete market dilution was dispositive.
Acquiring and keeping pirated copies has fared badly. In Bartz v. Anthropic, the court drew a line through the middle of the same conduct: training was transformative, but building and maintaining a permanent, general-purpose digital library of pirated books was not protected by fair use. That distinction is what produced the settlement — a non-reversionary fund of $1.5 billion covering roughly 482,000 works on a court-approved Works List, granted final approval on July 20, 2026 and described by the court as the largest copyright class action settlement in history.
The new complaint aims squarely at the second category. Its subject is not primarily what Llama learned; it is how Meta got the books.
What the Complaint Alleges
The factual core is acquisition by torrent. Plaintiffs allege Meta obtained more than 267 terabytes of books and journal articles from the shadow libraries LibGen and Anna’s Archive — a volume the complaint characterizes as many times the entire print collection of the Library of Congress — and then reproduced those works repeatedly in the course of training.
The willfulness allegations are the part that separates this filing from its predecessors. According to the complaint, Meta spent January through April 2023 negotiating dataset licenses and discussed raising its licensing budget to as much as $200 million, up from roughly $17 million. In early April 2023 the licensing effort stopped abruptly. The decision was escalated to Zuckerberg, after which the business-development team received verbal instructions to halt licensing.
Variety reported the complaint’s allegation that Zuckerberg “personally authorized and actively encouraged” the infringement — which is why he is named individually rather than only as an officer of the company.
AAP President Maria Pallante said Meta “made calculated decisions to enrich itself with literary properties that it did not create.” Hachette CEO David Shelley framed it as a matter of first principles: “Copyright is the bedrock of all creative industries.”
The pleaded claim is willful copyright infringement under the Copyright Act. Relief sought includes damages and injunctive relief — notably destruction of all infringing copies in Meta’s possession.
Why the Willfulness Framing Carries Weight
Willfulness is not rhetorical here. Under the Copyright Act, statutory damages for registered works rise substantially where infringement is found willful, and a documented decision to abandon a licensing negotiation in favor of acquiring the same material from piracy repositories is close to the paradigm case. That the alternative was priced — a $200 million budget under active discussion — makes the “we could not have licensed this” argument considerably harder to run.
It also complicates a fair-use defense at the acquisition stage. Bartz already indicated that maintaining a pirated library sits outside fair use even where downstream training is transformative. A record showing the pirated route was chosen over an available licensed one speaks directly to good faith, and to the fourth fair-use factor’s concern with the market for licensing.
What This Actually Means for Individual Authors
Four practical implications, stated without more optimism than the record supports.
A class action is the delivery mechanism, and class definition is where your interest lives. Bartz paid out against a court-approved list of specific works. Whether your titles fall inside this proposed class — and on what schedule — will be determined by class certification and any works list that follows, not by the strength of the allegations. Certification is contested in cases of this size, and it is the first real gate.
Registration timing governs what a claim is worth. Statutory damages and attorney’s fees generally require registration before infringement, or within the narrow window after publication. Authors relying on the existence of copyright without registration are relying on actual damages, which are hard to prove for an individual title in a training corpus. This is the least glamorous and most consequential item on the list.
Your publisher may control the claim. Whether you or your publisher has standing to assert infringement of a given title depends on what your contract granted. Where a publishing agreement assigned or exclusively licensed the relevant rights, the publisher is typically the party who can sue — which is part of why five publishers, rather than five thousand authors, are the named plaintiffs. Reading your grant clause against the scope of what a license actually transfers is the way to know which side of that line you are on.
Training itself remains largely unresolved in authors’ favor. Kadrey was decided for Meta on the training question and the case continues on an amended pleading after the court permitted amendment in March 2026. Chhabria’s market-dilution reasoning — his observation that no other use has anything near LLM training’s potential to flood the market with competing works — is a theory with real potential, but it is a theory that has so far cost plaintiffs a summary judgment rather than won them one. The reliable liability hook is still the pirated copy.
Reading the Case Going Forward
The signal to watch is not the size of the damages number in the complaint. It is whether the acquisition-versus-training distinction hardens into a rule. If it does, the operative question for every AI developer becomes provenance — where the corpus came from and whether it was paid for — rather than the philosophical status of machine learning. That is a question with documentary answers, which is exactly why this complaint reads the way it does.
For authors, the near-term work is unglamorous: know your registration status, know what your contract granted, and watch class certification rather than headlines. Those three facts determine whether a case like this reaches you at all.
This article is editorial and informational, not legal advice. Consult a licensed attorney about your specific situation.
