Zuckerberg Personally Named in Publishers' Copyright Lawsuit Over Llama Training Data
Five publishing houses and author Scott Turow sued Meta and CEO Mark Zuckerberg in federal court in Manhattan on Tuesday, alleging the company illegally trained its Llama AI system on millions of copyrighted works without authorization. The class action accuses Zuckerberg of personally authorizing and actively encouraging the infringement — naming a CEO individually is an escalation beyond the standard corporate complaint.
The Allegation
The suit argues that Meta’s Llama training datasets included copyrighted books obtained without license. The “personally authorized and actively encouraged” framing is significant: plaintiffs are arguing this was not a rogue engineering decision but a deliberate business choice made at the top level, which creates a legal theory for piercing the corporate shield on damages.
Scott Turow, whose legal thrillers have sold tens of millions of copies, joins five major publishers in the action. The class structure means other rights holders could join if certified.
Why This Matters More Than Prior Cases
Previous copyright suits against AI companies — including earlier cases involving training data — have largely targeted corporate entities. Naming Zuckerberg personally raises the stakes significantly. If courts accept the “personally authorized” framing, it opens executives to liability exposure that cannot be discharged through corporate restructuring and creates a deterrent that settlement of prior suits did not.
Meta’s legal exposure from training data lawsuits has been a persistent overhang since the first Llama models shipped in open-weights form. The open-weights distribution compounds the issue: the allegedly infringing training outputs are now in the hands of tens of thousands of downstream users.
The Broader Pattern
This suit lands the same week a Stanford study confirmed that finetuning can cause GPT-4o, Gemini 2.5 Pro, and DeepSeek to reproduce up to 90% of copyrighted books verbatim — a finding that has already been covered here. The publishers’ lawsuit and that research together establish a reinforcing legal and empirical narrative that AI training data practices constitute systemic, not incidental, infringement.
Meta has not publicly responded to this specific suit. The company has previously argued that training on publicly available data constitutes fair use — an argument that has not yet been tested at trial in any major AI copyright case.
The case is filed in the Southern District of New York.