17 Sources
[1]
Zuckerberg's YouTube Defense Sparks Debate in Meta's AI Copyright Battle
Zuckerberg's Unconventional Defense: Fair Use or Flawed Logic in AI Training Debate? Meta CEO Mark Zuckerberg recently defended his company's use of a copyrighted e-book dataset in a deposition for the ongoing AI copyright case, Kadrey v. The case is part of a massive lawsuit that involves AI
[2]
In AI copyright case, Zuckerberg turns to YouTube for his defense | TechCrunch
Meta CEO Mark Zuckerberg appears to have used YouTube and its battle to take down pirated content to defend his own company's use of a data set containing copyrighted e-books to train AI models, newly released snippets of his deposition reveals. The deposition, which was part of a complaint
[3]
Mark Zuckerberg Allegedly Trained AI Models on Copyrighted Works
LibGen is a link aggregator that provides access to copyrighted works Meta is facing a copyright lawsuit over allegedly using copyrighted works to train its artificial intelligence (AI) models. The lawsuit was filed by multiple complainants that also include several bestselling authors. The
[4]
Mark Zuckerberg gave Meta's Llama team the OK to train on copyrighted works, filing claims
Counsel for plaintiffs in a copyright lawsuit filed against Meta allege that Meta CEO Mark Zuckerberg gave the green light to the team behind the company's Llama AI models to use a data set of pirated ebooks and articles for training. The case, Kadrey v. Meta, is one of many against tech giants
[5]
Zuckerberg Appeared to Know Meta Trained AI on Pirated Library
What Happened With California's Water Supply During the Wildfires? The AI rush has brought with it thorny questions of copyright and ownership of data as tech companies train bots like ChatGPT on existing texts, but it seems Meta largely brushed these aside as they worked to integrate such tools
[6]
Zuckerberg Knowingly Used Pirated Data to Train Meta AI, Authors Allege - Decrypt
Mark Zuckerberg approved using pirated books to train Meta AI, even after his own team warned the material was illegally obtained, a group of authors allege in a recent court filing. The allegations come from a copyright infringement lawsuit filed by a group of authors including the comedian Sarah
[7]
Zuckerberg approved Meta's use of 'pirated' books to train AI models, authors claim
Sarah Silverman and others file court case claiming CEO approved use of dataset despite warnings Mark Zuckerberg approved Meta's use of "pirated" versions of copyright-protected books to train the company's artificial intelligence models, a group of authors has alleged in a US court
[8]
Lawsuit says Mark Zuckerberg approved Meta's use of pirated materials to train Llama AI
Meta allegedly used LibGen for AI training and even stripped copyright information from its materials. Meta knowingly used pirated materials to train its Llama AI models -- with the blessing of company chief Mark Zuckerberg -- according to an ongoing copyright lawsuit against the company. As
[9]
Mark Zuckerberg named in lawsuit over Meta's use of pirated books for AI training
A group of authors, including Ta-Nehisi Coates and Sarah Silverman, alleged in a court filing that Meta CEO Mark Zuckerberg approved "Meta's torrenting and processing of pirated copyrighted works" to train the company's AI models. The California filing, which was made public on Wednesday, claims
[10]
Lawsuit Alleges Mark Zuckerberg Gave Permission for Meta to Train AI on Stolen Content
An amended lawsuit against Meta alleges that the company's CEO, Mark Zuckerberg, approved training Meta's Llama AI models using copyrighted content. The initial lawsuit was filed in 2023, alleging that Meta had trained its AI using copyrighted content. However, courts have favored Meta's position
[11]
Meta AI copyright case: Who's liable for using open-source data?
A group of authors in the US -- Richard Kadrey, Sarah Silverman and Christopher Golden claim that Meta torrented and processed their books to train its artificial intelligence models. The authors had originally filed a case against Meta's copyright infringement of their works back in 2023. In their
[12]
Meta knew it used pirated books to train AI, authors say
(Reuters) - Meta Platforms used pirated versions of copyrighted books to train its artificial intelligence systems with approval from its CEO Mark Zuckerberg, a group of authors alleged in newly disclosed court papers. Ta-Nehisi Coates, comedian Sarah Silverman and other authors suing Meta for
[13]
Court docs allege Meta trained AI model using LibGen
Did Zuck's definition of 'free expression' just get even broader? Meta allegedly downloaded material from an online source that's been sued for breaching copyright, because it wanted the material to train its AI models, according to a new court filing. The accusation was made in a document [PDF]
[14]
Meta accused of training its AI using pirated content from torrents
A new day, a new controversy around artificial intelligence. This time, Meta has been accused of using pirated content from torrents to train its large language model (LLM) Llama, which powers Meta AI. The case was one of the first copyright lawsuits filed against a tech company for training
[15]
Meta Secretly Trained Its AI on a Notorious Russian 'Shadow Library,' Newly Unredacted Court Docs Reveal
One of the most important AI copyright legal battles just took a major turn. Meta just lost a major fight in its ongoing legal battle with a group of authors suing the company for copyright infringement over how it trained its artificial intelligence models. Against the company's wishes, a court
[16]
Facebook Apparently Trained Its AI by Torrenting Pirated Books Stolen From Authors
And Zuckerberg personally approved the piracy, according to these documents. Newly unredacted court documents allege that Meta, formerly Facebook, knowingly used pirated books obtained from the online archive Library Genesis to train its AI models, Wired reports. Submitted in an ongoing lawsuit
[17]
Meta trained Llama on copyrighted material, new filing claims
Meta requested that a 'large portion' of the new filing be redacted, the judge however denied the request. In a new filing, the counsel for the trio of authors suing Meta claim that the company allowed its artificial intelligence (AI) large language model Llama to commit copyright infringement on
Share
Copy Link
Meta CEO Mark Zuckerberg defends the use of copyrighted e-books to train AI models, comparing it to YouTube's content moderation challenges. The case raises questions about fair use in AI development.

In a high-profile lawsuit, Meta faces allegations of using copyrighted materials to train its AI models without proper authorization. The case, Kadrey v. Meta, involves bestselling authors Sarah Silverman and Ta-Nehisi Coates as plaintiffs, challenging the tech giant's practices in AI development
1
.During a deposition, Meta CEO Mark Zuckerberg drew a controversial parallel between Meta's use of copyrighted e-books and YouTube's content moderation challenges. He argued that, like YouTube, which may temporarily host pirated content, it's not always rational to completely avoid using certain datasets in AI training
2
.Zuckerberg stated, "So would I want to have a policy against people using YouTube because some of the content may be copyrighted? No. There are cases where having such a blanket ban might not be the right thing to do"
2
.At the heart of the lawsuit is Meta's alleged use of LibGen, a controversial "links aggregator" providing access to copyrighted works. Court filings suggest that Zuckerberg approved the use of LibGen for training Meta's Llama AI models, despite internal concerns about legal implications
3
.Plaintiffs' counsel alleges that Meta attempted to conceal its use of copyrighted materials. According to the filings, Meta engineer Nikolay Bashlykov wrote a script to remove copyright information from ebooks in LibGen. The company is also accused of stripping copyright markers from science journal articles and source metadata in the training data
4
.The lawsuit also claims that Meta torrented the LibGen dataset, potentially engaging in another form of copyright infringement by participating in the distribution of copyrighted materials. This decision allegedly raised concerns among some Meta research engineers
4
.Related Stories
Meta's primary defense rests on the fair use doctrine, arguing that using text to statistically model language and generate original expression falls under permissible use of copyrighted material. However, the recently unsealed documents appear to challenge this argument
5
.This case is part of a larger debate surrounding AI companies' use of copyrighted works for training. The outcome could set a precedent for how fair use is interpreted in the context of AI development, potentially affecting the entire tech industry's approach to AI training data
1
.As the AI industry continues to grapple with these legal and ethical challenges, the resolution of this case may have far-reaching implications for the future of AI development and copyright law in the digital age.
Summarized by
Navi
[1]
[4]
[5]
08 Feb 2025•Technology
11 Mar 2025•Policy and Regulation

05 May 2026•Policy and Regulation

1
Technology

2
Policy and Regulation

3
Technology
