11 Sources
[1]
Court filings show Meta paused efforts to license books for AI training | TechCrunch
New court filings in an AI copyright case against Meta add credence to earlier reports that the company "paused" discussions with book publishers on licensing deals to supply some of its generative AI models with training data. The filings are related to the case Kadrey v. Meta Platforms -- one of
[2]
Meta purportedly trained its AI on more than 80TB of pirated content and then open-sourced Llama for the greater good
Court filings suggest Meta took steps to unsuccessfully mask its AI training activities Meta is facing a class-action lawsuit alleging copyright infringement and unfair competition over the training of its AI model, Llama. According to court documents released by vx-underground, Meta allegedly
[3]
Meta faces lawsuit for training AI with pirated books
In a recent lawsuit, Meta has been accused of using pirated books to train its AI models, with CEO Mark Zuckerberg's approval. As per Ars Technica, the lawsuit filed by authors including Ta-Nehisi Coates and Sarah Silverman in a California federal court, cite internal Meta communications indicating
[4]
Meta's Llama AI in hot water: Alleged copyright theft leads to class action lawsuit, accused of pirating 82TB of books for AI training
Meta is being sued for allegedly using pirated books from shadow libraries to train its AI models, despite internal concerns and ethical warnings. The company reportedly downloaded 81.7TB of data through torrents and concealed its involvement to bypass copyright laws.Meta's LLaMA AI is in the
[5]
Meta used pirated books to train its AI models, and there are emails to prove it
Facepalm: A group of authors has sued Meta, alleging that the company used unauthorized copies of their books to train its generative AI models. While Meta has denied any wrongdoing, newly unsealed messages suggest that executives and engineers were well aware of their actions - and that they were
[6]
Meta staff torrented nearly 82TB of pirated books for AI training -- court records reveal copyright violations
Facebook parent-company Meta is currently fighting a class action lawsuit alleging copyright infringement and unfair competition, among others, with regards to how it trained LLaMA. According to an X (formerly Twitter) post by vx-underground, court records reveal that the social media company used
[7]
Meta accused of pirating 82 terabytes of books from 'shadow libraries' to train AI
TL;DR: Leaked court documents indicate that Meta is accused of illegally downloading 82 terabytes of data to train its artificial intelligence systems. For sophisticated AI chatbots to exist, they need to be trained on large swaths of data, but where things get murky is when the big question is
[8]
Unredacted Meta emails reveal scale of book piracy for AI training
Disclaimer: This content generated by AI & may have errors or hallucinations. Edit before use. Read our Terms of use Newly uncovered emails gave fresh impetus to a copyright case against Meta, raised by book authors who claim the tech giant utilised pirated books to train its AI models. Authors
[9]
Meta accused of downloading torrents of 81.7TB of pirated books to train its Llama AI models
Meta has been accused of torrenting an astonishing 81.7TB of pirated books to train its Llama AI models according to a new lawsuit filed in the US District Court for the Northern District of California. The social networking giant has been accused of illegally torrenting copyrighted materials from
[10]
Court documents show not only did Meta torrent terabytes of pirated books to train AI models, employees wouldn't stop emailing each other about it: 'Torrenting from a corporate laptop doesn't feel right'
First reported by Ars Technica, the copyright case against Facebook parent company Meta over its use of authors' work to train large language models has unearthed some embarrassing dirty laundry in discovery. Dozens of emails, allegedly between Meta employees, discuss torrenting massive amounts of
[11]
81.7 terabytes of data: Meta downloads entire digital libraries without permission - Softonic
The information revealed suggests that among the 81.7 terabytes of downloaded data, at least 35.7 terabytes include library books Meta, the company of Mark Zuckerberg that was previously called Facebook, is facing serious accusations in the Kadrey case against the company, in which it is accused
Share
Copy Link
Meta is embroiled in a lawsuit alleging the company used pirated books to train its AI models, including Llama. Internal communications reveal ethical concerns and attempts to conceal the practice.
Meta, the parent company of Facebook, is facing a class-action lawsuit over allegations that it used pirated books to train its artificial intelligence models, including the popular Llama series. Court filings and internal communications have revealed that Meta allegedly downloaded and used vast amounts of copyrighted material without proper authorization
1
.According to the lawsuit, Meta is accused of downloading nearly 82 terabytes of pirated books from shadow libraries such as Anna's Archive, Z-Library, and LibGen
2
. The plaintiffs, including bestselling authors Sarah Silverman and Ta-Nehisi Coates, allege that Meta infringed upon their copyrights and potentially harmed their livelihoods3
.Unsealed court documents reveal that some Meta employees raised ethical concerns as early as 2022. One researcher explicitly stated, "I don't think we should use pirated material," while another employee commented that "torrenting from a corporate laptop doesn't feel right"
4
.Despite these internal warnings, Meta allegedly took steps to conceal its activities. Employees discussed ways to prevent Meta's infrastructure from being directly linked to the downloads, including using servers outside of Facebook's main network in what was referred to as "stealth mode"
5
.Meta has defended its practices by invoking the "fair use" doctrine, asserting that using publicly available materials to train AI tools is legal in certain cases. The company argues that it uses text to statistically model language and generate original expression
3
.Related Stories
This case is part of a larger trend of legal challenges against tech companies developing AI technologies. OpenAI and Nvidia have also faced similar accusations regarding their use of copyrighted materials for AI training
2
.U.S. District Judge Vince Chhabria has dismissed some claims but allowed the authors to amend their complaint to include new allegations, including those related to the removal of copyright management information
3
. The outcome of this lawsuit could have significant implications for the tech industry, particularly concerning the use of copyrighted materials in AI training.Summarized by
Navi
[2]
[3]
11 Mar 2025•Policy and Regulation

05 May 2026•Policy and Regulation

10 Mar 2026•Policy and Regulation

1
Technology

2
Policy and Regulation

3
Technology
