2 Sources
[1]
Meta allegedly used pirated books to train AI. Australian authors have objected, but US courts may decide if this is 'fair use'
Companies developing AI models, such as OpenAI and Meta, train their systems on enormous datasets. These consist of text from newspapers, books (often sourced from unauthorised repositories), academic publications and various internet sources. The material includes works that are copyrighted. The
[2]
Meta allegedly used pirated books to train AI -- US courts may decide if this is 'fair use'
Companies developing AI models, such as OpenAI and Meta, train their systems on enormous datasets. These consist of text from newspapers, books (often sourced from unauthorized repositories), academic publications and various internet sources. The material includes works that are copyrighted. The
Share
Copy Link
Meta faces legal challenges for allegedly using pirated books to train AI, raising questions about copyright infringement and fair use in the AI industry. The case highlights growing tensions between tech companies and content creators.

Meta, the parent company of Facebook and Instagram, is facing serious allegations of using pirated books to train its artificial intelligence (AI) models. According to recent reports, Meta allegedly utilized LibGen, an illegal book repository, to access copyrighted material for AI training purposes
1
2
. This revelation has ignited a fierce debate about the ethics and legality of using copyrighted content in AI development.LibGen, created by Russian scientists in 2008, hosts over 7 million books and 81 million research papers, making it one of the world's largest repositories of pirated work
1
2
. The Atlantic magazine's allegations suggest that Meta's use of this unauthorized database for AI training could have far-reaching implications for the publishing industry and individual authors.The controversy has sparked multiple legal challenges against Meta and other AI companies. A group of authors, including Michael Chabon, Ta-Nehisi Coates, and Sarah Silverman, have filed a lawsuit against Meta for copyright infringement
1
2
. Court documents allege that Meta CEO Mark Zuckerberg approved the use of the LibGen dataset despite knowing it contained pirated material.At the heart of these legal battles is the question of whether mass data scraping for AI training constitutes "fair use"
1
2
. AI companies argue that their use of copyrighted works is transformative and falls under the fair use doctrine. However, when AI systems can reproduce content that closely mimics an author's style or regenerates substantial portions of copyrighted material, it raises legitimate concerns about infringement.The ongoing legal challenges have significant implications for both the publishing industry and AI companies. Authors and publishers are increasingly concerned about losing control over their intellectual property and the potential devaluation of their work
1
2
. The average median full-time income for authors in the United States was just over $20,000 in 2023, highlighting the precarious financial situation many writers face1
2
.Related Stories
In response to these challenges, organizations like the Australian Society of Authors (ASA) are calling for government regulation of AI
1
2
. They propose that AI companies should be required to obtain permission before using copyrighted work and provide fair compensation to writers. The ASA also advocates for clear labeling of AI-generated content and transparency regarding the use of copyrighted works in AI training.As the legal battles unfold, the outcome will likely shape the future relationship between AI development and copyright law. The industry is grappling with finding a balance between fostering innovation and protecting the rights and livelihoods of content creators. The resolution of these cases may set important precedents for how AI companies can ethically and legally use copyrighted material in the development of their technologies.
Summarized by
Navi
[1]
15 Jul 2025•Policy and Regulation

25 Jun 2025•Policy and Regulation

06 Aug 2025•Policy and Regulation

1
Technology

2
Technology

3
Policy and Regulation
