6 Sources
[1]
Researchers suggest OpenAI trained AI models on paywalled O'Reilly books | TechCrunch
OpenAI has been accused by many parties of training its AI on copyrighted content sans permission. Now a new paper by an AI watchdog organization makes the serious accusation that the company increasingly relied on non-public books it didn't license to train more sophisticated AI models. AI models
[2]
OpenAI's models 'memorized' copyrighted content, new study suggests | TechCrunch
A new study appears to lend credence to allegations that OpenAI trained at least some of its AI models on copyrighted content. OpenAI is embroiled in suits brought by authors, programmers, and other rights-holders who accuse the company of using their works -- books, codebases, and so on -- to
[3]
Study suggests OpenAI isn't waiting for copyright exemption
GPT-4o likely trained on O'Reilly books without permission, figures appear to show Tech textbook tycoon Tim O'Reilly claims OpenAI mined his publishing house's copyright-protected tomes for training data and fed it all into its top-tier GPT-4o model without permission. This comes as the
[4]
An AI Watchdog accused OpenAI of using copyrighted books without permission
An artificial intelligence watchdog is accusing OpenAI of training its default ChatGPT model on copyrighted book content without permission. In a new paper published this week, the AI Disclosures Project alleges that OpenAI likely trained its GPT-4o model using non-public material from O'Reilly
[5]
OpenAI might have trained its AI on stolen books
OpenAI is facing accusations of training its AI models on copyrighted material without permission, as a new paper alleges the company used paywalled books from O'Reilly Media to train its GPT-4o model. The AI Disclosures Project, a nonprofit co-founded by Tim O'Reilly and Ilan Strauss, published
[6]
Researchers Claim OpenAI Trained Its AI Models on Copyrighted Content
GPT-4o was said to show the highest recognition of copyrighted content OpenAI might have trained its artificial intelligence (AI) models on copyrighted content, according to a research paper. A recently published paper from the non-profit organisation AI Disclosures Project, the San
Share
Copy Link
A new study by the AI Disclosures Project suggests that OpenAI may have used paywalled O'Reilly Media books to train its GPT-4o model without proper licensing, raising concerns about copyright infringement and the need for transparency in AI training data sources.

A new study by the AI Disclosures Project, a nonprofit co-founded by Tim O'Reilly and Ilan Strauss, has accused OpenAI of training its GPT-4o model on copyrighted O'Reilly Media books without permission
1
. The research, which used a method called DE-COP, suggests that OpenAI's latest model demonstrates strong recognition of paywalled O'Reilly book content compared to earlier models1
.The researchers used 13,962 paragraph excerpts from 34 O'Reilly books to probe GPT-4o, GPT-3.5 Turbo, and other OpenAI models
1
. The study found that GPT-4o "recognized" far more paywalled O'Reilly book content than older models, even after accounting for potential confounding factors1
.This accusation comes amid ongoing debates about AI companies' use of copyrighted material for training purposes. OpenAI has been advocating for looser restrictions on developing models using copyrighted data
2
. The company has some content licensing deals in place but faces several lawsuits over its training data practices1
.A separate study by researchers from the University of Washington, the University of Copenhagen, and Stanford proposed a new method for identifying training data "memorized" by models
2
. This study suggested that GPT-4 showed signs of having memorized portions of popular fiction books and New York Times articles2
.Related Stories
The findings highlight the need for increased transparency regarding pre-training data sources and the development of formal licensing frameworks for AI content training
3
. There are concerns that failure to adequately compensate creators could lead to a decline in internet content quality and diversity3
.OpenAI has been seeking higher-quality training data and has hired experts in various domains to fine-tune its models' outputs
1
. The company has also urged the US government to relax copyright restrictions to facilitate AI model training3
.As the AI industry grapples with these issues, some companies are introducing measures to protect copyrighted material. For instance, Cloudflare has developed an AI-powered system designed to deter unauthorized web scraping
3
.Summarized by
Navi
[3]
[5]
24 Oct 2024•Technology

09 Jul 2026•Policy and Regulation

05 Sept 2026•Policy and Regulation
1
Science and Research

2
Policy and Regulation

3
Technology