6 Sources
[1]
Former OpenAI Employee Condemns the Company's Data Scraping Practices
An artificial intelligence researcher who worked at OpenAI as recently as August says the company violates copyright law. Part of Suchir Balaji's job was to gather enormous amounts of data for OpenAI's GPT-4 multimodal AI but at the time he treated it as a research project and didn't think that
[2]
OpenAI Whistleblower Disgusted That His Job Was to Vacuum Up Copyrighted Data to Train Its Models
"If you believe what I believe, you have to just leave the company." A former OpenAI researcher is blowing the whistle on the company's AI training practices, alleging that OpenAI violated copyright law to train its AI models -- and arguing that OpenAI's current business model stands to upend the
[3]
Former OpenAI researcher says the company broke copyright law
Suchir Balaji spent nearly four years as an artificial intelligence researcher at OpenAI. Among other projects, he helped gather and organize the enormous amounts of internet data the company used to build its online chatbot, ChatGPT. At the time, he did not carefully consider whether the company
[4]
Former OpenAI Staffer Says the Company Is Breaking Copyright Law and Destroying the Internet
Is OpenAI breaking U.S. copyright law? A former employee of the company says yes. A former researcher at the OpenAI has come out against the company's business model, writing, in a personal blog, that he believes the company is not complying with U.S. copyright law. That makes him one of a growing
[5]
Former OpenAI Researcher Says Company Broke Copyright Law
Cade Metz has written about artificial intelligence for 15 years. Suchir Balaji spent nearly four years as an artificial intelligence researcher at OpenAI. Among other projects, he helped gather and organize the enormous amounts of internet data the company used to build its online chatbot,
[6]
Former OpenAI researcher says the company broke copyright law
In August, Balaji, 25, left OpenAI because he no longer wanted to contribute to technologies that he believed would bring society more harm than benefit. "If you believe what I believe, you have to just leave the company," he said during a recent series of interviews with The New York Times.Suchir
Share
Copy Link
Suchir Balaji, a former OpenAI employee, speaks out against the company's data scraping practices, claiming they violate copyright law and pose a threat to the internet ecosystem.

Suchir Balaji, a 25-year-old artificial intelligence researcher who worked at OpenAI for nearly four years, has come forward with serious allegations against the company's data practices. Balaji, who left OpenAI in August 2024, claims that the company's use of copyrighted data to train its AI models violates copyright law and poses a significant threat to the internet ecosystem
1
2
.During his time at OpenAI, Balaji was involved in gathering and organizing vast amounts of internet data used to build products like ChatGPT. Initially, he viewed his work as part of a research project, assuming that using any internet data, copyrighted or not, was acceptable in that context
3
.However, Balaji's perspective changed dramatically after the release of ChatGPT in late 2022. He realized that what was once a closed-door research project had transformed into a commercialized product, raising serious ethical and legal concerns
2
.Balaji argues that OpenAI's data scraping practices do not meet the criteria for fair use, a legal doctrine that allows limited use of copyrighted material without permission
1
. He contends that while the outputs of AI models like ChatGPT aren't exact copies of the inputs, they are also not fundamentally novel, potentially infringing on copyrights4
.OpenAI, however, maintains that its use of publicly available data is protected by fair use principles and is critical for U.S. competitiveness
5
. The company is currently facing several lawsuits related to copyright infringement, including a high-profile case brought by The New York Times2
.Balaji expresses deep concern about the sustainability of OpenAI's business model for the internet ecosystem. He argues that AI technologies like ChatGPT are "destroying the commercial viability of the individuals, businesses and internet services that created the digital data used to train these AI systems"
5
.This sentiment is echoed by others in the tech industry, with a growing chorus of voices questioning the legitimacy and ethics of AI companies' data-hoovering practices
4
.Related Stories
The controversy surrounding AI training practices has led to increased calls for government intervention. Bradley Hulbert, an intellectual property lawyer, suggests that "it is time for Congress to step in" given the rapid evolution of AI technology
2
.As the debate intensifies, the AI industry faces mounting pressure to address these concerns. OpenAI's transition from a non-profit research organization to a commercial entity has only heightened scrutiny of its practices
2
5
.Balaji's whistleblowing adds to the ongoing discussion about the future of AI development and its impact on society. While he initially joined the AI industry believing in its potential to solve major global challenges, he now sees the technology causing more harm than good
1
4
.As lawsuits pile up and former insiders speak out, the AI industry may be forced to reckon with its data practices and their long-term consequences for creativity, innovation, and the digital landscape as a whole.
Summarized by
Navi
[2]
[4]
02 Apr 2025•Technology

09 Jul 2026•Policy and Regulation

30 Nov 2024•Policy and Regulation

1
Science and Research

2
Policy and Regulation

3
Technology