15 Sources
[1]
Reddit blocks Internet Archive to end sneaky AI scraping
Reddit is now blocking the Internet Archive (IA) from indexing popular Reddit threads after allegedly catching sneaky AI firms -- restricted from scraping Reddit -- instead simply scraping data from IA's archived content. Where before IA's Wayback Machine dependably archived Reddit pages,
[2]
Reddit will block the Internet Archive
Reddit says that it has caught AI companies scraping its data from the Internet Archive's Wayback Machine, so it's going to start blocking the Internet Archive from indexing the vast majority of Reddit. The Wayback Machine will no longer be able to crawl post detail pages, comments, or profiles;
[3]
Reddit blocks the Internet Archive from crawling its data - here's why
Publishers (and others) are suing AI companies for copyright infringement. Reddit is defending its privacy from AI companies that are taking roundabout approaches to scraping its content. The social media platform, known as a resource where users can post anonymously and find information about
[4]
Reddit Is Blocking Internet Archive to Halt Free Scraping of User Data
(Credit: Thomas Fuller/SOPA Images/LightRocket via Getty Images) Reddit is limiting access to the Internet Archive after finding out that AI companies have used its Wayback Machine to scrape user data for free, The Verge reports. The Internet Archive is a nonprofit digital library that preserves
[5]
Reddit is restricting its availability to the Internet Archive's Wayback Machine
The Internet Archive's Wayback Machine is the latest victim of Reddit's crackdown on data access. The company has begun to place new restrictions on what the archive site will be able to access in a move that will significantly limit the Wayback Machine's ability to preserve information from
[6]
Reddit Is Blocking the Wayback Machine From Archiving Posts
Reddit is limiting the Wayback Machine from indexing most of its site over concerns of unauthorized AI scraping. Reddit is blocking the Internet Archive’s Wayback Machine from indexing most of its site, after discovering that AI companies were scraping its data from the digital time
[7]
Reddit blocks non-profit Wayback Machine from archiving the site - 9to5Mac
The Internet Archive's Wayback Machine is one of the most valuable free services available on the web, ensuring that important sources of information are protected from the vicissitudes of fate and tech companies. Until recently, the archive was able to capture the entirety of Reddit, but that is
[8]
Reddit is blocking Wayback Machine from archiving users' posts
Reddit will reportedly block the Internet Archive's Wayback Machine from saving users' posts. The social media platform states that the measure is intended to stop AI companies from scraping archived comments to train their algorithms. Or at least, prevent them from doing so without paying up. As
[9]
The internet is about to get a little worse as Reddit moves to block the Internet Archive so AI companies can't scrape its content
Google and OpenAI can scrape Reddit's content, but they paid for it. The internet, which was once a useful thing, is about to become a little less so: A new report from The Verge says Reddit is going to start blocking the Wayback Machine from indexing most of its content. The Wayback Machine,
[10]
It's About to Get Harder to Read Old Reddit Threads, and You Can Blame AI
Reddit and the Internet Archive are still in talks about the decision. With more and more AI showing up in Google searches as of late, I've been leaning extra hard on that one magic word that makes the internet work: Reddit. It's got its problems, but appending "Reddit" to a search is still the
[11]
AI data wars push Reddit to block the Wayback Machine
As the battle to train artificial intelligence models becomes more intense and Reddit's rich content library becomes more valuable, the social media giant has taken steps to block the Internet Archive from indexing its pages. While the Wayback Machine has historically recorded all Reddit pages,
[12]
Reddit says its blocking the Internet Archive to stop sneaky AI scrapers accessing its content - SiliconANGLE
Reddit says its blocking the Internet Archive to stop sneaky AI scrapers accessing its content Reddit Inc. said today it has decided to block the Internet Archive from indexing its popular web forums in order to prevent sneaky artificial intelligence firms from scraping its content for training
[13]
Reddit locks out Wayback machine to stop AI from scraping old posts
Reddit has restricted the Internet Archive's Wayback Machine from extensively capturing its content due to concerns over unauthorized AI data scraping. The platform will now allow only the homepage to be archived, aiming to protect user privacy and control content use. This move highlights the
[14]
Reddit Restricts Wayback Machine's Access To Only Its Homepage
Reddit has stopped the Internet Archive's Wayback Machine from indexing most of its content, saying that AI companies have been using it to bypass licensing fees and scrape user data, as per a report by The Verge. From now on, the Wayback Machine will not be able to archive posts, comments, or
[15]
Reddit Cuts Off Wayback Machine Over AI Data-Scraping Concerns
The action ends a long practice of regularly preserving public pages by the Wayback Machine for research and historical reasons. The shift is targeted directly at preventing artificial intelligence firms from circumventing its licensing agreements. While the Internet Archive is a 'good-faith
Share
Copy Link
Reddit has implemented restrictions on the Internet Archive's Wayback Machine to prevent AI companies from scraping user data without permission, sparking debates about data privacy and AI training practices.
In a significant move to protect user data and enforce its platform policies, Reddit has implemented restrictions on the Internet Archive's Wayback Machine. This decision comes after the discovery that AI companies were using the archive to circumvent Reddit's data scraping restrictions
1
.
Source: SiliconANGLE
Under the new policy, the Wayback Machine will only be allowed to archive Reddit's homepage, effectively limiting its ability to preserve the platform's vast content ecosystem. The restrictions prevent the archiving of post detail pages, comments, user profiles, and subreddit pages
2
.Reddit spokesperson Tim Rathschmidt stated that the company became "aware of instances where AI companies violate platform policies, including ours, and scrape data from the Wayback Machine"
3
. This move is part of Reddit's broader strategy to control access to its data, especially in the context of AI training.The Internet Archive, a non-profit digital library, has been an essential resource for researchers and historians. The restrictions will significantly limit its ability to preserve Reddit's content, potentially impacting future studies on online culture and digital forensics
3
.
Source: 9to5Mac
Reddit has been actively pursuing data licensing agreements with AI companies. It has struck deals with Google and OpenAI, allowing them to use Reddit's content for AI training in exchange for substantial fees
4
. The platform's approach underscores the growing value of user-generated content in the AI era.Related Stories
This incident highlights the ongoing tensions between AI companies, content platforms, and copyright holders. Several publishers and creators have filed lawsuits against AI firms for alleged copyright infringement, challenging the notion of "fair use" in AI training
3
.
Source: ZDNet
Reddit's decision raises questions about the future of data access for AI training. As platforms become more protective of their data, AI companies may need to reassess their data acquisition strategies and potentially negotiate more licensing agreements
5
.The Internet Archive has expressed a willingness to continue discussions with Reddit about this matter. Mark Graham, director of the Wayback Machine, stated that they have a "longstanding relationship with Reddit" and hope to find an amicable solution
4
.Summarized by
Navi
[1]
[2]
15 Apr 2026•Policy and Regulation

01 May 2026•Policy and Regulation

05 Jun 2025•Policy and Regulation

1
Technology

2
Technology

3
Policy and Regulation
