7 Sources
[1]
AI bots strain Wikimedia as bandwidth surges 50%
On Tuesday, the Wikimedia Foundation announced that relentless AI scraping is putting strain on Wikipedia's servers. Automated bots seeking AI model training data for LLMs have been vacuuming up terabytes of data, growing the foundation's bandwidth used for downloading multimedia content by 50
[2]
AI crawlers cause Wikimedia Commons bandwidth demands to surge 50% | TechCrunch
The Wikimedia Foundation, the umbrella organization of Wikipedia and a dozen or so other crowdsourced knowledge projects, said on Wednesday that bandwidth consumption for multimedia downloads from Wikimedia Commons has surged by 50% since January 2024. The reason, the outfit wrote in a blog post
[3]
AI data scrapers are an existential threat to Wikipedia
Wikipedia is one of the greatest knowledge resources ever assembled, containing crowdsourced contributions from millions of humans worldwide - and it faces a growing threat from artificial intelligence developers. The non-profit Wikimedia Foundation, which operates Wikipedia, says since January
[4]
Wikipedia Faces Flood of AI Bots That Are Eating Bandwidth, Raising Costs
Wikipedia is paying the price for the AI boom: The online encyclopedia is grappling with rising costs from bots scraping its articles to train AI models, which is straining the site's bandwidth. On Tuesday, the nonprofit that hosts Wikipedia warned that "automated requests for our content have
[5]
Wikimedia Foundation bemoans AI bot bandwidth burden
Crawlers snarfing long-tail content for training and whatnot cost us a fortune Web-scraping bots have become an unsupportable burden for the Wikimedia community due to their insatiable appetite for online content to train AI models. Representatives from the Wikimedia Foundation, which oversees
[6]
Wikipedia is struggling with voracious AI bot crawlers
The Wikimedia Foundation is getting pummeled by crawlers, which could cause issues for actual readers. Wikimedia has seen a 50 percent increase in bandwidth used for downloading multimedia content since January 2024, the foundation said in an update. But it's not because human readers have
[7]
Wikipedia servers are struggling under pressure from AI scraping bots
Editor's take: AI bots have recently become the scourge of websites dealing with written content or other media types. From Wikipedia to the humble personal blog, no one is safe from the network sledgehammer wielded by OpenAI and other tech giants in search of fresh content to feed their AI
Share
Copy Link
The Wikimedia Foundation reports a 50% increase in bandwidth consumption due to AI bots scraping content, causing technical and financial strain on their infrastructure.

The Wikimedia Foundation, the organization behind Wikipedia and other crowdsourced knowledge projects, has reported a significant increase in bandwidth consumption. Since January 2024, the foundation has experienced a 50% surge in bandwidth usage for multimedia downloads from Wikimedia Commons
1
. This surge is primarily attributed to automated bots scraping content for AI model training, rather than increased human traffic.The foundation's infrastructure, designed to handle sudden spikes in human traffic during high-interest events, is struggling to cope with the unprecedented volume of bot-generated traffic. Wikimedia's internal data reveals that bots account for 65% of the most expensive requests to its core infrastructure, despite making up only 35% of total pageviews
2
.This asymmetry in resource consumption is due to the nature of bot behavior. Unlike human users who tend to access popular and frequently cached content, bots indiscriminately crawl obscure and less-accessed pages. This forces Wikimedia's core datacenters to serve content directly, bypassing caching systems designed for predictable human browsing patterns
1
.The situation is further complicated by the sophisticated tactics employed by some AI-focused crawlers. Many of these bots ignore robots.txt directives, spoof browser user agents to appear as human visitors, and rotate through residential IP addresses to avoid blocking
1
. This cat-and-mouse game has forced Wikimedia's Site Reliability team into a perpetual state of defense, diverting resources from supporting contributors, users, and technical improvements.This issue is not unique to Wikimedia. Similar challenges are being faced across the open-source community and the broader internet. Other platforms like Fedora's Pagure repository, GNOME's GitLab instance, and Read the Docs have implemented various measures to combat excessive bot access and reduce bandwidth costs
1
.Related Stories
In response to these challenges, the Wikimedia Foundation is developing a "Responsible Use of Infrastructure" plan. This initiative aims to identify and filter access from AI bot scrapers, potentially requiring authentication for high-volume scraping and API use
4
.The foundation is also exploring systemic approaches under a new initiative called WE5: Responsible Use of Infrastructure. This raises critical questions about guiding developers toward less resource-intensive access methods and establishing sustainable boundaries while preserving openness
1
.The challenge lies in bridging the gap between open knowledge repositories and commercial AI development. Many companies rely on open knowledge to train commercial models but don't contribute to the infrastructure making that knowledge accessible. This creates a technical imbalance that threatens the sustainability of community-run platforms
1
.As the Wikimedia Foundation aptly states, "Our content is free, our infrastructure is not."
5
This situation calls for better coordination between AI developers and resource providers, potentially through dedicated APIs, shared infrastructure funding, or more efficient access patterns. Without such practical collaboration, the very platforms that have enabled AI advancement may struggle to maintain reliable service.Summarized by
Navi
[1]
[3]
[5]
10 Nov 2025•Business and Economy

17 Oct 2025•Technology

19 Jun 2025•Technology

1
Technology

2
Technology

3
Policy and Regulation
