2 Sources
[1]
Training AI on Mastodon posts? That idea's extinct
Such rules could be tricky to enforce in the Fediverse, though Mastodon is the latest platform to push back against AI training, updating its terms and conditions to ban the use of user content for large language models (LLMs). "We want to make it clear," the federated platform stated in an email
[2]
Mastodon's New Terms Block AI Scraping, But Gaps Still Remain
Open-source social media platform Mastodon has changed its rules to block companies from using user posts to train AI systems. The new policy, which takes effect from July 1, specifically bans scraping data from its main server, namely 'Mastodon.social'. The updated terms of service now clearly
Share
Copy Link
Mastodon updates its terms to prohibit AI training on user content, but the decentralized nature of the platform poses challenges in enforcing these rules across the entire Fediverse.
In a significant move to protect user privacy, Mastodon, the open-source social media platform, has updated its terms and conditions to prohibit the use of user content for training large language models (LLMs). The new policy, set to take effect from July 1, 2025, specifically bans the scraping of data from its main server, Mastodon.social
1
.
Source: The Register
"We want to make it clear that training LLMs on the data of Mastodon users on our instances is not permitted," the platform stated in an email to users
1
. This decision aligns Mastodon with other platforms like Bluesky, which have expressed similar intentions to protect user data from AI training1
.While the move is welcomed by privacy advocates, the decentralized nature of Mastodon poses significant challenges in enforcing these rules across the entire Fediverse. The new terms apply only to Mastodon's own instances, not the wider network of independently managed servers
2
.Eugen Rochko, founder of Mastodon, acknowledged the difficulty in enforcing such restrictions on a platform that prides itself on decentralization and openness. While it's possible to deploy a file to block AI crawlers, the effectiveness of this measure relies on the compliance of those behind the bots
1
.Mastodon's policy change comes amid growing concerns about AI companies using public online content without permission to train their models. This issue has sparked legal actions, such as Reddit's lawsuit against Anthropic for allegedly scraping user-generated content in violation of contractual terms
1
.The debate around AI training data has intensified, with platforms like Reddit signing data-sharing deals with AI companies like OpenAI and Google, while simultaneously taking legal action against others for unauthorized use of their data
1
.Related Stories
Along with the AI training ban, Mastodon has introduced other significant policy updates:
2
.2
.2
.While Mastodon's efforts to protect user data are commendable, the limited scope of the policy highlights the challenges faced by decentralized platforms in presenting a unified front against AI data harvesting. Unless more servers in the Fediverse adopt similar rules, AI companies may still have access to vast amounts of user-generated content
2
.As the debate over AI training data continues, Mastodon's policy change serves as a significant milestone in the ongoing struggle between open, decentralized platforms and the need for user data protection in the age of AI.
Summarized by
Navi
[1]
1
Science and Research

2
Technology

3
Policy and Regulation
