6 Sources
[1]
Your Bluesky posts might be training AI
Bluesky is grappling with a significant privacy issue after one million public posts were scraped from its platform for AI training, according to a 404Media report. The dataset, compiled by machine learning librarian Daniel van Strien from the AI company Hugging Face, was intended for use in
[2]
One million public Bluesky posts scraped for AI training
Bluesky is already facing its first major AI scrape, despite the stance of its owners that it will never train generative AI on user data. Reported by 404Media on Nov. 26, one million public Bluesky posts -- complete with identifying user information -- were crawled and then uploaded to AI company
[3]
In Bluesky, they also use your data to train AI, even though the company claims not to do so. - Softonic
A researcher published a dataset with scraped data from Bluesky users, but has already deleted it. Bluesky, the decentralized social network, is at the center of controversy following the recent publication of a dataset on Hugging Face, a community platform for artificial intelligence. According
[4]
Bluesky's open API means anyone can scrape your data for AI training | TechCrunch
Per a report by 404 Media, a machine learning librarian at AI firm Hugging Face pulled 1 million public posts from Bluesky via its Firehose API for machine learning research, pushing the dataset to a public repository. Daniel van Strien later removed the data due to the controversy that ensued,
[5]
Think Twice Before Joining Bluesky
Bluesky offers an open API, allowing its data to be used for training AI models. Since the election day in the US, Bluesky, a microblogging alternative to X (formerly Twitter), has been rapidly gaining popularity. The user base has doubled since September to reach 20 million by November 20. The
[6]
Bluesky dataset for AI training removed from Hugging Face
The creator of the dataset issued an apology to concerned users in a post on Bluesky. A dataset of 1m Bluesky posts that was uploaded to machine learning platform Hugging Face earlier this week has been removed. On 26 November, Daniel van Strien, a machine learning librarian at Hugging Face,
Share
Copy Link
Bluesky, a decentralized social network, faces privacy concerns after one million public posts were scraped for AI training, highlighting the platform's vulnerability to data collection despite its stance against using user data for AI.

Bluesky, a decentralized social network gaining popularity as an alternative to X (formerly Twitter), has found itself at the center of a privacy controversy. A dataset containing one million public posts from the platform was scraped and uploaded to AI company Hugging Face, intended for machine learning research
1
. This incident has raised significant concerns about user privacy and consent in the age of AI.The dataset was compiled by Daniel van Strien, a machine learning librarian at Hugging Face. It included not only the text of posts but also users' decentralized identifiers (DIDs) and metadata
2
. Van Strien's intention was to use this data for research related to natural language processing, social media analysis, and content moderation.The core issue stems from Bluesky's open and decentralized nature, built on the Authenticated Transfer (AT) Protocol. The platform's Firehose API provides an aggregated stream of public data updates, making it vulnerable to external scrapers
3
. This openness, while beneficial for third-party developers, has exposed a significant privacy risk for users.Bluesky users did not provide explicit permission for their posts to be used in this manner. The platform has stated that it will never train generative AI on user data, but it acknowledges that it cannot enforce this policy outside its systems
4
. Bluesky is now exploring ways to enable users to communicate their consent preferences to external parties, though enforcement will ultimately depend on outside developers.Following the backlash, van Strien removed the dataset from Hugging Face and issued a public apology, acknowledging the breach of transparency and consent in his data collection approach
5
. This incident has sparked a broader discussion about data privacy and the ethical use of public information for AI training.Related Stories
This controversy serves as a reminder that content shared publicly on platforms like Bluesky is accessible to external entities. As Bluesky continues to grow, surpassing 20 million users, it will likely face increasing scrutiny regarding its data protection measures and user privacy policies.
The incident highlights a growing trend in the tech industry where user-generated content is being used to train AI models. Other platforms like X and Meta have updated their terms of service to allow for such practices, while LinkedIn has introduced an opt-out option for users who don't want their data used for AI training.
As AI technology advances, the balance between open platforms, user privacy, and ethical AI training becomes increasingly complex. Bluesky's situation underscores the need for clearer policies, user consent mechanisms, and industry-wide standards for the responsible use of public data in AI development.
Summarized by
Navi
[1]
[3]
16 Nov 2024•Technology

29 Mar 2026•Technology

18 Oct 2024•Technology

1
Science and Research

2
Technology

3
Policy and Regulation
