3 Sources
[1]
Reddit keeps weird DMCA lawsuit against web scraper alive despite Google's loss
On Friday, a judge largely denied a motion to dismiss from a web scraper, SerpApi, which is accused of conspiring with Perplexity AI to illegally scrape copyrighted Reddit content from Google search results. In his opinion, US District Judge Paul A. Engelmayer said that at this early stage, Reddit has plausibly pleaded that there was a conspiracy, with SerpApi providing a product to circumvent Google access controls and Perplexity AI paying for it. Engelmayer's decision came less than two weeks after another court dismissed a similar action raised by Google, finding that the company had not proven that rights holders, such as Reddit, had ever authorized the search engine to prevent the scraping of protected content. Google told Ars that it planned to amend its complaint to keep its lawsuit alive, but SerpApi told Ars that Google and Reddit were both trying to "use the DMCA to wall off the open Internet by retroactively claiming control over content that they didn't author and don't own." For both Google and Reddit, the mission is to first prove that SerpApi and Perplexity AI conspired to access snippets of works covered by the Copyright Act that were, second, protected by a technological measure effectively controlling access, and that, third, the defendants circumvented that technology. Google failed on the first prong, giving SerpApi a rare win at such an early stage, but it may strengthen Google's arguments that Engelmayer agreed with Reddit that it was plausible the company had authorized Google to use anti-circumvention technology to block malicious scraping. And it's likely upsetting to web scrapers like SerpApi that Engelmayer thinks Reddit can make that case, even though Google's technology was invented more than a year after Google and Reddit struck their licensing deal. According to Engelmayer, it would be impractical to expect partners to update licensing deals every time a company rolls out new security methods. Additionally, Engelmayer found that "the Google Decision is not to the contrary" of Reddit's case because, unlike Google, Reddit went "beyond the bare allegation" that Google used to broadly claim that it generally "has licenses to display copyrighted content." Instead, Reddit argued that its licensing agreement with Google directly prohibits certain uses of Reddit data that are now being accessed due to the circumvention methods employed by malicious web scrapers. Specifically, Reddit argued that when it licenses content to partners like Google, its partners agree to delete posts that Reddit flags when users remove content. According to Reddit, "millions of posts" are deleted monthly, and unsanctioned efforts like SerpApi's partnership with Perplexity AI make it impossible for Reddit to protect its promise to users to honor content removals. And allowing deleted posts to fester in Perplexity AI's answer engine allegedly harms Reddit's reputation, as well as its profits, Reddit successfully argued. Reddit cheers; SerpApi prepares to fight If Reddit wins the fight, the popular online discussion forum could be in a better position to force all AI scrapers to enter into licensing agreements. Reddit has asked the court for an injunction blocking SerpApi and Perplexity AI access to both Reddit and Google websites, another injunction stopping circumvention of Google SearchGuard, and a third stopping SerpApi and Reddit from using previously scraped data. A Reddit spokesperson celebrated the ruling against the motion to dismiss in a statement provided to Ars. "Today's ruling brings us one step closer to holding bad actors accountable," Reddit's spokesperson said. "Reddit supports responsible access to public content, but we oppose companies that bypass our protections, ignore our rules, and profit off our communities without permission. Redditors create some of the most valuable human conversations on the Internet. We intend to protect them." The fight is seemingly far from over, though, with Engelmayer noting that SerpApi and Perplexity AI may prove through discovery that Reddit never authorized Google to protect its content in search results. SerpApi may also strengthen its defense if it can prove that all publicly accessible content in Google search results is not protected by the Copyright Act, a footnote in Engelmayer's opinion suggested. Last week, a DMCA expert with Public Knowledge, Meredith Rose, told Ars that Google and Reddit seemed to be "sort of grasping at whatever tool is available" in the face of the sudden, continuous rise of AI scraping over the past three years. And Reddit in particular appeared weirdly positioned in its DMCA claim because "the judge in the Google case said, 'Well, in order to have standing to bring a lawsuit under the DMCA, you can be the copyright owner or the exclusive licensee or the person who is deploying and manufacturing the technological protection measure at issue,'" Rose told Ars. And confusingly, "Reddit is none of those things." But Rose did acknowledge that DMCA rulings seemed to be more about "vibes," suggesting that it may be hard to predict a winner or loser in this fight just yet. This week wasn't a total loss for SerpApi and Perplexity AI, which did manage to get Reddit's unjust enrichment and unfair competition claims tossed, since they were both preempted by the Copyright Act. Asked for comment, Jeff Homrig, a lawyer for SerpApi, told Ars that "we remain confident in our position. The court has decided to hear the facts; the facts are on our side. SerpApi accesses public search results, not Reddit's platform, and public information does not become protected because a platform wants to charge for it. We look forward to making that case." Perplexity AI did not immediately respond to Ars' request for comment. Advance Publications, which owns Ars Technica parent Condé Nast, is the largest shareholder in Reddit.
[2]
Reddit's AI copyright lawsuit against Perplexity can move forward.
A judge rejected Perplexity's motion to dismiss the lawsuit, which accuses the AI startup and three data-scraping services of vacuuming up Reddit's content without permission. Ben Lee, Reddit's chief legal officer, says in a statement to The Verge that the ruling "brings us one step closer to holding bad actors accountable:" Reddit supports responsible access to public content, but we oppose companies that bypass our protections, ignore our rules, and profit off our communities without permission.
[3]
Perplexity AI loses bid to toss Reddit lawsuit over data scraping
July 31 (Reuters) - A Manhattan federal judge on Friday rejected most of Perplexity AI's bid to dismiss a lawsuit brought by Reddit (RDDT.N), opens new tab accusing it of violating U.S. copyright law by scraping the online discussion platform's data to train Perplexity's AI-powered search engine. U.S. District Judge Paul Engelmayer said in his decision, opens new tab that Reddit could continue pressing its claims that Perplexity and three data scrapers unlawfully circumvented protective measures to steal content for AI training. Among other things, Engelmayer said that Reddit had standing to sue for Perplexity's alleged misuse of its users' content. "Reddit's suit claims the right to control access to public web pages it doesn't own, using a security tool it didn't build, on behalf of users it hasn't asked," a Perplexity spokesperson said. "We're going to defend the open internet, and we're going to win." "Today's ruling brings us one step closer to holding bad actors accountable," a Reddit spokesperson said. "Reddit supports responsible access to public content, but we oppose companies that bypass our protections, ignore our rules, and profit off our communities without permission." The case is one of many filed by content owners including authors, music labels and news outlets against tech companies over the alleged misuse of copyrighted material to train AI systems. Reddit filed a separate data-scraping lawsuit against Anthropic last year that is still ongoing in California state court. Reddit, which features thousands of interest-based "subreddit" web communities, said in the lawsuit against Perplexity that the platform is the most commonly cited source for AI-generated answers to user questions. It has licensed its content to Google, OpenAI and others for their AI training. Reddit said that Lithuania-based Oxylabs, Russia-based AWMProxy and Texas-based SerpApi -- which it also sued -- scraped Reddit data from billions of search results without permission and that Perplexity, which does not have a license to Reddit content, worked with at least one of them to obtain its material. SerpApi attorney Jeff Homrig of Weil Gotshal & Manges said in a statement that the company "accesses public search results, not Reddit's platform, and public information does not become protected because a platform wants to charge for it." Spokespeople for Oxylabs did not immediately respond to a request for comment, and AWMProxy could not immediately be reached for comment. Reddit asked the court for unspecified monetary damages and an order blocking Perplexity from using its data. Perplexity has denied the allegations. Engelmayer on Friday dismissed some of Reddit's secondary claims, but advanced its claims that Perplexity unlawfully scraped its data and engaged in a conspiracy with the data-scraping companies. The case is Reddit Inc v. SerpApi LLC, U.S. District Court for the Southern District of New York, No. 1:25-cv-08736. For Reddit: Reid Bolton, Matthew Ford and William Gohl of Bartlit Beck For Perplexity: Eric MacMichael, Sharif Jacob, Benjamin Rothstein and Christina Lee of Keker Van Nest & Peters Reddit sues Perplexity for scraping data to train AI system Reporting by Blake Brittain in Washington Our Standards: The Thomson Reuters Trust Principles., opens new tab * Suggested Topics: * Litigation * Data Privacy * Intellectual Property Blake Brittain Thomson Reuters Blake Brittain reports on intellectual property law, including patents, trademarks, copyrights and trade secrets, for Reuters Legal. He has previously written for Bloomberg Law and Thomson Reuters Practical Law and practiced as an attorney.
Share
Copy Link
A Manhattan federal judge rejected Perplexity AI's motion to dismiss Reddit's DMCA lawsuit over data scraping. Reddit accuses Perplexity and third-party scrapers like SerpApi of bypassing protections to illegally scrape copyrighted Reddit content from Google search results for AI training data.
A Manhattan federal judge has rejected most of Perplexity AI's bid to dismiss a Reddit lawsuit accusing the AI startup of violating copyright law through unauthorized scraping of Reddit's content
3
. US District Judge Paul A. Engelmayer ruled on Friday that Reddit has plausibly demonstrated that Perplexity AI and data scraping services conspired to illegally scrape copyrighted Reddit content from Google search results1
. The decision allows Reddit to continue pressing claims that Perplexity and three data scrapers—SerpApi, Oxylabs, and AWMProxy—unlawfully circumvented protective measures to access content for AI training data3
.
Source: The Verge
The ruling comes less than two weeks after another court dismissed a similar action raised by Google, finding insufficient proof that rights holders like Reddit had authorized the search engine to prevent scraping
1
. Judge Engelmayer distinguished Reddit's case by noting the company went beyond bare allegations, arguing its licensing agreement with Google directly prohibits certain uses of Reddit data now being accessed through anti-circumvention measures employed by third-party scrapers1
. Reddit successfully argued that when it licenses content to partners like Google, those partners agree to delete posts flagged when users remove content—a promise that millions of monthly post deletions depend upon1
.Judge Engelmayer ruled that Reddit has standing to sue for Perplexity's alleged misuse of its users' content, despite the platform not owning the copyrighted material itself
3
. This addresses a key question raised by DMCA experts about whether Reddit, as neither the copyright owner nor the creator of Google's protective technology, could bring such claims1
. The decision suggests courts may accept that platforms can enforce protections on behalf of their user communities when licensing agreements are in place.
Source: Ars Technica
Reddit has asked the court for multiple forms of relief: an injunction blocking SerpApi and Perplexity AI access to both Reddit and Google websites, another injunction stopping circumvention of Google SearchGuard, and a third preventing the defendants from using previously scraped data
1
. The platform also seeks unspecified monetary damages3
. Ben Lee, Reddit's chief legal officer, stated the ruling brings the company closer to holding bad actors accountable, emphasizing that Reddit supports responsible access to public content but opposes companies that bypass protections and profit off communities without permission2
.Perplexity AI responded defiantly, stating that Reddit's suit claims the right to control access to public web pages it doesn't own, using a security tool it didn't build, on behalf of users it hasn't asked
3
. The company vowed to defend the open internet and predicted victory. SerpApi similarly argued that Google and Reddit are trying to use the DMCA to wall off the open Internet by retroactively claiming control over content they didn't author and don't own1
. SerpApi's attorney stated the company accesses public search results, not Reddit's platform, and that public information doesn't become protected simply because a platform wants to charge for it3
.Related Stories
If Reddit wins this AI copyright lawsuit, the platform could be better positioned to force all AI scrapers to enter licensing agreements
1
. Reddit, which features thousands of interest-based subreddit communities, has already licensed its content to Google, OpenAI, and others for AI training3
. The lawsuit claims Reddit is the most commonly cited source for AI-generated answers to user questions, making control over this data particularly valuable3
. This case joins many filed by content owners including authors, music labels, and news outlets against tech companies over alleged misuse of copyrighted material for AI systems3
.Judge Engelmayer noted that SerpApi and Perplexity AI may still prove through discovery that Reddit never authorized Google to protect its content in Google search results
1
. SerpApi may also strengthen its defense by proving that publicly accessible content in search results isn't protected by the Copyright Act1
. The fight appears far from over, with fundamental questions about who controls public web content and how anti-circumvention measures apply to AI training still unresolved. Reddit has also filed a separate data-scraping lawsuit against Anthropic that remains ongoing in California state court3
.Summarized by
Navi
22 Oct 2025•Policy and Regulation

05 Jun 2025•Policy and Regulation

01 Aug 2024
