Reddit's DMCA lawsuit against Perplexity survives dismissal, testing whether platform data licensing deals create anti-scraping moats
A federal judge denied SerpApi's motion to dismiss Reddit's lawsuit accusing the web scraper of conspiring with Perplexity AI to harvest copyrighted Reddit content from Google search results. The ruling advances a novel legal theory: that Reddit's licensing deal with Google implicitly authorized anti-circumvention protections, giving Reddit standing under the DMCA even though Google's parallel case was dismissed. If the theory holds, every platform with a big-tech data deal could gain DMCA-level control over who accesses its content, potentially making paid licensing mandatory for AI training data.

When Reddit signed a content licensing agreement with Google, it got something neither company negotiated for: a potential federal legal weapon against every AI scraper that pulls Reddit content without paying.
That weapon is the Digital Millennium Copyright Act. A Manhattan federal judge allowed Reddit to proceed with core claims under Section 1201 of the DMCA, which prohibits circumvention of technological measures designed to control access to copyrighted works 1.
The decision advances a legal theory with implications far beyond this case. If a private licensing agreement between a platform and a tech partner can implicitly create anti-circumvention protections under federal law, then every platform with a big-tech data deal gains DMCA-level control over who accesses its content for AI training. Voluntary licensing stops being a business choice and becomes a de facto access requirement, not because Congress passed a law, but because a judge let a private contract reshape the rules.
The case is Reddit, Inc. v. SerpApi, LLC; Oxylabs UAB; AWMProxy; and Perplexity AI, Inc., filed October 22, 2025, in the Southern District of New York 2. Reddit alleges that SerpApi and two other intermediaries, proxy-network operators Oxylabs and AWMProxy, scraped Reddit content at what the complaint calls industrial scale by pulling it from Google search results, then sold the data to Perplexity AI
2. Perplexity AI surfaces Reddit content in its answer engine, providing summaries and citations from threads Reddit never licensed to the company
2.
What makes the ruling surprising is that Google filed a nearly identical lawsuit and lost. A separate court dismissed Google's parallel action less than two weeks before Engelmayer's ruling, finding that Google had not proven that rights holders like Reddit had authorized the search engine to prevent the scraping of protected content 3.
Reddit survived where Google failed because it went "beyond the bare allegation" that Google used to broadly claim it generally has licenses to display copyrighted content 3. Reddit argued that its licensing agreement with Google directly prohibits certain uses of Reddit data. When Reddit licenses content to partners like Google, those partners agree to delete posts that Reddit flags when users remove content. According to Reddit, millions of posts are deleted monthly, and unsanctioned scraping operations make it impossible for the platform to honor those removals
3.
There is a second wrinkle. Google's anti-bot technology, which SerpApi is accused of circumventing, was invented more than a year after Google and Reddit struck their licensing deal. Engelmayer ruled that it would be impractical to expect partners to update licensing deals every time a company rolls out new security methods 3. Under that logic, a licensing deal signed today could retroactively cover security measures that do not exist yet. The anti-scraping protections are not negotiated between the parties. They are inferred by a court from the existence of a commercial relationship.
The downstream angle compounds the risk for AI companies. Perplexity AI never directly scraped Reddit. It bought data from intermediaries who did. Perplexity AI argued that the DMCA's anti-circumvention provision does not carry secondary liability, and that purchasing data someone else collected is not itself a violation 2. Reddit countered that knowingly buying data derived from circumvention and monetizing it is not meaningfully different from doing the circumventing yourself
2. If Reddit's theory survives, a vendor's assurance that training data was sourced legally stops being sufficient. Buyers would need chain-of-title documentation showing where data originated and whether any technical protection measure was bypassed upstream
2.
Not everyone is convinced the theory will hold. Meredith Rose of Public Knowledge noted that Reddit is neither the copyright owner, nor an exclusive licensee, nor the party deploying the technological protection measure at issue 3.
Jeff Homrig, a lawyer for SerpApi, was more direct. "SerpApi accesses public search results, not Reddit's platform," he told Ars, "and public information does not become protected because a platform wants to charge for it" 3.
The case now moves to discovery, where SerpApi and Perplexity AI may prove that Reddit never authorized Google to protect its content in search results 3. Engelmayer also noted in a footnote that SerpApi could strengthen its defense by proving that publicly accessible content in Google search results is not protected by the Copyright Act
3. Reddit did lose its unjust enrichment and unfair competition claims, which the court found were preempted by the Copyright Act
3.
For AI builders, the practical takeaway is this: training data pulled from the open web is not necessarily safe just because it passed through a search engine. If courts treat a licensing deal as an implicit access control, the legal status of publicly available content may depend on private agreements you have never seen and cannot inspect.
For data-rich platforms, the incentive is clear. A partnership with a major tech company now serves a second function beyond revenue: a potential DMCA enforcement tool against competitors who scrape the same content. The licensing deal signed to monetize data could also be the barrier that keeps rivals away from it.
Congress has not passed a law requiring AI companies to license training data. But if this legal theory survives discovery and trial, a judge may have effectively built one through private contract law instead. The question is whether the open web stays open, or whether every data deal quietly builds a toll booth around it.
References
Cite this story
ProvenBrief (2026). "Reddit's DMCA lawsuit against Perplexity survives dismissal, testing whether platform data licensing deals create anti-scraping moats." ProvenBrief. https://provenbrief.com/story/reddit-s-dmca-lawsuit-against-perplexity-survives-dismissal-testing-whether-plat
Free to quote and link with attribution. Republishing in full or AI-training use requires a license.
Get the next brief in your inbox
One weekly email. Every claim verified against primary sources before we hit send.
This story
WordsProduced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our editorial standards.