07.05.2025

Scraping: Reddit vs Anthropic

Reddit sued Anthropic for scraping its platform to train Claude without permission or payment — a first-of-its-kind test of content-licensing law.

Scraping: Reddit vs AnthropicScraping: Reddit vs Anthropic

Reddit has filed a lawsuit against Anthropic — the developer of the AI assistant Claude — over the unauthorized use of user-generated content to train its neural networks.

This is one of the first legal cases where a digital platform owner formally challenges the commercial use of user-posted content without licensing or compensation. The lawsuit raises fundamental questions about who owns online content and under what conditions it may be used for machine learning.

Reddit's lawsuit against Anthropic is primarily about the unlawful use of public user content, not unauthorized access to personal data, which is a secondary question here. In this case, "scraping" refers specifically to publicly posted materials (texts, comments, and so on) created by Reddit users.

What is Scraping?

Scraping is the automated collection of data from websites using specialized software. It typically involves large volumes of automated queries, which place a significant load on servers.

In this case, Anthropic used scraping to gather a large volume of content from Reddit in order to train its language model, Claude. The core issue is that the data was used to build a commercially successful product that has already attracted interest from investors and tech companies. This was not research.

Legal and Regulatory Risks

Reddit brings five main arguments against Anthropic:

  • Breach of the user agreement — Reddit's terms of service explicitly prohibit automated collection and commercial use of content without separate permission. Anthropic used the platform's materials to train its AI system and then sold access to it to clients.
  • Unjust enrichment — Anthropic gained significant benefit by using content that Reddit monetizes through licensing agreements with other AI companies, without paying anything for it.
  • Trespass to chattels — the scraping bots created excessive load on Reddit's servers and could degrade service quality for regular users.
  • Interference with contractual relationships — Anthropic was aware of Reddit's confidentiality obligations to its users, yet knowingly violated them and continued collecting data after receiving official notices.
  • Unfair competition — Anthropic claimed it had stopped scraping at Reddit's demand, but continued unauthorized collection and harmed the platform's reputation.

Ethical and Privacy Issues

Scraping raises serious questions about a user's control over their own content: deleted posts may remain in training datasets. Licensed partnerships with OpenAI and Google use APIs where content removal is synchronized. With illegal scraping, users have no such control. This raises questions about digital content rights and about the possible processing of personal information, if any made it into the collected set. AI developers are increasingly expected to behave ethically: to respect the terms under which content is used, to honor users' choices, and to be transparent about data collection.

Historical Precedent: HiQ vs LinkedIn

The case of HiQ Labs vs LinkedIn (2017–2022) set important legal reference points for scraping. HiQ analyzed labor-market trends and routinely collected public LinkedIn profiles for its analytics.

The positions that emerged from that case affirmed that scraping is permissible where there is no explicit prohibition and no substantial harm to the platform. They also established that the CFAA (Computer Fraud and Abuse Act) cannot be used as a catch-all tool against the collection of publicly accessible data.

The Uniqueness of This Case

This lawsuit is a fundamentally new approach to regulating scraping in the training of AI models. Past lawsuits were built on the CFAA. Reddit instead builds its case on the violation of licensing terms and the unlawful commercial use of content. The company emphasizes that Anthropic knowingly ignored technical restrictions (such as robots.txt) and continued collecting data even after official warnings.

The absence of any CFAA claim shows an evolution in legal tactics. Instead of arguing over unauthorized access to systems (as in HiQ vs LinkedIn), Reddit protects economic interests and demands compliance with its terms of service and confidentiality. This approach could set a new precedent for the relationship between content-owning platforms and AI developers.

What Changed by July 2026

The case has moved from a freshly announced complaint to live, precedent-setting litigation, and the court rejected Anthropic's first attempt to escape it. Reddit filed the suit in the Superior Court of California, County of San Francisco, on June 4, 2025, bringing the same five claims described above: breach of contract, unjust enrichment, trespass to chattels, interference with contractual relationships, and unfair competition. According to the complaint, Anthropic scraped Reddit content, including deleted posts, to train Claude and kept collecting after saying it had stopped.

Anthropic removed the case to federal court and argued that the federal Copyright Act preempts Reddit's state-law claims. On March 30, 2026, Judge Trina Thompson of the U.S. District Court for the Northern District of California disagreed and remanded the case to state court. The court held that Reddit's contract and tort claims impose obligations "qualitatively different" from copyright: the user agreement's limits on methods of access and its protection of Reddit's technical infrastructure fall outside what copyright governs. That ruling matters in its own right — it shows that platforms can pursue AI scrapers on contract and unfair-competition grounds, without being funneled into copyright law.

The practical takeaway is that an anti-scraping clause in the terms of service is a standalone lever that works independently of copyright. For a platform, this means protection is not limited to copyright: the terms of use should explicitly prohibit automated collection and model training, backed by technical measures (robots.txt, rate-limiting) and notices to the infringer — that creates grounds for breach-of-contract and unfair-competition claims. For those collecting data, it is important to see the other side: a fair use argument answers only the copyright question and does not dispose of a contract claim. Breach of the terms of service is a separate risk zone, and fair use does not remove it.
Gennady Kurdiumov, Co-Founder, Futura Digital

As of July 2026, the case is proceeding in San Francisco Superior Court: there is no settlement and no trial on the merits yet. The close of fact discovery is set for January 18, 2027, and the next case management conference is scheduled for December 17, 2026.

Conclusion

The Reddit vs Anthropic lawsuit adds to the emerging legal framework that protects digital content in the age of AI. There are already precedents where courts sided with data owners against the unauthorized use of their materials to train models. Licensing agreements are increasingly becoming a mandatory part of working with training data — as seen in Reddit's deals with OpenAI and Google.

The outcome of Reddit vs Anthropic could become a pivotal precedent in AI regulation and help establish a new balance between innovation in machine learning and the rights of content owners.

Behind the dispute over scraping lies a larger question — the future of AI governance and digital rights in a world where data has become the fuel of technological progress.

This material was updated in July 2026 by the Futura team.

Discuss
the Task

Speak to our team

Speak to our team. Tell us about your task –

we’ll help you with it in any jurisdiction.

Tell us about your task –
we’ll help you with it in any jurisdiction.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

We use cookies to improve your experience.