Anthropic
Anthropic is the operator of Claude, an assistant that reads the live web and cites what it reads. Web search reached every Claude plan worldwide on 27 May 2025[1], and answers built on retrieved pages carry direct citations to those pages[2]. Anthropic runs three separately named crawlers with different purposes and publishes their source addresses, which makes its traffic unusually easy for a site owner to identify. Its training data is the subject of two suits, one from Reddit and one from a class of authors settled for $1.5 billion[3].
Claude and the live web
Claude answers from a retrieved page rather than from memory alone once web search is on, and that search is available on every plan in every market from 27 May 2025[1]. Where an answer uses retrieved material, Claude prints direct citations so the reader can check the source[2]. For a publisher this puts Claude in the same category as ChatGPT search and Perplexity: a surface where a page can be quoted and linked without a click on a results list.
Crawlers
Three agents do three jobs. ClaudeBot collects web content that may contribute to training Anthropic's models, so blocking it in robots.txt is a statement about future training sets[4]. Claude-User fetches a page only when a person's question requires it, so blocking it removes the site from live answers[5]. Claude-SearchBot builds and improves the search index behind those answers, analysing pages for relevance and accuracy[6]. Anthropic publishes a machine-readable list of the source addresses these crawlers use at claude.com/crawling/bots.json[7]. A request claiming to be ClaudeBot from an address outside that list is not Anthropic.
Litigation over training data
Reddit sued Anthropic in June 2025, alleging that its models were trained on Reddit user data taken in breach of the site's user agreement, pleaded as unlawful and unfair business acts[8]. In a separate class action brought by authors, Bartz v. Anthropic, Anthropic agreed in September 2025 to pay $1.5 billion over the use of pirated book copies for training, described at the time as the largest copyright resolution in US history[3]. The settlement pays class members $3,000 a book plus interest[9], and Judge Araceli Martinez-Olguin granted it final approval in July 2026, with some authors and publishers opting out to sue separately[10].
Why it matters for AI visibility
Blocking decisions here are not one decision. ClaudeBot governs training sets[4]; Claude-User is the fetch that happens because someone asked a question, so disallowing it removes the site from live answers[5]. Because answers built on retrieved pages carry direct citations[2], a page Claude can fetch is a page a reader can be sent to. The published address list[7] makes verification mechanical: a request naming ClaudeBot from an address outside it should be treated as forged, not as Anthropic traffic to allow.
References (10)
- Claude can now search the web Archive
- Claude can now search the web Archive
- Anthropic - Wikipedia Archive
- Does Anthropic crawl data from the web, and how can site owners block the crawler? Archive
- Does Anthropic crawl data from the web, and how can site owners block the crawler? Archive
- Does Anthropic crawl data from the web, and how can site owners block the crawler? Archive
- Anthropic support: does Anthropic crawl data from the web, and how can site owners block the crawler Archive
- Anthropic - Wikipedia Archive
- Anthropic - Wikipedia Archive
- Anthropic - Wikipedia Archive
Last updated 2026-09-05. Written and maintained by Baseline Labs.