Skip to content

Search crawlers

Concept in AI search and SEO
Volatile, last checked 2026-09-04 This page carries measured figures that move quickly.Reviewed 2026-09-04

Search crawlers fetch pages in bulk to build an index that an answer engine retrieves from at query time[1]. They are the class an AI answer usually cites: OpenAI runs OAI-SearchBot for ChatGPT's search features[2], Anthropic runs Claude-SearchBot[3], and Perplexity runs PerplexityBot to surface and link sites in its results[4]. What separates them from a training scraper is what happens to the copy: an index entry can be retrieved and linked back, a training corpus is absorbed into model weights, and the two are controlled by different robots.txt tokens[5].

Separate tokens for index and training

Google publishes different robots.txt tokens for the two jobs. GoogleOther is the generic crawler its product teams use to fetch publicly accessible content[6], while Google-Extended is a standalone token deciding whether crawled content may train future Gemini models[5]. A site can stay in the index and out of the training corpus. OpenAI documents the same split, running OAI-SearchBot for search and a separate token for model training[2]. See training scrapers and robots.txt.

Robots.txt handling

Google states that its common crawlers always respect robots.txt rules for automatic crawls[7], and the protocol assumes exactly that: a crawler reads the file and decides what to follow[1]. Obedience across the wider population is uneven, and it went unmeasured for most of the protocol's life: the 2025 paper that first tested it at scale, on anonymised institutional web logs, records that the evidence before it was anecdotal[8]. That study, of 130 self-declared bots over 40 days, found compliance falls as a directive gets stricter, and that AI search crawlers often do not request robots.txt at all[9], so a Disallow line is a request rather than a control.

What being indexed buys

Index membership is the precondition for citation. OpenAI states that a site opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though it can still appear as a navigational link[2], Anthropic describes Claude-SearchBot as analysing content to improve the relevance of search responses[3], and Perplexity asks webmasters to allow PerplexityBot so pages can be surfaced and linked[4]. Blocking the search token of an engine removes the page from the pool its answers are drawn from, which is a different decision from refusing training. See ChatGPT Search, answer engines and AI citation sources.

Volumes

Search crawling remains the largest share of automated fetching. Cloudflare recorded Googlebot rising from 30% to 50% of crawler requests between May 2024 and May 2025[10], with GPTBot going from 2.2% to 7.7% on a 305% rise in requests[11]. Across 2025 Googlebot originated 4.5% of HTML requests on the Cloudflare network against 4.2% for every other AI bot combined[12]. Retrieval is a minority of declared AI bot purpose: a vendor crawler panel for January to July 2026 put search or answer retrieval at 11.1% of AI bot requests against 46.5% for training[13]. See Cloudflare and AI crawlers.

Bots in this class

The Baseline Labs bot registry (v1) tracks 33 of these. Each one has its own entry with its user agent, its robots.txt token and how to verify it.

The full bot encyclopedia lists every class side by side.

Seen by Baseline

Requests, 30 days
...
Last 7 days
...
Bots seen
...

Requests from bots in this class across sites tracked by Baseline, last 30 days. Aggregated over - sites, never reported per site.

References (13)
  1. Web crawler - Wikipedia
    Journalism Retrieved 2026-09-04
  2. Overview of OpenAI Crawlers - OpenAI developer docs
    Documentation Retrieved 2026-09-04
  3. Does Anthropic crawl data from the web, and how can site owners block the crawler? - Claude support
    Documentation Retrieved 2026-09-04
  4. PerplexityBot - Perplexity developer docs
    Documentation Retrieved 2026-09-04
  5. Google common crawlers - Google Search Central
    Documentation Retrieved 2026-09-04
  6. Google common crawlers - Google Search Central
    Documentation Retrieved 2026-09-04
  7. Overview of Google crawlers and fetchers - Google Search Central
    Documentation Retrieved 2026-09-04
  8. Scrapers selectively respect robots.txt directives: evidence from a large-scale empirical study
    Academic Published 2025-05-27 Retrieved 2026-09-04
  9. Scrapers selectively respect robots.txt directives - arXiv 2505.21733 (ACM IMC 2025)
    Academic Published 2025-05-27 Retrieved 2026-09-04
  10. From Googlebot to GPTBot: who's crawling your site in 2025 - Cloudflare Blog
    Vendor documentation Published 2025-07-01 Retrieved 2026-09-04 Volatile, last checked 2026-09-04
  11. From Googlebot to GPTBot: who's crawling your site in 2025 - Cloudflare Blog
    Vendor documentation Published 2025-07-01 Retrieved 2026-09-04 Volatile, last checked 2026-09-04
  12. Cloudflare Radar 2025 Year in Review - Cloudflare Blog
    Vendor documentation Retrieved 2026-09-04 Volatile, last checked 2026-09-04
  13. Crawl-to-refer ratio: AI crawlers and LLM bots - Seomator crawler panel
    Vendor documentation Published 2026-07-21 Retrieved 2026-09-04 Single-source Volatile, last checked 2026-09-04

Last updated 2026-09-04. Written and maintained by Baseline Labs.

George
Online
0%