Skip to content

Training scrapers

Concept in AI search and SEO
Volatile, last checked 2026-09-04 This page carries measured figures that move quickly.Reviewed 2026-09-04

Training scrapers collect text in bulk for model training rather than for an index a live answer retrieves from. OpenAI documents GPTBot as the crawler that makes its foundation models more useful and safe[1], and says a Disallow for it signals that a site's content should not be used in training[2]. Anthropic's ClaudeBot collects content that could contribute to training[3], and Common Crawl's CCBot feeds a public archive gathered since 2008 that most open corpora are cut from[4][5]. Blocking this class is a separate decision from blocking a search crawler.

Opting out

The larger operators publish a token that controls training alone. Google-Extended decides whether crawled content may train future Gemini models[6], and Apple is explicit that Applebot-Extended does not itself crawl anything: it is a preference read at training time, and a page that disallows it still appears in search results[7]. That separation is the point. A site can refuse the training corpus and keep the index entry an answer engine cites from. See robots.txt.

Apple's opt-out control, documented in its own support pages: Applebot-Extended governs training use only, and refusing it does not remove a site from Siri or Spotlight results.
Apple's opt-out control, documented in its own support pages: Applebot-Extended governs training use only, and refusing it does not remove a site from Siri or Spotlight results.Captured 2026-09-08 from support.apple.com. Signed out, no cookies.

Common Crawl

Common Crawl is the largest shared input. CCBot is a Nutch-based crawler built on Apache Hadoop[4], feeding a corpus of petabytes gathered since 2008[5]. Common Crawl publishes a robots.txt block that stops the crawler from crawling a site[8], which acts forward, on crawls not yet taken. The withdrawal figures below are therefore measured against a corpus, C4, rather than against a live crawler[9].

What opting out costs

Refusals have moved fast and unevenly. An audit of 14,000 web domains found robots.txt restrictions added in a single year fully restricted more than 5% of all C4 tokens and more than 28% of its most actively maintained sources[9], with terms of service covering 45% of the corpus[10]. Among news sites, 60.0% of reputable outlets disallow at least one AI crawler against 9.1% of misinformation sites[11], a gap that skews what a model reads. Blocking is not free: publishers who block generative AI bots see reduced site traffic against those who do not[12]. Compliance is partial in any case, since bots obey less as a directive gets stricter[13].

The 2025 default flip

Blocking became the default rather than an act. On 1 July 2025 Cloudflare changed the default on its network to block AI crawlers unless they pay creators for their content[14], and blocking among reputable news sites had already climbed from 23% in September 2023 to nearly 60% by May 2025[15]. The counterpart is paid access rather than free access, covered under content licensing deals and AI crawlers.

Bots in this class

The Baseline Labs bot registry (v1) tracks 73 of these. Each one has its own entry with its user agent, its robots.txt token and how to verify it.

The full bot encyclopedia lists every class side by side.

Seen by Baseline

Requests, 30 days
...
Last 7 days
...
Bots seen
...

Requests from bots in this class across sites tracked by Baseline, last 30 days. Aggregated over - sites, never reported per site.

Data

Share of the C4 training corpus placed off limits, 2024 auditAll C4 tokens, robots.txt5%Critical sources, robots.txt28%C4, terms of service45%
Share of the C4 training corpus placed off limits, 2024 audit. Figures in %.Source: arxiv.org
Show the numbers (3)
Point%
All C4 tokens, robots.txt5
Critical sources, robots.txt28
C4, terms of service45
References (15)
  1. Overview of OpenAI Crawlers - OpenAI developer docs
    Documentation Retrieved 2026-09-04
  2. Overview of OpenAI Crawlers - OpenAI developer docs
    Documentation Retrieved 2026-09-04
  3. Does Anthropic crawl data from the web, and how can site owners block the crawler? - Claude support
    Documentation Retrieved 2026-09-04
  4. Frequently asked questions - Common Crawl
    Documentation Retrieved 2026-09-04
  5. Overview - Common Crawl
    Documentation Retrieved 2026-09-04
  6. Google common crawlers - Google Search Central
    Documentation Retrieved 2026-09-04
  7. About Applebot - Apple Support
    Documentation Retrieved 2026-09-04
  8. Frequently asked questions - Common Crawl
    Documentation Retrieved 2026-09-04
  9. Consent in Crisis: the rapid decline of the AI data commons - arXiv 2407.14933
    Academic Published 2024-07-20 Retrieved 2026-09-04
  10. Consent in Crisis: the rapid decline of the AI data commons - arXiv 2407.14933
    Academic Published 2024-07-20 Retrieved 2026-09-04
  11. Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web - arXiv 2510.10315
    Academic Published 2025-10-11 Retrieved 2026-09-04
  12. Strategic Response of News Publishers to Generative AI - arXiv 2512.24968
    Academic Published 2025-12-31 Retrieved 2026-09-04 Single-source
  13. Scrapers selectively respect robots.txt directives - arXiv 2505.21733 (ACM IMC 2025)
    Academic Published 2025-05-27 Retrieved 2026-09-04
  14. Content Independence Day: no AI crawl without compensation - Cloudflare Blog
    Vendor documentation Published 2025-07-01 Retrieved 2026-09-04
  15. Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web - arXiv 2510.10315
    Academic Published 2025-10-11 Retrieved 2026-09-04 Volatile, last checked 2026-09-04

Last updated 2026-09-04. Written and maintained by Baseline Labs.

George
Online
0%