OpenAI
The crawler OpenAI uses to gather training data for its GPT models. One of the highest-volume AI crawlers on the web. Blocking it keeps your content out of future model training runs, but has no effect on whether ChatGPT can cite you in search answers - that is OAI-SearchBot.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.2; +https://openai.com/gptbot
- User agent
gptbot- Robots token
GPTBot- Feeds
- GPT model training corpora
- Robots.txt
- Honours robots.txt
Anthropic
The crawler Anthropic uses to collect training data for Claude models. Anthropic documents the token and honours robots.txt, but publishes no IP ranges, so a claimed ClaudeBot hit cannot be verified against an official list - anyone can wear the name.
Mozilla/5.0; compatible; ClaudeBot/1.0; +claudebot@anthropic.com
- User agent
claudebot- Robots token
ClaudeBot- Feeds
- Claude model training corpora
- Robots.txt
- Honours robots.txt
- Verify
- No IP list published - identity cannot be verified
Anthropic
An older Anthropic token that predates the current ClaudeBot naming. Sites still declare it in robots.txt as a belt-and-braces training opt-out, and it occasionally appears in logs from older tooling. Treat it as a legacy alias of the ClaudeBot family.
- User agent
anthropic-ai- Robots token
anthropic-ai- Feeds
- Historical Anthropic data collection
- Robots.txt
- Unclear
Anthropic
Another legacy Anthropic token seen in community registries and older robots.txt files. Not documented on the current Anthropic crawler page; kept here because it is still widely declared and occasionally observed.
- User agent
claude-web- Robots token
Claude-Web- Feeds
- Historical Anthropic data collection
- Robots.txt
- Unclear
Google
Not a crawler - a control token. Google crawls with ordinary Googlebot and reads Google-Extended in robots.txt to decide whether that content may train Gemini models or ground Gemini answers. Disallowing it does not change your Google Search crawling or ranking.
- User agent
- Robots-only token - never appears in logs
- Robots token
Google-Extended- Feeds
- Gemini training and grounding permissions
- Robots.txt
- Honours robots.txt
- Verify
- Robots-only token - never appears in logs
Google
Crawls sites at the request of Google Cloud customers building Vertex AI Search applications. It only fetches domains a customer has configured, so a hit means someone is building an AI search product over your content.
- User agent
google-cloudvertexbot- Robots token
Google-CloudVertexBot- Feeds
- Vertex AI Search customer indexes
- Robots.txt
- Honours robots.txt
Apple
The Apple control token, mirroring Google-Extended: it never crawls, it only tells Apple whether pages Applebot already fetched may train Apple foundation models. Disallow it to opt out of training while keeping Siri and Spotlight visibility.
- User agent
- Robots-only token - never appears in logs
- Robots token
Applebot-Extended- Feeds
- Apple foundation model training permissions
- Robots.txt
- Honours robots.txt
- Verify
- Robots-only token - never appears in logs
Meta
The primary Meta AI crawler, collecting training data for Llama models and content for Meta AI products. Introduced in 2024 as the successor to FacebookBot for AI collection. Meta documents robots.txt support, with some independent reports of inconsistency.
meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
- User agent
meta-externalagent- Robots token
meta-externalagent- Feeds
- Llama training and Meta AI
- Robots.txt
- Honours robots.txt
Meta
The older Meta crawler, documented as training language models for products like speech recognition. Largely superseded by meta-externalagent for AI collection but still active and still worth a robots.txt line if you are opting out of Meta training.
- User agent
facebookbot- Robots token
FacebookBot- Feeds
- Meta language model training
- Robots.txt
- Honours robots.txt
Meta
Born as the link-preview fetcher for shares on Facebook and Messenger, and most of its traffic is still that. Meta has also acknowledged using it for other collection, which is why it appears in AI bot lists despite predating the AI era by a decade.
- User agent
facebookexternalhit- Robots token
facebookexternalhit- Feeds
- Link previews, plus unspecified Meta collection
- Robots.txt
- Ignores robots.txt Preview fetches fire on user shares and do not consult robots.txt
Amazon
Crawls sites that AWS Bedrock customers configure as data sources for their generative AI applications. Like Google-CloudVertexBot, a hit means a specific company is ingesting your content into their own AI product, not Amazon itself.
- User agent
bedrockbot- Robots token
bedrockbot- Feeds
- AWS Bedrock customer knowledge bases
- Robots.txt
- Honours robots.txt
ByteDance
The ByteDance training crawler, gathering data for its LLM efforts including Doubao. Notorious for two things: extremely aggressive crawl volume, and widespread independent reports that it ignores robots.txt entirely. The bot most often named in server-load complaints.
Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36; compatible; Bytespider; https://zhanzhang.toutiao.com/
- User agent
bytespider- Robots token
Bytespider- Feeds
- ByteDance LLM training
- Robots.txt
- Ignores robots.txt Claims compliance; widely reported not to honour it
ByteDance
A sibling ByteDance token seen alongside Bytespider. ByteDance publishes no documentation for it, and it is understood to gather TikTok-branded content for the same LLM training pipeline.
- User agent
tiktokspider- Robots token
TikTokSpider- Feeds
- ByteDance LLM training
- Robots.txt
- Unclear
DeepSeek
The DeepSeek training crawler. Community trackers report it does not reliably honour robots.txt, and DeepSeek publishes no crawler documentation or IP ranges.
- User agent
deepseekbot- Robots token
DeepSeekBot- Feeds
- DeepSeek model training
- Robots.txt
- Ignores robots.txt Reported non-compliance; no vendor documentation
xAI
A token attributed to xAI data collection for Grok. No crawler documentation exists from xAI at all - these tokens come from community directories and are the least verified entries in this encyclopedia. Grok is widely believed to rely on collection that does not announce itself.
- User agent
xai-bot xai-grok- Robots token
xAI-Bot- Feeds
- Grok model training, attributed
- Robots.txt
- Unclear No vendor documentation exists for any xAI token
Common Crawl
The Common Crawl Foundation crawler, building the open web archive that seeded the training sets of nearly every large language model - GPT-3 famously included. Blocking CCBot is the single broadest training opt-out on the web, because dozens of labs consume the corpus downstream.
CCBot/2.0 (https://commoncrawl.org/faq/)
- User agent
ccbot- Robots token
CCBot- Feeds
- The Common Crawl open corpus, used by many AI labs
- Robots.txt
- Honours robots.txt
Allen Institute
The Allen Institute for AI crawler, collecting data for open language models such as OLMo and the Dolma corpus. A rare case where the training data, the models and the crawler policy are all published openly.
- User agent
ai2bot- Robots token
AI2Bot- Feeds
- Open model training (OLMo, Dolma)
- Robots.txt
- Honours robots.txt
Diffbot
The structured-extraction crawler behind the Diffbot Knowledge Graph API. It parses pages into structured records and sells that data to companies, including AI developers. Compliance is documented as discretionary depending on the product.
Mozilla/5.0; compatible; Diffbot/0.1; +http://www.diffbot.com
- User agent
diffbot- Robots token
Diffbot- Feeds
- The Diffbot Knowledge Graph, resold as data
- Robots.txt
- Partial Documented as discretionary per product
Webz.io
The Webz.io crawler, harvesting forums, reviews and news into datasets sold for research and AI training. Appears in logs under the older omgili name.
- User agent
omgili- Robots token
omgilibot- Feeds
- Webz.io datasets, sold for AI training
- Robots.txt
- Honours robots.txt
ImageSift
The ImageSift crawler from Hive, indexing images for reverse-image search and image datasets. The image-side counterpart of the text scrapers on this page.
- User agent
imagesift- Robots token
ImagesiftBot- Feeds
- Image search and image datasets
- Robots.txt
- Honours robots.txt
Firecrawl
The fetcher of Firecrawl, a scrape-to-markdown API popular with LLM application builders. Traffic is driven by whatever its customers are building, so volume and targets vary wildly.
- User agent
firecrawlagent- Robots token
FirecrawlAgent- Feeds
- Firecrawl customer scraping jobs
- Robots.txt
- Honours robots.txt
Crawl4AI
The default signature of Crawl4AI, an open-source crawling library aimed at LLM data pipelines. Not one operator but thousands of independent users of the same tool.
- User agent
crawl4ai- Robots token
Crawl4AI- Feeds
- Assorted LLM data pipelines
- Robots.txt
- Unclear Behaviour depends entirely on each operator
Apify
A signature of the Apify scraping platform, whose hosted actors collect web data for customers including AI teams. Like Scrapy and Crawl4AI, one name covering many operators.
- User agent
apifybot- Robots token
ApifyBot- Feeds
- Apify customer scraping jobs
- Robots.txt
- Unclear Behaviour depends on each customer's actor
Zyte
The default user agent of Scrapy, the most widely used open-source scraping framework. Anything from academic research to industrial AI data collection ships under this name, so it is a tool signature rather than an operator identity.
- User agent
scrapy- Robots token
Scrapy- Feeds
- Assorted scraping projects
- Robots.txt
- Unclear Honours robots.txt by default; operators can disable it
Semrush
The Semrush crawler serving ContentShake, its AI writing tool. Included because its fetches feed AI content generation, though Semrush is an SEO platform rather than a lab.
- User agent
semrushbot-ocob- Robots token
SemrushBot-OCOB- Feeds
- ContentShake AI writing tool
- Robots.txt
- Honours robots.txt