Skip to content

The AI bot encyclopedia

Registry page in AI search and SEO

Every AI crawler, assistant fetcher, training scraper and browsing agent we track: who operates it, what it feeds, whether it honours robots.txt, and how to verify it. The page is generated from the same registry that classifies live traffic in our AI traffic tracker, so it cannot drift from what the product actually recognises.

The 150 tokens below fall into four classes, and the class is the part that decides what to do about a hit. Each one has its own article: search crawlers, AI assistants, training scrapers and AI agents. Every bot also has its own entry at /wiki/bots/{id}, which is the address to link at.

A user agent is a claim, not a proof: anything can wear any name. Where an operator publishes IP ranges we link them; where none exist we say so.

Seen by Baseline

Assistants
...
Search crawlers
...
Training scrapers
...
Agents
...

Requests from bots in each class across sites tracked by Baseline, last 30 days. Aggregated over - sites, never reported per site.

Assistants - the demand signal

ChatGPT-User

Assistant
OpenAI

Fires when a ChatGPT user asks about a specific page and the assistant goes to read it live. Each hit represents a real person, mid-conversation, being shown your content - the clearest demand signal in this table.

Mozilla/5.0 AppleWebKit/537.36; compatible; ChatGPT-User/1.0; +https://openai.com/bot
User agent
chatgpt-user
Robots token
ChatGPT-User
Feeds
Live answers inside ChatGPT conversations
Robots.txt
Partial Acts on direct user requests; OpenAI says it is not used for automated crawling

Claude-User

Assistant
Anthropic

Fires when a Claude user asks about a page and Claude fetches it live during the conversation. Like ChatGPT-User, every hit is a real person being read your content by an assistant - a demand signal, not a crawl.

Mozilla/5.0; compatible; Claude-User/1.0; +Claude-User@anthropic.com
User agent
claude-user
Robots token
Claude-User
Feeds
Live answers inside Claude conversations
Robots.txt
Honours robots.txt

Google-NotebookLM

Assistant
Google

Fetches a page when a NotebookLM user adds it as a source. Strictly user-initiated - one URL, one user, one notebook - which is why Google files it under user-triggered fetchers rather than crawlers.

User agent
google-notebooklm
Robots token
Google-NotebookLM
Feeds
NotebookLM user notebooks
Robots.txt
Ignores robots.txt User-triggered fetchers ignore robots.txt by design

Gemini-Deep-Research

Assistant
Google

The fetcher behind Gemini Deep Research runs. When a user commissions a research report, this agent reads dozens of pages to compile it - user-initiated, but at research volume rather than single-page volume.

User agent
gemini-deep-research
Robots token
Gemini-Deep-Research
Feeds
Gemini Deep Research reports
Robots.txt
Ignores robots.txt User-triggered fetchers ignore robots.txt by design

GoogleAgent-URLContext

Assistant
Google

Fetches a URL a user has explicitly passed to Gemini as context for their prompt. The narrowest of the Google fetchers: it reads exactly the page the user pasted.

User agent
googleagent-urlcontext
Robots token
GoogleAgent-URLContext
Feeds
Gemini URL-context requests
Robots.txt
Ignores robots.txt User-triggered fetchers ignore robots.txt by design

Perplexity-User

Assistant
Perplexity

Fetches a page when a Perplexity user asks about it directly. Perplexity documents that these user-initiated fetches generally do not consult robots.txt, on the argument that the request comes from a human, not a crawler.

Mozilla/5.0 AppleWebKit/537.36; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user
User agent
perplexity-user
Robots token
Perplexity-User
Feeds
Live answers inside Perplexity
Robots.txt
Ignores robots.txt By design for user-initiated fetches, per Perplexity docs

Meta-ExternalFetcher

Assistant
Meta

The Meta user-initiated fetcher: retrieves a specific link when a Meta AI user asks about it. Meta documents that it may bypass robots.txt because the fetch is performed on behalf of a person.

User agent
meta-externalfetcher
Robots token
Meta-ExternalFetcher
Feeds
Live answers inside Meta AI
Robots.txt
Ignores robots.txt Documented as able to bypass robots.txt for user requests

MistralAI-User

Assistant
Mistral

Fetches pages to ground and cite answers in Le Chat, the Mistral assistant. It is user-initiated and documented, and it honours robots.txt - the European entry in the assistant-fetcher class.

Mozilla/5.0; compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots
User agent
mistralai-user
Robots token
MistralAI-User
Feeds
Live answers inside Le Chat
Robots.txt
Honours robots.txt

cohere-ai

Assistant
Cohere

The Cohere token, covering both user-initiated fetches and a training-data crawler variant. Documentation is thin compared to the other major labs, and compliance behaviour is not independently established.

User agent
cohere-ai cohere-training-data-crawler
Robots token
cohere-ai
Feeds
Cohere models and user requests
Robots.txt
Unclear

Grok-DeepSearch

Assistant
xAI

A community-reported token attributed to Grok DeepSearch research runs. Unverified by any xAI publication.

User agent
grok-deepsearch
Robots token
Grok-DeepSearch
Feeds
Grok DeepSearch, attributed
Robots.txt
Unclear No vendor documentation exists for any xAI token

DuckAssistBot

Assistant
DuckDuckGo

Fetches source pages for DuckAssist, the AI answer feature in DuckDuckGo search. Low volume and well behaved, consistent with the DuckDuckGo privacy posture.

DuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html)
User agent
duckassistbot
Robots token
DuckAssistBot
Feeds
DuckAssist AI answers
Robots.txt
Honours robots.txt

kagi-fetcher

Assistant
Kagi

Fetches pages for the Kagi Assistant when a subscriber asks about them. Small, paid-search-engine volume.

User agent
kagi-fetcher
Robots token
kagi-fetcher
Feeds
Kagi Assistant answers
Robots.txt
Unclear

Search crawlers - the citation pool

OAI-SearchBot

Search crawler
OpenAI

Builds the index behind ChatGPT search. A site that blocks OAI-SearchBot cannot appear as a cited source in ChatGPT answers with search, regardless of what it does about GPTBot. For visibility in AI answers, this is the OpenAI token that matters.

Mozilla/5.0 AppleWebKit/537.36; compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot
User agent
oai-searchbot
Robots token
OAI-SearchBot
Feeds
ChatGPT search index and citations
Robots.txt
Honours robots.txt

Claude-SearchBot

Search crawler
Anthropic

Builds the index Claude uses to ground search answers. Blocking it removes your site from the pool Claude can cite when answering with search - the Anthropic counterpart of OAI-SearchBot.

User agent
claude-searchbot
Robots token
Claude-SearchBot
Feeds
Claude search index and citations
Robots.txt
Honours robots.txt

GoogleOther

Search crawler
Google

The generic Google crawler for research, development and product uses outside Search. It shares infrastructure and IP ranges with Googlebot but its fetches do not feed Search ranking. Image and video variants carry their own tokens.

User agent
googleother
Robots token
GoogleOther
Feeds
Google research and non-Search products
Robots.txt
Honours robots.txt

Applebot

Search crawler
Apple

The Apple crawler behind Siri and Spotlight suggestions. Since 2024 its crawl also feeds Apple Intelligence features unless a site disallows Applebot-Extended - the crawl and the AI-use permission are split across the two tokens.

Mozilla/5.0 AppleWebKit/605.1.15; compatible; Applebot/0.1; +http://www.apple.com/go/applebot
User agent
applebot
Robots token
Applebot
Feeds
Siri, Spotlight, and Apple Intelligence
Robots.txt
Honours robots.txt

PerplexityBot

Search crawler
Perplexity

Builds the index behind Perplexity answers. Blocking it removes you from the source pool Perplexity cites. Perplexity publishes official IP ranges after earlier third-party reports of undeclared fetching, so claimed hits can now be verified.

Mozilla/5.0 AppleWebKit/537.36; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot
User agent
perplexitybot
Robots token
PerplexityBot
Feeds
Perplexity answer index and citations
Robots.txt
Honours robots.txt

meta-webindexer

Search crawler
Meta

The newer Meta crawler building a search index for Meta AI answers - the Meta counterpart of OAI-SearchBot. Blocking it affects whether Meta AI can cite you; blocking meta-externalagent affects whether Llama trains on you.

meta-webindexer/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
User agent
meta-webindexer
Robots token
meta-webindexer
Feeds
Meta AI search index
Robots.txt
Unclear

Amazonbot

Search crawler
Amazon

The Amazon crawler feeding Alexa answers and, per Amazon, the training of its AI services. A steady mid-volume presence in most server logs. Registry class is crawler because its index serves live Alexa responses, though it doubles as a training source.

Mozilla/5.0; compatible; Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot
User agent
amazonbot
Robots token
Amazonbot
Feeds
Alexa answers and Amazon AI training
Robots.txt
Honours robots.txt

GrokBot

Search crawler
xAI

A community-reported token attributed to a Grok search index crawler. As with all xAI entries, there is no official documentation to verify against.

User agent
grokbot
Robots token
GrokBot
Feeds
Grok answer grounding, attributed
Robots.txt
Unclear No vendor documentation exists for any xAI token
You.com

The You.com crawler, indexing for its AI search engine and assistant products.

User agent
youbot
Robots token
YouBot
Feeds
You.com search and assistant answers
Robots.txt
Honours robots.txt

Timpibot

Search crawler
Timpi

The crawler of Timpi, a decentralised search index project whose data is also positioned for LLM training.

User agent
timpibot
Robots token
Timpibot
Feeds
The Timpi decentralised index
Robots.txt
Unclear

Bravebot

Search crawler
Brave

The Brave Search index crawler. Notable because Brave sells its independent index through the Brave Search API, which several AI companies buy for grounding - so a Bravebot crawl can surface in other products' AI answers.

User agent
bravebot
Robots token
Bravebot
Feeds
Brave Search index, resold to AI products via API
Robots.txt
Honours robots.txt

PetalBot

Search crawler
Huawei

The Huawei crawler behind Petal Search, now also feeding Huawei AI features. High volume across European sites owing to the Petal Search footprint on Huawei devices.

User agent
petalbot
Robots token
PetalBot
Feeds
Petal Search and Huawei AI
Robots.txt
Honours robots.txt
Exa

The Exa crawler, building a neural search index designed to be queried by AI applications rather than people. Exa sells retrieval to agent and RAG builders, so its index is infrastructure for other AI products.

User agent
exabot
Robots token
ExaBot
Feeds
The Exa neural search API
Robots.txt
Unclear

TavilyBot

Search crawler
Tavily

Crawls for Tavily, a search API built specifically for AI agents and RAG pipelines. A hit suggests your content is being served into other companies' AI applications at answer time.

User agent
tavilybot
Robots token
TavilyBot
Feeds
The Tavily agent search API
Robots.txt
Unclear

Training scrapers - the corpus builders

OpenAI

The crawler OpenAI uses to gather training data for its GPT models. One of the highest-volume AI crawlers on the web. Blocking it keeps your content out of future model training runs, but has no effect on whether ChatGPT can cite you in search answers - that is OAI-SearchBot.

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.2; +https://openai.com/gptbot
User agent
gptbot
Robots token
GPTBot
Feeds
GPT model training corpora
Robots.txt
Honours robots.txt

ClaudeBot

Training scraper
Anthropic

The crawler Anthropic uses to collect training data for Claude models. Anthropic documents the token and honours robots.txt, but publishes no IP ranges, so a claimed ClaudeBot hit cannot be verified against an official list - anyone can wear the name.

Mozilla/5.0; compatible; ClaudeBot/1.0; +claudebot@anthropic.com
User agent
claudebot
Robots token
ClaudeBot
Feeds
Claude model training corpora
Robots.txt
Honours robots.txt
Verify
No IP list published - identity cannot be verified

anthropic-ai

Training scraper
Anthropic

An older Anthropic token that predates the current ClaudeBot naming. Sites still declare it in robots.txt as a belt-and-braces training opt-out, and it occasionally appears in logs from older tooling. Treat it as a legacy alias of the ClaudeBot family.

User agent
anthropic-ai
Robots token
anthropic-ai
Feeds
Historical Anthropic data collection
Robots.txt
Unclear

Claude-Web

Training scraper
Anthropic

Another legacy Anthropic token seen in community registries and older robots.txt files. Not documented on the current Anthropic crawler page; kept here because it is still widely declared and occasionally observed.

User agent
claude-web
Robots token
Claude-Web
Feeds
Historical Anthropic data collection
Robots.txt
Unclear

Google-Extended

Training scraper
Google

Not a crawler - a control token. Google crawls with ordinary Googlebot and reads Google-Extended in robots.txt to decide whether that content may train Gemini models or ground Gemini answers. Disallowing it does not change your Google Search crawling or ranking.

User agent
Robots-only token - never appears in logs
Robots token
Google-Extended
Feeds
Gemini training and grounding permissions
Robots.txt
Honours robots.txt
Verify
Robots-only token - never appears in logs

Google-CloudVertexBot

Training scraper
Google

Crawls sites at the request of Google Cloud customers building Vertex AI Search applications. It only fetches domains a customer has configured, so a hit means someone is building an AI search product over your content.

User agent
google-cloudvertexbot
Robots token
Google-CloudVertexBot
Feeds
Vertex AI Search customer indexes
Robots.txt
Honours robots.txt

Applebot-Extended

Training scraper
Apple

The Apple control token, mirroring Google-Extended: it never crawls, it only tells Apple whether pages Applebot already fetched may train Apple foundation models. Disallow it to opt out of training while keeping Siri and Spotlight visibility.

User agent
Robots-only token - never appears in logs
Robots token
Applebot-Extended
Feeds
Apple foundation model training permissions
Robots.txt
Honours robots.txt
Verify
Robots-only token - never appears in logs

meta-externalagent

Training scraper
Meta

The primary Meta AI crawler, collecting training data for Llama models and content for Meta AI products. Introduced in 2024 as the successor to FacebookBot for AI collection. Meta documents robots.txt support, with some independent reports of inconsistency.

meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
User agent
meta-externalagent
Robots token
meta-externalagent
Feeds
Llama training and Meta AI
Robots.txt
Honours robots.txt

FacebookBot

Training scraper
Meta

The older Meta crawler, documented as training language models for products like speech recognition. Largely superseded by meta-externalagent for AI collection but still active and still worth a robots.txt line if you are opting out of Meta training.

User agent
facebookbot
Robots token
FacebookBot
Feeds
Meta language model training
Robots.txt
Honours robots.txt

facebookexternalhit

Training scraper
Meta

Born as the link-preview fetcher for shares on Facebook and Messenger, and most of its traffic is still that. Meta has also acknowledged using it for other collection, which is why it appears in AI bot lists despite predating the AI era by a decade.

User agent
facebookexternalhit
Robots token
facebookexternalhit
Feeds
Link previews, plus unspecified Meta collection
Robots.txt
Ignores robots.txt Preview fetches fire on user shares and do not consult robots.txt

bedrockbot

Training scraper
Amazon

Crawls sites that AWS Bedrock customers configure as data sources for their generative AI applications. Like Google-CloudVertexBot, a hit means a specific company is ingesting your content into their own AI product, not Amazon itself.

User agent
bedrockbot
Robots token
bedrockbot
Feeds
AWS Bedrock customer knowledge bases
Robots.txt
Honours robots.txt

Bytespider

Training scraper
ByteDance

The ByteDance training crawler, gathering data for its LLM efforts including Doubao. Notorious for two things: extremely aggressive crawl volume, and widespread independent reports that it ignores robots.txt entirely. The bot most often named in server-load complaints.

Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36; compatible; Bytespider; https://zhanzhang.toutiao.com/
User agent
bytespider
Robots token
Bytespider
Feeds
ByteDance LLM training
Robots.txt
Ignores robots.txt Claims compliance; widely reported not to honour it

TikTokSpider

Training scraper
ByteDance

A sibling ByteDance token seen alongside Bytespider. ByteDance publishes no documentation for it, and it is understood to gather TikTok-branded content for the same LLM training pipeline.

User agent
tiktokspider
Robots token
TikTokSpider
Feeds
ByteDance LLM training
Robots.txt
Unclear

DeepSeekBot

Training scraper
DeepSeek

The DeepSeek training crawler. Community trackers report it does not reliably honour robots.txt, and DeepSeek publishes no crawler documentation or IP ranges.

User agent
deepseekbot
Robots token
DeepSeekBot
Feeds
DeepSeek model training
Robots.txt
Ignores robots.txt Reported non-compliance; no vendor documentation
xAI

A token attributed to xAI data collection for Grok. No crawler documentation exists from xAI at all - these tokens come from community directories and are the least verified entries in this encyclopedia. Grok is widely believed to rely on collection that does not announce itself.

User agent
xai-bot xai-grok
Robots token
xAI-Bot
Feeds
Grok model training, attributed
Robots.txt
Unclear No vendor documentation exists for any xAI token
Common Crawl

The Common Crawl Foundation crawler, building the open web archive that seeded the training sets of nearly every large language model - GPT-3 famously included. Blocking CCBot is the single broadest training opt-out on the web, because dozens of labs consume the corpus downstream.

CCBot/2.0 (https://commoncrawl.org/faq/)
User agent
ccbot
Robots token
CCBot
Feeds
The Common Crawl open corpus, used by many AI labs
Robots.txt
Honours robots.txt
Allen Institute

The Allen Institute for AI crawler, collecting data for open language models such as OLMo and the Dolma corpus. A rare case where the training data, the models and the crawler policy are all published openly.

User agent
ai2bot
Robots token
AI2Bot
Feeds
Open model training (OLMo, Dolma)
Robots.txt
Honours robots.txt
Diffbot

The structured-extraction crawler behind the Diffbot Knowledge Graph API. It parses pages into structured records and sells that data to companies, including AI developers. Compliance is documented as discretionary depending on the product.

Mozilla/5.0; compatible; Diffbot/0.1; +http://www.diffbot.com
User agent
diffbot
Robots token
Diffbot
Feeds
The Diffbot Knowledge Graph, resold as data
Robots.txt
Partial Documented as discretionary per product

omgilibot

Training scraper
Webz.io

The Webz.io crawler, harvesting forums, reviews and news into datasets sold for research and AI training. Appears in logs under the older omgili name.

User agent
omgili
Robots token
omgilibot
Feeds
Webz.io datasets, sold for AI training
Robots.txt
Honours robots.txt

ImagesiftBot

Training scraper
ImageSift

The ImageSift crawler from Hive, indexing images for reverse-image search and image datasets. The image-side counterpart of the text scrapers on this page.

User agent
imagesift
Robots token
ImagesiftBot
Feeds
Image search and image datasets
Robots.txt
Honours robots.txt

FirecrawlAgent

Training scraper
Firecrawl

The fetcher of Firecrawl, a scrape-to-markdown API popular with LLM application builders. Traffic is driven by whatever its customers are building, so volume and targets vary wildly.

User agent
firecrawlagent
Robots token
FirecrawlAgent
Feeds
Firecrawl customer scraping jobs
Robots.txt
Honours robots.txt

Crawl4AI

Training scraper
Crawl4AI

The default signature of Crawl4AI, an open-source crawling library aimed at LLM data pipelines. Not one operator but thousands of independent users of the same tool.

User agent
crawl4ai
Robots token
Crawl4AI
Feeds
Assorted LLM data pipelines
Robots.txt
Unclear Behaviour depends entirely on each operator

ApifyBot

Training scraper
Apify

A signature of the Apify scraping platform, whose hosted actors collect web data for customers including AI teams. Like Scrapy and Crawl4AI, one name covering many operators.

User agent
apifybot
Robots token
ApifyBot
Feeds
Apify customer scraping jobs
Robots.txt
Unclear Behaviour depends on each customer's actor
Zyte

The default user agent of Scrapy, the most widely used open-source scraping framework. Anything from academic research to industrial AI data collection ships under this name, so it is a tool signature rather than an operator identity.

User agent
scrapy
Robots token
Scrapy
Feeds
Assorted scraping projects
Robots.txt
Unclear Honours robots.txt by default; operators can disable it

SemrushBot-OCOB

Training scraper
Semrush

The Semrush crawler serving ContentShake, its AI writing tool. Included because its fetches feed AI content generation, though Semrush is an SEO platform rather than a lab.

User agent
semrushbot-ocob
Robots token
SemrushBot-OCOB
Feeds
ContentShake AI writing tool
Robots.txt
Honours robots.txt

Agents - the ones that act

ChatGPT Agent

Agent
OpenAI

The robots.txt token for ChatGPT agent mode, where the assistant operates a browser to complete tasks. Published for declaration only - agent traffic presents as a browser session, so this token can be disallowed but never observed in a server log.

User agent
Robots-only token - never appears in logs
Robots token
ChatGPT Agent
Feeds
ChatGPT agent-mode browsing
Robots.txt
Honours robots.txt

Operator

Agent
OpenAI

The robots.txt token for Operator, the OpenAI browsing agent that clicks, fills forms and transacts on behalf of a user. Like ChatGPT Agent it is declarable in robots.txt but does not appear as a distinct user agent in logs.

User agent
Robots-only token - never appears in logs
Robots token
Operator
Feeds
Operator agent browsing
Robots.txt
Honours robots.txt

Claude-Code

Agent
Anthropic

The agentic coding tool from Anthropic, fetching documentation and pages while a developer works. Traffic is user-directed and bursty - a hit usually means a developer pointed their session at your docs.

User agent
claude-code
Robots token
Claude-Code
Feeds
Claude Code agent sessions
Robots.txt
Unclear

GoogleAgent-Mariner

Agent
Google

Project Mariner, the Google browsing agent that navigates sites and performs multi-step tasks on behalf of a user. Early-stage and low-volume, but a preview of what agentic traffic will look like at Google scale.

User agent
googleagent-mariner
Robots token
GoogleAgent-Mariner
Feeds
Project Mariner agent tasks
Robots.txt
Unclear

NovaAct

Agent
Amazon

Amazon Nova Act, the browsing agent from Amazon AGI Labs that navigates and acts on websites for users. Early-stage agentic traffic.

User agent
novaact
Robots token
NovaAct
Feeds
Nova Act agent sessions
Robots.txt
Unclear

Also tracked

The rest of the registry - 94 more tokens we classify in live traffic, from niche answer engines to open-source scraping signatures.

BotOperatorClassUser agentRobots token
GoogleOther-ImageGoogleSearch crawlergoogleother-imageGoogleOther-Image
GoogleOther-VideoGoogleSearch crawlergoogleother-videoGoogleOther-Video
Google-FirebaseGoogleTraining scrapergoogle-firebaseGoogle-Firebase
Google-Gemini-CLIGoogleAgentgoogle-gemini-cliGoogle-Gemini-CLI
Google-AgentGoogleAgentgoogle-agentGoogle-Agent
Amzn-SearchBotAmazonSearch crawleramzn-searchbotAmzn-SearchBot
amazon-kendraAmazonSearch crawleramazon-kendraamazon-kendra
Amzn-UserAmazonAssistantamzn-userAmzn-User
amazon-QBusinessAmazonAssistantamazon-qbusinessamazon-QBusiness
AmazonBuyForMeAmazonAgentamazonbuyformeAmazonBuyForMe
Ai2Bot-DolmaAllen InstituteTraining scraperai2bot-dolmaAi2Bot-Dolma
AI2Bot-DeepResearchEvalAllen InstituteAssistantai2bot-deepresearchevalAI2Bot-DeepResearchEval
PhindBotPhindAssistantphindbotPhindBot
AndibotAndiSearch crawlerandibotAndibot
LinerBotLinerAssistantlinerbotLinerBot
iAskBotiAskSearch crawleriaskbotiAskBot
iaskspideriAskSearch crawleriaskspideriaskspider
AzureAI-SearchBotMicrosoftSearch crawlerazureai-searchbotAzureAI-SearchBot
PanguBotHuaweiTraining scraperpangubotPanguBot
YandexAdditionalYandexTraining scraperyandexadditionalYandexAdditional
YandexAdditionalBotYandexTraining scraperyandexadditionalbotYandexAdditionalBot
YiyanBotBaiduAssistantyiyanbotYiyanBot
TongyiBotAlibabaAssistanttongyibotTongyiBot
Kimi-UserMoonshot AIAssistantkimi-userKimi-User
ChatGLM-SpiderZhipu AITraining scraperchatglm-spiderChatGLM-Spider
WRTNBotWrtnAssistantwrtnbotWRTNBot
SBIntuitionsBotSB IntuitionsTraining scrapersbintuitionsbotSBIntuitionsBot
CotoyogiROISTraining scrapercotoyogiCotoyogi
ICC-CrawlerNICTTraining scrapericc-crawlerICC-Crawler
omgiliWebz.ioTraining scraper-omgili
Webzio-ExtendedWebz.ioTraining scraperwebzio-extendedWebzio-Extended
imageSpiderUnknownTraining scraperimagespiderimageSpider
LinkupBotLinkupSearch crawlerlinkupbotLinkupBot
QueritBotQueritSearch crawlerqueritbotQueritBot
Querit-SearchBotQueritSearch crawlerquerit-searchbotQuerit-SearchBot
Cloudflare-AutoRAGCloudflareSearch crawlercloudflare-autoragCloudflare-AutoRAG
AddSearchBotAddSearchSearch crawleraddsearchbotAddSearchBot
AIWebIndexLyrenthSearch crawleraiwebindexAIWebIndex
AnomuraDireqtSearch crawleranomuraAnomura
Aranet-SearchBotUnknownSearch crawleraranet-searchbotAranet-SearchBot
Channel3BotUnknownSearch crawlerchannel3botChannel3Bot
ZanistaBotUnknownSearch crawlerzanistabotZanistaBot
atlassian-botAtlassianSearch crawleratlassian-botatlassian-bot
ApifyWebsiteContentCrawlerApifyTraining scraperapifywebsitecontentcrawlerApifyWebsiteContentCrawler
CrawlspaceCrawlspaceTraining scrapercrawlspaceCrawlspace
BrightbotBright DataTraining scraperbrightbotBrightbot
Mozilla-TabstackMozillaTraining scrapermozilla-tabstackMozilla-Tabstack
TerraCottaCeramic AITraining scraperterracottaTerraCotta
ShapBotParallelTraining scrapershapbotShapBot
Shap-UserParallelAssistantshap-userShap-User
PanscientPanscientTraining scraperpanscientPanscient
CragCrawlerCragSoftwareTraining scrapercragcrawlerCragCrawler
HenkBotUnknownTraining scraperhenkbotHenkBot
KunatoCrawlerKunatoTraining scraperkunatocrawlerKunatoCrawler
NagetBotNagetTraining scrapernagetbotNagetBot
WARDBotWEBSPARKTraining scraperwardbotWARDBot
MyCentralAIScraperBotUnknownTraining scrapermycentralaiscraperbotMyCentralAIScraperBot
newsaiUnknownTraining scrapernewsainewsai
AgentTimesThe Agent TimesTraining scraperagenttimesAgentTimes
img2datasetimg2datasetTraining scraperimg2datasetimg2dataset
LAIONDownloaderLAIONTraining scraperlaiondownloaderLAIONDownloader
laion-huggingface-processorLAIONTraining scraperlaion-huggingface-processorlaion-huggingface-processor
FriendlyCrawlerUnknownTraining scraperfriendlycrawlerFriendlyCrawler
Poseidon Research CrawlerPoseidon ResearchTraining scraperposeidon research crawlerPoseidon Research Crawler
ISSCyberRiskCrawlerISS-CorporateTraining scraperisscyberriskcrawlerISSCyberRiskCrawler
Factset_spyderbotFactSetTraining scraperfactset_spyderbotFactset_spyderbot
VelenPublicWebCrawlerVelenTraining scrapervelenpublicwebcrawlerVelenPublicWebCrawler
Kangaroo BotUnknownTraining scraperkangaroo botKangaroo Bot
Datenbank CrawlerDatenbankTraining scraperdatenbank crawlerDatenbank Crawler
netEstate Imprint CrawlernetEstateTraining scrapernetestate imprint crawlernetEstate Imprint Crawler
Sidetrade indexer botSidetradeTraining scrapersidetrade indexer botSidetrade indexer bot
SemrushBot-SWASemrushTraining scrapersemrushbot-swaSemrushBot-SWA
AwarioAwarioTraining scraperawarioAwario
Echobot BotEchoboxTraining scraperechobotEchobot Bot
EchoboxBotEchoboxTraining scraperechoboxbotEchoboxBot
YaKMeltwaterTraining scraper-YaK
aiHitBotaiHitTraining scraperaihitbotaiHitBot
QuillBotQuillBotTraining scraperquillbotQuillBot
Linguee BotLingueeTraining scraperlinguee botLinguee Bot
ThinkbotThinkbotTraining scraperthinkbotThinkbot
KlaviyoAIBotKlaviyoAssistantklaviyoaibotKlaviyoAIBot
Poggio-CitationsPoggioAssistantpoggio-citationsPoggio-Citations
QualifiedBotQualifiedAssistantqualifiedbotQualifiedBot
BuddyBotBuddyBotLearningAssistantbuddybotBuddyBot
GeistHaus-PageFetcherGeistHausAssistantgeisthaus-pagefetcherGeistHaus-PageFetcher
bigsur.aiBig Sur AIAssistantbigsur.aibigsur.ai
wpbotQuantumCloudAssistantwpbotwpbot
UseAIUnknownAssistant-UseAI
Manus-UserButterfly EffectAgentmanus-userManus-User
TwinAgentTwinAgenttwinagentTwinAgent
CursorCursorAgentcursorCursor
DevinCognitionAgentdevinDevin
TraeByteDanceAgenttraeTrae
opencodeopencodeAgentopencodeopencode
No bots match that filter.

Common questions

What is the difference between an AI crawler and an AI assistant fetch?
A crawler builds an index ahead of time - it visits on its own schedule, and its crawl decides whether an engine can cite you later. An assistant fetch happens live: a person asked an AI about a page, and the AI went to read it during the conversation. Assistant hits are a demand signal; crawler hits are supply-side plumbing.
Do AI bots respect robots.txt?
The documented crawlers from the major labs - GPTBot, ClaudeBot, OAI-SearchBot, Google-Extended, CCBot - honour it. User-initiated fetchers such as Perplexity-User and the Google user-triggered fetchers bypass it by design, on the argument that a human requested the page. And some training scrapers, most notoriously Bytespider, are widely reported to ignore it entirely. Each entry above carries its documented status.
Should I block AI training crawlers?
It is a legitimate publisher choice with a real trade-off. Blocking training scrapers keeps your content out of future model weights but changes nothing about today's AI answers. Blocking search crawlers such as OAI-SearchBot or PerplexityBot removes you from the pool those engines cite - which is usually the opposite of what a business wants. Decide per class, not per blanket rule.
How do I see which AI bots actually visit my site?
Server logs are the ground truth, and most analytics tools never show bots at all. Our AI Traffic Tracker classifies every hit against this registry via a lightweight reporting channel - see the install page - and the crawler log analyzer does the same for an uploaded access log.

Compiled from the Baseline Labs bot registry v1 (2026-08-03) - 150 bots tracked, 56 profiled above. The registry also powers the machine-readable agents endpoint our reporting channels poll. Spotted a bot we are missing? Tell us.

George
Online
0%