AI assistant fetchers
AI assistant fetchers retrieve one page at a time, live, because a person asked the assistant something. OpenAI documents ChatGPT-User as serving user actions and says it is not used for crawling the web in an automatic fashion[1], and Google keeps user-triggered fetchers in a category apart from its crawlers[2]. The class matters to site owners for one reason: some operators hold that robots.txt, written for crawlers, does not govern a fetch a human asked for[3][4]. A hit from this class is also the clearest demand signal a site gets, because a real person asked a question the page answered. An audit of about 37,000 commercial recommendation runs frames the products behind these fetches as recommendation engines rather than search engines, answering a buying question by nominating brands instead of returning a list of links[5].
The robots.txt argument
Two operators write the exemption into their own documentation. OpenAI says that because ChatGPT-User actions are initiated by a user, robots.txt rules may not apply[3]. Perplexity is blunter: since a user requested the fetch, Perplexity-User generally ignores robots.txt rules[4]. The reasoning is that the file governs systematic collection, and a single retrieval on a person's behalf resembles that person opening the page. Anthropic reaches the opposite conclusion for the same behaviour, documenting Claude-User as the agent that visits a site when someone asks Claude a question[6] while stating that its bots honour do-not-crawl directives in robots.txt[7]. See robots.txt.
The dispute
The argument became a public fight in August 2025, when Cloudflare reported that Perplexity modified its user agent and rotated source ASNs to keep fetching domains that had disallowed it[8]. Perplexity's reply rested on the class distinction, arguing that a modern AI assistant answering one question works differently from traditional web crawling[9]. Publishers read the same fetch as uncompensated access to content, which is what the block-by-default and licensing moves described under AI crawlers respond to.
How much traffic
The class is small but the fastest growing. Cloudflare recorded ChatGPT-User requests up 2,825% between May 2024 and May 2025, to a 1.3% share of crawler requests[10], and put the growth of AI user-action crawling across 2025 at more than 15 times[11]. It remains a thin slice of the total: a vendor crawler panel for January to July 2026 attributes 2.5% of AI bot requests to user-triggered action, against 46.5% for training[12]. Because each hit stands for one question asked, this class is the traffic a site can read as demand rather than harvesting. See AI visibility measurement and AI referral traffic.
Why it matters for AI visibility
This is the smallest class by volume and the most informative one. It is 2.5% of AI bot requests against 46.5% for training[12], but every hit stands for a question a person actually asked, which makes it the closest thing to a demand signal in a log file. Blocking it is a decision about being answerable rather than about training: Anthropic documents Claude-User as the agent that visits a site when someone asks Claude a question[6]. The audit framing is that these assistants now act as recommendation engines[5], so a site invisible to this class is absent from the recommendation, not merely uncrawled.
Bots in this class
The Baseline Labs bot registry (v1) tracks 30 of these. Each one has its own entry with its user agent, its robots.txt token and how to verify it.
- AI2Bot-DeepResearchEvalAllen Institute
- amazon-QBusinessAmazon
- Amzn-UserAmazon
- bigsur.aiBig Sur AI
- BuddyBotBuddyBotLearning
- ChatGPT-UserOpenAI
- Claude-UserAnthropic
- cohere-aiCohere
- DuckAssistBotDuckDuckGo
- GeistHaus-PageFetcherGeistHaus
- Gemini-Deep-ResearchGoogle
- Google-NotebookLMGoogle
- GoogleAgent-URLContextGoogle
- Grok-DeepSearchxAI
- kagi-fetcherKagi
- Kimi-UserMoonshot AI
- KlaviyoAIBotKlaviyo
- LinerBotLiner
- Meta-ExternalFetcherMeta
- MistralAI-UserMistral
- Perplexity-UserPerplexity
- PhindBotPhind
- Poggio-CitationsPoggio
- QualifiedBotQualified
- Shap-UserParallel
- TongyiBotAlibaba
- UseAIUnknown
- wpbotQuantumCloud
- WRTNBotWrtn
- YiyanBotBaidu
The full bot encyclopedia lists every class side by side.
Seen by Baseline
Requests from bots in this class across sites tracked by Baseline, last 30 days. Aggregated over - sites, never reported per site.
References (12)
- Overview of OpenAI Crawlers - OpenAI developer docs
- Overview of Google crawlers and fetchers - Google Search Central
- Overview of OpenAI Crawlers - OpenAI developer docs
- PerplexityBot - Perplexity developer docs
- Prominence-Stratified Failure Modes in Retrieval-Augmented Commercial Recommendation: A 37,000-Run Audit
- Does Anthropic crawl data from the web, and how can site owners block the crawler? - Claude support
- Does Anthropic crawl data from the web, and how can site owners block the crawler? - Claude support
- Perplexity is using stealth, undeclared crawlers to evade no-crawl directives - Cloudflare Blog
- Perplexity AI ignores no-crawling rules on websites, crawls them anyway - Malwarebytes
- From Googlebot to GPTBot: who's crawling your site in 2025 - Cloudflare Blog
- Cloudflare Radar 2025 Year in Review - Cloudflare Blog
- Crawl-to-refer ratio: AI crawlers and LLM bots - Seomator crawler panel
Last updated 2026-09-05. Written and maintained by Baseline Labs.