Cloudflare splits AI bots in three on 15 September
Until this summer the only question Cloudflare asked you about AI crawlers was yes or no. From 15 September there are three, because Cloudflare has stopped treating the bot that fetches a page to answer a live question as the same animal as the bot that swallows your archive to train a model. Search stays allowed by default. Training and Agent get blocked by default on any page that carries ads, for a wider set of customers than most of them realise. Do you know what your own setting says today?
One switch becomes three
Cloudflare's argument for dropping the binary is that the label keeps moving. As its 1 July announcement puts it, "we could debate the cutoff for what qualifies as 'AI' today, just to find that the standard changes tomorrow", so the new classification asks what a bot is doing on your site rather than which company sent it. The full taxonomy runs to eleven behaviours, from Feed Fetching to Security Testing. Three of them are the ones every customer, including Free, can now set for themselves.
| Behaviour | What Cloudflare means by it | Default from 15 September |
|---|---|---|
| Search | Any behaviour that collects or indexes your content so it can answer questions about it later. Cloudflare says site owners should expect referral traffic or other equitable compensation in return. | Allowed |
| Agent | Automation acting in real time on a person's behalf to get something done now. Cloudflare's examples are chat fetch bots such as ChatGPT-User, and browser-use agents driving Chrome. | Blocked on pages that display ads |
| Training | A crawler taking your content to train or fine-tune a model, so that your data is permanently absorbed into the model's architecture. | Blocked on pages that display ads |
Two smaller changes travel with it. AI search is no longer a category of its own: Cloudflare now says there is no meaningful distinction between "AI Search" and traditional search, and both are simply Search. And being a Verified bot has stopped meaning allowed. Verification now only makes a bot eligible for whatever its category permits, so the category you set is the thing that decides access.
Who inherits a setting they never chose
Read the announcement and you get a narrow change. The blog post says that on 15 September, "for all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default". New domains only. Nothing to do for anyone already running a site.
The press release published the same day is broader. It applies the new defaults to new customers and to new sites created by existing customers, and then adds a sentence the blog post does not carry: "On September 15, 2026 these changes will also be made for all existing free customers that have not changed their settings by September 15, 2026 in their dashboard." TechCrunch reported the same three groups on the day. If you are a small business on a free Cloudflare plan who has never opened the bot settings, the press release is the document that describes you, and it is the one worth acting on.
One limit matters before anyone panics. The new default does not switch off AI access to a whole site. It applies to pages that Cloudflare's automated detection says display ads, and Cloudflare's reasoning is that an ad marks a page a site owner meant a human to land on. A brochure site with no advertising inherits nothing from that clause. A publisher, a recipe site or anyone running display inventory is squarely in it.
Googlebot counts as a training bot
This is the part that can cost you real money, and it has almost nothing to do with AI. Googlebot, Applebot and BingBot crawl for search and for training under one user agent. Cloudflare resolves that by taking the strictest rule that applies, and says so plainly: "multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service)".
Note the second half of that sentence. The one-click "Block AI bots" toggle that a great many small sites switched on during 2025 is being deprecated on 15 September, and until now its documentation carried an explicit carve-out: "This option excludes mixed-purpose bots that are used both for Training and for Search." The toggle changes meaning underneath the people who flicked it. Same switch, same position, and from 15 September it takes Googlebot with it.
There is already an early sighting. Search Engine Journal reported on 4 August a site owner describing exactly this behaviour: "When I set AI Training = Block, both Googlebot and Bingbot start receiving HTTP 403 responses when trying to fetch my sitemap. As soon as I disable the AI Training block, the sitemap is accessible again." That is one report rather than a confirmed incident, and Google's John Mueller asked to look into it privately. We argued last year that blocking every AI bot was already costing brands visibility, and the mixed-crawler rule raises the price of the blunt version. Check your own logs rather than assuming either way.
Nobody has published which bot is which
The awkward bit for anyone trying to act on this is that Cloudflare has not published a per-bot map of the new categories. The only bots named anywhere in the announcement are ChatGPT-User, given as an example of an Agent, Gemini and Claude driving Chrome as browser-use agents, and Googlebot, Applebot and BingBot as the mixed Search plus Training case. Where GPTBot, ClaudeBot, PerplexityBot, Google-Extended or Bytespider land is not stated. Cloudflare's own bot reference table was last updated in April 2026 and still uses the previous labels: AI Crawler, AI Search, AI Assistant, Search Engine. Anyone telling you confidently that GPTBot is Training has inferred it.
Site owners have not waited for the official map. A read of Cloudflare Radar's parsed robots.txt files by TechnologyChecker found publishers already sorting one company's bots into different piles: OpenAI's GPTBot disallowed 2.33 times for every allow, while OpenAI's OAI-SearchBot is allowed slightly more often than it is blocked. Anthropic shows the same shape, with ClaudeBot running 2.39 to 1 against and Claude-User roughly at parity, and Apple's training opt-out agent Applebot-Extended sits at 3.24 to 1 against. Those are raw domain counts from a single-day snapshot of 4,223 files taken on 27 July 2026, not a share of the web, so read them as a pattern rather than a measurement. The pattern is the interesting part: people are blocking training and allowing answering from the same vendor, by hand, because until now the infrastructure gave them no way to say it once.
Which leaves you doing the mapping yourself, per bot, from what each operator publishes about its own crawlers. That is the job our bot encyclopedia exists to do: who runs each user agent, what it fetches for, and how to allow or block it without taking your search crawler down with it.
Why Cloudflare thinks this is fair
The case rests on waste and on imbalance. On waste, Cloudflare says over 50% of crawl traffic from AI crawlers is spent re-fetching unchanged pages, which is your bandwidth paying for a page the crawler already has. On imbalance, it points at the ratio between pages crawled and visitors sent back.
Be careful with those ratios, because the current ones are second-hand. Cloudflare's own July 2026 post still quotes the 2025 range of 118:1 up to nearly 50,000:1. The 2026 figures in circulation come from third parties reading Cloudflare Radar, and the clearest published set is SEOmator's, for the rolling 28-day window ending 21 July 2026.
Two caveats belong beside every one of those numbers. Cloudflare's own methodology note says the ratios may overstate extraction, because referrals from native apps carry no referrer header and so never reach the denominator. And they move fast. SEOmator's reading has Anthropic falling from roughly 57,000:1 to roughly 2,300:1 in six months, driven by referrals finally appearing rather than by any drop in crawling. A ratio quoted without its date is close to meaningless.
Underneath the fairness argument sits a competitive one. Cloudflare says the largest search engine "has access to about 2X more information than leading AI companies because they make it difficult for customers to remain discoverable without also being used for AI", and Matthew Prince names the hope behind it: that the proposed default changes encourage mixed-use crawlers to separate search from agent use and training. With more than 20% of web domains sitting behind Cloudflare, that is a lever aimed at Google as much as at OpenAI. Cloudflare also says automated agents and bots now drive more than half of all web requests, which is its framing rather than a settled figure, and we went through what the bot-share numbers do and do not support in the web went majority-machine.
One honest gap sits under all of it. The tooling that shows your own crawl-to-referral ratio per operator, and whether allowing Search is buying you anything, is an Enterprise feature, so the small business being asked to make this call cannot see the evidence the argument is built on.
What you actually do
1. Find out what you are set to. In the Cloudflare dashboard the path is Security Settings, then Configure AI bot policies. Each of Search, Agent and Training offers three options: block on all pages, block on pages with ads, or allow. Write down what you find before you change anything. Plenty of these settings were made by an agency or a previous developer.
2. Treat the legacy toggle as a decision to remake. If "Block AI bots" is on, it is currently sparing Googlebot, Applebot and BingBot and it will stop doing so. Replace it with an explicit setting per category rather than letting it lapse into one.
3. Decide by behaviour rather than by brand. Search is the category that sends people back, and allowing it is the default for a reason. Agent is a genuine judgement call: those are fetches with a human waiting, so blocking them means a customer asking an assistant about you gets an answer built without your page. Training is the one where refusing costs you least today.
4. Check what your robots.txt is already saying. Run curl https://yoursite.com/robots.txt. If Cloudflare's managed robots.txt is enabled you will see content signals declaring that crawling for search is fine and crawling for training is not, and since July that block also carries use=reference, which asks machines to index, excerpt and link back rather than reproduce. Make sure the file and the dashboard are not saying different things.
5. Opt out in writing if you want no change. Cloudflare says a site owner who wants to keep their current configuration can mark that in Security settings at any time before 15 September, which confirms they want no changes to Training crawlers that also crawl for Search.
6. Watch your logs after the date. Look for 403 responses to Googlebot and Bingbot on sitemap and feed URLs, and check Search Console for crawl errors that week. A search crawler locked out of an ad-bearing page shows up as ranking drift a month later.
Questions people ask about this
Sources: Cloudflare's 1 July 2026 announcement and press release, its developer documentation on blocking AI bots, verified bots and the AI Crawl Control bot reference, its Attribution business insights post, TechCrunch, Search Engine Journal, SEOmator and TechnologyChecker, drawn from publicly available reports published in 2025 and 2026. Crawl-to-referral ratios are third-party readings of Cloudflare Radar, dated where quoted.
