Skip to main content
All articles

Which AI bots should you block? Search bots vs. training bots

By Mohamed Hadri

Blocking AI bots sounds like a single switch, but each big AI company runs several bots with different jobs. One builds a search index, one opens a page when a user asks about it, and one collects pages that may be used to train future models. Block the search bot and your site can drop out of AI answers. Block only the training bot and you stay in their search results, because the companies document the two as separate choices.

This matters more since 15 September 2026, when Cloudflare changed what its AI bot settings do. If your site runs behind Cloudflare, the word “Block” in that menu no longer means what it did.

Three jobs, three kinds of bots

  • Search bots build the index that AI search draws on when it answers and cites sources. Blocking them is how a site disappears from those answers. OpenAI says sites that opt out of OAI-SearchBot are not shown in ChatGPT search answers.
  • User bots open a page at the moment someone asks the assistant about it, and they do not crawl on their own. OpenAI says robots.txt rules may not apply to ChatGPT-User because a person started the request, and Google says the same of its user-triggered fetchers.
  • Training bots collect pages that may be used to train future models. Blocking them keeps your content out of that training, and the companies treat it as a separate decision from search.

Here is how four of them split the work, from each company’s own documentation:

CompanySearchUser requestsTraining
OpenAIOAI-SearchBotChatGPT-UserGPTBot
AnthropicClaude-SearchBotClaude-UserClaudeBot
PerplexityPerplexityBotPerplexity-UserNone listed
GoogleGooglebotGoogle-AgentGoogle-Extended

Two details are easy to miss. Perplexity says PerplexityBot is not used to collect content for training AI models, so it has no training bot for you to block. And Google-Extended does not crawl anything itself. It is a name you use in robots.txt to tell Google whether it may use your pages to train Gemini models and to ground answers in Gemini Apps. Google says it has no effect on your inclusion or ranking in Google Search, but blocking it does stop Gemini Apps from grounding their answers in your pages.

What Cloudflare changed

Cloudflare sits in front of a large share of the web, so for many sites its settings decide which bots get through. Three changes this year matter:

  • 1 July 2026: every Cloudflare plan, the free one included, got separate controls for three kinds of AI traffic: Search, Agent and Training.
  • 21 August 2026: Bot Preference Sync, which keeps your robots.txt in step with the choices you make in the dashboard. Its lines are added above whatever your file already says.
  • 15 September 2026: “Block” now also stops crawlers that serve both search and training, such as Googlebot, Bingbot and Applebot, so blocking them takes you out of their search as well. The setting that stops training and keeps search is now called Disallow AI Training. The old one-click “Block AI bots” option and the managed robots.txt feature are being retired.

If you had turned on the old “Block AI bots” option, Cloudflare moved you to the new settings for you: Search allowed, Training set to Disallow AI Training, and Agents blocked on pages that show ads. Its advice to existing customers is that in almost every case there is nothing to do. The risk sits in the new menu: pick Block to keep AI out, and you now keep Google, Bing and Apple search out too.

Cloudflare’s own numbers show what most sites choose: under 1% of its sites block search bots, while 17% block AI training in some way.

Check your site in four steps

  1. Read your robots.txt. Open it in a browser at the root of your site, for example example.com/robots.txt. A group with User-agent: * followed by Disallow: / closes the whole site to every bot that follows the file. A group for a search bot such as OAI-SearchBot or PerplexityBot with Disallow: / takes you out of that assistant’s search. If you use Cloudflare, look here too, because your dashboard choices can end up in this file.
  2. Open the AI settings in Cloudflare, if you use it. Search should say Allow. If you don’t want your content used for training, choose Disallow AI Training rather than Block. Agents are AI tools acting for a person, such as an assistant opening your page because a user asked about it. Block them only if you have a reason: Anthropic says turning its user bot away may reduce your visibility when people search through Claude.
  3. Check your firewall too. A firewall rule, a bot protection mode or a security plugin can turn a bot away with an error before it ever reads robots.txt. OpenAI and Perplexity both publish the IP addresses their search bots use and ask sites to allow them. If your security logs show those bots getting errors, add a rule that lets them in. Anthropic, for its part, warns that blocking its IP addresses is not a reliable way to keep its bots out, and asks you to use robots.txt for that.
  4. If a platform hosts your site, such as an online store builder, you may not control the firewall or robots.txt at all. Ask the platform how it treats AI search bots and whether you can change it.

A robots.txt that keeps search and says no to training

This example leaves search and user bots alone and blocks the three training bots from the table. The last group states the same wishes in a Content-Signal line, a format Cloudflare published in 2025 with three values: search, ai-input (using your pages in AI answers) and ai-train. Note that search=yes on its own does not cover AI-written answers; that is what ai-input is for.

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Allow: /

Sitemap: https://www.example.com/sitemap.xml

Three things to know before you copy it. A bot that finds a group with its own name follows only that group (RFC 9309), so if you keep private paths out under User-agent: *, repeat them in any group you add. OpenAI says its search systems can take about 24 hours to pick up a change. And if you would rather not block anyone, keep only the Content-Signal line: it states your wishes without closing the door, which is what we do on this site.

What our free scan checks

The free Mariy AI scan reads your robots.txt as part of its checks. It shows whether the file blocks your whole site, whether it blocks any AI bot by name, whether it names AI bots at all, and whether it carries a Content-Signal line. It does not count a block on training-only bots against you: blocking GPTBot, ClaudeBot, Applebot-Extended or CCBot only says you do not want your pages used for training. It does count a block on Google-Extended, because blocking it also stops Gemini Apps from using your pages in their answers, so the sample above shows up as a partial result. If you blocked it on purpose, that result is expected.

The scan cannot see your firewall. It fetches your site from our own servers, whose addresses differ from those of the OpenAI and Anthropic bots, so a rule that stops the real bots can go unnoticed. For that part, step 3 above is the check.

Sources