Skip to content
AI Visibility

Should you allow AI crawlers on your site?

Published on

Yes for crawlers that serve search, if you want to be cited by AIs; it is a free choice for those collecting training data. OpenAI, Anthropic, and Google separate these two uses: at OpenAI, a site that blocks OAI-SearchBot is not shown in ChatGPT search responses, whereas blocking GPTBot only concerns training. Our measurements point in the same direction: none of the sites cited by ChatGPT blocked OAI-SearchBot, but 3% of them blocked GPTBot.

For what type of question? It all depends on the crawler. A search crawler determines whether your site can appear in responses. A training crawler determines whether your content can be used to train future models. The two settings are independent.

Two families of crawlers, two different decisions

OpenAI documents three crawlers. OAI-SearchBot is used to display sites in ChatGPT search results; a site that excludes it is not shown in its search responses, even though it may still appear as a navigational link. GPTBot collects content that can be used to train its models. ChatGPT-User intervenes for certain user actions. OpenAI specifies that each setting is independent: a site can allow OAI-SearchBot to appear in search while blocking GPTBot.

Anthropic makes the same distinction. ClaudeBot collects content for training; blocking it excludes the site's future content from training data. Claude-SearchBot improves search result quality, and Claude-User accesses sites when a user asks a question; blocking them can reduce the site's visibility in responses. At Google, Google-Extended determines whether content can be used to train future Gemini models, and Google specifies that it has no effect on a site's inclusion in Google Search or its ranking.

What we measured in robots.txt files

We read the robots.txt file of each site returned by search on ChatGPT, Claude, and Gemini across a dozen business sectors queried in French, on the day of collection, September 10, 2026. First observation: the vast majority of sites do not mention AI crawlers. OAI-SearchBot is only mentioned by 5 to 6% of sites, and GPTBot by 11 to 16%, depending on whether they were cited or not.

Second observation: less than 1% of the sites returned by ChatGPT search blocked OAI-SearchBot, and none of the sites it cited, which aligns with OpenAI's documentation. On the other hand, blocking training did not prevent citation: 3% of the sites cited by ChatGPT blocked GPTBot, compared to 8% of the sites it returned without citing. At Claude, 5% of the cited sites blocked ClaudeBot, compared to 4% of non-cited sites; at Gemini, 2% of the cited sites blocked Google-Extended.

What these numbers mean

Blocking a training crawler does not shut the door on responses. Sites that refuse to let their content train GPT, Claude, or Gemini are still found and cited by these AIs when they search the web, because other crawlers, or the search engine itself, are reading the pages at that moment.

Conversely, blocking a search crawler removes the site from the responses of the respective AI. For ChatGPT, the finding is clear: among the sites it cited, we found none that blocked OAI-SearchBot. Allowing this crawler is not enough to be cited; according to OpenAI, blocking it is enough to not be shown in its search responses.

What this changes for a business

Check your robots.txt file. If you want to be cited, let search crawlers through: OAI-SearchBot and ChatGPT-User for ChatGPT, Claude-SearchBot and Claude-User for Claude, and Googlebot for Google and its AI responses. Also, make sure that a firewall or bot protection service is not blocking them even before the file is read.

For training crawlers, such as GPTBot, ClaudeBot, and Google-Extended, the decision depends on your policy regarding the use of your content, not your visibility in responses: our measurements do not show that blocking them prevents being cited. It is a choice to be made knowingly, not a technical setting to be left to chance.

What we do not know

We read the robots.txt files on the day of collection, without verifying whether the crawlers actually respect these files or whether a firewall was blocking them. Claude-SearchBot and Claude-User were not among the crawlers read. A missing file was counted as not mentioning any crawler.

Sources

  1. OpenAI — OpenAI crawlers, .
  2. Anthropic — Anthropic crawlers and site blocking, .
  3. Google Search Central — common Google crawlers, .
  4. FaireDuBruit measurements, from September 4 to 16, 2026 — see "How we measured".

Related questions

All questions about visibility in AI · Our offer