Blog · GEO · 25 Sep 2026
Blog
GEO

Managing AI crawlers:block training,stay visible

GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended: each name stands for a different purpose. Learn what each crawler does, how to set rules for it in robots.txt and why a firewall often overrides your decision without telling you.

Get in touch
Enquiry
Nikolai Schöbel und Jeremias Burger, Co-Founder Scalableloops

Let's talk about your project.

First we check whether the project fits your business model. Then you get a proposal with phases and effort.

Have your AI crawler settings reviewed or call: +49 151 1576 5566
Blog · GEO

AI crawlers are managed in robots.txt one by one, according to their purpose: you can block training crawlers such as GPTBot or ClaudeBot without shutting out the search crawlers, OAI-SearchBot, Claude-SearchBot and PerplexityBot, that decide whether you show up in AI search. The major providers now separate crawlers that collect content for model training, crawlers that build a search index and fetches triggered by a user in a chat. Block everything that sounds like AI and you disappear from the answers your customers read when they ask for a supplier. Leave everything open and your content is available for training without anyone asking you. The sensible path lies in between, and it takes only a few lines to put in place.

In brief
  • OpenAI, Anthropic and Perplexity run separate, named crawlers for training, search and user-triggered fetches.
  • Our own survey: none of 22 small and mid-sized business websites blocks the search crawlers of ChatGPT, Claude or Perplexity in the robots.txt.
  • Blocking GPTBot keeps you visible in ChatGPT search as long as OAI-SearchBot is allowed.
  • Google-Extended only governs use for Gemini, not Google Search and not AI Overviews.
  • Firewalls and bot protection often block AI crawlers while robots.txt shows nothing of it.
Published 25 Sep 2026Nikolai Schöbel and Jeremias Burger9 min read
Nikolai SchöbelJeremias Burger

Nikolai Schöbel and Jeremias Burger

Co-founders of Scalableloops GmbH. Nikolai Schöbel leads online marketing and AI strategy, Jeremias Burger the AI architecture. Both build AI systems and train teams on them in their own agency work.

On this page
  1. What are AI crawlers?
  2. Which AI crawlers exist and what are they used for?
  3. What does Google-Extended control, and what does it leave alone?
  4. Should you block AI crawlers or let them in?
  5. What does a robots.txt for AI crawlers look like?
  6. Do small and mid-sized business websites block AI crawlers at all?
  7. Why do AI crawlers fail to reach your site even though robots.txt allows them?
  8. Where does robots.txt reach its limits with AI crawlers?
  9. Frequently asked questions
  10. How to get your AI crawler rules in order
  11. Where the information on this page comes from
Definition

What are AI crawlers?

AI crawlers are automated programs that AI providers use to fetch publicly accessible web pages. Technically they work much like Googlebot: they request pages, read the content and identify themselves with a name of their own, the user agent. That name is what lets you give each crawler its own rules in your robots.txt.

What sets them apart is purpose. Some crawlers collect content that may be used to train future language models. Others build a search index from which ChatGPT, Claude or Perplexity pull and link current sources when someone asks a question. A third group only fetches a page when a user asks for it in a chat. For your visibility in AI answers, the second group matters most. How that visibility comes about overall is covered in our article on AI SEO.

Overview

Which AI crawlers exist and what are they used for?

The overview below summarises what OpenAI, Anthropic, Perplexity and Google say about their crawlers in their own documentation (as of 25 September 2026). Providers update these pages from time to time, so check the original before you change your robots.txt.

OpenAI states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, while disallowing GPTBot indicates that content should not be used to train its foundation models. Anthropic draws the same line: blocking ClaudeBot excludes future content from training, and blocking Claude-SearchBot prevents indexing for search, which Anthropic says may reduce visibility in search results. Perplexity says it does not run a training crawler; PerplexityBot exists to surface and link websites in Perplexity's search results.

CrawlerProviderPurposeToken in robots.txt
GPTBotOpenAITraining foundation modelsGPTBot
OAI-SearchBotOpenAISearch index for ChatGPT searchOAI-SearchBot
ChatGPT-UserOpenAIFetch when a user asks ChatGPT for a pageChatGPT-User
ClaudeBotAnthropicTraining Claude modelsClaudeBot
Claude-SearchBotAnthropicIndexing for search in ClaudeClaude-SearchBot
Claude-UserAnthropicFetch on behalf of a Claude userClaude-User
PerplexityBotPerplexitySearch index, no training according to the providerPerplexityBot
Perplexity-UserPerplexityFetch on behalf of a userPerplexity-User
Google-ExtendedGoogleUse for Gemini training and grounding, not a crawler of its ownGoogle-Extended
Google

What does Google-Extended control, and what does it leave alone?

Google-Extended is a token in robots.txt, not a separate crawler. Google says it has no user agent string of its own; crawling happens with the existing Google crawlers. With Google-Extended you decide whether content Google already crawls may be used to train future Gemini models and for grounding in the Gemini apps.

Google states explicitly that Google-Extended does not affect a site's inclusion in Google Search and is not used as a ranking signal. That applies to AI Overviews and AI Mode as well, because both are part of Google Search. If you want to limit what appears there, Google points to the usual Search controls: nosnippet, data-nosnippet, max-snippet or noindex. There is no switch that applies to the AI features alone.

In practice this means blocking Google-Extended costs you no visibility in Google Search, whereas blocking Googlebot removes you from Search altogether. How Google assembles its AI summaries is explained in our article on Google AI Overviews.

Decision

Should you block AI crawlers or let them in?

There is no single answer, but the question splits neatly into two separate decisions. The first is about training: may your content feed into future models? The second is about visibility: do you want to appear as a source in AI search answers? Because providers run separate crawlers for each, you can answer the two independently.

For businesses that want to be found through AI answers, search is the more important of the two. The overview shows which stance fits which goal.

  1. 01

    01

    Visible, training allowed

    Allow every crawler. This fits when your content is meant to promote you publicly anyway and you see no reason to keep it out of training.

  2. 02

    02

    Visible, training excluded

    Block GPTBot, ClaudeBot and Google-Extended, allow OAI-SearchBot, Claude-SearchBot and PerplexityBot. This fits when you want to protect your texts but still be named in AI search.

  3. 03

    03

    Protect specific areas

    Block individual directories for all AI crawlers, such as customer areas, dealer price lists or internal documents, and leave the rest open. For truly confidential content robots.txt is not enough, as explained below.

  4. 04

    04

    Opt out completely

    Block every AI crawler. This is a deliberate choice with a cost: your business is missing as a source in AI search answers when customers ask there for a supplier.

Setup

What does a robots.txt for AI crawlers look like?

Rules come in groups: a User-agent line names the crawler, followed by Disallow and Allow. For the stance "visible, training excluded" it looks like this:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

One detail is easy to miss. Under the RFC 9309 standard and Google's documentation, a crawler follows only the group that matches its name most specifically and ignores all other groups, including User-agent: *. If you block certain directories for all crawlers and also add a group for OAI-SearchBot, those blocks have to be repeated inside the OAI-SearchBot group. Otherwise the crawler may go exactly where you meant to keep it out.

Changes take a little time: OpenAI says it can take around 24 hours for its systems to reflect an updated robots.txt. Note that robots.txt is not the same as llms.txt; what that second file can do is covered in our article on llms.txt.

Our own survey

Do small and mid-sized business websites block AI crawlers at all?

None of the websites we examined block the crawlers that ChatGPT, Claude or Perplexity use to read web pages for their search. Crawlers are programs that fetch web pages automatically. On 26 Sep 2026 we checked the robots.txt of 22 websites of small and mid-sized German businesses. The robots.txt is a small file on a website that tells crawlers which pages they may fetch. We evaluated it according to the RFC 9309 standard. Put simply: if the file names a crawler, its instructions apply to that crawler, otherwise the general ones do. And the most specific rule takes precedence.

17 of the 22 websites have a robots.txt. Only two of them name an AI crawler at all, meaning a program that AI services use to read web pages. Only one website blocks one of these crawlers: CCBot, the crawler of the Common Crawl archive. GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, PerplexityBot and Google-Extended are allowed on all 22 websites. In most cases this is simply because the robots.txt does not mention them.

So if an AI service does not fetch your pages, this survey suggests the robots.txt is probably not the reason. That is why it pays to look at your firewall and bot protection. More on that in the next section.

One more finding: six of the 22 websites offer an llms.txt. This is a file that gives AI systems a short overview of the most important content on a website. One of these files is generated automatically by an SEO plugin for WordPress. What this file does and does not do is explained in our article on llms.txt. The sample is not representative. It comes from our client base and from businesses whose website we reviewed for a proposal. We look after some of these websites ourselves.

CrawlerPurposeBlocked in robots.txt
GPTBotTraining (OpenAI)on none of 22
OAI-SearchBotSearch in ChatGPTon none of 22
ChatGPT-UserFetch on user requeston none of 22
ClaudeBotTraining (Anthropic)on none of 22
Claude-SearchBotSearch in Claudeon none of 22
PerplexityBotSearch in Perplexityon none of 22
Google-ExtendedUse for Geminion none of 22
CCBotCommon Crawl web archiveon one of 22
Firewall

Why do AI crawlers fail to reach your site even though robots.txt allows them?

robots.txt is only a request to the crawler. Whether it reaches your pages at all is decided earlier, by your firewall, bot protection and hosting, and their defaults increasingly lean towards blocking. Cloudflare, which says around 20 percent of web traffic runs through its network, has blocked AI crawlers by default for newly added domains since July 2025; access then requires explicit permission.

The trouble starts when the person responsible for visibility does not know about that setting. robots.txt looks fine, yet the search crawler from ChatGPT or Claude is turned away at the firewall, and nothing in robots.txt reveals it. Ask your agency or host which bot rules are active and whether they distinguish between training and search crawlers. Your server access logs are revealing too: if OAI-SearchBot, Claude-SearchBot or PerplexityBot show up there with error codes, the cause is usually technical rather than a matter of content.

Whether opening access makes a difference only shows when you measure it with repeated queries in the AI systems. How to go about that is described in our article on how to measure AI visibility.

From our own work: A user-agent test alone is not enough for this diagnosis. At an industrial company we sent eight crawler user agents to the website, and all eight were rejected, including Googlebot. So the block was not aimed at AI but hit crawler user agents in general. A request that merely claims to be a crawler may be rejected by good bot protection anyway. Whether the real crawlers get through only shows in the server logs or in the settings of the bot protection itself.

20 %

of web traffic runs through Cloudflare by its own account, and it has blocked AI crawlers by default for new domains since July 2025.

Cloudflare, press release, July 2025
Limits

Where does robots.txt reach its limits with AI crawlers?

robots.txt is not access control. The RFC 9309 standard states plainly that its rules are not a form of access authorization, and Google notes that the file cannot enforce crawler behaviour. Anything genuinely confidential belongs behind a password, not behind a line in robots.txt.

Then there are fetches triggered by users. OpenAI's documentation says that robots.txt rules may not apply to ChatGPT-User because a person initiated the action; according to OpenAI, ChatGPT-User does not determine whether content can appear in ChatGPT search. Perplexity says Perplexity-User generally ignores robots.txt. Anthropic, by contrast, says it respects robots.txt and offers a separate block for Claude-User.

Declared crawlers are not free of dispute either. In August 2025 Cloudflare accused Perplexity of switching to undeclared crawlers with rotating addresses when blocked; Perplexity rejected the accusations. The takeaway: robots.txt is the right tool to state your wishes to reputable providers. If you must make sure content is not fetched at all, you need technical barriers. For a view of your settings in the wider context of AI visibility, talk to our GEO agency.

Frequently asked questions

Frequently asked questions

Do many companies block AI crawlers?

In our survey of 22 websites of small and mid-sized businesses on 26 Sep 2026, none blocked the search crawlers of ChatGPT, Claude or Perplexity in the robots.txt, and only one blocked the archive crawler CCBot. The sample comes from our client base and is not representative.

What happens if I block GPTBot?

According to OpenAI, disallowing GPTBot indicates that your content should not be used to train its foundation models. Your visibility in ChatGPT search does not depend on it; that is OAI-SearchBot's job.

Will I disappear from Google if I block Google-Extended?

No. Google states that Google-Extended does not affect inclusion in Google Search and is not a ranking signal. AI Overviews are part of Google Search and are governed by the usual controls such as nosnippet or noindex.

Which AI crawlers should I allow to be mentioned in AI answers?

The search crawlers: OAI-SearchBot for ChatGPT search, Claude-SearchBot for Claude and PerplexityBot for Perplexity. OpenAI states explicitly that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers.

Is blocking ClaudeBot enough to keep out all Anthropic crawlers?

No. Anthropic runs ClaudeBot, Claude-SearchBot and Claude-User as separate crawlers with their own names. Each one needs its own rule in robots.txt.

Do all AI crawlers follow robots.txt?

Not every fetch does. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores robots.txt, because both fetches are triggered by users. robots.txt is also never a form of access control.

How quickly does a robots.txt change take effect?

OpenAI says it can take around 24 hours for its systems to reflect an updated robots.txt. If opening access still has no effect, check your firewall and bot protection.

What now

How to get your AI crawler rules in order

  1. 01

    Set your goal

    Decide separately whether your content may be used for training and whether you want to be visible in AI search.

  2. 02

    Review robots.txt

    Open your robots.txt, check the groups for GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot and Google-Extended, and make sure general blocks are repeated in every named group.

  3. 03

    Check the firewall

    Ask your host or agency which bot rules are active in the firewall and CDN, and whether they turn away AI search crawlers.

  4. 04

    Measure the effect

    Check your access logs to see whether the search crawlers get through, and ask the AI systems your customers' questions repeatedly.

Changing robots.txt takes minutes. The harder part is the decision before it: what your content is worth as training material, and what you are prepared to give up in AI visibility.

or call: +49 151 1576 5566

Further reading
Sources

Where the information on this page comes from

Projekt-Detail

    Got a project in mind?

    We reply personally. First a use-case check, then an architecture proposal.

    Start your inquiry