Managing AI crawlers:block training,stay visible
GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended: each name stands for a different purpose. Learn what each crawler does, how to set rules for it in robots.txt and why a firewall often overrides your decision without telling you.
Get in touch- What are AI crawlers?
- Which AI crawlers exist and what are they used for?
- What does Google-Extended control, and what does it leave alone?
- Should you block AI crawlers or let them in?
- What does a robots.txt for AI crawlers look like?
- Do small and mid-sized business websites block AI crawlers at all?
- Why do AI crawlers fail to reach your site even though robots.txt allows them?
- Where does robots.txt reach its limits with AI crawlers?
- Frequently asked questions
- How to get your AI crawler rules in order
- Where the information on this page comes from

Let's talk about your project.
First we check whether the project fits your business model. Then you get a proposal with phases and effort.
Have your AI crawler settings reviewed or call: +49 151 1576 5566AI crawlers are managed in robots.txt one by one, according to their purpose: you can block training crawlers such as GPTBot or ClaudeBot without shutting out the search crawlers, OAI-SearchBot, Claude-SearchBot and PerplexityBot, that decide whether you show up in AI search. The major providers now separate crawlers that collect content for model training, crawlers that build a search index and fetches triggered by a user in a chat. Block everything that sounds like AI and you disappear from the answers your customers read when they ask for a supplier. Leave everything open and your content is available for training without anyone asking you. The sensible path lies in between, and it takes only a few lines to put in place.
- OpenAI, Anthropic and Perplexity run separate, named crawlers for training, search and user-triggered fetches.
- Our own survey: none of 22 small and mid-sized business websites blocks the search crawlers of ChatGPT, Claude or Perplexity in the robots.txt.
- Blocking GPTBot keeps you visible in ChatGPT search as long as OAI-SearchBot is allowed.
- Google-Extended only governs use for Gemini, not Google Search and not AI Overviews.
- Firewalls and bot protection often block AI crawlers while robots.txt shows nothing of it.
On this page
- What are AI crawlers?
- Which AI crawlers exist and what are they used for?
- What does Google-Extended control, and what does it leave alone?
- Should you block AI crawlers or let them in?
- What does a robots.txt for AI crawlers look like?
- Do small and mid-sized business websites block AI crawlers at all?
- Why do AI crawlers fail to reach your site even though robots.txt allows them?
- Where does robots.txt reach its limits with AI crawlers?
- Frequently asked questions
- How to get your AI crawler rules in order
- Where the information on this page comes from
What are AI crawlers?
AI crawlers are automated programs that AI providers use to fetch publicly accessible web pages. Technically they work much like Googlebot: they request pages, read the content and identify themselves with a name of their own, the user agent. That name is what lets you give each crawler its own rules in your robots.txt.
What sets them apart is purpose. Some crawlers collect content that may be used to train future language models. Others build a search index from which ChatGPT, Claude or Perplexity pull and link current sources when someone asks a question. A third group only fetches a page when a user asks for it in a chat. For your visibility in AI answers, the second group matters most. How that visibility comes about overall is covered in our article on AI SEO.
Which AI crawlers exist and what are they used for?
The overview below summarises what OpenAI, Anthropic, Perplexity and Google say about their crawlers in their own documentation (as of 25 September 2026). Providers update these pages from time to time, so check the original before you change your robots.txt.
OpenAI states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, while disallowing GPTBot indicates that content should not be used to train its foundation models. Anthropic draws the same line: blocking ClaudeBot excludes future content from training, and blocking Claude-SearchBot prevents indexing for search, which Anthropic says may reduce visibility in search results. Perplexity says it does not run a training crawler; PerplexityBot exists to surface and link websites in Perplexity's search results.
| Crawler | Provider | Purpose | Token in robots.txt |
|---|---|---|---|
| GPTBot | OpenAI | Training foundation models | GPTBot |
| OAI-SearchBot | OpenAI | Search index for ChatGPT search | OAI-SearchBot |
| ChatGPT-User | OpenAI | Fetch when a user asks ChatGPT for a page | ChatGPT-User |
| ClaudeBot | Anthropic | Training Claude models | ClaudeBot |
| Claude-SearchBot | Anthropic | Indexing for search in Claude | Claude-SearchBot |
| Claude-User | Anthropic | Fetch on behalf of a Claude user | Claude-User |
| PerplexityBot | Perplexity | Search index, no training according to the provider | PerplexityBot |
| Perplexity-User | Perplexity | Fetch on behalf of a user | Perplexity-User |
| Google-Extended | Use for Gemini training and grounding, not a crawler of its own | Google-Extended |
What does Google-Extended control, and what does it leave alone?
Google-Extended is a token in robots.txt, not a separate crawler. Google says it has no user agent string of its own; crawling happens with the existing Google crawlers. With Google-Extended you decide whether content Google already crawls may be used to train future Gemini models and for grounding in the Gemini apps.
Google states explicitly that Google-Extended does not affect a site's inclusion in Google Search and is not used as a ranking signal. That applies to AI Overviews and AI Mode as well, because both are part of Google Search. If you want to limit what appears there, Google points to the usual Search controls: nosnippet, data-nosnippet, max-snippet or noindex. There is no switch that applies to the AI features alone.
In practice this means blocking Google-Extended costs you no visibility in Google Search, whereas blocking Googlebot removes you from Search altogether. How Google assembles its AI summaries is explained in our article on Google AI Overviews.
Should you block AI crawlers or let them in?
There is no single answer, but the question splits neatly into two separate decisions. The first is about training: may your content feed into future models? The second is about visibility: do you want to appear as a source in AI search answers? Because providers run separate crawlers for each, you can answer the two independently.
For businesses that want to be found through AI answers, search is the more important of the two. The overview shows which stance fits which goal.
- 01
01
Visible, training allowedAllow every crawler. This fits when your content is meant to promote you publicly anyway and you see no reason to keep it out of training.
- 02
02
Visible, training excludedBlock GPTBot, ClaudeBot and Google-Extended, allow OAI-SearchBot, Claude-SearchBot and PerplexityBot. This fits when you want to protect your texts but still be named in AI search.
- 03
03
Protect specific areasBlock individual directories for all AI crawlers, such as customer areas, dealer price lists or internal documents, and leave the rest open. For truly confidential content robots.txt is not enough, as explained below.
- 04
04
Opt out completelyBlock every AI crawler. This is a deliberate choice with a cost: your business is missing as a source in AI search answers when customers ask there for a supplier.
What does a robots.txt for AI crawlers look like?
Rules come in groups: a User-agent line names the crawler, followed by Disallow and Allow. For the stance "visible, training excluded" it looks like this:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
One detail is easy to miss. Under the RFC 9309 standard and Google's documentation, a crawler follows only the group that matches its name most specifically and ignores all other groups, including User-agent: *. If you block certain directories for all crawlers and also add a group for OAI-SearchBot, those blocks have to be repeated inside the OAI-SearchBot group. Otherwise the crawler may go exactly where you meant to keep it out.
Changes take a little time: OpenAI says it can take around 24 hours for its systems to reflect an updated robots.txt. Note that robots.txt is not the same as llms.txt; what that second file can do is covered in our article on llms.txt.
Do small and mid-sized business websites block AI crawlers at all?
None of the websites we examined block the crawlers that ChatGPT, Claude or Perplexity use to read web pages for their search. Crawlers are programs that fetch web pages automatically. On 26 Sep 2026 we checked the robots.txt of 22 websites of small and mid-sized German businesses. The robots.txt is a small file on a website that tells crawlers which pages they may fetch. We evaluated it according to the RFC 9309 standard. Put simply: if the file names a crawler, its instructions apply to that crawler, otherwise the general ones do. And the most specific rule takes precedence.
17 of the 22 websites have a robots.txt. Only two of them name an AI crawler at all, meaning a program that AI services use to read web pages. Only one website blocks one of these crawlers: CCBot, the crawler of the Common Crawl archive. GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, PerplexityBot and Google-Extended are allowed on all 22 websites. In most cases this is simply because the robots.txt does not mention them.
So if an AI service does not fetch your pages, this survey suggests the robots.txt is probably not the reason. That is why it pays to look at your firewall and bot protection. More on that in the next section.
One more finding: six of the 22 websites offer an llms.txt. This is a file that gives AI systems a short overview of the most important content on a website. One of these files is generated automatically by an SEO plugin for WordPress. What this file does and does not do is explained in our article on llms.txt. The sample is not representative. It comes from our client base and from businesses whose website we reviewed for a proposal. We look after some of these websites ourselves.
| Crawler | Purpose | Blocked in robots.txt |
|---|---|---|
| GPTBot | Training (OpenAI) | on none of 22 |
| OAI-SearchBot | Search in ChatGPT | on none of 22 |
| ChatGPT-User | Fetch on user request | on none of 22 |
| ClaudeBot | Training (Anthropic) | on none of 22 |
| Claude-SearchBot | Search in Claude | on none of 22 |
| PerplexityBot | Search in Perplexity | on none of 22 |
| Google-Extended | Use for Gemini | on none of 22 |
| CCBot | Common Crawl web archive | on one of 22 |
Why do AI crawlers fail to reach your site even though robots.txt allows them?
robots.txt is only a request to the crawler. Whether it reaches your pages at all is decided earlier, by your firewall, bot protection and hosting, and their defaults increasingly lean towards blocking. Cloudflare, which says around 20 percent of web traffic runs through its network, has blocked AI crawlers by default for newly added domains since July 2025; access then requires explicit permission.
The trouble starts when the person responsible for visibility does not know about that setting. robots.txt looks fine, yet the search crawler from ChatGPT or Claude is turned away at the firewall, and nothing in robots.txt reveals it. Ask your agency or host which bot rules are active and whether they distinguish between training and search crawlers. Your server access logs are revealing too: if OAI-SearchBot, Claude-SearchBot or PerplexityBot show up there with error codes, the cause is usually technical rather than a matter of content.
Whether opening access makes a difference only shows when you measure it with repeated queries in the AI systems. How to go about that is described in our article on how to measure AI visibility.
From our own work: A user-agent test alone is not enough for this diagnosis. At an industrial company we sent eight crawler user agents to the website, and all eight were rejected, including Googlebot. So the block was not aimed at AI but hit crawler user agents in general. A request that merely claims to be a crawler may be rejected by good bot protection anyway. Whether the real crawlers get through only shows in the server logs or in the settings of the bot protection itself.
of web traffic runs through Cloudflare by its own account, and it has blocked AI crawlers by default for new domains since July 2025.
Where does robots.txt reach its limits with AI crawlers?
robots.txt is not access control. The RFC 9309 standard states plainly that its rules are not a form of access authorization, and Google notes that the file cannot enforce crawler behaviour. Anything genuinely confidential belongs behind a password, not behind a line in robots.txt.
Then there are fetches triggered by users. OpenAI's documentation says that robots.txt rules may not apply to ChatGPT-User because a person initiated the action; according to OpenAI, ChatGPT-User does not determine whether content can appear in ChatGPT search. Perplexity says Perplexity-User generally ignores robots.txt. Anthropic, by contrast, says it respects robots.txt and offers a separate block for Claude-User.
Declared crawlers are not free of dispute either. In August 2025 Cloudflare accused Perplexity of switching to undeclared crawlers with rotating addresses when blocked; Perplexity rejected the accusations. The takeaway: robots.txt is the right tool to state your wishes to reputable providers. If you must make sure content is not fetched at all, you need technical barriers. For a view of your settings in the wider context of AI visibility, talk to our GEO agency.
Frequently asked questions
Do many companies block AI crawlers?
In our survey of 22 websites of small and mid-sized businesses on 26 Sep 2026, none blocked the search crawlers of ChatGPT, Claude or Perplexity in the robots.txt, and only one blocked the archive crawler CCBot. The sample comes from our client base and is not representative.
What happens if I block GPTBot?
According to OpenAI, disallowing GPTBot indicates that your content should not be used to train its foundation models. Your visibility in ChatGPT search does not depend on it; that is OAI-SearchBot's job.
Will I disappear from Google if I block Google-Extended?
No. Google states that Google-Extended does not affect inclusion in Google Search and is not a ranking signal. AI Overviews are part of Google Search and are governed by the usual controls such as nosnippet or noindex.
Which AI crawlers should I allow to be mentioned in AI answers?
The search crawlers: OAI-SearchBot for ChatGPT search, Claude-SearchBot for Claude and PerplexityBot for Perplexity. OpenAI states explicitly that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers.
Is blocking ClaudeBot enough to keep out all Anthropic crawlers?
No. Anthropic runs ClaudeBot, Claude-SearchBot and Claude-User as separate crawlers with their own names. Each one needs its own rule in robots.txt.
Do all AI crawlers follow robots.txt?
Not every fetch does. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores robots.txt, because both fetches are triggered by users. robots.txt is also never a form of access control.
How quickly does a robots.txt change take effect?
OpenAI says it can take around 24 hours for its systems to reflect an updated robots.txt. If opening access still has no effect, check your firewall and bot protection.
How to get your AI crawler rules in order
- 01
Set your goal
Decide separately whether your content may be used for training and whether you want to be visible in AI search.
- 02
Review robots.txt
Open your robots.txt, check the groups for GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot and Google-Extended, and make sure general blocks are repeated in every named group.
- 03
Check the firewall
Ask your host or agency which bot rules are active in the firewall and CDN, and whether they turn away AI search crawlers.
- 04
Measure the effect
Check your access logs to see whether the search crawlers get through, and ask the AI systems your customers' questions repeatedly.
Changing robots.txt takes minutes. The harder part is the decision before it: what your content is worth as training material, and what you are prepared to give up in AI visibility.
Google spam updateGoogle Spam Update: What Can Your Numbers Really Tell You?
LLM costsLLM Cost Optimization: When a Model Switch Pays Off
Detect AI-written textDetect AI-written text: what AI detectors get wrong
AI agentsAgentic AI Explained: What It Means for Your Business
Google AI ModeGoogle AI Mode: What It Means for Your Website
AI agentsWhat Is an AI Agent, and How Is It Different From a Chatbot?
Structured DataStructured Data: What It Really Does for AI Search
AI AssistantAI Assistant for Business: Types, Uses and Data Protection
GEOE-E-A-T: Trust Signals for Google and AI Search
Local AILocal AI for Business: What “Local” Really Means
AI for SMEsAI for SMEs: How to Introduce AI Step by Step
Perplexity SEOPerplexity SEO: How to Get Your Site Cited as a Source
ChatGPT SEOChatGPT SEO: How to Get Your Business Found in ChatGPT
llms.txtllms.txt: What the File Does and When It Pays Off
AI SEOAI SEO: What Changes Compared to Traditional SEO
AI OverviewsGoogle AI Overviews: How Google Picks Its Sources
GEO vs AEO vs LLMOGEO vs AEO vs LLMO: The AI Search Terms Explained
Measure AI VisibilityMeasure AI Visibility: Method, Metrics and Limits
ChatGPT AdsChatGPT Ads: How to Advertise on ChatGPT in Germany
AI Literacy ObligationAI Literacy Obligation: What Article 4 Requires Since 2026
AI DisclosureAI Chatbot Disclosure: Article 50 in Practice
AI Text WatermarkAI Text Watermark: What Claude Marks and What It Doesn't
ShopwareShopware Plugin Development: Buy or Build? A Guide
Where the information on this page comes from
- Scalableloops survey: robots.txt of 22 small and mid-sized business websites, evaluated according to RFC 9309 (Robots Exclusion Protocol)26 Sep 2026
- OpenAI: Overview of OpenAI Crawlersaccessed 25 Sep 2026
- Search Engine World: Tracking OpenAI ChatGPT Botsaccessed 25 Sep 2026
- Search Engine Journal: OpenAI Says Robots.txt May Not Apply To ChatGPT's Fetch Botaccessed 25 Sep 2026
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?accessed 25 Sep 2026
- Search Engine Journal: Anthropic's Claude Bots Make Robots.txt Decisions More Granularaccessed 25 Sep 2026
- Perplexity: Perplexity Crawlersaccessed 25 Sep 2026
- Search Roost: PerplexityBot robots.txt guideaccessed 25 Sep 2026
- Google Search Central: Google's common crawlers (Google-Extended)accessed 25 Sep 2026
- Google Search Central: AI Features and Your Websiteaccessed 25 Sep 2026
- Anglera: Google-Extended, what it controls and what it does notaccessed 25 Sep 2026
- IETF: RFC 9309, Robots Exclusion Protocolaccessed 25 Sep 2026
- Google Search Central: How Google interprets the robots.txt specificationaccessed 25 Sep 2026
- Google Search Central: Introduction to robots.txtaccessed 25 Sep 2026
- Cloudflare: Cloudflare Just Changed How AI Crawlers Scrape the Internet-at-Largeaccessed 25 Sep 2026
- Transparency Coalition: Cloudflare to block AI crawlers by defaultaccessed 25 Sep 2026
- Cloudflare: Perplexity is using stealth, undeclared crawlers to evade website no-crawl directivesaccessed 25 Sep 2026
- Intelligent CISO: Cloudflare accuses Perplexity of using stealth crawlersaccessed 25 Sep 2026
- AppleInsider: Perplexity defensive over ignoring robots.txtaccessed 25 Sep 2026


