If you want to appear in AI search, allow the search crawlers, such as Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot, PerplexityBot and Applebot. Treat model training as a separate decision: operators document GPTBot, ClaudeBot, Google-Extended and Applebot-Extended separately from their search crawlers. Fetchers triggered by a user's request may not follow robots.txt.
Who this guide is for
Owners, trustees, practice managers and developers of UK charity, clinic and small business websites who want to make a deliberate choice about AI crawlers, and anyone editing robots.txt or Cloudflare bot settings.
Who this is not for
Anyone trying to keep content private: robots.txt is not access control, so use a login. It is not a guide to licensing content to AI companies either, and nothing here guarantees that a search product will cite you.
What is the difference between search, training and user-triggered crawlers?
They do different jobs, so they need separate decisions. Search crawlers build the index that a search or answer product draws on. Training crawlers collect content that may be used to train models. User-triggered fetchers visit a page because a person asked a question.
- Search crawlers. OpenAI says OAI-SearchBot “is used to surface websites in search results in ChatGPT’s search features”. Perplexity says PerplexityBot is “designed to surface and link websites in search results”. Anthropic uses Claude-SearchBot. Googlebot and Bingbot feed Google and Bing, including the AI features built into them.
- Training crawlers and control tokens. GPTBot, ClaudeBot and CCBot fetch pages. Google-Extended and Applebot-Extended do not. Google says Google-Extended “doesn’t have a separate HTTP request user agent string”, and Apple says Applebot-Extended “does not crawl webpages”. They are robots.txt tokens that tell the company how content already fetched by Googlebot or Applebot may be used.
- User-triggered fetchers. ChatGPT-User, Claude-User and Perplexity-User fetch a page in response to a user’s question. Because a person started the request, OpenAI and Perplexity say robots.txt may not apply.
Some crawlers serve more than one purpose. Amazon says Amazonbot improves its products and services and “may be used to train Amazon AI models”. Meta describes meta-externalagent as crawling for “training foundation AI models or improving products by indexing content directly”. Apple says Applebot’s data powers search and may also help train its foundation models. For these, a block can affect more than training.
Which crawler does what, and what does blocking it change?
The table summarises each operator’s own documentation, checked on 25 September 2026. The fourth column answers one narrow question: does blocking this token take you out of that company’s search or answer product? Where the documentation does not say, we say so.
| Crawler | Operator | What it’s for | Does blocking it remove you from that company’s search/answers? | robots.txt token | Official docs |
|---|---|---|---|---|---|
| OAI-SearchBot | OpenAI | Surfaces websites in search results in ChatGPT’s search features. | Yes. OpenAI says opted-out sites are not shown in ChatGPT search answers, though they can still appear as navigational links. | OAI-SearchBot | OpenAI crawlers |
| GPTBot | OpenAI | Crawls content that may be used to train OpenAI’s generative AI foundation models. | No. OpenAI says each setting is independent: you can allow OAI-SearchBot and disallow GPTBot. | GPTBot | OpenAI crawlers |
| ChatGPT-User | OpenAI | Visits a page when a ChatGPT or custom GPT user’s request needs it. Not used for automatic crawling. | Not reliably. OpenAI says robots.txt rules may not apply to these user-initiated requests. | ChatGPT-User | OpenAI crawlers |
| Googlebot | Crawls for Google Search and all its features, plus Google Images, Video, News and Discover. | Yes. Google says Googlebot’s robots.txt rules are the control for how your site is crawled for Search, which is where AI Overviews and AI Mode sit. | Googlebot | Google common crawlers | |
| Google-Extended | A control token, not a separate crawler. Governs use of content Google crawls for training future Gemini models and for grounding in Gemini Apps and Grounding with Google Search on Vertex AI. | No for Google Search. Google says it does not affect inclusion in Google Search and is not a ranking signal. | Google-Extended | Google common crawlers | |
| ClaudeBot | Anthropic | Collects web content to help improve Anthropic’s generative AI models. | Anthropic describes it as a training control: restricting it signals that future content should be excluded from training datasets. | ClaudeBot | Anthropic help article |
| Claude-SearchBot | Anthropic | Navigates the web to improve search result quality for users. | Likely. Anthropic says disabling it may reduce your visibility and accuracy in user search results. | Claude-SearchBot | Anthropic help article |
| Claude-User | Anthropic | Fetches pages when a Claude user asks a question. | Anthropic says disabling it stops retrieval for user queries and may reduce visibility in user-directed web search. | Claude-User | Anthropic help article |
| PerplexityBot | Perplexity | Surfaces and links websites in Perplexity search results. Perplexity says it is not used to crawl content for AI foundation models. | Likely. Perplexity recommends allowing it to ensure your site appears in search results. | PerplexityBot | Perplexity crawlers |
| Perplexity-User | Perplexity | Visits a page when a user asks Perplexity a question, and may link it in the answer. | Not reliably. Perplexity says it generally ignores robots.txt because a user requested the fetch. | Perplexity-User | Perplexity crawlers |
| Applebot | Apple | Powers search in Spotlight, Siri and Safari. Apple says the data may also help train its foundation models. | Apple ties Applebot’s data to search in Spotlight, Siri and Safari, so blocking it withholds your pages from that source. | Applebot | About Applebot |
| Applebot-Extended | Apple | A control token that does not crawl. Lets you opt out of training Apple’s foundation models. | No. Apple says pages that disallow it can still be included in search results. | Applebot-Extended | About Applebot |
| Bingbot | Microsoft | Bing’s standard crawler. Bing says Bing and Copilot search rely on the same crawling and indexing foundation. | Very likely. Bing lists blocking Bingbot in robots.txt among practices to avoid if you want to be eligible for Bing and Copilot results. | bingbot | Bing crawlers |
| CCBot | Common Crawl | Builds Common Crawl’s free, open repository of web crawl data, which anyone can use. | Not applicable. Common Crawl does not describe a search or answer product of its own. | CCBot | CCBot |
| meta-externalagent | Meta | Crawls for uses such as training foundation AI models or improving products by indexing content directly. | Not documented as Meta AI’s search crawler. Meta links citations in Meta AI’s responses to Meta-WebIndexer. | meta-externalagent | Meta web crawlers |
| Meta-WebIndexer | Meta | Improves Meta AI search result quality. | Likely. Meta says allowing it helps Meta cite and link to your content in Meta AI’s responses. | meta-webindexer | Meta web crawlers |
| Amazonbot | Amazon | Improves Amazon products and services. Amazon says it may be used to train Amazon AI models. | Not documented as Amazon’s search agent. Amazon names Amzn-SearchBot for search experiences such as Alexa. | Amazonbot | About Amazonbot |
| Amzn-SearchBot | Amazon | Improves search in Amazon products and services. Amazon says it does not crawl for generative AI model training. | Likely. Amazon says allowing it makes your content eligible for search experiences such as Alexa. | Amzn-SearchBot | About Amazonbot |
Two rows are not in most templates but belong next to their training counterparts. Meta says allowing Meta-WebIndexer “helps us cite and link to your content in Meta AI’s responses”. Amazon says allowing Amzn-SearchBot makes your content “eligible to appear in search experiences such as Alexa”.
What does Google-Extended control, and what does it not?
Google-Extended controls whether content Google crawls may be used to train future Gemini models and for grounding in Gemini Apps and Vertex AI. It does not control Google Search, and Google does not point to it for AI Overviews or AI Mode.
Google describes Google-Extended as a “standalone product token” for “training future generations of Gemini models that power Gemini Apps and Vertex AI API for Gemini and for grounding in Gemini Apps and Grounding with Google Search on Vertex AI”. It adds: “Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.”
Googlebot is different. Google says rules for Googlebot affect Google Search, “including Discover and all Google Search features”. Its page on AI Overviews and AI Mode says “robots.txt directives for Googlebot is the control for site owners to manage access to how their sites are crawled for Search”. To limit what those features show, Google points to nosnippet, data-nosnippet, max-snippet and noindex. These also affect your ordinary search results.
Two practical points follow. Blocking Google-Extended is not a pure training switch, because Google says it also covers grounding in Gemini Apps. And to be shown as a supporting link in AI Overviews or AI Mode, Google says a page “must be indexed and eligible to be shown in Google Search with a snippet”.
How do Bingbot and Microsoft Copilot fit together?
Bing says “Bing and Copilot search experiences rely on the same core crawling, indexing, and ranking foundation as traditional search”. Bing’s list of crawlers has no separate Copilot crawler, so Bingbot access is what matters.
Bing lists Bingbot, AdIdxBot, BingPreview, MicrosoftPreview and BingVideoPreview. Its Webmaster Guidelines list “Blocking Bingbot in your robots.txt file” among the things to avoid. They also say “robots.txt controls crawl access, not indexing”. Use noindex when a URL should not appear in Bing search, Copilot or grounding API results.
For finer control, Bing says noarchive “prevents content from being used in Copilot responses and grounding results”. nocache “limits Copilot to using only the URL, title, and snippet”. We did not find a separate Microsoft training token in the Bing documentation we checked, so we do not list one.
What should robots.txt look like if you want AI search but not training?
Put the search crawlers in one group and the training tokens in another. The example below is a starting point for that common choice. It is not a recommendation for every site, so adapt the paths and check each operator’s current documentation first.
# EXAMPLE ONLY: "appear in AI search, opt out of model training".
# Check each operator's current documentation before you use it.
# Search and answer crawlers: allowed.
# Repeat any path rules you need in every group.
User-agent: Googlebot
User-agent: Bingbot
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: Applebot
User-agent: Meta-WebIndexer
User-agent: Amzn-SearchBot
Allow: /
Disallow: /admin/
# Training crawlers and training control tokens: disallowed.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
User-agent: Amazonbot
Disallow: /
# Everyone else, including user-triggered fetchers.
User-agent: *
Allow: /
Disallow: /admin/
Sitemap: https://www.example.org/sitemap.xml Before you copy it, check these points:
- Named groups replace the catch-all group. Google says a crawler follows the group with the most specific matching user agent, and “user agent specific groups and global groups (*) are not combined”. If your
User-agent: *group blocks/admin/, repeat that line in each named group. - Some blocks do more than stop training. Google-Extended also covers grounding in Gemini Apps. Amazonbot and meta-externalagent have broader stated purposes. Amazon also supports a
noarchiverobots meta tag, which it describes as “do not use the page for model training”. - Apple falls back to Googlebot. Apple says that if your rules don’t mention Applebot but do mention Googlebot, Applebot follows the Googlebot rules. Name Applebot if you want it treated differently.
- User-triggered fetchers are left under
*. Anthropic says its bots honour robots.txt. OpenAI and Perplexity say their user-triggered fetchers may not. - Changes are not instant. OpenAI says search can take about 24 hours to reflect a robots.txt update. Perplexity and Amazon give similar timescales. Google says it generally caches robots.txt for up to 24 hours.
Google says user-agent values are case-insensitive but paths are case-sensitive. That is why bingbot and Bingbot both match, while /Admin/ and /admin/ do not.
Does robots.txt actually stop AI crawlers?
No. Robots.txt is a public request. Reputable operators say they honour it, but nothing in the file enforces it.
The standard, RFC 9309, describes rules “that crawlers are requested to honor” and says “these rules are not a form of access authorization”. Google says robots.txt instructions “cannot enforce crawler behavior to your site; it’s up to the crawler to obey them”. It recommends other methods, such as password protection, for anything that must stay private. Cloudflare says the same: “robots.txt compliance is voluntary.”
User-triggered fetchers are the clearest gap. OpenAI says robots.txt “may not apply” to ChatGPT-User. Perplexity says Perplexity-User “generally ignores robots.txt rules”. Meta says Meta-ExternalFetcher “may bypass robots.txt rules”. If you need to enforce a block, it has to happen at your server, firewall or CDN.
What does Cloudflare block by default, and what are Content Signals?
It depends on the setting and on when the domain was added. Cloudflare’s documentation describes three separate things: AI bot blocking policies, a managed robots.txt, and Content Signals.
AI bot policies (enforcement)
Cloudflare groups AI bots into Search, Agent and Training, with options to block on all pages, block on pages with ads, or allow. It says that from 15 September 2026, new domains get updated defaults: bots classified as Training or Agent “will be blocked on pages that display ads, and Search will remain allowed”. Mixed-purpose crawlers that combine Search and Training “will also be blocked by all configurations to block AI training”. The older “Block AI bots” option is marked as deprecating on the same date. AI Crawl Control also lets you allow or block individual crawlers and reports robots.txt violations.
Managed robots.txt (a preference)
When you turn this setting on, Cloudflare adds its own rules to the start of your existing robots.txt. Its documented example adds Content-signal: search=yes, ai-train=no, use=reference for all user agents. It then disallows Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent. Cloudflare says the feature is available on all plans.
Content Signals
Content Signals are robots.txt lines with three categories: search, ai-input and ai-train. Cloudflare’s policy text says search “does not include providing AI-generated search summaries”. ai-input covers uses such as retrieval-augmented generation and grounding. Free-plan domains that have no robots.txt of their own and do not use the managed setting display the policy text, which defines the categories without setting preferences. Cloudflare is also testing an optional use field. It notes that Google Search Console may report “Syntax not understood” for these lines, but says it has observed no impact on crawling or SEO.
If your site sits behind Cloudflare, check Security Settings and AI Crawl Control before assuming robots.txt decides access. An Allow line does not override a block at the network edge.
Do you need an llms.txt file?
No, not for Google. llms.txt is an optional proposal, and Google says Google Search ignores it.
The llms.txt site describes it as “a proposal to standardise on using an /llms.txt file to provide information to help agents use a website”. Google’s guide to its generative AI search features says you don’t need “new machine readable files, AI text files, markup, or Markdown” to appear in Google Search, “as Google Search itself doesn’t use them”. It adds that creating llms.txt files for other services “will neither harm nor help your site’s visibility or rankings in Google Search”.
Some other tools may read llms.txt, so publishing one is a reasonable choice if you maintain it. Treat it as an extra, not a substitute for crawlable, indexable pages.
How can you check which crawlers actually visit your site?
Look at your server or CDN access logs, then verify anything that matters. A user-agent string proves nothing on its own. Bing warns that user agent strings “are easy to spoof”.
- Export recent access logs from your host or CDN and search for the tokens in the table.
- Don’t look for Google-Extended or Applebot-Extended. They are control tokens, so their traffic appears as Googlebot or Applebot.
-
Verify the source. Google supports a reverse DNS lookup that resolves to googlebot.com, google.com or googleusercontent.com, confirmed by a forward lookup, or a match against its published IP ranges. Bing offers a Verify Bingbot tool. Apple uses reverse DNS in
*.applebot.apple.comand publishes IP ranges. OpenAI, Perplexity, Anthropic and Amazon publish IP ranges. Common Crawl runs CCBot on dedicated IP ranges with reverse DNS. - Check the status codes. An allowed crawler receiving 403 responses or challenges points to a firewall or bot setting, not robots.txt.
- If a blocked crawler keeps visiting, check whether it is a user-triggered fetcher, a spoofed user agent or a genuine violation.
What should you do next?
Write down your policy first: which search products you want to appear in, and whether you accept training use by each operator. Then make robots.txt and any CDN settings match it.
- Run our free robots.txt checker. It tests a URL against your declared rules for Googlebot and a set of named AI crawlers. It cannot see firewall blocks or confirm indexing.
- Read how to assess an AI search citation before treating access as a route to being cited.
- Use our AI search access guide to check each product in turn.
-
Recheck the documentation from time to time. Crawler names change. Anthropic’s current article lists ClaudeBot, Claude-SearchBot and Claude-User, and does not mention older tokens such as
anthropic-aiorClaude-Webthat still appear in many templates.
Sources and further reading: what did we check?
Every factual claim above comes from these official pages. Each was checked on 25 September 2026.
- Overview of OpenAI crawlers (OpenAI). Checked 25 September 2026.
- Google’s common crawlers (Google). Checked 25 September 2026.
- AI features and your website (Google Search Central). Checked 25 September 2026.
- Optimizing your website for generative AI features on Google Search (Google Search Central). Checked 25 September 2026.
- Introduction to robots.txt (Google Search Central). Checked 25 September 2026.
- How Google interprets the robots.txt specification (Google). Checked 25 September 2026.
- Verifying Google crawler requests (Google). Checked 25 September 2026.
- Does Anthropic crawl data from the web, and how can site owners block the crawler? (Anthropic). Checked 25 September 2026.
- Perplexity crawlers (Perplexity). Checked 25 September 2026.
- About Applebot (Apple). Checked 25 September 2026.
- Which crawlers does Bing use? (Microsoft Bing). Checked 25 September 2026.
- Bing Webmaster Guidelines (Microsoft Bing). Checked 25 September 2026.
- CCBot (Common Crawl). Checked 25 September 2026.
- Common Crawl FAQ (Common Crawl). Checked 25 September 2026.
- Meta web crawlers (Meta). Checked 25 September 2026.
- About Amazonbot (Amazon). Checked 25 September 2026.
- robots.txt setting (managed robots.txt and Content Signals) (Cloudflare). Checked 25 September 2026.
- Block AI Bots (Cloudflare). Checked 25 September 2026.
- Manage AI crawlers (AI Crawl Control) (Cloudflare). Checked 25 September 2026.
- RFC 9309: Robots Exclusion Protocol (IETF / RFC Editor). Checked 25 September 2026.
- The /llms.txt file (proposal) (llmstxt.org). Checked 25 September 2026.