For most businesses: allow OAI-SearchBot, Claude-SearchBot and PerplexityBot so AI search can cite you, allow the user-request agents, and choose separately for training crawlers (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot). Google-Extended does not affect Google Search or AI Overviews. Then check that your firewall or CDN is not blocking these bots anyway.
Which AI bots matter in 2026?
Each major AI company now separates training, search and user-requested fetching into different user agents. The table summarises the official documentation.
| Company | User agent | Purpose | If you block it |
|---|---|---|---|
| OpenAI | OAI-SearchBot | Surfaces sites in ChatGPT search results | Your pages are no longer eligible for ChatGPT search answers |
| OpenAI | GPTBot | Content that may be used to train models | Excluded from future training; search unaffected |
| OpenAI | ChatGPT-User | User-initiated visits from ChatGPT and GPTs | OpenAI says robots.txt may not apply |
| Anthropic | Claude-SearchBot / Claude-User / ClaudeBot | Search quality / user retrieval / training | Less search visibility / no retrieval on request / excluded from training |
| Perplexity | PerplexityBot / Perplexity-User | Surfacing sites in Perplexity / user-triggered fetch | Not surfaced in Perplexity answers / no fetch on request |
| Google-Extended | Use in Gemini models and grounding | No effect on Google Search, AI Overviews or rankings |
Sources: OpenAI, Anthropic, Perplexity and Google. OpenAI also runs OAI-AdsBot, which only checks the landing pages of ads submitted to ChatGPT.
Three robots.txt policies you can copy
Policy A — visible everywhere. The simplest option if you want maximum reach: no AI-specific rules. Your existing rules for private paths still apply.
User-agent: *
Allow: /
Disallow: /admin/
Policy B — visible in AI search, no training. The most common choice for content businesses in 2026.
# AI search and user-requested visits: allowed
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /
Disallow: /admin/
# Model training: not allowed
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
Disallow: /
Policy C — no AI crawlers. Disallow every agent above. Expect to disappear from ChatGPT, Claude and Perplexity search answers. Keep Googlebot allowed: Google Search, including AI Overviews, depends on it.
Remember that a crawler follows only its most specific matching group — repeat your private paths in each group, as in Policy B.
Why check your firewall and CDN too?
Because security rules can block what robots.txt allows. In our 25 September 2026 check of 100 yachting websites in Miami and South Florida, only 2 of the 78 robots.txt files blocked an AI crawler — but 4 sites returned an HTTP 403 error to a standard browser request. Bot protection that aggressive can turn away legitimate crawlers as well. See our Miami yachting index for the sample.
- Review your CDN or firewall settings for AI-bot blocking options.
- Check server logs for the user agents above and their response codes.
- Verify real bots by IP: OpenAI publishes ranges at
openai.com/searchbot.json,openai.com/gptbot.jsonandopenai.com/chatgpt-user.json; Perplexity documents its own.
What can robots.txt not do?
It is a published request, not an access control. Well-behaved crawlers follow it; it does not remove content already collected, and user agents can be spoofed. OpenAI says robots.txt changes take about 24 hours to reach its systems; Perplexity says up to 24 hours.
It is also not a ranking tool. An llms.txt file does not replace robots.txt and is not used by Google Search — we explain what it can and cannot do in llms.txt and GEO.
Sources and editorial scope
Agent names and behaviours are taken from each company’s documentation as of September 2026 and can change. The robots.txt policies are templates: adapt the private paths to your site and test them before publishing.
- OpenAI — Overview of OpenAI crawlers
- Claude Help Center — Does Anthropic crawl data from the web?
- Perplexity — Perplexity crawlers
- Google — Google’s common crawlers (Google-Extended)
- Google Search Central — AI features and your website
- Arrow AI — Miami & South Florida yachting website readiness index
The Arrow read
Open the right doors, then earn the citation.
Access is a prerequisite, not a strategy. Once search crawlers can reach you, what decides a citation is how clearly your pages answer the question. Our free audit checks robots.txt, firewall responses and content clarity together.