Short answer

For most businesses: allow OAI-SearchBot, Claude-SearchBot and PerplexityBot so AI search can cite you, allow the user-request agents, and choose separately for training crawlers (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot). Google-Extended does not affect Google Search or AI Overviews. Then check that your firewall or CDN is not blocking these bots anyway.

Which AI bots matter in 2026?

Each major AI company now separates training, search and user-requested fetching into different user agents. The table summarises the official documentation.

CompanyUser agentPurposeIf you block it
OpenAIOAI-SearchBotSurfaces sites in ChatGPT search resultsYour pages are no longer eligible for ChatGPT search answers
OpenAIGPTBotContent that may be used to train modelsExcluded from future training; search unaffected
OpenAIChatGPT-UserUser-initiated visits from ChatGPT and GPTsOpenAI says robots.txt may not apply
AnthropicClaude-SearchBot / Claude-User / ClaudeBotSearch quality / user retrieval / trainingLess search visibility / no retrieval on request / excluded from training
PerplexityPerplexityBot / Perplexity-UserSurfacing sites in Perplexity / user-triggered fetchNot surfaced in Perplexity answers / no fetch on request
GoogleGoogle-ExtendedUse in Gemini models and groundingNo effect on Google Search, AI Overviews or rankings

Sources: OpenAI, Anthropic, Perplexity and Google. OpenAI also runs OAI-AdsBot, which only checks the landing pages of ads submitted to ChatGPT.

Three robots.txt policies you can copy

Policy A — visible everywhere. The simplest option if you want maximum reach: no AI-specific rules. Your existing rules for private paths still apply.

User-agent: *
Allow: /
Disallow: /admin/

Policy B — visible in AI search, no training. The most common choice for content businesses in 2026.

# AI search and user-requested visits: allowed
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /
Disallow: /admin/

# Model training: not allowed
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
Disallow: /

Policy C — no AI crawlers. Disallow every agent above. Expect to disappear from ChatGPT, Claude and Perplexity search answers. Keep Googlebot allowed: Google Search, including AI Overviews, depends on it.

Remember that a crawler follows only its most specific matching group — repeat your private paths in each group, as in Policy B.

Why check your firewall and CDN too?

Because security rules can block what robots.txt allows. In our 25 September 2026 check of 100 yachting websites in Miami and South Florida, only 2 of the 78 robots.txt files blocked an AI crawler — but 4 sites returned an HTTP 403 error to a standard browser request. Bot protection that aggressive can turn away legitimate crawlers as well. See our Miami yachting index for the sample.

  • Review your CDN or firewall settings for AI-bot blocking options.
  • Check server logs for the user agents above and their response codes.
  • Verify real bots by IP: OpenAI publishes ranges at openai.com/searchbot.json, openai.com/gptbot.json and openai.com/chatgpt-user.json; Perplexity documents its own.

What can robots.txt not do?

It is a published request, not an access control. Well-behaved crawlers follow it; it does not remove content already collected, and user agents can be spoofed. OpenAI says robots.txt changes take about 24 hours to reach its systems; Perplexity says up to 24 hours.

It is also not a ranking tool. An llms.txt file does not replace robots.txt and is not used by Google Search — we explain what it can and cannot do in llms.txt and GEO.

Sources and editorial scope

Agent names and behaviours are taken from each company’s documentation as of September 2026 and can change. The robots.txt policies are templates: adapt the private paths to your site and test them before publishing.

The Arrow read

Open the right doors, then earn the citation.

Access is a prerequisite, not a strategy. Once search crawlers can reach you, what decides a citation is how clearly your pages answer the question. Our free audit checks robots.txt, firewall responses and content clarity together.