Block or allow AI crawlers? A robots.txt guide for 2026

Which AI bots train models and which put you in AI answers, ready-to-use robots.txt snippets for each policy, and why robots.txt alone is not enough.

Updated October 11, 2026 · 7 min read

“AI crawlers” are not one thing. Each major AI company runs separate bots for training, for search and for fetching pages on a user’s request. Blocking all of them is rarely what a business wants.

The bots, by purpose

OperatorTrainingSearch (citations)User-triggered
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBotClaude-SearchBotClaude-User
Perplexity—PerplexityBotPerplexity-User
GoogleGoogle-Extended (token)Googlebot—

Policy 1: stay visible in AI answers, opt out of training

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: *
Allow: /

OAI-SearchBot, Claude-SearchBot, PerplexityBot and the user-triggered agents keep access. Google-Extended does not affect Google Search ranking.

Policy 2: block all AI

User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
Disallow: /

Consecutive User-agent lines share one group. Note that user-triggered fetchers may not follow robots.txt in every case — check each operator’s documentation.

Policy 3: allow everything

No AI-specific rules needed. Make sure your CDN agrees (next section).

robots.txt is not the whole story

robots.txt is a request, read only by crawlers that choose to honour it. The opposite problem is more common: robots.txt allows AI crawlers, but the CDN or firewall refuses them. A 2026 scan of the top 2,000 domains found about 9% of sites with permissive robots.txt whose edge still refused AI crawlers — most of them behind Cloudflare. A robots-only checker reports those sites as fine.

The AI crawler checker tests both: what robots.txt says and what each crawler actually receives.

Block AI Crawlers with robots.txt (or Allow Them): 2026 Guide