“AI crawlers” are not one thing. Each major AI company runs separate bots for training, for search and for fetching pages on a user’s request. Blocking all of them is rarely what a business wants.
The bots, by purpose
| Operator | Training | Search (citations) | User-triggered |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Perplexity | — | PerplexityBot | Perplexity-User |
| Google-Extended (token) | Googlebot | — |
Policy 1: stay visible in AI answers, opt out of training
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: *
Allow: /OAI-SearchBot, Claude-SearchBot, PerplexityBot and the user-triggered agents keep access. Google-Extended does not affect Google Search ranking.
Policy 2: block all AI
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
Disallow: /Consecutive User-agent lines share one group. Note that user-triggered fetchers may not follow robots.txt in every case — check each operator’s documentation.
Policy 3: allow everything
No AI-specific rules needed. Make sure your CDN agrees (next section).
robots.txt is not the whole story
robots.txt is a request, read only by crawlers that choose to honour it. The opposite problem is more common: robots.txt allows AI crawlers, but the CDN or firewall refuses them. A 2026 scan of the top 2,000 domains found about 9% of sites with permissive robots.txt whose edge still refused AI crawlers — most of them behind Cloudflare. A robots-only checker reports those sites as fine.
The AI crawler checker tests both: what robots.txt says and what each crawler actually receives.