| Operator | OpenAI |
|---|---|
| Purpose | Collecting data to train OpenAI models |
| robots.txt token | GPTBot |
| Respects robots.txt | Yes, per OpenAI documentation. |
| If you block it | Your content is excluded from future training data. No effect on ChatGPT search visibility. |
User-agent string
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.3; +https://openai.com/gptbot)Version numbers inside user-agents change over time. Match on the token (GPTBot), never on the full string.
robots.txt examples
Block GPTBot
User-agent: GPTBot
Disallow: /Allow GPTBot, block a folder
User-agent: GPTBot
Disallow: /private/
Allow: /A group for a specific crawler replaces the User-agent: * group for that crawler — rules are not combined. Repeat any rules from * that should still apply.
How to verify GPTBot
OpenAI publishes the IP ranges used by GPTBot.
The user-agent alone proves nothing — it is the first thing scrapers copy. See how to verify crawlers.
Is GPTBot blocked on your site?
Your robots.txt may allow GPTBot while your CDN or firewall refuses it. Test the real response with the AI crawler checker.
Official documentation: platform.openai.com