Anyone can send User-Agent: Googlebot. Scrapers do it to slip past bot protection. To decide whether a request really came from Google, check the IP address — not the user-agent.
Method 1: reverse and forward DNS
- Run a reverse DNS lookup on the IP.
- The hostname must end in
googlebot.com,google.comorgoogleusercontent.com. - Run a forward lookup on that hostname. It must resolve back to the same IP.
$ host 66.249.66.1
1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
$ host crawl-66-249-66-1.googlebot.com
crawl-66-249-66-1.googlebot.com has address 66.249.66.1The forward step matters: anyone controlling reverse DNS for their own IPs can make it say “googlebot.com”.
Method 2: Google’s IP range files
Google publishes machine-readable JSON lists of IP ranges for its common crawlers, special-case crawlers and user-triggered fetchers. Match the request IP against these ranges — faster than DNS for high-volume log processing. Refresh the lists regularly; they change.
Common mistakes
- Allowlisting by user-agent. It lets every fake Googlebot in.
- Forgetting Google-InspectionTool. Search Console’s URL Inspection and the Rich Results Test fetch with a different user-agent token. If you allow only “Googlebot”, your own tests look blocked while real crawling works.
- Blocking all US data centers. Googlebot crawls mostly from the United States.
The same applies to other crawlers
Bing documents the same reverse-DNS check (search.msn.com). OpenAI, Anthropic and Perplexity publish IP ranges for their crawlers. Our crawler directory links to each operator’s documentation.