Googlebot

Google’s main crawler for Search. Most sites are crawled primarily by Googlebot Smartphone (mobile-first indexing), with Googlebot Desktop as a secondary crawler.

OperatorGoogle
PurposeSearch indexing (also classified as AI training by Cloudflare)
robots.txt tokenGooglebot
Respects robots.txtYes. Uses the most specific user-agent group and the longest matching rule.
If you block itPages drop out of Google Search. Since 15 Sept 2026, Cloudflare’s “Block” for AI training also blocks Googlebot.

User-agent string

Googlebot Smartphone

Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0.7339.80 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)

Googlebot Desktop

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/140.0.7339.80 Safari/537.36

Version numbers inside user-agents change over time. Match on the token (Googlebot), never on the full string.

robots.txt examples

Block Googlebot

User-agent: Googlebot
Disallow: /

Allow Googlebot, block a folder

User-agent: Googlebot
Disallow: /private/
Allow: /

A group for a specific crawler replaces the User-agent: * group for that crawler — rules are not combined. Repeat any rules from * that should still apply.

How to verify Googlebot

Reverse DNS must end in googlebot.com, google.com or googleusercontent.com and resolve forward to the same IP; or match Google’s published IP range files.

The user-agent alone proves nothing — it is the first thing scrapers copy. See how to verify crawlers.

Is Googlebot blocked on your site?

Your robots.txt may allow Googlebot while your CDN or firewall refuses it. Test the real response with the Googlebot checker.

Official documentation: developers.google.com

Googlebot user agent, robots.txt & how to verify · BotCanary