Cloudflare allows verified search crawlers by default. When Googlebot starts receiving 403s anyway, the cause is almost always a setting someone changed — often without thinking about SEO. Nothing changes in robots.txt, the CMS or the code, so the SEO team does not see it until traffic drops.
Since 15 September 2026: Cloudflare applies the Block option for AI training to mixed-use crawlers — Googlebot, Bingbot and Applebot. Blocking AI training with “Block” now blocks Google Search too. Use Disallow AI Training instead.
Symptoms
- Search Console → Settings → Crawl stats shows a rise in 403 or “other client error” responses.
- Pages move to “Blocked due to access forbidden (403)” in the Page indexing report.
- Sitemaps show “Couldn’t fetch”.
- Merchant Center starts disapproving products because landing pages can’t be fetched.
- Organic traffic falls with no deploy, no core update and no robots.txt change.
The four usual causes
1. AI Crawl Control set to “Block” for Training
Cloudflare classifies Googlebot as a crawler with several purposes: search indexing and AI training. Since 15 September 2026, a crawler is judged by its most restrictive purpose, so blocking training also blocks Googlebot. Reports of Googlebot and Bingbot getting 403 on sitemaps after enabling a training block appeared on Reddit in August 2026, a few weeks before the official date.
Fix: Security → AI Crawl Control → Training → Disallow AI Training. It publishes a no-training preference in robots.txt and keeps search crawling allowed. For Google specifically, a Disallow for the Google-Extended token opts you out of Gemini training without affecting Search.
2. Bot Fight Mode
Bot Fight Mode challenges traffic Cloudflare classifies as automated. Verified search crawlers are normally exempt, but crawlers outside the verified list, monitoring tools and API clients get challenged. On the Free plan it cannot be bypassed with a WAF Skip rule. Site owners have reported traffic declines 30–45 days after enabling it, because challenges are logged as “managed challenge” rather than “block” and are easy to overlook.
3. A custom WAF rule matching the user-agent
Rules like http.user_agent contains "bot" or country blocks that include the United States catch Googlebot. Google crawls mostly from US IP addresses.
4. Rate limiting
Aggressive rate limits return 429 to Googlebot during crawl spikes. Google slows down and eventually crawls less.
How to confirm it is really Googlebot being blocked
Tools that send Googlebot’s user-agent from their own servers — including ours — are not Google. Cloudflare can tell the difference and may block the impostor while letting the real Googlebot in. So confirm with real data:
- Cloudflare → Security → Events. Filter by Verified bot or by user-agent containing Googlebot. The Service column names the setting that acted (e.g. “Block AI training crawlers”, “Bot fight mode”).
- Search Console → URL Inspection → Test live URL. A failed live test with “Blocked due to access forbidden” is confirmation.
- Crawl stats → By response. Compare the share of 403/429 before and after the change.
Fix without disabling protection
- Prefer Cloudflare’s built-in verified-bot handling over custom user-agent rules.
- If you need a custom exception, match on
cf.client.bot(verified bot) rather than on the user-agent string, which anyone can fake. - Exclude
/robots.txtand sitemap URLs from challenges and rate limits. - After the fix, request indexing for key URLs. Recovery is not instant: in reported cases it took about two weeks for Google to re-crawl and restore rankings.
Catch it the day it happens
The expensive part of this problem is the delay between a settings change and someone noticing. Run the Cloudflare SEO checker now, and turn on monitoring to get an alert when Googlebot or AI crawlers start getting blocked.