Can Google and AI crawlers actually reach your site?
robots.txt says yes. Your CDN may say no. Check what Googlebot, Bingbot, ChatGPT, Claude and Perplexity really get when they request your pages.
Every layer between a crawler and your content
One report covering the rules you wrote and the blocks you didn’t know about.
Search engine access
Googlebot Smartphone & Desktop, Bingbot and Applebot — status, block pages and what content they receive.
AI crawler access
AI search crawlers, user-triggered assistants and training bots, each judged separately.
CDN & WAF detection
Recognises Cloudflare, Fastly, Akamai, Imperva, Sucuri, DataDome and Wordfence blocks — even behind HTTP 200.
robots.txt, noindex, sitemap
Per-crawler robots.txt evaluation, meta robots and X-Robots-Tag, sitemap fetched as Googlebot.
Cloudflare settings
Connect a read-only token and we read AI Crawl Control and Bot Fight Mode directly — no guessing.
Multiple locations
The same checks from several countries, so a regional block does not stay invisible.
Most accidental blocks happen at the edge
A 2026 scan of the top 2,000 domains found about 9% of sites that allow AI crawlers in robots.txt but refuse them at the CDN. A robots-only checker reports them as fine.
Googlebot, GPTBot…
Requests your page
CDN · WAF · bot protection
AI Crawl Control, Bot Fight Mode, firewall rules, rate limits. Can refuse before your server sees anything.
Security plugins
Wordfence, mod_security, hosting firewalls
robots.txt · noindex
Only now do your own rules apply
Find out the day it breaks — not a month later
A plugin update, a new firewall rule, one toggle in Cloudflare. Rankings don’t fall for weeks, so nobody connects the drop to the change. BotCanary re-checks your sites on a schedule and alerts you in Telegram when a crawler gets locked out.
2 new problems detected
🔴 Googlebot is blocked together with AI crawlers on Cloudflare
🟠 The sitemap declared in robots.txt cannot be fetched by Googlebot
Questions
Isn’t checking robots.txt enough?
No. robots.txt is what you ask crawlers to do. The CDN, firewall and security plugins decide what actually happens. The most common accidental block is a site that allows a crawler in robots.txt and refuses it at the edge.
Can you send real Googlebot traffic?
Nobody outside Google can. We use the exact user-agents from our own servers and include a control bot to tell “the site verifies Googlebot” apart from “the site blocks bots”. When only the CDN, Search Console or logs can confirm a block, the report says so.
What changed at Cloudflare in September 2026?
Since 15 September 2026, choosing “Block” for AI training in Cloudflare also blocks mixed-use crawlers such as Googlebot, Bingbot and Applebot. “Disallow AI Training” keeps search access.
Is the check free?
One-off checks are free and need no account. Monitoring with alerts is a paid plan, paid in crypto, with no recurring charges.
Check your site in 10 seconds
Free, no account. 14 crawlers, one report.