How it works
Honest crawler testing: what we can see from outside, what we can’t, and how we tell them apart.
We read robots.txt like Google does
Fetch status first: a 5xx or timeout means Google stops crawling. Then each crawler is matched to its most specific user-agent group, and your URL is evaluated with longest-match rules and wildcards.
We request the page as 14 different clients
A normal Chrome visitor, an honest unknown bot (the control), Googlebot Smartphone and Desktop, Bingbot, Applebot, and the AI crawlers from OpenAI, Anthropic and Perplexity — each with its official user-agent.
We identify who answered
Response headers and block pages reveal the CDN or firewall: Cloudflare, Fastly, CloudFront, Akamai, Imperva, Sucuri, DataDome, Wordfence. Challenge pages are detected even when they return HTTP 200.
We compare the answers
Did the crawler get the same content as the visitor? Did the control bot pass where Googlebot failed? Does every AI crawler fail the same way? Patterns separate fake-bot protection from real blocks.
We tell you how sure we are
Every finding is labelled Confirmed (read directly from your files and headers), Likely (a strong pattern) or Needs confirmation (only your CDN, Search Console or logs can prove it).
Optional: we read your Cloudflare settings
With a read-only token we fetch AI Crawl Control and Bot Fight Mode configuration and name the exact setting behind a block.