Crawlability: can the crawler fetch the page? Indexability: once fetched, is the page allowed and eligible to be stored and shown in results?
| Signal | Affects | Example |
|---|---|---|
| robots.txt Disallow | Crawlability | Googlebot never fetches the page |
| 403 / CDN challenge | Crawlability | Fetch fails at the edge |
| 5xx on robots.txt | Crawlability (whole site) | Google pauses crawling |
| noindex meta / X-Robots-Tag | Indexability | Fetched, then excluded |
| Canonical to another URL | Indexability | Another URL is indexed instead |
The classic conflict
Disallowing a page in robots.txt and adding noindex does not work: the crawler can’t fetch the page, so it never sees the noindex. The URL can still be indexed from links.
Check both for any page with the Googlebot checker.