Crawlability vs indexability

Crawlability is whether a crawler can fetch a page; indexability is whether it may be added to the index. How robots.txt, noindex, 403s and canonicals affect each.

Crawlability: can the crawler fetch the page? Indexability: once fetched, is the page allowed and eligible to be stored and shown in results?

SignalAffectsExample
robots.txt DisallowCrawlabilityGooglebot never fetches the page
403 / CDN challengeCrawlabilityFetch fails at the edge
5xx on robots.txtCrawlability (whole site)Google pauses crawling
noindex meta / X-Robots-TagIndexabilityFetched, then excluded
Canonical to another URLIndexabilityAnother URL is indexed instead

The classic conflict

Disallowing a page in robots.txt and adding noindex does not work: the crawler can’t fetch the page, so it never sees the noindex. The URL can still be indexed from links.

Check both for any page with the Googlebot checker.

Crawlability vs Indexability: The Difference (with Examples)