AI Crawlability Checklist: 25 Things to Check on Your Website
A practical preflight checklist for making public pages accessible to the AI crawlers you intentionally support.
AI Crawlability Checklist: 25 Things to Check on Your Website
Use this checklist when launching a new site, publishing an important page, or auditing AI search accessibility.
Access and HTTP
- The target URL is publicly reachable.
- The canonical URL returns a successful HTTP response.
- There are no redirect loops.
- Redirect chains are intentional and short.
- Important pages do not return intermittent 5xx errors.
- Rate limiting does not unnecessarily reject legitimate crawlers.
- Public pages do not require authentication.
robots.txt
- robots.txt is available at the site root.
- Important public paths are not accidentally Disallowed.
- Crawler-specific rules are intentional.
- Wildcard rules have been reviewed.
- AI crawler policies match your business goals.
- robots.txt is not being used as a security mechanism.
Security and infrastructure
- WAF rules do not accidentally block legitimate crawlers.
- Bot-management challenges have been reviewed.
- CDN policies do not create crawler-only failures.
- Geo restrictions do not unintentionally block intended access.
- Server logs can identify relevant crawler requests.
Rendering and content
- Important content is present in accessible HTML.
- Critical information does not depend on fragile client-side execution.
- Headings describe the document structure clearly.
- Internal links expose important pages.
- Canonical URLs are consistent.
- Indexing directives match the publishing strategy.
- Critical pages are re-tested after infrastructure changes.
How to prioritize failures
Not every warning deserves the same response.
A 403 on your most important public guide is high priority.
A minor metadata warning on a low-value page is usually lower priority.
Use a simple framework:
Critical: crawler cannot access the target page.
High: access is intermittent or dependent on fragile infrastructure.
Medium: discovery or rendering is unnecessarily complex.
Low: informational warning with no demonstrated access impact.
Build an ongoing process
AI crawlability should not be a one-time audit.
Re-check when you:
- change robots.txt
- migrate hosting
- change CDN providers
- deploy a new WAF policy
- introduce authentication
- rebuild your frontend
- change rendering architecture
- launch important content
Final takeaway
A good AI crawlability checklist turns a vague AI-search concern into concrete engineering checks.
Run the checklist on your highest-value URLs first, fix actual blockers, and re-test.