Paste a domain and see which AI crawlers your robots.txt blocks or allows - split into the retrieval bots that power citations (where blocking costs you visibility) and the training bots you may deliberately block.
Most robots.txt mistakes in AI search come from blocking the wrong bot. Training crawlers such as GPTBot, ClaudeBot, and CCBot collect data for model training; retrieval crawlers such as OAI-SearchBot, ChatGPT-User, PerplexityBot, and Claude-User fetch pages for live answers and citations. Blocking training bots is a rights decision; blocking retrieval bots removes you from AI answers. This check shows which is which for your domain.
Free and anonymous. Rate-limited to 20 checks per hour per IP. We fetch only /robots.txt.
Crawler access is one of three agent-readiness pillars. Next: run the agent readiness check on a key page, or run the full free AI visibility audit to see whether engines actually cite you.
It is not a yes-or-no question - it is three questions. Training crawlers harvest content for model training and send nothing back. Retrieval crawlers fetch pages so engines can cite them in live answers. User-triggered agents fetch a page because a person asked about it.
If you sell content itself - journalism, research, courses - reserving training rights can be rational. If you sell products or services, blocking retrieval bots mostly means the engines recommend your competitors instead of you.
GPTBot collects training data for OpenAI's models. OAI-SearchBot indexes pages for ChatGPT's live search answers, and ChatGPT-User fires only when a person asks ChatGPT to open a page. Blocking one has no effect on the others - and per OpenAI's own documentation, sites opted out of OAI-SearchBot are not shown in ChatGPT search answers at all.
The same split exists elsewhere: Anthropic runs ClaudeBot for training and Claude-User and Claude-SearchBot for retrieval, and Google-Extended is a training-only token that never touches Google Search ranking.
This check reads your robots.txt policy. The free audit asks the engines themselves - real buyer prompts across ChatGPT, Claude, Perplexity, Gemini, and Google AI - and shows where you are cited and where you are invisible.
No. Googlebot is a separate crawler, Google states that Google-Extended does not affect Search inclusion or ranking, and publisher-network analyses reviewed by Playwire found no ranking impact from blocking GPTBot. The cost of blocking is paid in AI answers, not in blue links - which is exactly why the decision deserves to be deliberate.
Because robots.txt is only one layer. CDN bot rules, WAF defaults, and firewall lists block AI user agents even when robots.txt allows them - research from ziptie.dev puts accidental CDN-level blocking near a quarter of B2B SaaS and e-commerce sites.
The tell is in your server logs: if OAI-SearchBot, PerplexityBot, or Claude-User never appear, or appear only as 403s, something upstream of robots.txt is turning them away.
Crawler-policy corrections ship as copy-paste artifacts in the Fix Engine - see how a fix package works.
It reads policy, not outcomes. An open robots.txt does not mean engines cite you, and a perfect file cannot repair extraction or grounding problems.
The free AI visibility audit measures the outcome side: whether engines actually cite your domain on buyer prompts, which sources they lean on instead, and which layer - access, extraction, entity clarity, or source coverage - is costing you the citation.
No. Separate them by role first. Blocking retrieval and user-triggered bots removes you from AI answers, while blocking training bots is a content-rights decision that does not affect Google rankings. For products and services, blocking retrieval mostly means competitors get recommended instead.
Reputable operators honor it, but the file is voluntary - it expresses preference, not enforcement. Enforcement lives in CDN and WAF rules, and those same rules frequently block AI bots by accident, so check both layers.
The retrieval and user-triggered set: OAI-SearchBot, ChatGPT-User, PerplexityBot, Claude-User, and Claude-SearchBot. Decide separately, per your content strategy, on the training set: GPTBot, ClaudeBot, Google-Extended, and CCBot.
No. Google-Extended is a training control token, not a search crawler, and Google states it does not affect inclusion or ranking in Search. AI Overviews are governed by the normal Search controls such as nosnippet and noindex.
After any CDN, WAF, robots.txt, or platform migration - and on a recurring schedule, because AI companies add and rename crawlers. Anthropic, for example, consolidated earlier bot names into ClaudeBot.
Explore the other free checks: