RankEcho
Free tool - no account

AI crawler policy check

Paste a domain to parse what its current robots.txt declares for RankEcho's built-in AI bot list. This is a policy heuristic, not a test of CDN or WAF enforcement, genuine bot identity, indexing, AI impressions, or citations.

Read the result as a robots.txt hypothesis. This checker uses a simplified, root-level robots.txt model. Some built-in labels group user-action and control tokens for diagnostic convenience; those labels are RankEcho terminology, not provider definitions. Current provider roles and robots caveats are documented below. Confirm any decision against the provider source and the exact target URL.

Fetch-failure limitation. Open the JSON report and inspect robotsStatus. If it is 0 or 500-599, no robots.txt result is available. Ignore every allowance label and score because the missing response may otherwise appear as if all listed crawlers are allowed. RFC 9309 treats an unreachable robots file as complete disallow initially. Retest and inspect the origin or edge failure.

Free and anonymous. Rate-limited to 20 checks per hour per IP. We fetch only /robots.txt.

Full heuristic report (JSON)

Important: labels such as "retrieval readiness" and "citation-critical" are RankEcho heuristic labels, not provider-defined eligibility or outcome measures. Verify the exact robots rule, genuine request, edge event, response body, indexing state, and later AI outcome separately.

Next: follow the current crawler-role and verification guide, or inspect one page with the page-readiness heuristic.

What does this crawler policy check actually test?

It fetches /robots.txt and applies RankEcho's built-in policy model. It does not request a representative content URL as an authenticated provider crawler, evaluate your Cloudflare account, inspect WAF rules, read server logs, or ask an AI engine whether the page was selected.

Always inspect robotsStatus in the JSON report. Status 0 or a 5xx is a fetch failure, not evidence that no robots.txt exists. No policy conclusion is available from that response; discard the allowance labels and score until the file can be fetched reliably.

The parser is useful for finding a possible root-policy conflict. RFC 9309 also defines exact-group merging, most-specific path matching, wildcard and end-anchor behavior, and URL-level evaluation. Review those cases manually before changing production policy.

Which bot roles should I separate?

Search discovery, model training, user-triggered actions, and use-control tokens are different jobs. OpenAI separates OAI-SearchBot, GPTBot, and ChatGPT-User, and says robots rules may not apply to ChatGPT-User. Perplexity says Perplexity-User generally ignores robots.txt. Anthropic says its current ClaudeBot, Claude-SearchBot, and Claude-User tokens honor robots.txt.

Googlebot and normal Search controls govern Google AI Overviews and AI Mode. Google-Extended is a control token for other Gemini and Vertex uses, has no separate HTTP user agent, and does not control Google Search inclusion or ranking.

Policy is not proof of access.

Free audits check Perplexity and Gemini once per prompt. Paid and trialing accounts add ChatGPT, Claude, and Google AI Overviews when configured for the account, for up to 5 engines. The audit records observed answers; it does not turn a robots rule into a citation promise.

How do I verify a real crawler request?

Evaluate the exact bot and URL, inspect CDN and WAF events, verify identity with official IP or reverse-DNS guidance, and reconcile edge and origin logs. Then inspect status, final URL, headers, canonical, directives, and useful HTML body. A bot-shaped User-Agent can be spoofed, and a 200 can still contain a challenge or empty shell.

Why can Cloudflare disagree with robots.txt?

Robots.txt records a preference. Cloudflare AI Crawl Control and WAF rules enforce requests at the edge, and another WAF rule can still block a crawler that AI Crawl Control allows. Cloudflare documents separate Search, Agent, and Training presets; inspect the zone's current choices, rule order, security event, and origin request instead of assuming a universal default. The updated defaults scheduled for September 15, 2026 were not yet effective on this September 1 review date.

Use the Cloudflare diagnostic workflow for the full evidence ladder.

What this check cannot conclude

It cannot conclude that a genuine crawler reached a page, that the page was indexed, that an AI surface recorded an impression, or that any answer will cite it. An allowed rule is one input into some providers' discovery process; provider selection and answer variation remain outside the tool.

Primary sources reviewed

Checked 2026-09-01: OpenAI bots, Anthropic crawlers, Perplexity crawlers, Google AI features, Cloudflare AI Crawl Control, Cloudflare AI bot policies, Cloudflare managed robots.txt, Cloudflare WAF ordering, RFC 9309, and Google crawler verification.

Frequently asked questions

Does this tool test Cloudflare or my WAF?

No. It fetches /robots.txt. Inspect Cloudflare Security Events, AI Crawl Control, rule order, and origin logs separately.

Does an allowed result mean the bot can reach my site?

No. It means the parser found an allowing policy under its model. Edge enforcement, genuine identity, target-path matching, response content, and indexing are separate.

Does blocking GPTBot remove a site from ChatGPT Search?

OpenAI documents GPTBot for potential training and OAI-SearchBot for ChatGPT search. The controls are independent.

Does ChatGPT-User obey robots.txt?

OpenAI says robots.txt rules may not apply to ChatGPT-User because it handles user-triggered requests. Do not use this parser's label as a reachability verdict for that token.

Does Google-Extended control AI Overviews?

No. Google says normal Search controls through Googlebot govern AI Overviews and AI Mode. Google-Extended does not affect Search inclusion or ranking.

Explore the other free checks:

Turn the diagnosis into a shipped change: use the RankEcho Resources library for Find, Fix, and Prove playbooks, checklists, and templates.