Home / Resources / Find / AI crawler blocked by robots.txt: issue diagnosis
AI Search Intelligence

AI crawler blocked by robots.txt: issue diagnosis

The short answer

Record this issue only when the production robots.txt response contains a rule that blocks the exact URL for the documented crawler role under RFC 9309 matching. Preserve the response, collection time, user-agent group, matched path, and provider purpose. A copied User-Agent request is only a path test, and a policy block does not by itself explain indexing, retrieval, ranking, a mention, recommendation, or citation.

Scope box: what does this page cover?

One production robots.txt response, one exact URL, and one documented crawler role at a recorded time.

A matching disallow rule supports a policy diagnosis only; it does not prove why an answer omitted a page.

In scopeOut of scope
Declared robots.txt policyCrawler identity from a copied User-Agent
RFC 9309 matchingCDN or WAF enforcement
Provider-specific crawler rolesIndexing or retrieval
Exact-path verificationRanking or citation causation

What is the definition and observable symptom?

The issue is a declared-policy conflict: the current production robots.txt selects a documented crawler group and its most-specific matching rule disallows the exact public URL the team intends that crawler to access. The observable symptom is the served policy and its deterministic match, not a missing answer-system citation.

Name the provider role precisely. OpenAI, Anthropic, Perplexity, Bing, and Google document different search, training, user-triggered, or Search controls. A block for one role cannot be generalized to every role or product.

Why might the issue matter?

A disallow rule can prevent a compliant crawler from requesting the covered path for the documented purpose. That makes the policy relevant when the intended public surface depends on that crawler role.

Permission is only one boundary. A page may still fail at DNS, authentication, CDN, WAF, redirect, response, rendering, canonical, indexing, retrieval, or selection stages; an intentional block may also be the correct security or licensing decision.

What is the exact trigger and detection logic?

Create the issue only when every required check is supported and attached. The linked RankEcho checker is a preliminary root-policy heuristic; it does not evaluate the exact target URL, group merging, wildcard or end-anchor behavior, or the most-specific rule, so complete those RFC checks separately.

  • Fetch the root-level production robots.txt at a recorded UTC time and preserve status, redirects, headers, and body.
  • Select the provider-documented crawler identity and purpose that match the intended surface; do not substitute a training or user-triggered agent.
  • Apply RFC 9309 group merging and most-specific path matching to one exact canonical URL.
  • Record the selected group, matched allow or disallow rule, match length, and final policy result.
  • Require a blocking result for the exact URL. A missing file, an allowed result, or an unrelated disallow is not this issue.

How should severity and affected scope be recorded?

Severity follows the verified business scope and intended crawler purpose, not the emotional weight of the phrase ‘AI crawler blocked’.

LevelUse whenAffected scope to record
ReviewThe block is intentional, ambiguous, or limited to a non-target roleCrawler role, rule, and rationale
BoundedOne intended public page or directory is blockedExact canonical URLs and matched rule
BroadThe selected group blocks all intended public pathsHost, group, representative URLs, and exclusions

What are false positives and non-issues?

Do not create this issue for an intentional policy decision, a training bot when only search access is in scope, a user-triggered fetcher whose documented robots behavior differs, an unrelated path rule, or a checker that ignored group merging and specificity.

A 403, challenge page, or empty response can be a CDN, WAF, authentication, or application issue rather than robots.txt. A curl command with a copied User-Agent does not verify genuine provider identity. A missing citation without a matching policy block is not this issue.

What does a worked example look like?

Illustrative synthetic example—this is not customer data or evidence of provider behavior. At 09:00 UTC a team fetches https://example.org/robots.txt and preserves the response. The file has a group for a documented example search crawler with ‘Disallow: /docs/’ and no more-specific allow. RFC matching classifies https://example.org/docs/widget as disallowed for that role.

The issue record names only that URL, policy version, role, rule, and time. It retains WAF enforcement, indexing, retrieval, prompt fit, entity clarity, and answer variability as separate questions. It does not claim the block caused the synthetic brand's absence from an answer.

Which primary sources bound the diagnosis?

Use RFC 9309 for robots matching and current provider documentation for crawler names, purposes, and control boundaries. The dated claim-level records below show what was checked and the interpretation limit.

Provider pages can change. Recheck them at implementation time, record the date, and avoid extending one provider's rules to another.

What are the correct fix options?

If the block is intentional, document it and close the issue without changing policy. If it is unintended, change only the selected group and public path needed for the approved purpose, preserving private and unrelated exclusions. If the robots policy allows the URL, investigate reachability, edge enforcement, indexing, extraction, source coverage, prompt fit, and entity clarity instead.

How is deployment validation performed?

After an approved change, fetch the production policy again, attach status, redirects, headers, body, and deployment version, then re-run matching for the same crawler role and exact URL. Separately verify DNS, final response, redirect chain, CDN/WAF events, authentication, and useful HTML.

Passing these checks validates the deployment and removes this diagnosed policy barrier. It does not validate indexing, ranking, retrieval, or any answer outcome.

How is the stable-panel retest and proof artifact recorded?

Preserve the pre-change prompt panel and raw observations, record the deployment time, then repeat the same prompts on the same declared surfaces or configured adapters under the same matching and exclusion rules. Report planned, unavailable, failed, and observed cells separately.

Label later movement as observed after the shipment. No movement and adverse movement remain reportable results, and neither movement nor stasis proves that robots.txt was the cause.

Who owns the page, review, tests, and corrections?

Author and accountable publisher: Abiot Y. Derbie, RankEcho founder. Published 2026-09-02; updated 2026-09-02. Independent reviewer: unassigned. No independent external review is claimed.

Tested scope: Static content contract, link targets, RFC/provider evidence mapping, and synthetic issue example; no live provider crawl was performed for this page. Material provider and standards statements use the dated source records below. Examples are synthetic, not customer results. Send corrections through /contact; material corrections should update the visible date and version history.

Sources reviewed

Material crawler-role and control claims below were checked against primary provider documentation and the robots standard. Access settings affect eligibility and reachability; they do not guarantee indexing, ranking, an AI impression, or a citation.

9 claim-level source records
Checked 2026-09-01 · Primary-source technical review · Confidence is recorded per claim.
Claim reviewedOfficial sourceReview record
RFC 9309 defines robots.txt matching, including merging multiple groups for the same user agent and using the most specific matching rule.Robots Exclusion Protocol, RFC 9309Checked 2026-09-01 · IETF standards-track RFC published September 2022 · Primary-standard review; simple policy checkers may not implement every URL-level matching case. · Confidence: High
OpenAI documents OAI-SearchBot for ChatGPT search, GPTBot for potential model training, and ChatGPT-User for user-triggered actions; the controls are independent and robots.txt rules may not apply to ChatGPT-User.OpenAI crawler documentationChecked 2026-09-01 · Current OAI-SearchBot, GPTBot, and ChatGPT-User documentation · Primary-source documentation review; no claim that a permitted bot will index, rank, or cite a page. · Confidence: High
Anthropic documents ClaudeBot for potential model training, Claude-SearchBot for search indexing, and Claude-User for user-directed retrieval; Anthropic says all three honor robots.txt.Anthropic crawler documentationChecked 2026-09-01 · Crawler taxonomy updated 2026-04-07 · Primary-source documentation review; crawler permission is treated as access policy, not citation eligibility. · Confidence: High
Perplexity documents PerplexityBot for search discovery and Perplexity-User for user-triggered fetching; Perplexity-User generally ignores robots.txt and is not a crawl or training bot.Perplexity crawler documentationChecked 2026-09-01 · Current PerplexityBot and Perplexity-User documentation · Primary-source documentation review; Perplexity recommends checking both the user agent and published IP ranges. · Confidence: High
Microsoft lists Bingbot as Bing's main web crawler. Permitting Bingbot supports Bing discovery but does not guarantee indexing, an impression, or a citation.Bing crawler documentationChecked 2026-09-01 · Current Bing crawler documentation · Primary-source documentation review; Bingbot access and index eligibility do not guarantee a Copilot citation or impression. · Confidence: High
Google says normal Search indexing and snippet controls govern eligibility for AI Overviews and AI Mode; there are no additional technical requirements or special AI schema files.Google Search AI features documentationChecked 2026-09-01 · AI features documentation updated 2025-12-10 · Primary-source documentation review; eligibility does not guarantee selection or presentation in an AI feature. · Confidence: High
Google documents Google-Extended as a robots.txt control for some Gemini Apps, Vertex AI grounding, and future model training uses; it has no separate HTTP user agent and does not affect Google Search inclusion or ranking.Google common crawler documentationChecked 2026-09-01 · Google-Extended documentation updated 2026-07-14 · Primary-source documentation review; Google-Extended is not represented as the control for Google Search AI features. · Confidence: High
A crawler user-agent string can be spoofed; Google documents reverse-DNS and published-IP methods for verifying genuine Google crawlers.Google crawler verification documentationChecked 2026-09-01 · Current verified-crawler guidance · Primary-source documentation review; a curl request that only changes User-Agent is a path test, not bot-identity verification. · Confidence: High
Cloudflare documents that a crawler allowed by AI Crawl Control can still be blocked by a custom WAF rule, so rule order and the matching security event must be reviewed.Cloudflare AI Crawl Control with WAFChecked 2026-09-01 · Current AI Crawl Control and WAF configuration guidance · Primary-source documentation review; the result depends on the zone's actual rules and request evidence. · Confidence: High

Frequently asked questions

Does a robots.txt block prove why an AI answer omitted my page?

No. It proves only the declared policy match for the documented crawler role and exact URL at the recorded time. Reachability, identity, indexing, retrieval, selection, and answer variability remain separate.

Can I test a crawler by changing curl's User-Agent?

That can test how the path responds to a string, but it cannot prove the request came from the provider. Use the provider's documented IP or DNS verification method where available and preserve genuine request evidence.

Should every AI crawler be allowed?

No. Access is a publisher policy and security decision. Identify the crawler's documented purpose, intended public scope, licensing and privacy requirements, then make the smallest approved change—or retain the block.

What happens after the policy is fixed?

Validate the served policy and delivery path, then repeat the registered prompt panel. Report any later outcome as a temporal observation, not a guaranteed or causal result.

Screen the root robots policy →
Last updated 2026-09-02 · RankEcho · Operated by Nexus Decision Systems LLC