Home / Research / Website AI readiness: what 5,542 business sites show
AI Search Intelligence

Website AI readiness: what 5,542 business sites show

The short answer

In RankEcho's July 2026 one-request screening of 5,542 responding business homepages, 33.7% had no detected structured data and 10.9% fell below an initial-HTML visible-text threshold. These indicators do not identify rendering architecture or measure indexing, AI impressions, citations, or causation. No row-level dataset is currently published.

What we found

Two aggregate signals were recorded only among the 5,542 domains that returned usable homepage HTML in the collection request.

The structured-data measure recorded presence, not correctness, coverage, or whether a provider used the markup. The rendering heuristic recorded that the initial response lacked the amount of visible homepage text expected by the screening rule; it did not run every provider's crawler or renderer.

The study did not observe search indexing, AI impressions, answer inclusion, citations, traffic, or outcomes after a change. The table therefore reports prevalence in this sample, not the impact of either signal.

SignalSites affectedShare of sample
No structured data of any kind1,86933.7%
Initial HTML below RankEcho's visible-text threshold60210.9%
Sample analysed5,542100%

Why structured data matters for AI answers

Structured data can make declared page entities and attributes machine-readable when it is accurate and matches visible content. Presence alone says nothing about whether the markup is valid, complete, trusted, or used by an answer system.

Google states that AI Overviews and AI Mode have no additional technical requirements or special AI schema. A page must meet normal Search eligibility and snippet controls, and even then appearance is not guaranteed. This study did not test whether structured data changed any provider outcome.

Treat missing markup as a review prompt, not a failure verdict. Add only schema that truthfully describes visible content and validate it under the requirements of the search feature or consumer you actually target.

What does the initial-HTML signal mean?

A useful initial HTML response improves compatibility with clients that do not execute the site's JavaScript. Other crawlers may render JavaScript, cache rendered output, or use a search index, so the screening signal cannot stand in for a provider-specific fetch.

The threshold measures extracted visible text in one initial response. It cannot distinguish server rendering from static generation or prerendering, and a sparse server-rendered page can fall below it. Inspect the title, canonical, headings, main answer, and directives, then test with the documented provider workflow.

How this was measured

7,126 business domains were harvested from public business listings across US regions. Each homepage was requested once with an identifying user agent, following redirects, with a nine-second timeout. Responses that returned usable HTML were screened for structured-data presence, extracted visible text in the initial HTML, crawler policy, and agent-guidance files.

1,584 domains did not return usable content and were excluded rather than counted. That group mixes refused requests, timeouts, redirect loops, and other failures that one request cannot separate. Every percentage on this page therefore uses 5,542 responding homepages as its denominator.

The analysis is a single July 2026 homepage snapshot. RankEcho currently publishes the aggregate counts and this method description, not the row-level domain sample, request records, screening code version, or raw classifications. Readers cannot independently reproduce the aggregates from a public artifact.

What this does not show

The population is local and small-business websites drawn from public business listings, weighted toward US regions. It is not a sample of B2B SaaS, ecommerce, publishers, or enterprise sites, and the figures should not be read as applying to them. No category comparison was performed.

Structured data presence was measured, not correctness. A site with invalid or contradictory markup still counts as having structured data. The initial-HTML threshold is a RankEcho heuristic from one response, not a rendering-architecture classifier or engine-eligibility test.

One request per homepage means transient failures are indistinguishable from persistent ones, which is part of why the excluded group is reported as excluded rather than characterised.

No AI provider impressions, citations, rankings, crawl frequency, or post-change outcomes were joined to these records. The results cannot establish that either indicator caused visibility or traffic differences.

What to do with this

Check your own high-value pages rather than assuming the sample applies. Separate robots policy, verified reachability, initial HTML, indexing directives, structured-data validity, and actual search or AI outcomes. RankEcho's tools provide heuristics and one-off observations, not provider certification.

Prioritize a correction when it fixes a verified user or crawler failure, then record the deployment and later outcomes. Do not promise citation movement from schema or rendering alone.

Sources reviewed

The prevalence figures on this page are first-party aggregate observations, not crawler-provider findings. The external records below bound how those signals may be interpreted. RankEcho has not published the row-level sample, so the aggregate counts are not independently reproducible from a public artifact.

2 claim-level source records
Checked 2026-09-01 · Aggregate-method and interpretation review · Confidence is recorded per claim.
Claim reviewedOfficial sourceReview record
Google says normal Search indexing and snippet controls govern eligibility for AI Overviews and AI Mode; there are no additional technical requirements or special AI schema files.Google Search AI features documentationChecked 2026-09-01 · AI features documentation updated 2025-12-10 · Primary-source documentation review; eligibility does not guarantee selection or presentation in an AI feature. · Confidence: High
RFC 9309 defines robots.txt matching, including merging multiple groups for the same user agent and using the most specific matching rule.Robots Exclusion Protocol, RFC 9309Checked 2026-09-01 · IETF standards-track RFC published September 2022 · Primary-standard review; simple policy checkers may not implement every URL-level matching case. · Confidence: High

Frequently asked questions

What percentage of websites have no structured data?

In this sample of 5,542 business websites analysed in July 2026, 33.7% carried no structured data of any kind on the homepage. The population is local and small-business sites harvested from public business listings across US regions, so the figure should not be generalised to B2B SaaS or enterprise sites.

Does structured data help with AI citations?

This study did not test that question. Structured data can make visible facts machine-readable, but Google says no special AI schema is required and eligibility does not guarantee appearance. Presence is not evidence of citation impact.

What did the initial-HTML screening measure?

It measured whether extracted visible text in one initial response met RankEcho's threshold. It did not identify server rendering, static generation, or prerendering, and it did not test provider eligibility or citation outcomes.

How was this study conducted?

7,126 business domains were harvested from public business listings across US regions. Each homepage was requested once with an identifying user agent and a nine-second timeout, then screened for structured-data presence, initial-HTML visible text, crawler policy, and agent-guidance files. 1,584 domains that did not return usable content were excluded, so every percentage uses 5,542 responding homepages.

Is the Website AI Readiness dataset public?

No row-level sample or request dataset is currently published. This page provides aggregate counts and method limits, so the figures are not independently reproducible from a public artifact.

How do I check my own site for these problems?

Use the crawler tool to parse robots policy and the page tool as a RankEcho heuristic, then verify genuine requests, indexing, and AI outcomes separately. Neither tool certifies provider eligibility or predicts citations.

See where AI ignores your brand — run a free audit →
Last updated 2026-09-01 · RankEcho · Operated by Nexus Decision Systems LLC