Home / Website AI readiness: what 5,542 business sites show
AI Search Intelligence

Website AI readiness: what 5,542 business sites show

The short answer

Across 5,542 business websites fetched and analysed in July 2026, 33.7% carry no structured data of any kind and 10.9% return no server-rendered content on the homepage. Both are prerequisites for being quoted in an AI answer: structured data tells an engine what the business is, and server-rendered HTML is what a crawler reads when it does not execute JavaScript. The discussion of AI visibility concentrates on crawler blocking, which is comparatively rare in this sample. The more common barrier is extraction: engines can reach these sites and still find little they can confidently attribute. The population is local and small-business websites harvested from public business listings across US regions, which is stated here because it bounds what the figures apply to.

What we found

Two signals stand out, both measured only on sites that returned a usable homepage.

A third of the sample publishes no machine-readable identity at all. No Organization markup, no LocalBusiness, no Product, no FAQPage. An engine reading these sites has to infer what the business is, what it sells, and where it operates from prose alone, and inference is where attribution gets vague or wrong.

One site in nine returns no server-rendered content on the homepage. The page looks complete in a browser because JavaScript assembles it, but a crawler that does not execute scripts receives a shell. For citation purposes that is close to an empty page.

SignalSites affectedShare of sample
No structured data of any kind1,86933.7%
No server-rendered homepage content60210.9%
Sample analysed5,542100%

Why structured data matters for AI answers

Structured data is the difference between a page an engine can quote and a page it has to interpret. Organization and LocalBusiness markup state the entity, its category, and its location as data rather than as sentences. Product and Offer markup state what is sold. FAQPage markup pairs questions with answers in a form that maps directly onto how people prompt an assistant.

None of it guarantees a citation, and any page claiming otherwise is overselling. What it does is remove ambiguity. When two sources say similar things and one of them is unambiguous about who it is, the unambiguous one is easier to attribute, and attribution is the mechanism by which a brand appears in an answer at all.

The practical point is cost. Adding schema that matches visible page content is an afternoon of work. Earning the third-party corroboration that decides competitive category prompts takes months. A third of these sites have not done the afternoon.

Why server-rendered content matters

Crawlers vary in whether they execute JavaScript, and the ones that do are slower and more selective about it. A page whose text exists only after client-side rendering is a gamble on which crawler arrives.

The check is quick: request the page and read the raw response rather than the rendered view. If the answer text, the headings, and the product description are absent from that response, they are absent for any crawler that does not run scripts.

This overlaps with the structured-data gap more often than not. Sites assembled entirely by client-side frameworks frequently inject their markup at runtime too, which means an engine sees neither the content nor the entity data.

How this was measured

7,126 business domains were harvested from public business listings across US regions. Each homepage was requested once with an identifying user agent, following redirects, with a nine-second timeout. Responses that returned usable HTML were analysed for structured data types, server-rendered content, crawler policy, and agent-guidance files.

1,584 domains did not return usable content and were excluded rather than counted. That group mixes refused requests, timeouts, and redirect loops, and we do not report a figure for it because those causes cannot be separated reliably from a single request per domain. Excluding them means every percentage on this page has the same denominator: sites that answered.

The analysis is a single snapshot of the homepage, taken in July 2026.

What this does not show

The population is local and small-business websites drawn from public business listings, weighted toward US regions. It is not a sample of B2B SaaS, ecommerce, publishers, or enterprise sites, and the figures should not be read as applying to them. Categories with heavier engineering investment would be expected to differ in both directions: better rendering, and not necessarily better structured data.

Structured data presence was measured, not correctness. A site with markup that contradicts its visible content counts as having structured data here, and in practice that markup can do more harm than none.

One request per homepage means transient failures are indistinguishable from persistent ones, which is part of why the excluded group is reported as excluded rather than characterised.

What to do with this

Check your own site rather than assuming the averages apply. Three checks cover what this study measured and take a few minutes: whether crawlers can reach the site, whether a key page returns server-rendered content, and whether that page carries structured data matching what a reader sees.

If any of the three fails, fix it before investing in content or outreach. Those are the cheapest interventions available and they gate everything downstream: a page an engine cannot read cannot be cited regardless of how good the writing is, and a page it cannot attribute is unlikely to be quoted with the brand attached.

Frequently asked questions

What percentage of websites have no structured data?

In this sample of 5,542 business websites analysed in July 2026, 33.7% carried no structured data of any kind on the homepage. The population is local and small-business sites harvested from public business listings across US regions, so the figure should not be generalised to B2B SaaS or enterprise sites.

Does structured data help with AI citations?

It removes ambiguity rather than guaranteeing anything. Markup that states the entity, category, offering, and question-answer pairs as data makes a page easier to attribute than one an engine has to interpret from prose. No markup produces a citation on its own, and any tool claiming otherwise is overselling.

What does server-rendered mean and why does it matter for AI?

Server-rendered means the page content is present in the HTML the server returns, rather than being assembled by JavaScript in the browser. Crawlers vary in whether they execute JavaScript, so content that only appears after rendering may be invisible to the crawler that arrives. In this sample, 10.9% of homepages returned no server-rendered content.

How was this study conducted?

7,126 business domains were harvested from public business listings across US regions. Each homepage was requested once with an identifying user agent and a nine-second timeout, then analysed for structured data, server-rendered content, crawler policy, and agent-guidance files. 1,584 domains that did not return usable content were excluded, so every percentage uses the same denominator of sites that answered.

How do I check my own site for these problems?

Three free checks cover what was measured here: a crawler policy check for access, an agent readiness check for rendering and structure on a specific page, and a citation check for whether an engine actually cites you. None require an account and each takes under a minute.

See where AI ignores your brand — run a free audit →
Last updated 2026-07-23 · RankEcho · Operated by Nexus Decision Systems LLC