Evidence hub · audit aggregate withheld

AI Visibility Benchmark: Method, Status, and Limits

One route, four different evidence jobs. This page identifies what RankEcho can publish now, what remains a technical screen, and what is still blocked by cohort and review requirements.

Publication status

RankEcho does not yet publish a general industry AI-visibility benchmark. Stored audits differ in prompt panels, configured providers, sample depth, observed denominators, and collection dates, so pooling them would imply comparability the current feed does not support. This hub publishes one dated citation-source study and the rules a future category benchmark must pass.

Public cross-audit aggregate
Withheld
Legacy 20-run display gate
met

The legacy overall gate is 20 completed audit records and the legacy industry-row gate is 5. These are preserved in runtime for compatibility. They are display rules, not evidence that records are unique domains, comparable, representative, or statistically sufficient.

Separate the instruments before comparing numbers

SurfaceObservation unit and frameWhat it can answerStatus
Q3 citation studyOne scheduled prompt × configured answer-surface × run cell; skipped cells remain disclosed.Cited-domain-set overlap inside the completed portion of that fixed panel.Published with CSV analysis export
Cross-audit aggregateOne completed audit run; repeat domains and unlike run configurations are not removed.Nothing public until a comparable cohort and review artifact pass the gate below.Withheld
AI-readiness directoryOne homepage-centred RankEcho-UA screen plus robots.txt and llms.txt checks, from a convenience and submission sample.RankEcho's observed technical-screen signals, not citations or recommendations.Separate directory
Rolling state reportThe evolving audit-run stream plus proof-loop summaries.Descriptive product observations; not the fixed Q3 panel or a census.Separate rolling report

The publishable fixed-panel asset

The State of AI Citations, Q3 2026 is the only citation analysis artifact distributed from this hub. Its design arithmetic is 48 predeclared prompts × 4 configured answer surfaces × 2 scheduled runs = 384 planned cells. Repository metadata reports the second run was scheduled 24 hours later; the CSV does not carry timestamps that independently verify that interval.

384
planned prompt × surface × run cells
368
cells marked non-skipped; 16 skipped (11 run 0, 5 run 1)
0.195
run-0 conditional pairwise Jaccard mean, n=181 same-prompt pairs where both had a publisher domain
0.671
mean repeat-run Jaccard similarity, n=178 matched non-skipped pairs; empty/empty is defined as 1

The configured surfaces were Claude through Anthropic's Messages API with its web-search tool, Perplexity API, Gemini API with Google Search grounding, Google AI Overviews through SerpApi. They are API and search integrations, not interchangeable consumer products. The panel had no OpenAI/ChatGPT observation, covered US-English commercial prompts, and scheduled only two runs on 2026-07-09 and 2026-07-10. A stability value of 0.671 is a mean set-similarity statistic; its 0.329 distance is not a literal percentage of individual citations that changed. The mean includes 23 Google-AI-Overview empty/empty pairs scored as similarity 1; the report's separate 0.617 Google-AI-Overview figure uses 15 source-bearing pairs. This is a bounded first-party probe with a normalized-domain export, not a raw-answer archive, market census, or claim about every provider, geography, category, or date.

Read the full method and findings · Download the CSV · CC BY 4.0

What the current audit stream actually contains

The executing aggregate reads up to the latest 5,000 rows marked done with a non-null scorecard. The unit is a completed audit run, not a unique domain. The projection used here carries industry and scorecard JSON, then skips malformed JSON; it supplies no domain key or finished timestamp for deduplication and calendar-window reporting.

  • Runs can mix sitewide and page scopes, custom and generated prompts, free and paid provider sets, sampling depths, repeats, skipped observations, and answer-reuse rules.
  • Industry is an inferred free-text label or “Other,” not a validated sampling stratum or stable public taxonomy.
  • The observed prompt denominator can vary because an audit can contain unavailable or skipped engine cells. No common cross-audit prompt-by-engine panel is enforced.

Executing metric meanings

Prompt-level brand appearance rateA prompt is positive when any successful engine result carries citedTarget. In the executing parser that flag can mean a target URL citation or a target-name text mention. Mention is not source citation.
Equal-run headline meanThe overall rate is the unweighted mean of each included audit run's rate. A two-prompt run and a ten-prompt run each receive one equal vote.
Non-owned all-URL occurrence shareThe current off-site calculation is one minus owned source occurrences divided by all surfaced URL occurrences. It can include competitor, infrastructure, and unrelated sources; it is not endorsement or third-party corroboration of the audited brand.
Legacy row thresholdsTwenty included runs unlock the legacy overall renderer and five runs unlock a legacy industry row. Threshold does not mean representative, comparable, reviewed, or reliable.
Worked synthetic denominator example · not customer data
RunPositive promptsObserved promptsRun rate
Audit A121 ÷ 2 = 50.0%
Audit B6106 ÷ 10 = 60.0%

The current equal-run mean is (50.0% + 60.0%) ÷ 2 = 55.0%. Pooling prompt counts would produce 7 ÷ 12 = 58.3%. The current aggregate uses the first calculation. Neither result is a population rate without a defined cohort and sampling frame.

The minimum gate for a future category benchmark

This is a pilot publication rule, not a universal statistical law. Every category cut must satisfy all rows together, and passing them still does not make a convenience sample representative.

GateMinimumRequired disclosure
Eligible population≥ 30 unique domains per categorySampling frame, eligibility, deduplication, repeat-domain handling
Prompt panel≥ 5 predeclared prompts per categoryFull prompt text, prompt class, additions and removals
Surface coverage≥ 3 named surfaces or modelsProvider surface, model/version where known, failures and skips
Repeated measurement≥ 3 dated windowsCollection dates, interval, answer-reuse and cache policy
Usable observations≥ 100 domain × prompt × surface observations after exclusionsPlanned, attempted, skipped, excluded, successful, and observed denominators
Publication artifactRequiredDownloadable aggregate, uncertainty, limitations, independent review status, version and correction log
Current status: blocked. The present cross-audit projection cannot demonstrate these conditions, so no overall rate or industry table is rendered here—even when its legacy row threshold is met.

Readiness is a separate technical screen

The AI-readiness directory uses a homepage-centred screen plus robots.txt and llms.txt requests made with RankEcho's user agent, not audit answers or a provider's crawler. Its current composite weights are 30% structured data, 25% crawl access, 25% machine readability, and 20% agent guidance. Bands begin at 85, 65, and 40.

The directory mixes harvested business-listing domains with submissions, counts successful captures, and does not expose a request-time freshness window on this hub. Category and category-by-metro display gates of 30 and 20 sites are coverage rules, not statistical sufficiency. Metro describes a harvest-source location, not AI-answer request geography or personalization.

Do not infer: readiness is not AI-answer visibility; a crawler or homepage signal is not indexing, citation, recommendation, traffic, or evidence that one technical feature caused an outcome.

Six inferences this hub rejects

Threshold ≠ representativeA minimum count does not repair a convenience sample or establish inference to a wider market.
Records ≠ unique domainsRepeat audits can create multiple rows for one site unless a cohort explicitly deduplicates them.
Mention ≠ source citationA brand named in answer text is a different event from a displayed or attributed source URL.
Non-owned URL ≠ corroborationA source outside the target domain can belong to a competitor, infrastructure layer, or unrelated publisher.
Aggregate difference ≠ causeChanges across unlike runs or dates do not prove that one page, fix, or signal caused the difference.
Readiness ≠ AI visibilityA RankEcho technical screen is not a provider's eligibility rule, ranking system, citation result, or recommendation outcome.

Choose the evidence surface that matches the job

Sources and provenance reviewed

These first-party sources support the public study, artifact, terminology, screening limits, and no-guarantee boundary. The audit-stream definitions above were also checked against the executing database selection, scorecard, citation parser, and aggregate code on 2026-09-02.

Open the five-record evidence ledger
  • The State of AI Citations, Q3 2026 The Q3 study scheduled 384 prompt-by-surface-by-run cells from 48 prompts, four configured answer surfaces, and two runs repository metadata reports as 24 hours apart; 368 cells are marked non-skipped and 16 skipped. Its prompt battery, method, limits, and normalized-domain CSV analysis export are public. Checked 2026-09-02 · Published 2026-07-10; fixed Q3 2026 study · First-party study review; the fixed panel is not merged with later customer audit runs or homepage screens.
  • Q3 2026 normalized-domain analysis export (CSV) The Q3 normalized-domain analysis export is downloadable as CSV under CC BY 4.0. It contains prompt ID, category, intent, adapter/surface key, run index, skip flag, distinct normalized domain, heuristic source type, and zero-based export order among the cell's distinct domains. It omits prompt text, raw answers, original URLs and titles, model/version, timestamps, skip reasons, answer length, and session or geography configuration. Checked 2026-09-02 · Published Q3 2026 CSV analysis export · Artifact review; normalized citation-domain rows are not a raw-answer archive, and availability does not make the US-English probe a census or universal provider benchmark.
  • RankEcho measurement methodology RankEcho's methodology treats citation, mention, recommendation, source mix, and re-test movement as distinct outcomes and rejects guaranteed placement or causal claims. Checked 2026-09-02 · Methodology updated August 4, 2026 · First-party methodology review; the hub additionally follows the executing audit code where legacy labels differ from these definitions.
  • Website AI Readiness study and limits The published website-readiness study reports a bounded homepage-screening sample and explicitly does not measure indexing, AI impressions, citations, traffic, or causal effects. Checked 2026-09-02 · July 2026 one-request homepage screening · First-party study review; the dated study and the evolving readiness directory are not represented as one population.
  • RankEcho AI visibility disclosure RankEcho's disclosure states that observed visibility can vary by engine, model, retrieval state, time, location, account state, and prompt wording, and that no citation or business outcome is guaranteed. Checked 2026-09-02 · Current public measurement disclosure · First-party disclosure review; a threshold crossing or later observation is not converted into a causal or outcome guarantee.

Benchmark FAQ

Which RankEcho benchmark data can I download today?

The Q3 2026 citation study has a downloadable normalized-domain analysis export: 48 prompts across four configured answer surfaces in two runs repository metadata reports as scheduled 24 hours apart. Of 384 planned cells, 368 are marked non-skipped and 16 skipped. The export omits prompt text, raw answers, URLs, model versions, timestamps, and skip reasons, and the cross-audit aggregate and readiness directory do not share this study's sample.

Why are cross-audit citation percentages withheld here?

The current feed contains completed audit runs with varying scopes, prompt panels, engine configurations, sampling depths, skipped observations, and repeat domains. It does not give this page a deduplicated cohort, common calendar window, or common observed denominator.

Does the 20-run legacy display threshold make an aggregate reliable?

No. Twenty runs and five runs per industry are legacy product display gates. They do not establish unique-domain coverage, representativeness, uncertainty, a common panel, or statistical validity, so crossing them does not publish results on this hub.

Does an AI-readiness score predict citations or recommendations?

No. The readiness directory is a homepage-centred screen plus robots.txt and llms.txt requests made with RankEcho's user agent. It is not a real-provider access test, an indexing observation, a citation measurement, a recommendation outcome, or evidence that one signal caused an AI answer.

What does location mean in the readiness directory?

Location describes the business-listing harvest frame used for some screened domains. It is not the location of an AI-answer request, proof of local-market coverage, or evidence about personalization.

What must happen before an industry benchmark is published?

At minimum, the cohort gate requires 30 eligible unique domains per category, 5 disclosed prompts per category, 3 named surfaces or models, 3 repeated windows, and 100 usable domain-by-prompt-by-surface observations after exclusions. Dates, missingness, exclusions, uncertainty, review status, and a downloadable artifact must also be disclosed. Passing the minimum is not a representativeness guarantee.

Review the measurement rules before using a number

The methodology defines what RankEcho intends to measure; this status page records where the current aggregate does and does not meet that standard.

Review the measurement methodology →Inspect the published Q3 study