AI Visibility Benchmark: Method, Status, and Limits
One route, four different evidence jobs. This page identifies what RankEcho can publish now, what remains a technical screen, and what is still blocked by cohort and review requirements.
RankEcho does not yet publish a general industry AI-visibility benchmark. Stored audits differ in prompt panels, configured providers, sample depth, observed denominators, and collection dates, so pooling them would imply comparability the current feed does not support. This hub publishes one dated citation-source study and the rules a future category benchmark must pass.
The legacy overall gate is 20 completed audit records and the legacy industry-row gate is 5. These are preserved in runtime for compatibility. They are display rules, not evidence that records are unique domains, comparable, representative, or statistically sufficient.
Separate the instruments before comparing numbers
| Surface | Observation unit and frame | What it can answer | Status |
|---|---|---|---|
| Q3 citation study | One scheduled prompt × configured answer-surface × run cell; skipped cells remain disclosed. | Cited-domain-set overlap inside the completed portion of that fixed panel. | Published with CSV analysis export |
| Cross-audit aggregate | One completed audit run; repeat domains and unlike run configurations are not removed. | Nothing public until a comparable cohort and review artifact pass the gate below. | Withheld |
| AI-readiness directory | One homepage-centred RankEcho-UA screen plus robots.txt and llms.txt checks, from a convenience and submission sample. | RankEcho's observed technical-screen signals, not citations or recommendations. | Separate directory |
| Rolling state report | The evolving audit-run stream plus proof-loop summaries. | Descriptive product observations; not the fixed Q3 panel or a census. | Separate rolling report |
The publishable fixed-panel asset
The State of AI Citations, Q3 2026 is the only citation analysis artifact distributed from this hub. Its design arithmetic is 48 predeclared prompts × 4 configured answer surfaces × 2 scheduled runs = 384 planned cells. Repository metadata reports the second run was scheduled 24 hours later; the CSV does not carry timestamps that independently verify that interval.
The configured surfaces were Claude through Anthropic's Messages API with its web-search tool, Perplexity API, Gemini API with Google Search grounding, Google AI Overviews through SerpApi. They are API and search integrations, not interchangeable consumer products. The panel had no OpenAI/ChatGPT observation, covered US-English commercial prompts, and scheduled only two runs on 2026-07-09 and 2026-07-10. A stability value of 0.671 is a mean set-similarity statistic; its 0.329 distance is not a literal percentage of individual citations that changed. The mean includes 23 Google-AI-Overview empty/empty pairs scored as similarity 1; the report's separate 0.617 Google-AI-Overview figure uses 15 source-bearing pairs. This is a bounded first-party probe with a normalized-domain export, not a raw-answer archive, market census, or claim about every provider, geography, category, or date.
Read the full method and findings · Download the CSV · CC BY 4.0
What the current audit stream actually contains
The executing aggregate reads up to the latest 5,000 rows marked done with a non-null scorecard. The unit is a completed audit run, not a unique domain. The projection used here carries industry and scorecard JSON, then skips malformed JSON; it supplies no domain key or finished timestamp for deduplication and calendar-window reporting.
- Runs can mix sitewide and page scopes, custom and generated prompts, free and paid provider sets, sampling depths, repeats, skipped observations, and answer-reuse rules.
- Industry is an inferred free-text label or “Other,” not a validated sampling stratum or stable public taxonomy.
- The observed prompt denominator can vary because an audit can contain unavailable or skipped engine cells. No common cross-audit prompt-by-engine panel is enforced.
Executing metric meanings
citedTarget. In the executing parser that flag can mean a target URL citation or a target-name text mention. Mention is not source citation.| Run | Positive prompts | Observed prompts | Run rate |
|---|---|---|---|
| Audit A | 1 | 2 | 1 ÷ 2 = 50.0% |
| Audit B | 6 | 10 | 6 ÷ 10 = 60.0% |
The current equal-run mean is (50.0% + 60.0%) ÷ 2 = 55.0%. Pooling prompt counts would produce 7 ÷ 12 = 58.3%. The current aggregate uses the first calculation. Neither result is a population rate without a defined cohort and sampling frame.
The minimum gate for a future category benchmark
This is a pilot publication rule, not a universal statistical law. Every category cut must satisfy all rows together, and passing them still does not make a convenience sample representative.
| Gate | Minimum | Required disclosure |
|---|---|---|
| Eligible population | ≥ 30 unique domains per category | Sampling frame, eligibility, deduplication, repeat-domain handling |
| Prompt panel | ≥ 5 predeclared prompts per category | Full prompt text, prompt class, additions and removals |
| Surface coverage | ≥ 3 named surfaces or models | Provider surface, model/version where known, failures and skips |
| Repeated measurement | ≥ 3 dated windows | Collection dates, interval, answer-reuse and cache policy |
| Usable observations | ≥ 100 domain × prompt × surface observations after exclusions | Planned, attempted, skipped, excluded, successful, and observed denominators |
| Publication artifact | Required | Downloadable aggregate, uncertainty, limitations, independent review status, version and correction log |
Readiness is a separate technical screen
The AI-readiness directory uses a homepage-centred screen plus robots.txt and llms.txt requests made with RankEcho's user agent, not audit answers or a provider's crawler. Its current composite weights are 30% structured data, 25% crawl access, 25% machine readability, and 20% agent guidance. Bands begin at 85, 65, and 40.
The directory mixes harvested business-listing domains with submissions, counts successful captures, and does not expose a request-time freshness window on this hub. Category and category-by-metro display gates of 30 and 20 sites are coverage rules, not statistical sufficiency. Metro describes a harvest-source location, not AI-answer request geography or personalization.
Six inferences this hub rejects
Choose the evidence surface that matches the job
Sources and provenance reviewed
These first-party sources support the public study, artifact, terminology, screening limits, and no-guarantee boundary. The audit-stream definitions above were also checked against the executing database selection, scorecard, citation parser, and aggregate code on 2026-09-02.
Open the five-record evidence ledger
- The State of AI Citations, Q3 2026 The Q3 study scheduled 384 prompt-by-surface-by-run cells from 48 prompts, four configured answer surfaces, and two runs repository metadata reports as 24 hours apart; 368 cells are marked non-skipped and 16 skipped. Its prompt battery, method, limits, and normalized-domain CSV analysis export are public. Checked 2026-09-02 · Published 2026-07-10; fixed Q3 2026 study · First-party study review; the fixed panel is not merged with later customer audit runs or homepage screens.
- Q3 2026 normalized-domain analysis export (CSV) The Q3 normalized-domain analysis export is downloadable as CSV under CC BY 4.0. It contains prompt ID, category, intent, adapter/surface key, run index, skip flag, distinct normalized domain, heuristic source type, and zero-based export order among the cell's distinct domains. It omits prompt text, raw answers, original URLs and titles, model/version, timestamps, skip reasons, answer length, and session or geography configuration. Checked 2026-09-02 · Published Q3 2026 CSV analysis export · Artifact review; normalized citation-domain rows are not a raw-answer archive, and availability does not make the US-English probe a census or universal provider benchmark.
- RankEcho measurement methodology RankEcho's methodology treats citation, mention, recommendation, source mix, and re-test movement as distinct outcomes and rejects guaranteed placement or causal claims. Checked 2026-09-02 · Methodology updated August 4, 2026 · First-party methodology review; the hub additionally follows the executing audit code where legacy labels differ from these definitions.
- Website AI Readiness study and limits The published website-readiness study reports a bounded homepage-screening sample and explicitly does not measure indexing, AI impressions, citations, traffic, or causal effects. Checked 2026-09-02 · July 2026 one-request homepage screening · First-party study review; the dated study and the evolving readiness directory are not represented as one population.
- RankEcho AI visibility disclosure RankEcho's disclosure states that observed visibility can vary by engine, model, retrieval state, time, location, account state, and prompt wording, and that no citation or business outcome is guaranteed. Checked 2026-09-02 · Current public measurement disclosure · First-party disclosure review; a threshold crossing or later observation is not converted into a causal or outcome guarantee.
Benchmark FAQ
Which RankEcho benchmark data can I download today?
The Q3 2026 citation study has a downloadable normalized-domain analysis export: 48 prompts across four configured answer surfaces in two runs repository metadata reports as scheduled 24 hours apart. Of 384 planned cells, 368 are marked non-skipped and 16 skipped. The export omits prompt text, raw answers, URLs, model versions, timestamps, and skip reasons, and the cross-audit aggregate and readiness directory do not share this study's sample.
Why are cross-audit citation percentages withheld here?
The current feed contains completed audit runs with varying scopes, prompt panels, engine configurations, sampling depths, skipped observations, and repeat domains. It does not give this page a deduplicated cohort, common calendar window, or common observed denominator.
Does the 20-run legacy display threshold make an aggregate reliable?
No. Twenty runs and five runs per industry are legacy product display gates. They do not establish unique-domain coverage, representativeness, uncertainty, a common panel, or statistical validity, so crossing them does not publish results on this hub.
Does an AI-readiness score predict citations or recommendations?
No. The readiness directory is a homepage-centred screen plus robots.txt and llms.txt requests made with RankEcho's user agent. It is not a real-provider access test, an indexing observation, a citation measurement, a recommendation outcome, or evidence that one signal caused an AI answer.
What does location mean in the readiness directory?
Location describes the business-listing harvest frame used for some screened domains. It is not the location of an AI-answer request, proof of local-market coverage, or evidence about personalization.
What must happen before an industry benchmark is published?
At minimum, the cohort gate requires 30 eligible unique domains per category, 5 disclosed prompts per category, 3 named surfaces or models, 3 repeated windows, and 100 usable domain-by-prompt-by-surface observations after exclusions. Dates, missingness, exclusions, uncertainty, review status, and a downloadable artifact must also be disclosed. Passing the minimum is not a representativeness guarantee.
Review the measurement rules before using a number
The methodology defines what RankEcho intends to measure; this status page records where the current aggregate does and does not meet that standard.
Review the measurement methodology →Inspect the published Q3 study