How to measure cross-engine citation agreement
In run 0 of RankEcho's dated Q3 2026 dataset, conditional mean cited-domain overlap was 19.5 percent across the four measured engines, and pooled overlap was 13.8 percent when comparisons with an absent source answer were assigned zero. The full study used 48 prompts and two dates; these cross-engine estimates describe the declared run-0 comparisons. They do not predict another engine's answer or reveal why any source was chosen. Report source overlap and brand-outcome agreement separately, with the prompt, engine, run, and denominator beside each estimate.
How often do AI engines cite the same sources?
Run 0 of RankEcho's Q3 2026 study recorded 19.5 percent conditional mean overlap between cited publisher-domain sets when both engines in a pair returned sources. When a comparison with an absent source answer was assigned zero overlap, the pooled mean was 13.8 percent.
The figures answer different descriptive questions. The conditional estimate summarizes overlap among pairs with sources on both sides. The pooled estimate also represents comparisons in which one side had no source answer. Neither figure measures what an engine read, why it selected a source, or how a different prompt would behave.
In the separate 12-prompt, five-engine example reported below, conditional source overlap was 13.5 percent across 72 eligible engine-pair comparisons, and no pair exceeded 23 percent. That is a single-site example, not external validation or a population estimate.
AirOps and Kevin Indig's 2026 State of AI Search reports repeat-answer brand-visibility figures. Its unit is brand visibility across repeated answers, not cited-domain overlap, so it is contextual evidence rather than a direct corroboration of RankEcho's source-overlap estimate.
These observations are a reason to report results by engine and metric rather than collapse them into one score. They do not identify how an engine assembled an answer or which change would alter it.
What is the difference between engines agreeing about sources and agreeing about you?
The measures answer separate questions. Source agreement compares whether observed answers cited the same publisher domains. Brand agreement compares whether the engines produced the same brand-citation outcome for the same prompt.
The two measures can diverge. Reporting both shows whether source-set overlap and brand-outcome agreement moved together in the observed panel. The gap is a descriptive comparison, not a diagnosis of why an engine produced an answer.
For a declared site panel, compute both from one observation per engine per prompt. Exclude attempts that never produced an answer rather than count them as disagreement, and filter infrastructure redirects so a grounding cache is not counted as a cited publisher.
What can a single visibility score obscure?
A single mean can conceal its component outcomes. In the 12-prompt, single-site example below, overall brand-outcome agreement was 67 percent. The aggregate alone does not show how those prompts were distributed.
It decomposes into three different observed groups.
Every engine cited the site on 3 of the 12 prompts - 25 percent.
No engine cited it on 5 prompts - 42 percent.
The engine outcomes differed on the remaining 4 prompts - 33 percent.
Five prompts therefore had agreement on no observed brand citation. Every figure in this example comes from the battery set out below and describes that site and run only.
The four prompts with differing engine outcomes warrant narrower inspection. The outputs alone cannot show whether the difference arose from access, retrieval, source availability, answer composition, or sampling variation.
A later run may decompose differently, so report the observation date and repeated runs rather than treat the percentages as permanent properties.
What can high source overlap with different brand outcomes show?
It shows that the observed cited-domain sets overlapped while the observed brand-citation outcomes differed. It does not show that each engine read the same pages, had access to the same candidate material, or omitted a brand because of extraction.
A team can inspect page-level features such as answer placement, passage length, table structure, and explicit attribution as content-quality checks. Any rewrite remains a hypothesis to evaluate with a matched re-test; the overlap pattern does not establish the cause.
This pattern can prioritize page inspection, but it cannot predict that editing the page will change an engine output.
What can low observed source overlap show?
It shows that the cited publisher-domain sets differed in the sampled answers. It does not prove that the engines read entirely different pools, that placement caused the difference, or that a rewrite cannot affect a later answer.
A useful follow-up is to inspect the returned source types and pages for the exact prompt: official documentation, roundups, directories, community discussions, or other sources. Accurate representation on a relevant third-party page can be tested as an off-site hypothesis, but no placement guarantees a citation.
Use source and brand agreement together to choose the next investigation, not to assign a cause from output overlap alone.
Which engine pairs differed in the 12-prompt example?
In the 12-prompt example below, pairwise brand-outcome agreement was 92 percent for Anthropic and OpenAI and 92 percent for OpenAI and Perplexity. It was 67 percent for Anthropic and Gemini and 75 percent for Gemini and OpenAI.
Within that battery, the pairs containing Gemini had lower brand-outcome agreement than the two 92-percent pairs. That supports engine-specific inspection for this example; it does not identify a cause or predict another run.
Google AI Overviews recorded zero brand citations in the example. A measurement record should distinguish an unavailable answer from an observed answer with no brand citation. An aggregate zero without that state is not evidence about eligibility or selection.
The data from one 12-prompt battery
Twelve prompts, five engines, no skipped runs, one observation per engine per prompt.
Per engine, the share of prompts where the brand was cited was: Perplexity 6 of 12 (50 percent), Anthropic 6 of 12 (50 percent), OpenAI 5 of 12 (42 percent), Gemini 4 of 12 (33 percent), and Google AI Overviews 0 of 12.
Pairwise, each figure is brand agreement first and source agreement second.
Anthropic and OpenAI: 92 percent brand agreement and 10 percent source agreement.
OpenAI and Perplexity: 92 percent and 11 percent.
Anthropic and Perplexity: 83 percent and 11 percent.
Gemini and Perplexity: 83 percent and 15 percent.
Gemini and OpenAI: 75 percent and 23 percent.
Anthropic and Gemini: 67 percent and 10 percent - the widest difference between the two measures among the listed pairs.
Brand-outcome agreement exceeds cited-domain overlap for every listed pair. That describes co-occurrence in this one battery; it does not show what pages the engines read, why their outcomes matched, or whether extraction caused the difference.
How do you measure this on your own site?
You need one observation per engine per prompt on the same battery, and you need to keep the three states apart: cited, not cited, and never ran. Collapsing the third into the second turns a provider outage into a visibility finding.
Then compute both numbers. Source agreement as the overlap between the publisher domains each pair of engines cited. Brand agreement as whether each pair reached the same verdict about you on the same prompt. Report the denominators, because a twelve-prompt battery makes pairwise estimates thin and a percentage without its sample size invites more confidence than it earns.
When comparing a site panel with the study, use the same definitions and show any configuration differences. Shared labels do not make panels with different prompts, engines, dates, or availability states directly comparable.
What this data cannot tell you
It cannot tell you why an engine chose what it chose. These are observations of outputs, not access to retrieval. Low overlap means the observed cited-domain sets differed; it does not establish the underlying candidate pools or why they differed.
The 384 study observations support exact descriptive estimates for that 48-prompt, four-engine design; they do not characterize an industry or other prompts and dates. A single 12-prompt site battery is narrower still. Both should be reported with their sample sizes and collection scope.
A figure quoted without its basis is ambiguous. The 19.5 percent here is conditional - both engines cited something - and uses run-0 comparisons with infrastructure redirects excluded. On the pooled basis, the same run gives 13.8 percent. State the basis with any agreement number, including this one.
The estimate may change across later runs, model versions, locales, or prompt wording. An agreement figure describes its declared panel and dates, not a permanent engine property.
Frequently asked questions
Measure the engines relevant to the audience and keep their results separate. Source overlap can identify observed similarities and differences between cited-domain sets, but it cannot reveal the reason or show whether page-level or off-site work will change a later answer.
Source agreement compares the cited publisher-domain sets in observed answers. Brand agreement compares whether each engine produced the same brand-citation outcome for a prompt. A gap between them is descriptive; it does not show that engines read the same material or that extraction caused an omission.
The aggregate needs its components. In the example above, 67 percent included 25 percent of prompts where every engine cited the brand, 42 percent where none did, and 33 percent where outcomes differed. The differing outcomes merit inspection, but they do not identify a cause or supply a decision threshold.
The example records zero brand citations for Google AI Overviews. Interpret that count alongside answer availability: an unavailable answer and an observed answer without a brand citation are different states. The aggregate zero alone does not explain eligibility or selection.
There is no universal minimum. A 12-prompt battery produces thin pairwise estimates, so report the prompt count, eligible comparisons, dates, and uncertainty. More observations improve the description of that panel but do not automatically establish representativeness or statistical confidence.
