Why one engine citing you predicts almost nothing about the others
Across 384 observations, the AI engines RankEcho measured overlapped on sources 19.5 percent of the time when both engines cited anything, and 13.8 percent when answers citing nothing are counted as non-overlap. Being cited by Perplexity therefore tells you very little about ChatGPT or Gemini, and a single visibility score averages away the number that would tell you what to fix.
How often do AI engines cite the same sources?
Rarely, and the honest answer needs two numbers rather than one. Across 384 observations, RankEcho measured 19.5 percent overlap between the publisher domains two engines cited - counting only the pairs where both engines cited something. Counting answers that cited nothing as non-overlap, the figure falls to 13.8 percent.
The distinction matters and is worth stating plainly. The first number describes engine behaviour: when two engines both reach for sources, how often do they reach for the same ones. The second describes a brand's experience: across everything that happened, including answers with no sources at all, how often did the two engines corroborate each other. Most published agreement figures do not say which they are reporting.
One live site measured 13.5 percent on the conditional basis across 72 engine pairs, with no pair above 23 percent - below the study's 19.5 percent, on the same definition.
Independent work lands in the same territory from a different direction. Analysis circulating in 2026 found that only around 11 percent of websites are cited by both ChatGPT and Perplexity, which is to say the overwhelming majority of citations on either engine are unique to it.
The practical consequence is that the phrase AI search visibility describes five different things. Optimising for the engine that already cites you may do nothing for the four that do not, because they are largely not reading the same web.
What is the difference between engines agreeing about sources and agreeing about you?
Two separate questions, and conflating them hides the more useful one. Source agreement asks whether the engines drew on the same pages. Brand agreement asks whether they reached the same conclusion about whether to cite you.
A site can score low on the first and high on the second, or the reverse, and each combination means something different. RankEcho measures both because the gap between them is the diagnosis - not the individual number.
Both are computed the same way for every site: one observation per engine per prompt, runs that never happened excluded rather than counted as disagreement, and infrastructure redirects filtered out so a grounding cache does not appear as a cited publisher.
Why does a single visibility score hide the problem?
Because averaging conceals the only part you can act on. Take a real 12-prompt battery with no skipped runs. Overall brand agreement came to 67 percent, which sounds like broad consensus.
It decomposes into three very different groups. Every engine cited the site on 3 prompts - 25 percent. No engine cited it on 5 - 42 percent. And they disagreed on the remaining 4 - 33 percent. Most of the apparent agreement is agreement that the brand does not exist.
The 33 percent is where the work is. On those prompts at least one engine already found the site citable, which means authority is not the obstacle and the pages losing are losing on treatment. A single score of 67 tells you none of that.
This is also why cross-engine consistency keeps showing up as a weak point elsewhere: AirOps has reported that only about 30 percent of brands stay visible across consecutive answers, and Semrush has found fewer than one brand in five is both frequently mentioned and consistently cited.
What does it mean when engines read the same pages but only some cite you?
It means the problem is extraction, and a rewrite can fix it. If source agreement is high, the engines are drawing on a common pool - so the ones ignoring you had access to the same material and chose something else from it.
That points at the page rather than at your reputation: an answer buried below three paragraphs of preamble, a section too long to quote without cutting mid-thought, figures in prose where a table would be liftable, a paragraph that never names the brand so quoting it credits nobody.
This is the good case. It is the one where editing the page changes the outcome.
What does it mean when they read different pages entirely?
It means the problem is placement, and a rewrite will not touch it. When source agreement is very low - the 13.5 percent case - each engine is answering from its own pool of pages. Winning all of them means being present in several different pools, not writing one better page.
That is off-site work: being named in the roundups and directories one engine favours, present in the community threads another draws on, referenced by the reference sites a third leans on. Slower, harder, and not something a content tool can do for you.
Saying so is more useful than selling a rewrite that cannot work. The split between the two metrics is what tells you which situation you are in, and it is the reason to measure both.
Which engine disagrees most?
On the site measured below, one engine stood apart. Pairwise brand agreement between Anthropic and OpenAI was 92 percent, and between OpenAI and Perplexity also 92 percent - but Anthropic and Gemini agreed only 67 percent of the time, and Gemini and OpenAI 75 percent.
In other words the engines disagreed with Gemini about this brand more than they disagreed with each other. That is a per-engine finding no aggregate score surfaces, and it is directly actionable: it names which engine to investigate.
One honest caveat on the table below. Google's AI Overviews shows zero, and part of that is a measurement artifact rather than a verdict - Overviews frequently returns no usable answer for a query, and an answer that never appeared cannot cite anybody. RankEcho now records that separately from a genuine absence.
The data from one 12-prompt battery
Twelve prompts, five engines, no skipped runs, one observation per engine per prompt.
Per engine, share of prompts where the brand was cited: Perplexity 6 of 12 (50 percent), Anthropic 6 of 12 (50 percent), OpenAI 5 of 12 (42 percent), Gemini 4 of 12 (33 percent), Google AI Overviews 0 of 12.
Pairwise, brand agreement against source agreement: Anthropic and OpenAI 92 percent against 10 percent. OpenAI and Perplexity 92 percent against 11 percent. Anthropic and Perplexity 83 percent against 11 percent. Gemini and Perplexity 83 percent against 15 percent. Gemini and OpenAI 75 percent against 23 percent. Anthropic and Gemini 67 percent against 10 percent.
Read the two columns together. Every pair agrees about the brand far more often than it agrees about sources - which means these engines are reaching similar conclusions from largely different reading. That is the extraction case, and it is the one worth acting on.
How do you measure this on your own site?
You need one observation per engine per prompt on the same battery, and you need to keep the three states apart: cited, not cited, and never ran. Collapsing the third into the second turns a provider outage into a visibility finding.
Then compute both numbers. Source agreement as the overlap between the publisher domains each pair of engines cited. Brand agreement as whether each pair reached the same verdict about you on the same prompt. Report the denominators, because a twelve-prompt battery makes pairwise estimates thin and a percentage without its sample size invites more confidence than it earns.
RankEcho does this per customer using the same definitions as the study above, so the number on your panel and the number in the research cannot drift apart.
What this data cannot tell you
It cannot tell you why an engine chose what it chose. These are observations of outputs, not access to retrieval. A low overlap says the pools differ; it does not say why they differ.
It cannot be extrapolated confidently from one site. 384 observations is enough to notice a pattern and not enough to characterise an industry, and a single 12-prompt battery is thinner still. Both are reported with their sample sizes for that reason.
And a figure quoted without its basis is close to meaningless. The 19.5 percent here is conditional - both engines cited something - measured on run 0 so repeat sampling cannot inflate it, with infrastructure redirects excluded. Quoted on the pooled basis the same data gives 13.8 percent. Any agreement number you read elsewhere, including ours, should say which it is.
And it will change. Engines re-crawl, re-index, and change models. An agreement figure describes a period, not a permanent property - which is the argument for measuring it repeatedly rather than once.
Frequently asked questions
Measure first. If source agreement across engines is high, the engines are reading similar pages and one set of page-level changes can move several of them. If it is low, each engine is answering from a different pool and winning them all is an off-site problem rather than a writing one.
Source agreement asks whether engines drew on the same pages. Brand agreement asks whether they reached the same verdict about citing you. High source agreement with low brand agreement means the engines read the same material and only some picked you out - an extraction problem you can fix by rewriting.
Not necessarily, because most of it can be agreement about your absence. In the battery above, 67 percent decomposed into 25 percent where every engine cited the brand, 42 percent where none did, and 33 percent where they split. The 33 percent is the actionable part and the headline number hides it.
Partly because it often returns no answer at all for a query, and an answer that never appeared cannot cite anyone. That is a measurement fact rather than a visibility verdict, and it should be recorded separately from a genuine absence - otherwise a missing answer is scored as a lost citation.
More than twelve. With a small battery the pairwise comparisons are thin, so a percentage can move several points on one prompt changing. Report the sample size alongside the figure and treat direction as more reliable than magnitude until the battery is larger.
