Home / Reports / State of AI Citations Q3 2026
Original research

The State of AI Citations, Q3 2026

0.195
cross-engine agreement (conditional)
18%
share held by the top 10 domains
0.67
citation stability after 24 hours
384
total observations
The short answer

We asked four AI engines the same 48 commercial questions, twice, 24 hours apart - 384 answers in all. They agree far less than the market assumes: across engines that both cited sources, only 20% of cited domains overlapped, and roughly four in five cited domains were unique to a single engine. Citations are not an oligarchy - the ten most-cited domains account for just 18% of all citations, and 81% point to ordinary independent sites, not big platforms. And AI answers drift: ask the same engine the same question a day later and about a third of its sources change - though how much depends heavily on which engine you ask.

Method (before findings)

We publish the method before the results because that is what separates a study from marketing. Everything here is reproducible: the full prompt battery is listed below and the raw per-citation data is a downloadable CSV.

  • Battery: 48 prompts across 8 commercial categories (SaaS, Ecommerce, Health, Finance, Legal, Home Services, Education, Travel), spanning four buyer intents (category, comparison, problem, brand). No prompt mentions RankEcho or any site we operate.
  • Engines: anthropic, perplexity, gemini, google_aio. Four engines - we did not have an OpenAI key at run time and say so plainly rather than implying coverage we lack.
  • Sampling: every prompt x every engine, run twice, 24h apart (July 9, 2026 and July 10, 2026). 384 observations. Run 0 measures the cross-sectional questions; run 1 exists to measure stability.
  • Statistic: cited-domain sets are compared with Jaccard similarity (|A∩B| / |A∪B|). Cross-engine agreement is the mean pairwise Jaccard per prompt; stability is the mean Jaccard between the two runs of the same prompt and engine.
  • Publishing rules: engine-infrastructure redirects (e.g. unresolved grounding URLs) are excluded from every domain statistic and reported separately. Skipped runs are excluded from denominators and their reasons are disclosed. Answers with no sources are kept as data, not dropped.

Finding 1 — The long tail dominates. There is no oligarchy.

The single most-cited domain is Reddit - but community sites are only 5% of all citations. The ten most-cited domains together account for just 18%; the rest is a very long tail. 81% of citations point to ordinary independent websites rather than to big platforms, review aggregators, or news outlets.

Share of AI citations by source type (all engines, run 0)
Independent sites81%Reddit5%YouTube4%News3%Social3%Scholarly2%Review sites1%

What it means: "get cited by the big platforms" is the wrong strategy. AI engines pull from a wide field of ordinary pages, which is good news for smaller sites - visibility is winnable without being a household name.

Finding 2 — The four engines barely agree with each other.

On identical prompts, the engines cite mostly different sources. Counting only cases where two engines both returned sources, mean overlap is 0.195 - about 20%. Counting an absent answer as non-overlap (the experience a brand actually has), it falls to 0.138. Either way, roughly four of five cited domains are unique to one engine. And it replicated: agreement was 0.195 on day one and 0.201 on day two.

Pairwise agreement between engines (conditional Jaccard, run 0)
anthropicanthropic0.1950.1860.180perplexityperplexity0.1950.1880.295geminigemini0.1860.1880.129google_aiogoogle_aio0.1800.2950.129
Pairwise Jaccard similarity of cited-domain sets. Darker = more agreement. Higher is more similar.

What it means: there is no single "AI visibility." Being cited by Perplexity tells you little about whether Gemini cites you. Any tool that reports one blended score is hiding this - visibility has to be measured per engine.

Finding 3 — Google's two AI surfaces agree with each other least of all.

The lowest-agreeing pair in the entire study is Gemini and Google's AI Overview at 0.129 - two systems from the same company, answering the same question, agreeing less than any cross-vendor pair. The highest pair is Google AI Overview and Perplexity (0.295), which makes sense: both lean on live web-search retrieval. So Google's AI Overview resembles Perplexity more than it resembles Google's own Gemini.

Caveat, stated up front: the AI Overview appeared for a minority of queries, so this pair rests on 15 comparisons. We flag it as a strong signal, not a settled fact, and will retest at larger scale next quarter.

Finding 4 — A third of citations change within 24 hours — but Gemini is far more volatile than the rest.

Asked the same question a day apart, engines returned overlapping-but-shifting sources: overall stability 0.671. But the average hides a wide split. Perplexity and Claude held about three-quarters of their sources steady; Gemini changed nearly two-thirds of its cited sources from one day to the next.

Citation stability after 24 hours, by engine (higher = steadier)
anthropic0.77perplexity0.77gemini0.36google_aio0.62

What it means: a single snapshot is not a measurement. If a third of citations turn over in a day - and most of an engine's, for Gemini - then a one-time "visibility score" is reporting a moment, not a position. This is the entire case for continuous tracking.

Finding 5 — Google shows an AI Overview for a minority of commercial queries — but cites generously when it does.

Google's AI Overview appeared for only 40% of our commercial prompts. On the rest, Google showed no AI Overview at all. But when it did appear, it cited about 7.7 sources - right in line with Perplexity and Claude. The often-quoted "Google AI Overview cites very little" is an averaging artifact: it is absent often, not stingy when present.

EngineAnswered with sourcesMean citations (when it did)
anthropic100%9.19
perplexity100%8.27
google_aio40%7.74
gemini83%7.25

Limitations

Read these before citing anything above.

  • 48 prompts across 8 categories is a probe, not a census. US-English, consumer-commercial intent only.
  • Two runs, 24 hours apart, one week. Stability over longer horizons is unmeasured.
  • Four engines. No OpenAI/ChatGPT this round - a real gap we will close next quarter.
  • Finding 3 (Gemini vs AI Overview) rests on 15 paired observations. Directionally strong, not definitive.
  • We measured which domains are cited, not whether a brand was named in prose without a citation. That "mentioned but not cited" question is deferred to Q4, when we will capture answer text.
  • AI engines change without notice. This is one snapshot of one week. That is precisely why we are publishing it as a dated, repeating series.

The data

Everything above comes from one downloadable file: every prompt, every engine, every cited domain, with its source-type classification and position, for both runs. Reuse it freely under CC BY 4.0 - a link back is all we ask. If you rerun the battery yourself, the prompt list below is the whole instrument.

Download the dataset (CSV)

The full 48-prompt battery
  1. best project management software for small teams
  2. best customer support helpdesk software
  3. Notion vs Asana for product teams
  4. Stripe vs PayPal for subscription billing
  5. how do I stop my SaaS trial users from churning
  6. is Linear worth it for issue tracking
  7. best ecommerce platform for a small business
  8. best running shoes for flat feet
  9. Shopify vs WooCommerce for a first store
  10. Dyson vs Shark vacuum which is better
  11. how do I reduce cart abandonment on my online store
  12. is Allbirds actually sustainable
  13. best online therapy platforms
  14. best free meditation apps
  15. BetterHelp vs Talkspace which is better
  16. creatine vs protein powder for beginners
  17. how do I fix my sleep schedule after night shifts
  18. is Noom effective for weight loss
  19. best high yield savings accounts
  20. best budgeting apps
  21. Vanguard vs Fidelity for index funds
  22. Roth IRA vs traditional IRA which should I choose
  23. how do I pay off credit card debt fastest
  24. is Wealthfront a good robo advisor
  25. best online will making services
  26. best LLC formation services
  27. LegalZoom vs Rocket Lawyer
  28. LLC vs S corp for a small business
  29. how do I trademark a business name myself
  30. is LegalZoom trustworthy for incorporation
  31. best home security systems without a contract
  32. best pest control companies
  33. SimpliSafe vs Ring alarm
  34. heat pump vs furnace for a cold climate
  35. how do I find a plumber I can actually trust
  36. is Angi worth using to find contractors
  37. best online coding bootcamps
  38. best language learning apps
  39. Coursera vs Udemy for career change
  40. Duolingo vs Babbel for Spanish
  41. how do I learn data analysis without a degree
  42. is Codecademy Pro worth the money
  43. best travel credit cards for beginners
  44. best travel insurance for international trips
  45. Airbnb vs hotels for a family trip
  46. Booking.com vs Expedia which is cheaper
  47. how do I find cheap flights last minute
  48. is Going (formerly Scott's Cheap Flights) worth it

Method note: this report front-loads its answer and marks up its own findings with Dataset and FAQ schema - the same extraction-friendly structure our product recommends. We eat our own cooking.

Frequently asked questions

How was cross-engine agreement calculated?

For each prompt, we took the set of domains each engine cited and computed the Jaccard similarity between every pair of engines, then averaged. The conditional figure (0.195) counts only pairs where both engines cited at least one source; the pooled figure (0.138) treats an absent answer as zero overlap, which reflects what a brand actually experiences across engines.

Why only four engines and not ChatGPT?

We did not have an OpenAI API key at run time, so the study covers Anthropic (Claude), Perplexity, Gemini, and Google's AI Overview. Rather than imply coverage we do not have, we state the gap openly. ChatGPT is the priority addition for the Q4 edition.

What does 'citation stability' mean?

We ran the identical battery twice, 24 hours apart, changing nothing in between, and measured how much each engine's cited-domain set for a given prompt overlapped between the two runs. An overall stability of 0.671 means roughly a third of cited sources changed within a day - and the per-engine range (0.36 for Gemini to 0.77 for Perplexity and Claude) is wide.

Is this enough data to draw firm conclusions?

It is a probe, not a census: 48 prompts, 8 categories, US-English, two runs over one week. The headline findings replicated across both runs, which is reassuring, but the smaller cuts - especially the Gemini-vs-AI-Overview pair at n=15 - are directional. We publish it as a dated, repeating series precisely so the picture sharpens over time.

Can I reuse the data?

Yes, under CC BY 4.0 - reuse freely with attribution. The full per-citation dataset is downloadable as CSV, and the complete prompt battery is published on the page so anyone can reproduce the study.

Run a no-card audit on your own site ->