The State of AI Citations, Q3 2026
We asked four AI engines the same 48 commercial questions, twice, 24 hours apart - 384 answers in all. They agree far less than the market assumes: across engines that both cited sources, only 20% of cited domains overlapped, and roughly four in five cited domains were unique to a single engine. Citations are not an oligarchy - the ten most-cited domains account for just 18% of all citations, and 81% point to ordinary independent sites, not big platforms. And AI answers drift: ask the same engine the same question a day later and about a third of its sources change - though how much depends heavily on which engine you ask.
Method (before findings)
We publish the method before the results because that is what separates a study from marketing. Everything here is reproducible: the full prompt battery is listed below and the raw per-citation data is a downloadable CSV.
- Battery: 48 prompts across 8 commercial categories (SaaS, Ecommerce, Health, Finance, Legal, Home Services, Education, Travel), spanning four buyer intents (category, comparison, problem, brand). No prompt mentions RankEcho or any site we operate.
- Engines: anthropic, perplexity, gemini, google_aio. Four engines - we did not have an OpenAI key at run time and say so plainly rather than implying coverage we lack.
- Sampling: every prompt x every engine, run twice, 24h apart (July 9, 2026 and July 10, 2026). 384 observations. Run 0 measures the cross-sectional questions; run 1 exists to measure stability.
- Statistic: cited-domain sets are compared with Jaccard similarity (|A∩B| / |A∪B|). Cross-engine agreement is the mean pairwise Jaccard per prompt; stability is the mean Jaccard between the two runs of the same prompt and engine.
- Publishing rules: engine-infrastructure redirects (e.g. unresolved grounding URLs) are excluded from every domain statistic and reported separately. Skipped runs are excluded from denominators and their reasons are disclosed. Answers with no sources are kept as data, not dropped.
Finding 1 — The long tail dominates. There is no oligarchy.
The single most-cited domain is Reddit - but community sites are only 5% of all citations. The ten most-cited domains together account for just 18%; the rest is a very long tail. 81% of citations point to ordinary independent websites rather than to big platforms, review aggregators, or news outlets.
What it means: "get cited by the big platforms" is the wrong strategy. AI engines pull from a wide field of ordinary pages, which is good news for smaller sites - visibility is winnable without being a household name.
Finding 2 — The four engines barely agree with each other.
On identical prompts, the engines cite mostly different sources. Counting only cases where two engines both returned sources, mean overlap is 0.195 - about 20%. Counting an absent answer as non-overlap (the experience a brand actually has), it falls to 0.138. Either way, roughly four of five cited domains are unique to one engine. And it replicated: agreement was 0.195 on day one and 0.201 on day two.
What it means: there is no single "AI visibility." Being cited by Perplexity tells you little about whether Gemini cites you. Any tool that reports one blended score is hiding this - visibility has to be measured per engine.
Finding 3 — Google's two AI surfaces agree with each other least of all.
The lowest-agreeing pair in the entire study is Gemini and Google's AI Overview at 0.129 - two systems from the same company, answering the same question, agreeing less than any cross-vendor pair. The highest pair is Google AI Overview and Perplexity (0.295), which makes sense: both lean on live web-search retrieval. So Google's AI Overview resembles Perplexity more than it resembles Google's own Gemini.
Caveat, stated up front: the AI Overview appeared for a minority of queries, so this pair rests on 15 comparisons. We flag it as a strong signal, not a settled fact, and will retest at larger scale next quarter.
Finding 4 — A third of citations change within 24 hours — but Gemini is far more volatile than the rest.
Asked the same question a day apart, engines returned overlapping-but-shifting sources: overall stability 0.671. But the average hides a wide split. Perplexity and Claude held about three-quarters of their sources steady; Gemini changed nearly two-thirds of its cited sources from one day to the next.
What it means: a single snapshot is not a measurement. If a third of citations turn over in a day - and most of an engine's, for Gemini - then a one-time "visibility score" is reporting a moment, not a position. This is the entire case for continuous tracking.
Finding 5 — Google shows an AI Overview for a minority of commercial queries — but cites generously when it does.
Google's AI Overview appeared for only 40% of our commercial prompts. On the rest, Google showed no AI Overview at all. But when it did appear, it cited about 7.7 sources - right in line with Perplexity and Claude. The often-quoted "Google AI Overview cites very little" is an averaging artifact: it is absent often, not stingy when present.
| Engine | Answered with sources | Mean citations (when it did) |
|---|---|---|
| anthropic | 100% | 9.19 |
| perplexity | 100% | 8.27 |
| google_aio | 40% | 7.74 |
| gemini | 83% | 7.25 |
Limitations
Read these before citing anything above.
- 48 prompts across 8 categories is a probe, not a census. US-English, consumer-commercial intent only.
- Two runs, 24 hours apart, one week. Stability over longer horizons is unmeasured.
- Four engines. No OpenAI/ChatGPT this round - a real gap we will close next quarter.
- Finding 3 (Gemini vs AI Overview) rests on 15 paired observations. Directionally strong, not definitive.
- We measured which domains are cited, not whether a brand was named in prose without a citation. That "mentioned but not cited" question is deferred to Q4, when we will capture answer text.
- AI engines change without notice. This is one snapshot of one week. That is precisely why we are publishing it as a dated, repeating series.
The data
Everything above comes from one downloadable file: every prompt, every engine, every cited domain, with its source-type classification and position, for both runs. Reuse it freely under CC BY 4.0 - a link back is all we ask. If you rerun the battery yourself, the prompt list below is the whole instrument.
The full 48-prompt battery
- best project management software for small teams
- best customer support helpdesk software
- Notion vs Asana for product teams
- Stripe vs PayPal for subscription billing
- how do I stop my SaaS trial users from churning
- is Linear worth it for issue tracking
- best ecommerce platform for a small business
- best running shoes for flat feet
- Shopify vs WooCommerce for a first store
- Dyson vs Shark vacuum which is better
- how do I reduce cart abandonment on my online store
- is Allbirds actually sustainable
- best online therapy platforms
- best free meditation apps
- BetterHelp vs Talkspace which is better
- creatine vs protein powder for beginners
- how do I fix my sleep schedule after night shifts
- is Noom effective for weight loss
- best high yield savings accounts
- best budgeting apps
- Vanguard vs Fidelity for index funds
- Roth IRA vs traditional IRA which should I choose
- how do I pay off credit card debt fastest
- is Wealthfront a good robo advisor
- best online will making services
- best LLC formation services
- LegalZoom vs Rocket Lawyer
- LLC vs S corp for a small business
- how do I trademark a business name myself
- is LegalZoom trustworthy for incorporation
- best home security systems without a contract
- best pest control companies
- SimpliSafe vs Ring alarm
- heat pump vs furnace for a cold climate
- how do I find a plumber I can actually trust
- is Angi worth using to find contractors
- best online coding bootcamps
- best language learning apps
- Coursera vs Udemy for career change
- Duolingo vs Babbel for Spanish
- how do I learn data analysis without a degree
- is Codecademy Pro worth the money
- best travel credit cards for beginners
- best travel insurance for international trips
- Airbnb vs hotels for a family trip
- Booking.com vs Expedia which is cheaper
- how do I find cheap flights last minute
- is Going (formerly Scott's Cheap Flights) worth it
Method note: this report front-loads its answer and marks up its own findings with Dataset and FAQ schema - the same extraction-friendly structure our product recommends. We eat our own cooking.
Frequently asked questions
For each prompt, we took the set of domains each engine cited and computed the Jaccard similarity between every pair of engines, then averaged. The conditional figure (0.195) counts only pairs where both engines cited at least one source; the pooled figure (0.138) treats an absent answer as zero overlap, which reflects what a brand actually experiences across engines.
We did not have an OpenAI API key at run time, so the study covers Anthropic (Claude), Perplexity, Gemini, and Google's AI Overview. Rather than imply coverage we do not have, we state the gap openly. ChatGPT is the priority addition for the Q4 edition.
We ran the identical battery twice, 24 hours apart, changing nothing in between, and measured how much each engine's cited-domain set for a given prompt overlapped between the two runs. An overall stability of 0.671 means roughly a third of cited sources changed within a day - and the per-engine range (0.36 for Gemini to 0.77 for Perplexity and Claude) is wide.
It is a probe, not a census: 48 prompts, 8 categories, US-English, two runs over one week. The headline findings replicated across both runs, which is reassuring, but the smaller cuts - especially the Gemini-vs-AI-Overview pair at n=15 - are directional. We publish it as a dated, repeating series precisely so the picture sharpens over time.
Yes, under CC BY 4.0 - reuse freely with attribution. The full per-citation dataset is downloadable as CSV, and the complete prompt battery is published on the page so anyone can reproduce the study.
