AI visibility audit checklist: 8 evidence-first steps
An AI visibility audit is a dated, repeatable observation of a fixed prompt-and-engine panel, not proof of why an answer appeared. For every eligible prompt-engine-repeat cell, record the exact prompt, run conditions, answer, cited URLs, brand mention, and named competitors; calculate rates only over completed eligible cells. Classify possible causes as hypotheses, ship one bounded fix, and repeat the same panel to measure movement without claiming causation.
What can an AI visibility audit establish?
This AI visibility audit checklist establishes what a defined set of systems returned under recorded conditions. It can show that a brand-owned URL was cited, that the brand was named without a link, that a tracked competitor appeared, or that a completed answer contained none of them.
It cannot establish why the system produced that answer. A citation is an observed attribution, not proof that the cited page caused every claim; a missing citation is not proof that a crawler was blocked; and a before/after change does not isolate your edit from index updates, answer variability, personalization, location, or provider changes.
- Observation: exact answer text, visible links, brand names, date, engine, prompt, and run status.
- Hypothesis: a testable explanation such as access, prompt fit, page evidence, source coverage, or answer variability.
- Decision: the smallest change justified by the evidence collected.
- Validation: the same scheduled cells repeated after the change, with misses and unavailable runs retained.
Step 1 — What decision will the audit support?
Write the decision before collecting answers. Examples include deciding which existing page to repair, whether a visibility loss is isolated to one engine, or whether a shipped access fix is associated with a changed result.
Do not combine unrelated decisions into one score. A category-discovery audit, a brand-accuracy audit, and a technical-access investigation need different prompts, labels, and success criteria.
- State the audience, market, language, geography, and decision owner.
- Define what counts as a brand mention, owned citation, competitor appearance, and unavailable run.
- Pre-register the primary metric and the minimum evidence needed to act.
- Record hypotheses separately from the observations used to test them.
Step 2 — How should you define the prompt panel?
Use prompts tied to the decision and preserve their exact wording. Include a deliberately bounded mix of category, comparison, alternative, problem, use-case, and proof questions only when each represents an audience need you can document.
A prompt list generated from a brand name is a starting hypothesis, not demand research. Validate candidate prompts against sales conversations, support questions, on-site search, paid-search terms, or other first-party evidence before treating them as buyer demand.
- Assign a stable prompt ID and retain the verbatim text.
- Label the intended audience, intent class, geography, and funnel stage.
- Keep exploratory prompts outside the proof panel until the next registered version.
- Version the panel instead of silently replacing prompts that perform poorly.
Step 3 — What is one eligible measurement cell?
Define one cell as one exact prompt asked to one named engine in one scheduled repeat. If six prompts run on two engines once, the planned denominator is 12 cells. If every cell is repeated three times, it is 36; do not average away run-level variability before retaining the raw records.
An eligible completed cell returned an answer that could be inspected under the registered conditions. Mark authentication failures, unavailable features, retrieval errors, and skipped regional surfaces separately. Exclude them from rate denominators, but publish their counts so a shrinking denominator cannot look like improvement.
- Planned cells = prompts × engines × scheduled repeats.
- Eligible cells = planned cells that returned an inspectable answer.
- Unavailable/error cells = reported separately and never relabelled as citation misses.
- The same eligibility rule must be applied to baseline and re-test windows.
Step 4 — What evidence should you retain for every cell?
Keep the answer, not just a derived score. Store the exact prompt, engine and surface label, run time, locale and account state when known, full response, visible cited URLs, detected brand strings, named tracked competitors, and the eligibility decision.
Classify a URL only after observing it. Labels such as owned site, documentation, review site, community, directory, or editorial source describe the page; they do not prove that the source caused the recommendation or that an unlinked source was not used.
- Preserve raw answer text and URLs before normalising domains.
- Distinguish a linked owned URL from an unlinked brand mention.
- Record the passage or answer position associated with each visible citation when available.
- Keep screenshots or exports for surfaces that cannot be reproduced later.
Step 5 — How do you calculate comparable rates?
Use the eligible completed cell as the denominator for cell-level rates and state the numerator beside every percentage. Do not describe a page as having an 80% citation rate without saying whether that means 8 of 10 cells, 80 of 100 prompts, or a provider-specific count.
Citation rate equals eligible cells with at least one visible link to a registered brand-owned domain divided by all eligible cells. Mention rate uses eligible cells that contain a registered brand name. A competitor-replacement label needs a written rule—for example, your brand absent while a tracked rival is explicitly presented as an option—and consistent human review.
- Owned citation rate = cells citing a registered owned domain ÷ eligible cells.
- Brand mention rate = cells naming the registered brand ÷ eligible cells.
- Competitor-replacement rate = cells meeting the pre-registered replacement rule ÷ eligible cells.
- Report engine-level numerators and denominators before any combined rate.
- Use percentage points, not percent change, when comparing two rates.
Step 6 — Which hypotheses survive disconfirming checks?
Turn each loss into competing explanations rather than assigning one cause. An access hypothesis weakens if the exact URL is indexed where relevant, publicly fetchable, and served successfully to a verified crawler in logs. A page-clarity hypothesis weakens if the requested facts are already explicit in public HTML. A third-party-source hypothesis weakens if observed cited sources already include your brand on comparable terms.
None of those checks proves the remaining hypothesis. They reduce avoidable guesswork and identify the smallest safe experiment. Keep an 'unknown or provider variability' option so the audit is not forced to invent a fix.
- Access hypothesis: inspect robots policy, edge controls, index status, responses, and verified request evidence.
- Prompt-fit hypothesis: verify the product actually serves the stated audience and use case.
- Page-evidence hypothesis: compare the requested claim with the exact public passage and its supporting source.
- External-evidence hypothesis: inspect observed cited pages without assuming every unlinked influence is visible.
- Variability hypothesis: repeat cells before treating a one-run miss as a stable loss.
Step 7 — Which bounded fix should you ship?
Prioritise a cell or prompt group by decision value, evidence strength, and reversibility—not by a score alone. The fix must correspond to the surviving hypothesis and name an owner, target URL, acceptance check, ship date, and re-test window.
Examples include correcting an actual crawler block, clarifying a verifiable answer in public HTML, repairing inconsistent product facts, or seeking factual inclusion in an observed source. Accurate structured data may support normal search understanding when it matches visible content, but Google documents no special AI schema and no markup guarantees selection.
- Find: preserve the losing cell and evidence that supports or weakens each hypothesis.
- Fix: change one bounded surface where feasible and record exactly what shipped.
- Prove: pre-register the same cells, eligibility rule, and comparison metric.
- Do not promise a citation, ranking, recommendation, or impression.
Step 8 — How should you re-test and report movement?
After the fix window, repeat the exact prompt-engine-repeat cells under the closest available conditions. Retain the full baseline and re-test answers, all eligibility decisions, the change log, and any provider or panel-version differences.
Report 'observed after the change' rather than 'caused by the change.' One bounded edit plus a stable panel is stronger evidence than an uncontrolled rewrite, but search indexes, retrieval systems, models, and cited pages can change during the same period. Replication across scheduled windows increases confidence without creating certainty.
- Publish baseline and re-test numerators, denominators, and unavailable counts.
- State the absolute percentage-point change and engine-level results.
- Keep negative, unchanged, and contradictory outcomes in the record.
- Label the result observed, suggestive, replicated, or inconclusive; never guaranteed.
What does a worked synthetic audit look like?
Suppose a team registers six prompts on two engines for one scheduled run: 12 planned cells. One engine does not expose an answer for one prompt, leaving 11 eligible completed cells and one unavailable cell. The brand has visible owned citations in 2 of 11 cells (18.2%), is named in 3 of 11 (27.3%), and meets the registered competitor-replacement rule in 5 of 11 (45.5%).
The team clarifies one existing comparison page and changes nothing else in the registered cohort. In the re-test, the same 11 cells are eligible and owned citations appear in 4 of 11 (36.4%): an observed increase of 18.2 percentage points. That is worth replicating, but it does not prove the page edit caused the two additional citations; an index refresh, answer variance, or source change could also explain them.
- Planned denominator: 12 cells; eligible denominator: 11; unavailable: 1.
- Baseline owned citations: 2/11; re-test owned citations: 4/11.
- Movement: +18.2 percentage points, not a 100% causal improvement claim.
- Next action: repeat the panel and inspect the two new citation records.
Which false positives and alternative explanations should you record?
Brand-name matching can confuse common words, subsidiaries, similarly named companies, quoted user text, navigation links, and a citation URL that redirects off the registered domain. A link beside a sentence may support only part of the sentence. Manual review should resolve ambiguous detections before the rate is final.
Changes can also come from provider experiments, account state, geography, personalisation, prompt interpretation, new web sources, index refreshes, cited-page edits, or ordinary stochastic variation. Report these as alternative explanations instead of selecting the most convenient story.
- Check exact brand aliases and domain ownership before counting.
- Separate 'mentioned', 'cited', and 'recommended' labels.
- Inspect redirects, citation placement, and duplicated URLs.
- Compare locale, account, engine label, feature availability, and run time.
- Retain an inconclusive outcome when the evidence conflicts.
What do the official provider sources say?
Google's AI-features documentation says normal Search eligibility and snippet controls apply to supporting links in AI Overviews and AI Mode, with no additional technical requirements or special AI schema; eligibility still does not guarantee appearance. Bing's February 2026 AI Performance announcement says its citation count is not placement, ranking, authority, or page importance and that grounding-query data is sampled.
OpenAI documents OAI-SearchBot for search, GPTBot for potential training, and ChatGPT-User for user-triggered actions as separate controls. An access audit must name the system and evidence being checked; permitting one documented bot does not prove indexing, answer inclusion, or citation. The claim-level source records below were checked on 2026-09-01.
How does the checklist hand off from Find to Fix to Prove?
Find ends with raw eligible-cell evidence and ranked hypotheses, not a list of asserted causes. Fix turns one surviving hypothesis into a bounded implementation with an acceptance check. Prove repeats the registered panel, preserves every outcome, and reports association with its limitations.
Use the free prompt and citation utilities for narrow discovery, then move to the product workflow only when you need a retained multi-engine audit and re-test history. A single pasted answer or generated prompt battery is not a completed AI visibility audit.
Sources reviewed
Provider eligibility and measurement claims below were checked against primary documentation. These records do not establish a universal selection formula, causation, or a guaranteed ranking, impression, recommendation, or citation.
3 claim-level source records
| Claim reviewed | Official source | Review record |
|---|---|---|
| Google says normal Search indexing and snippet controls govern eligibility for AI Overviews and AI Mode; there are no additional technical requirements or special AI schema files. | Google Search AI features documentation | Checked 2026-09-01 · AI features documentation updated 2025-12-10 · Primary-source documentation review; eligibility does not guarantee selection or presentation in an AI feature. · Confidence: High |
| OpenAI documents OAI-SearchBot for ChatGPT search, GPTBot for potential model training, and ChatGPT-User for user-triggered actions; the controls are independent and robots.txt rules may not apply to ChatGPT-User. | OpenAI crawler documentation | Checked 2026-09-01 · Current OAI-SearchBot, GPTBot, and ChatGPT-User documentation · Primary-source documentation review; no claim that a permitted bot will index, rank, or cite a page. · Confidence: High |
| Bing's AI Performance report counts observed citations and exposes sampled grounding queries, but Microsoft says citation count is not placement, ranking, authority, or page importance. | Bing Webmaster Blog: AI Performance | Checked 2026-09-01 · Public preview announced February 2026 · Primary-source documentation review; Bing metrics are treated as observations with their stated sampling and interpretation limits. · Confidence: High |
Frequently asked questions
It is a pre-defined record of prompts, engines, repeats, eligibility rules, raw answers, visible citations, brand outcomes, hypotheses, fixes, and like-for-like re-tests. It measures observed output; it does not prove why a system produced it.
Use all eligible completed prompt-engine-repeat cells for the stated panel and window. Show the numerator, denominator, and count of unavailable or errored cells; do not silently count unavailable surfaces as misses or remove completed misses.
There is no universal number. Use the smallest fixed panel that represents the documented decision, audience, and intent classes, then add scheduled repeats to measure variability. Version additions instead of changing the proof panel mid-test.
No. A like-for-like increase after one bounded change is evidence of movement, but provider updates, index changes, new sources, geography, account state, and answer variability remain alternative explanations. Replicate the result and keep the limitation visible.
No. Access is one prerequisite for some retrieval paths. Google states that Search eligibility does not guarantee appearance in AI features, and OpenAI documents separate search, training, and user-triggered bot controls.
Choose a cadence that matches the decision and likely refresh window, and re-test after a material change using the same conditions. More repeats help quantify variability, but no cadence guarantees a citation or recommendation.
