Prove: repeat the panel and report what moved
Use Prove after a dated, accepted fix has shipped. Re-run the same prompt wording on the same declared surfaces under the original outcome and exclusion rules; preserve cited, not cited, unavailable, and failed states; and publish raw numerators and denominators beside the comparison. Report movement, non-movement, or reversal as an observed temporal association—not proof that the fix caused the result or a guarantee that it will persist.
What is the Prove stage for?
Prove compares a registered baseline with a post-shipment panel. Its purpose is to make the sequence inspectable: what was observed before, exactly what shipped and when, what was observed afterward, which cells were unavailable or failed, and how the result should be interpreted.
A matched before-and-after can show that an outcome changed after a shipment. It cannot by itself isolate that shipment from provider updates, index changes, source changes, competitors, location, account state, or ordinary answer variation.
Which Prove resource should you use?
Choose the route based on what is missing from the proof record. Provider dashboards and RankEcho prompt observations use different units; keep each source's definitions and denominators intact rather than collapsing them into one universal rank.
| Proof need | Start with | Required output |
|---|---|---|
| No defensible pre-change panel exists | /learn/baseline-nobody-collects | A prospective baseline for the next change, labeled thin when observations are sparse |
| Repeat a shipped fix against its baseline | /product/proof-loop | Matched prompt cells, strict split time, outcomes, exclusions, and confidence limits |
| Continue observing a fixed panel | /product/citation-monitoring | Stable panel version, cadence, per-run evidence, and provider-specific outcomes |
| Interpret disagreement across surfaces | /learn/cross-engine-citation-agreement | Per-surface results and denominators instead of a misleading blended verdict |
| Inspect a public first-party outcome record | /proof-ledger | Dated observation sequence with non-movement and limits left visible |
| Check the definitions and causal limits | /methodology | Declared measurement unit, exclusions, comparison rules, and allowed conclusion |
What must a matched proof record contain?
A reader should be able to recompute the displayed rate and see every exception. Preserve cell-level evidence; a percentage without its planned and observed populations hides whether movement came from outcomes or availability.
- The inherited Find observation ID, panel version, verbatim prompts, declared surfaces, outcome rules, and exclusions.
- The exact Fix artifact, canonical target, pre-change version, accepted deployment version, and split timestamp.
- For each period: planned, available, successful, observed, unavailable, and failed cell counts.
- Separate raw outcomes for mention, owned-source link, other citation, recommendation, absence, unavailable, and failed—not a forced binary.
- Raw answers and visible source URLs needed to audit the classification.
- Numerators, denominators, absolute differences, collection windows, reversals, and confidence or thin-sample notes.
- Known concurrent changes and an explicit statement that temporal movement is not a causal estimate by itself.
Illustrative synthetic example: a matched increase with causal limits
This example is illustrative and synthetic. It is not customer data, a benchmark, a product result, or evidence that a fix caused an answer-system change.
Before shipment, a fixed panel has 12 planned cells: 11 return valid observed answers and one does not run. The target receives an owned-source citation in 3 of 11 observed cells. After the dated shipment, the same panel again has 12 planned cells: 11 valid observed answers and one unavailable cell. The target receives an owned-source citation in 6 of 11 observed cells.
The proof record reports an observed increase of 3 cited cells, from 3/11 to 6/11, while displaying both non-observed cells and the raw classifications. It does not say the shipment caused the increase. A provider update, a changed third-party source, competitor activity, or ordinary answer variation could also contribute, and later runs could reverse the result.
How do you decide what the result means?
Use the evidence state to choose the next workflow move. Preserve non-movement and reversal; deleting an inconvenient run changes the record rather than improving it.
| Observed state | Allowed conclusion | Next handoff |
|---|---|---|
| Matched movement with stable denominators | The registered outcome changed after shipment; cause and durability remain uncertain | Continue the panel and inspect alternative explanations in Find |
| No material movement | The bounded shipment was not followed by a clear change in this window | Return to Find before choosing another Fix |
| Reversal after an initial change | The earlier movement did not persist across the later observations | Keep every run visible and reassess variance and sources |
| Availability or failure changed | The periods are not directly comparable without a qualified denominator analysis | Repair the collection protocol; do not label the difference a visibility result |
| Prompt, surface, or rule changed | This is a new panel, not a matched before-and-after | Start a new prospective baseline in Find |
What is in scope, and what is not?
In scope are matched prompt panels, strict pre/post split points, cell-level raw evidence, explicit denominators and exclusions, provider-specific reporting kept in its own unit, transparent confidence limits, and publication of movement, non-movement, and reversal.
Out of scope are reconstructed baselines presented as prospective measurements, changed prompts treated as matched, unavailable cells counted as misses, provider-specific metrics merged into a universal score, selective removal of negative runs, causal attribution from timing alone, and guarantees about persistence, ranking, traffic, revenue, recommendation, mention, or citation.
What exactly leaves Prove?
The output is a proof record containing the inherited protocol and raw baseline; the exact shipped artifact and split timestamp; every post-shipment cell; planned, unavailable, failed, and observed counts; raw numerators, denominators, and differences; concurrent changes and limitations; and a decision to continue observation, return to Find, or test another bounded Fix.
A durable-looking result remains an observation under the registered conditions. Feed the full record—not only the favorable headline—back into Find so the next diagnosis benefits from both the movement and its uncertainty.
Where can you move in the workflow?
Stay in Prove while matched observations are accumulating. Return to Find when the result is ambiguous, unavailable cells break comparability, or a new gap appears. Move to Fix only after the next bounded hypothesis and target are supported.
| Stage | Use it when | Route |
|---|---|---|
| Find | The result needs a new diagnosis or prospective baseline | /resources/find |
| Fix | A new supported hypothesis has one bounded target | /resources/fix |
| Prove | The matched panel is still collecting or being reported | /resources/prove |
Sources reviewed
Provider eligibility and measurement claims below were checked against primary documentation. These records do not establish a universal selection formula, causation, or a guaranteed ranking, impression, recommendation, or citation.
2 claim-level source records
| Claim reviewed | Official source | Review record |
|---|---|---|
| Google's dedicated Generative AI performance report records link impressions from AI Overviews and AI Mode and groups them by page, country, date, and device. Its documentation does not expose a universal prompt rank or selection reason. | Google Search Console: Generative AI performance report | Checked 2026-09-02 · Worldwide rollout stated as August 31, 2026 · Primary-source review; a link impression is not represented as a quotation, endorsement, click, or causal explanation. · Confidence: High |
| Bing's AI Performance report exposes observed citations, cited pages, and sampled grounding queries. Microsoft cautions that citation count is not placement, ranking, authority, or page importance. | Bing Webmaster Blog: AI Performance | Checked 2026-09-02 · Public preview announced February 2026 · Primary-source review; Bing observations are not merged with Google link impressions or RankEcho prompt-engine cells. · Confidence: High |
Frequently asked questions
No. It documents a temporal sequence under matched rules. Provider updates, source changes, competitors, context, and ordinary answer variation remain alternative explanations unless a stronger design addresses them.
Do not reconstruct one and label it prospective. Start a dated baseline now for the next bounded change, disclose that the prior shipment lacks a valid before panel, and label a sparse baseline as thin.
No. Report planned, unavailable, failed, and observed cells separately in each period. If availability changes materially, qualify or stop the comparison rather than turning infrastructure state into a visibility result.
Publish the unchanged numerators and denominators, the collection windows, the shipped artifact, exclusions, and known limitations. Non-movement narrows the next diagnosis and is part of the evidence record.
When the declared window has been reported with every cell and limit visible and a next decision is recorded. Continued monitoring may still be appropriate because answer-system outcomes and source availability can change later.
