Home / Resources / Prove / Fixed-prompt retest: a before/after proof protocol
AI Search Intelligence

Fixed-prompt retest: a before/after proof protocol

The short answer

A fixed-prompt retest repeats the same versioned prompts on the same declared surfaces or configured adapters after one dated shipment. Give every prompt-surface-repeat combination a stable cell ID, preserve every raw observed answer and source URL, and compare outcomes over cells observed in both windows. Report planned, unavailable, failed, and observed counts plus per-window available-case rates separately as coverage context. Movement after a fix is temporal evidence, not proof that the fix caused it.

Scope box: what does this page cover?

One versioned prompt panel measured before and after one dated shipment under declared rules.

A later change is an observation after shipment, not proof that the shipment caused it.

In scopeOut of scope
Fixed prompts and surfacesAfter-only screenshots
Explicit denominatorsFailed runs counted as misses
Raw answer and URL evidenceCross-provider blended rank
Temporal before/after observationsCausal attribution

What method does a fixed-prompt retest use?

Register the prompt panel, surfaces or configured adapters, locale and account state when controllable, repeat count, schedule, outcome definitions, source-resolution rule, exclusions, and stable cell-key rule before the after window begins. Preserve the baseline rather than regenerating it after shipment.

Repeat that protocol after one dated, accepted deployment. A provider surface that is unavailable and a run that fails are reported as such; neither silently becomes a non-citation.

What must the baseline preserve?

A reviewer should be able to reconstruct every planned and observed cell.

  • Panel ID and version, stable cell ID, verbatim prompt, intent, audience, repeat index, and inclusion reason.
  • Declared surface or configured adapter, model or feature label when known, locale, account state, and UTC timestamp.
  • Planned, available, successful, observed, unavailable, and failed status with explicit reasons.
  • Raw answer, visible source URLs, brand-match rule, and separate mention, owned-link, other-citation, recommendation, competitor, and absence outcomes.
  • The linked Find record, fix target, before version, deployment time, accepted version, and credible alternatives.

Which inputs must remain fixed?

Keep prompt wording, panel version, declared surface or adapter, locale and account state when controllable, repeat policy, outcome definitions, URL ownership rules, brand and competitor matching, and exclusions unchanged across compared windows.

If any material field changes, create a new panel version and report it as a new series rather than quietly appending it to the old denominator. Provider model or product changes outside your control belong in the limitations record.

How are the metrics and denominators calculated?

For each window and surface, show planned, unavailable, failed, and observed cell counts. Treat the primary matched denominator as stable cell IDs with inspectable answers in both windows; calculate a paired outcome rate as outcome-positive common observed cells divided by common observed cells. Show outcome-positive observed cells divided by eligible observed cells for each window only as available-case coverage context.

RankEcho can have up to five configured adapters—perplexity, openai, anthropic, gemini, google_aio—but availability and semantics differ by adapter. Google AI Overviews is reported only when an overview is shown and retrievable; otherwise that cell is marked skipped/unavailable rather than counted as a citation miss.

FieldDefinitionDenominator rule
Observed cellAn inspectable answer was returned under the registered ruleEligible for that window's available-case context
Common observed cellThe same stable cell ID has an inspectable answer in both windowsPrimary matched before/after denominator
Unavailable cellThe declared feature or surface was not availableReport separately; exclude from observed rate
Failed cellThe attempt did not return an inspectable answerReport separately; exclude from observed rate
Paired owned-link rateCommon observed cells containing a resolved owned-source URL in the named windowOwned-link positive / common observed cells

What does a worked example show?

Illustrative synthetic example—this is not customer evidence or a benchmark. Six prompts × two surfaces × one repeat produce 12 planned cells in each window. Before deployment, one cell is unavailable and one fails, leaving 10 observed cells; after deployment, no cell is unavailable and one fails, leaving 11 observed cells. The twelve-cell ledger preserves every stable ID, status, outcome, and evidence pointer; nine stable IDs were observed in both windows.

Within those nine common cells, one gain, one persisted positive, and seven persisted negatives move owned links from one of nine (11.1%) to two of nine (22.2%); there are zero reversals. A reversal would be a common cell that changes from positive before to negative after: it remains in the common denominator but counts as positive only in the before numerator. The three non-common cells retain their explicit failure or unavailability details. The 10.0% and 18.2% available-case rates use different observed populations and remain coverage context only. The report does not say the fix caused the movement.

What are false positives, noise, and alternative explanations?

A text mention is not automatically an owned link, a visible URL is not automatically a recommendation, and a provider-reported link impression is not interchangeable with a prompt-panel citation. Redirects, aliases, and third-party profiles can also be misclassified without a frozen ownership rule.

Model and index updates, retrieval changes, feature availability, source changes, competitors, location, device, account state, prompt sensitivity, and ordinary output variability can explain before/after differences. Repeats and unchanged comparison prompts can improve interpretation but do not establish causality.

What deployment validation is required before the after window?

Confirm the exact target, saved before version, deployed version or hash, UTC shipment time, owner, acceptance result, and rollback status. For the robots pilot, validate the served policy for the exact crawler role and URL, then inspect delivery controls separately.

Do not begin the after window while the artifact is unverified or while a rollback changes exposure. If timing changes, record the new boundary before classifying observations.

How should the proof report state the result?

State the registered protocol, stable cell-key rule, before and after windows, shipment, per-window coverage counts, common-cell denominator, outcome definitions, observed difference, raw-evidence location, uncertainty statement, concurrent changes, confounders, and decision. Keep no-movement, adverse, unavailable, and failed evidence visible.

Use language such as ‘owned-link observations increased after the shipment under this registered panel’. Avoid ‘the fix made the provider cite us’, ranking claims, extrapolation beyond the sample, or combining provider metrics into an unsupported score.

When should the team repeat, stop, or return to Find?

Continue only through the predeclared window or cadence. Stop and label the result insufficient if evidence is missing, the panel drifted, or too few cells were observed. Return to Find when results challenge the hypothesis, a different gap becomes plausible, or the protocol no longer represents the intended buyer job.

Do not keep sampling until a favorable result appears. Register any extended window or revised panel as a new version before collecting it.

Who owns the page, review, tests, and corrections?

Author and accountable publisher: Abiot Y. Derbie, RankEcho founder. Published 2026-09-02; updated 2026-09-02. Independent reviewer: unassigned. No independent external review is claimed.

Tested scope: Static proof-method contract, metric and denominator examples, product adapter labels, and dated provider guidance; no live retest or causal analysis was performed for this page. Material provider and standards statements use the dated source records below. Examples are synthetic, not customer results. Send corrections through /contact; material corrections should update the visible date and version history.

Sources reviewed

Provider eligibility and measurement claims below were checked against primary documentation. These records do not establish a universal selection formula, causation, or a guaranteed ranking, impression, recommendation, or citation.

5 claim-level source records
Checked 2026-09-01 · Primary-source diagnostic review · Confidence is recorded per claim.
Claim reviewedOfficial sourceReview record
Google says normal Search indexing and snippet controls govern eligibility for AI Overviews and AI Mode; there are no additional technical requirements or special AI schema files.Google Search AI features documentationChecked 2026-09-01 · AI features documentation updated 2025-12-10 · Primary-source documentation review; eligibility does not guarantee selection or presentation in an AI feature. · Confidence: High
OpenAI documents OAI-SearchBot for ChatGPT search, GPTBot for potential model training, and ChatGPT-User for user-triggered actions; the controls are independent and robots.txt rules may not apply to ChatGPT-User.OpenAI crawler documentationChecked 2026-09-01 · Current OAI-SearchBot, GPTBot, and ChatGPT-User documentation · Primary-source documentation review; no claim that a permitted bot will index, rank, or cite a page. · Confidence: High
Google says its generative AI Search features use core Search systems: a page must be indexed, snippet-eligible, and included in Search generative AI features in Search Console. It requires no special AI file, content chunking, or AI-specific structured data, and eligibility does not guarantee display.Google guide to generative AI Search optimizationChecked 2026-09-01 · Google Search guidance updated July 10, 2026 · Primary-source documentation review; this is a Google Search eligibility and optimization boundary, not a universal answer-engine formula or ranking guarantee. · Confidence: High
Google's Generative AI performance report counts link impressions from AI Overviews and AI Mode and groups them by page, country, date, or device. It does not document prompt, ranking, citation-cause, or selection-formula fields.Google Search Console: Generative AI performance reportChecked 2026-09-01 · Worldwide rollout stated as August 31, 2026 · Primary-source documentation review; report visibility can be absent with insufficient impressions or exclusion, property and page aggregation can differ, and these link impressions remain distinct from other systems' metrics. · Confidence: High
Bing's AI Performance report counts observed citations and exposes sampled grounding queries, but Microsoft says citation count is not placement, ranking, authority, or page importance.Bing Webmaster Blog: AI PerformanceChecked 2026-09-01 · Public preview announced February 2026 · Primary-source documentation review; Bing metrics are treated as observations with their stated sampling and interpretation limits. · Confidence: High

Frequently asked questions

Why must the prompts remain fixed?

Changing the questions changes the sample. Fixed wording and a versioned panel make the compared windows inspectable; material changes start a new series.

Do unavailable runs count as citation misses?

No. Report unavailable and failed cells separately. Exclude them from available-case outcome rates and from the primary paired denominator, which contains only stable cell IDs observed in both windows. Show every denominator so readers can see availability changes.

Does before/after movement prove the fix worked?

No. It shows temporal movement under the registered protocol. Provider, index, retrieval, source, competitor, context, and ordinary variability remain plausible explanations.

Can Google and Bing AI metrics be combined with prompt-panel rates?

Not as though they were the same unit. Report each provider's documented metric separately, then describe any directional agreement with its limits.

Start the fixed-prompt retest →
Last updated 2026-09-02 · RankEcho · Operated by Nexus Decision Systems LLC