Home / Resources / Prove / Answer volatility: measure repeat-to-repeat variation
AI Search Intelligence

Answer volatility: measure repeat-to-repeat variation

The short answer

Answer volatility is repeat-to-repeat disagreement in a predeclared answer-feature signature for the same prompt, surface, window, and controlled context. For each group, compare every unordered pair of eligible observed repeats; divide unequal signatures by all comparable pairs, report n/N with per-surface counts, and keep unavailable or failed runs outside the denominator. A group with fewer than two observed repeats and an overall zero denominator are not available, not 0%. This is a descriptive sample measure, not a provider property or a causal test of a page change. Site traffic is a separate outcome and is not required.

What is the scope of answer volatility?

This method measures disagreement in one declared output-feature construct within a finite registered sample. It does not automatically measure prose diversity, citation rate, accuracy, ranking, share of voice, or general model reliability.

Define the decision and feature signature first. A result describes only the recorded prompt, surface, window, contexts, repeats, and rubric.

Which context fields must remain fixed?

Retain prompt and cell ID, verbatim prompt, service and surface, window, locale, account or session state, exposed model label, tool or search mode, time policy, repeat index, and collection procedure.

Record provider-driven changes as limitations. Identical visible prompts do not establish that rewritten searches, retrieval context, model state, or source candidates were identical.

ContextService, surface, and access routeControlled fieldsWindows and repeat planSignature and collection rule
CTX-ASynthetic Answer Service · Surface A · fictional consumer web surfaceen-US · fictional signed-out session · model not exposed · fresh session for every repeat2026-08-31T15:00:00Z/2026-08-31T15:10:00Z · 2026-09-01T15:00:00Z/2026-09-01T15:10:00Z · 3 registered repeats per windownorthstar-g02-ie-v1 · Submit the exact prompt once per fresh fictional session and preserve the complete answer before coding.
CTX-BSynthetic Answer Service · Surface B · fictional search answer surfaceen-US · fictional signed-out session · model not exposed · fresh session for every repeat2026-08-31T15:20:00Z/2026-08-31T15:30:00Z · 2026-09-01T15:20:00Z/2026-09-01T15:30:00Z · 3 registered repeats per windownorthstar-g02-ie-v1 · Submit the exact prompt once per fresh fictional session and preserve the complete answer or run-state evidence before coding.

How is an answer-feature signature registered?

Define every feature before reading outcomes, including positive, negative, and unclassifiable rules. Preserve the raw answer and evidence pointer for every classification, and calibrate human coding under the same rubric.

In the fictional G02 record, I=1 means the correct 12-month standard-condition interval is present and E=1 means the correct after-shock exception is present. A visible observed answer coded 00 contains neither required fact; it is not a missing run.

How do unavailable and failed runs affect the denominator?

Only classified observed repeats create comparable pairs. Never convert collection state into answer content.

Run stateSignaturePair treatment
Observed and classifiable00, 01, 10, or 11 under the frozen rubricEligible for within-group unordered pairs
Observed but unclassifiableNoneReport separately; do not force a signature
UnavailableNoneReport the reason; exclude from answer pairs
FailedNoneReport the error; exclude from answer pairs
PlannedNoneNot yet observed; report separately

What is the manual pairwise volatility formula?

Within each prompt-surface-window group containing k classified observed repeats, enumerate all unordered pairs i<j. Pair volatility equals unequal-signature pairs divided by k(k-1)/2 comparable pairs. Publish n/N, percentage, group results, and the pooled numerator and denominator.

A group with fewer than two classified observed repeats contributes no pairs and is reported as insufficient. If the pooled denominator is zero, report the volatility rate as not available; 0/0 is undefined and must not be displayed as 0%. No repeat count or decision threshold is universal.

Which worked synthetic cell ledger produces the result?

SYNTHETIC EXAMPLE — Northstar S1, its pages, sources, surfaces, and observations are fictional; this is not RankEcho data, customer data, provider evidence, or a benchmark.

The six stable G02 cells use two generic surfaces and three repeats only for this example. V06 changes from unavailable to failed and has no answer signature in either window; it never becomes a negative or a comparable pair.

CellSurface and repeatFrozen contextBaseline state and signatureBaseline evidenceLater state and signatureLater evidence
V01Surface A · R1CTX-AObserved · 10synthetic-evidence/V01/baseline-answerObserved · 11synthetic-evidence/V01/later-answer
V02Surface A · R2CTX-AObserved · 00synthetic-evidence/V02/baseline-answerObserved · 10synthetic-evidence/V02/later-answer
V03Surface A · R3CTX-AObserved · 10synthetic-evidence/V03/baseline-answerObserved · 11synthetic-evidence/V03/later-answer
V04Surface B · R1CTX-BObserved · 10synthetic-evidence/V04/baseline-answerObserved · 11synthetic-evidence/V04/later-answer
V05Surface B · R2CTX-BObserved · 10synthetic-evidence/V05/baseline-answerObserved · 11synthetic-evidence/V05/later-answer
V06Surface B · R3CTX-BUnavailable · no signaturesynthetic-evidence/V06/baseline-unavailableFailed · no signaturesynthetic-evidence/V06/later-failure

How does the worked synthetic arithmetic reconcile?

Baseline Surface A has signatures 10, 00, 10: two of three unordered pairs disagree. Surface B has 10, 10: zero of one pair disagrees. The pooled result is 2/4 = 50.0%.

Later Surface A has 11, 10, 11: two of three pairs disagree. Surface B has 11, 11: zero of one pair disagrees. The later pooled result also equals 2/4 = 50.0%.

WindowSurface ASurface BPooled within-surface pairs
Baseline2/3 = 66.7%0/1 = 0.0%2/4 = 50.0%
Later2/3 = 66.7%0/1 = 0.0%2/4 = 50.0%
Reconciliation3 + 1 comparable pairsIncluded above4 total pairs in each window
Zero-denominator ruleFewer than two observed repeats = insufficientFewer than two observed repeats = insufficientNo pairs = not available

How do completeness and volatility remain separate?

V01–V05 are observed in both windows. Complete 11 signatures move from 0/5 = 0.0% at baseline to 4/5 = 80.0% later, a descriptive +80.0 percentage-point change, while pair volatility remains 2/4 = 50.0% in both windows.

Report this as later completeness movement under the fictional registered cells. Do not say the answer became stable or that the content edit caused the movement; context, query rewriting, retrieval, system changes, source changes, and sampling remain plausible explanations.

How does this manual measure differ from RankEcho product outputs?

RankEcho does not currently publish this feature-signature pair-volatility metric. Product engine-by-run observations, the fixed-prompt retest, product citation outputs, and this companion manual coding use different records and units.

Do not claim that the Proof Loop, scorecard, or citation monitoring reproduces the signature formula. The current proof-report template is a broader before-and-after framework; it does not compute or store feature-signature pair volatility. Attach a companion cell, signature, and unordered-pair ledger keyed by cell ID, then use the proof report only for the surrounding protocol, limitations, and decision narrative.

Which deterministic validation checks complete Prove?

Validate the declared sample measure without extrapolating it to a provider population or waiting for site traffic.

CheckPass conditionEvidence
Cell identityCell IDs and prompt-surface-window-repeat keys are uniqueRegistered cell ledger
RubricI and E rules were frozen before classificationVersioned feature definitions
Run stateObserved signatures and unavailable or failed nulls are consistentRaw answer and run evidence
Pair enumerationEvery unordered pair stays within one surface and windowDerived pair list
ArithmeticPer-surface and pooled n/N values reconcile exactlyComputed projection
MissingnessInsufficient groups and zero denominator render not availableState counts and display check
InterpretationCompleteness, product metrics, causality, and population claims remain separatePublished limitations

When should the volatility protocol repeat or stop?

Repeat only under a registered cadence or a clearly versioned new design. Stop and mark the comparison insufficient when the rubric drifts, evidence is missing, or contexts cannot be compared.

Do not keep sampling until a favorable value appears. Site traffic is a separate outcome and is not required to calculate, report, or publish the registered result.

Who maintains this answer-volatility guide?

Written and maintained by Abiot Y. Derbie. Published 2026-09-07; sources checked 2026-09-07; last updated 2026-09-07. The worked ledger is synthetic.

Recheck provider surface behavior, context fields, feature definitions, source validity, and measurement limits before applying the method. Send corrections through the contact page.

Sources reviewed

Provider eligibility and measurement claims below were checked against primary documentation. These records do not establish a universal selection formula, causation, or a guaranteed ranking, impression, recommendation, or citation.

5 claim-level source records
Checked 2026-09-07 · Primary-source diagnostic review · Confidence is recorded per claim.
Claim reviewedOfficial sourceReview record
Google says AI Overviews and AI Mode may use query fan-out, may use different models and techniques, and can show varying responses and links; AI Overviews do not trigger for every query, and eligibility does not guarantee crawling, indexing, or serving.Google Search: AI features and your websiteChecked 2026-09-07 · Google Search documentation last updated December 10, 2025 · Primary-source documentation check; this does not expose all derived queries, quantify prompt demand, prescribe one page per fan-out query, explain a particular result, or generalize to another provider. · Confidence: High
OpenAI says ChatGPT Search may rewrite a user question into one or more targeted partner queries; location and enabled memory can affect that process; web-search responses may include citations, which can be incomplete, outdated, or incorrect.OpenAI Help: searching the web with ChatGPTChecked 2026-09-07 · Current consumer ChatGPT Search guidance · Primary-source documentation check; the guidance does not expose every rewritten query, define a ranking formula, guarantee placement, or describe API and non-Search surfaces universally. · Confidence: High
OpenAI says generative models may produce different outputs from the same input and recommends scoped, task-specific, logged, repeated or continuous evaluation rather than generic or impressionistic assessment.OpenAI API: evaluation best practicesChecked 2026-09-07 · Current OpenAI developer evaluation guidance · Developer and API evaluation guidance does not define consumer ChatGPT Search behavior, an AI-search volatility metric, a repeat count, or a pass threshold. · Confidence: High
NIST notes that prompt sensitivity and heterogeneous contexts can widen gaps between benchmark and real-world use, calls for documented measurement limits and context, and recommends reviewing sources and citations during ongoing monitoring.NIST AI 600-1: Generative AI ProfileChecked 2026-09-07 · NIST AI 600-1, July 2024 · Government risk-management guidance; it is not a search-provider protocol, causal model, volatility estimate, or prescribed repeat count. · Confidence: High
The arXiv preprint reports output variation from repeated identical prompts and settings in a sentiment-analysis experiment, and separately recommends retaining model, version, provider, access route, query date, verbatim prompts and settings, every run's inputs and raw outputs, and the number of draws for reproducibility.arXiv:2607.24372 preprintChecked 2026-09-07 · Preprint version checked September 7, 2026 · The empirical domain is sentiment analysis, not AI search; it cannot estimate AI-search volatility, validate the synthetic rate, or prescribe repeats or thresholds. · Confidence: Medium

Frequently asked questions

Does one changed answer prove volatility?

It establishes one discordant observed pair under the registered signature, not a general property of the provider or model.

Do failed runs count as disagreement?

No. Unavailable and failed runs have no answer signature and remain outside the comparable-pair denominator.

Is a 00 signature the same as no answer?

No. A 00 signature is an eligible observed answer containing neither registered fact; no answer has no signature.

How many repeats are enough?

There is no universal number. Choose and disclose a design appropriate to the decision, and report thin or zero denominators plainly.

Does reduced volatility prove a fix worked?

No. It describes repeat-pair disagreement in the registered sample. Site traffic is a separate outcome and is not required, and temporal movement alone does not establish causation.

Open the proof report framework →
Last updated 2026-09-07 · RankEcho · Operated by Nexus Decision Systems LLC