Answer volatility: measure repeat-to-repeat variation
Answer volatility is repeat-to-repeat disagreement in a predeclared answer-feature signature for the same prompt, surface, window, and controlled context. For each group, compare every unordered pair of eligible observed repeats; divide unequal signatures by all comparable pairs, report n/N with per-surface counts, and keep unavailable or failed runs outside the denominator. A group with fewer than two observed repeats and an overall zero denominator are not available, not 0%. This is a descriptive sample measure, not a provider property or a causal test of a page change. Site traffic is a separate outcome and is not required.
What is the scope of answer volatility?
This method measures disagreement in one declared output-feature construct within a finite registered sample. It does not automatically measure prose diversity, citation rate, accuracy, ranking, share of voice, or general model reliability.
Define the decision and feature signature first. A result describes only the recorded prompt, surface, window, contexts, repeats, and rubric.
Which context fields must remain fixed?
Retain prompt and cell ID, verbatim prompt, service and surface, window, locale, account or session state, exposed model label, tool or search mode, time policy, repeat index, and collection procedure.
Record provider-driven changes as limitations. Identical visible prompts do not establish that rewritten searches, retrieval context, model state, or source candidates were identical.
| Context | Service, surface, and access route | Controlled fields | Windows and repeat plan | Signature and collection rule |
|---|---|---|---|---|
| CTX-A | Synthetic Answer Service · Surface A · fictional consumer web surface | en-US · fictional signed-out session · model not exposed · fresh session for every repeat | 2026-08-31T15:00:00Z/2026-08-31T15:10:00Z · 2026-09-01T15:00:00Z/2026-09-01T15:10:00Z · 3 registered repeats per window | northstar-g02-ie-v1 · Submit the exact prompt once per fresh fictional session and preserve the complete answer before coding. |
| CTX-B | Synthetic Answer Service · Surface B · fictional search answer surface | en-US · fictional signed-out session · model not exposed · fresh session for every repeat | 2026-08-31T15:20:00Z/2026-08-31T15:30:00Z · 2026-09-01T15:20:00Z/2026-09-01T15:30:00Z · 3 registered repeats per window | northstar-g02-ie-v1 · Submit the exact prompt once per fresh fictional session and preserve the complete answer or run-state evidence before coding. |
How is an answer-feature signature registered?
Define every feature before reading outcomes, including positive, negative, and unclassifiable rules. Preserve the raw answer and evidence pointer for every classification, and calibrate human coding under the same rubric.
In the fictional G02 record, I=1 means the correct 12-month standard-condition interval is present and E=1 means the correct after-shock exception is present. A visible observed answer coded 00 contains neither required fact; it is not a missing run.
How do unavailable and failed runs affect the denominator?
Only classified observed repeats create comparable pairs. Never convert collection state into answer content.
| Run state | Signature | Pair treatment |
|---|---|---|
| Observed and classifiable | 00, 01, 10, or 11 under the frozen rubric | Eligible for within-group unordered pairs |
| Observed but unclassifiable | None | Report separately; do not force a signature |
| Unavailable | None | Report the reason; exclude from answer pairs |
| Failed | None | Report the error; exclude from answer pairs |
| Planned | None | Not yet observed; report separately |
What is the manual pairwise volatility formula?
Within each prompt-surface-window group containing k classified observed repeats, enumerate all unordered pairs i<j. Pair volatility equals unequal-signature pairs divided by k(k-1)/2 comparable pairs. Publish n/N, percentage, group results, and the pooled numerator and denominator.
A group with fewer than two classified observed repeats contributes no pairs and is reported as insufficient. If the pooled denominator is zero, report the volatility rate as not available; 0/0 is undefined and must not be displayed as 0%. No repeat count or decision threshold is universal.
Which worked synthetic cell ledger produces the result?
SYNTHETIC EXAMPLE — Northstar S1, its pages, sources, surfaces, and observations are fictional; this is not RankEcho data, customer data, provider evidence, or a benchmark.
The six stable G02 cells use two generic surfaces and three repeats only for this example. V06 changes from unavailable to failed and has no answer signature in either window; it never becomes a negative or a comparable pair.
| Cell | Surface and repeat | Frozen context | Baseline state and signature | Baseline evidence | Later state and signature | Later evidence |
|---|---|---|---|---|---|---|
| V01 | Surface A · R1 | CTX-A | Observed · 10 | synthetic-evidence/V01/baseline-answer | Observed · 11 | synthetic-evidence/V01/later-answer |
| V02 | Surface A · R2 | CTX-A | Observed · 00 | synthetic-evidence/V02/baseline-answer | Observed · 10 | synthetic-evidence/V02/later-answer |
| V03 | Surface A · R3 | CTX-A | Observed · 10 | synthetic-evidence/V03/baseline-answer | Observed · 11 | synthetic-evidence/V03/later-answer |
| V04 | Surface B · R1 | CTX-B | Observed · 10 | synthetic-evidence/V04/baseline-answer | Observed · 11 | synthetic-evidence/V04/later-answer |
| V05 | Surface B · R2 | CTX-B | Observed · 10 | synthetic-evidence/V05/baseline-answer | Observed · 11 | synthetic-evidence/V05/later-answer |
| V06 | Surface B · R3 | CTX-B | Unavailable · no signature | synthetic-evidence/V06/baseline-unavailable | Failed · no signature | synthetic-evidence/V06/later-failure |
How does the worked synthetic arithmetic reconcile?
Baseline Surface A has signatures 10, 00, 10: two of three unordered pairs disagree. Surface B has 10, 10: zero of one pair disagrees. The pooled result is 2/4 = 50.0%.
Later Surface A has 11, 10, 11: two of three pairs disagree. Surface B has 11, 11: zero of one pair disagrees. The later pooled result also equals 2/4 = 50.0%.
| Window | Surface A | Surface B | Pooled within-surface pairs |
|---|---|---|---|
| Baseline | 2/3 = 66.7% | 0/1 = 0.0% | 2/4 = 50.0% |
| Later | 2/3 = 66.7% | 0/1 = 0.0% | 2/4 = 50.0% |
| Reconciliation | 3 + 1 comparable pairs | Included above | 4 total pairs in each window |
| Zero-denominator rule | Fewer than two observed repeats = insufficient | Fewer than two observed repeats = insufficient | No pairs = not available |
How do completeness and volatility remain separate?
V01–V05 are observed in both windows. Complete 11 signatures move from 0/5 = 0.0% at baseline to 4/5 = 80.0% later, a descriptive +80.0 percentage-point change, while pair volatility remains 2/4 = 50.0% in both windows.
Report this as later completeness movement under the fictional registered cells. Do not say the answer became stable or that the content edit caused the movement; context, query rewriting, retrieval, system changes, source changes, and sampling remain plausible explanations.
How does this manual measure differ from RankEcho product outputs?
RankEcho does not currently publish this feature-signature pair-volatility metric. Product engine-by-run observations, the fixed-prompt retest, product citation outputs, and this companion manual coding use different records and units.
Do not claim that the Proof Loop, scorecard, or citation monitoring reproduces the signature formula. The current proof-report template is a broader before-and-after framework; it does not compute or store feature-signature pair volatility. Attach a companion cell, signature, and unordered-pair ledger keyed by cell ID, then use the proof report only for the surrounding protocol, limitations, and decision narrative.
Which deterministic validation checks complete Prove?
Validate the declared sample measure without extrapolating it to a provider population or waiting for site traffic.
| Check | Pass condition | Evidence |
|---|---|---|
| Cell identity | Cell IDs and prompt-surface-window-repeat keys are unique | Registered cell ledger |
| Rubric | I and E rules were frozen before classification | Versioned feature definitions |
| Run state | Observed signatures and unavailable or failed nulls are consistent | Raw answer and run evidence |
| Pair enumeration | Every unordered pair stays within one surface and window | Derived pair list |
| Arithmetic | Per-surface and pooled n/N values reconcile exactly | Computed projection |
| Missingness | Insufficient groups and zero denominator render not available | State counts and display check |
| Interpretation | Completeness, product metrics, causality, and population claims remain separate | Published limitations |
When should the volatility protocol repeat or stop?
Repeat only under a registered cadence or a clearly versioned new design. Stop and mark the comparison insufficient when the rubric drifts, evidence is missing, or contexts cannot be compared.
Do not keep sampling until a favorable value appears. Site traffic is a separate outcome and is not required to calculate, report, or publish the registered result.
Who maintains this answer-volatility guide?
Written and maintained by Abiot Y. Derbie. Published 2026-09-07; sources checked 2026-09-07; last updated 2026-09-07. The worked ledger is synthetic.
Recheck provider surface behavior, context fields, feature definitions, source validity, and measurement limits before applying the method. Send corrections through the contact page.
Sources reviewed
Provider eligibility and measurement claims below were checked against primary documentation. These records do not establish a universal selection formula, causation, or a guaranteed ranking, impression, recommendation, or citation.
5 claim-level source records
| Claim reviewed | Official source | Review record |
|---|---|---|
| Google says AI Overviews and AI Mode may use query fan-out, may use different models and techniques, and can show varying responses and links; AI Overviews do not trigger for every query, and eligibility does not guarantee crawling, indexing, or serving. | Google Search: AI features and your website | Checked 2026-09-07 · Google Search documentation last updated December 10, 2025 · Primary-source documentation check; this does not expose all derived queries, quantify prompt demand, prescribe one page per fan-out query, explain a particular result, or generalize to another provider. · Confidence: High |
| OpenAI says ChatGPT Search may rewrite a user question into one or more targeted partner queries; location and enabled memory can affect that process; web-search responses may include citations, which can be incomplete, outdated, or incorrect. | OpenAI Help: searching the web with ChatGPT | Checked 2026-09-07 · Current consumer ChatGPT Search guidance · Primary-source documentation check; the guidance does not expose every rewritten query, define a ranking formula, guarantee placement, or describe API and non-Search surfaces universally. · Confidence: High |
| OpenAI says generative models may produce different outputs from the same input and recommends scoped, task-specific, logged, repeated or continuous evaluation rather than generic or impressionistic assessment. | OpenAI API: evaluation best practices | Checked 2026-09-07 · Current OpenAI developer evaluation guidance · Developer and API evaluation guidance does not define consumer ChatGPT Search behavior, an AI-search volatility metric, a repeat count, or a pass threshold. · Confidence: High |
| NIST notes that prompt sensitivity and heterogeneous contexts can widen gaps between benchmark and real-world use, calls for documented measurement limits and context, and recommends reviewing sources and citations during ongoing monitoring. | NIST AI 600-1: Generative AI Profile | Checked 2026-09-07 · NIST AI 600-1, July 2024 · Government risk-management guidance; it is not a search-provider protocol, causal model, volatility estimate, or prescribed repeat count. · Confidence: High |
| The arXiv preprint reports output variation from repeated identical prompts and settings in a sentiment-analysis experiment, and separately recommends retaining model, version, provider, access route, query date, verbatim prompts and settings, every run's inputs and raw outputs, and the number of draws for reproducibility. | arXiv:2607.24372 preprint | Checked 2026-09-07 · Preprint version checked September 7, 2026 · The empirical domain is sentiment analysis, not AI search; it cannot estimate AI-search volatility, validate the synthetic rate, or prescribe repeats or thresholds. · Confidence: Medium |
Frequently asked questions
It establishes one discordant observed pair under the registered signature, not a general property of the provider or model.
No. Unavailable and failed runs have no answer signature and remain outside the comparable-pair denominator.
No. A 00 signature is an eligible observed answer containing neither registered fact; no answer has no signature.
There is no universal number. Choose and disclose a design appropriate to the decision, and report thin or zero denominators plainly.
No. It describes repeat-pair disagreement in the registered sample. Site traffic is a separate outcome and is not required, and temporal movement alone does not establish causation.
