AI SEO Testing: Build a Treatment and Control Plan
Assign the page or source cluster that receives a change, then use fixed prompts to observe it. Labeling half your questions ‘control’ does not isolate them if they can retrieve the same edited page. Map likely source overlap, pair comparable units, record the assignment and freeze the collection rules first. The plan below is a tested design exercise with no collected AI answers or claimed visibility effect.
Copy the complete pilot plan
A fictional documentation pilot with eight source clusters, four paired blocks and sixteen assigned prompts. Adapt the source map, actual provider context, windows and review criteria before use. All result fields remain unobserved.
{
"schemaVersion": 1,
"id": "fictional-docs-pilot-v1",
"purpose": "pilot-feasibility",
"hypothesis": "Adding supported setup prerequisites to selected documentation may change the frequency of owned-documentation links in the frozen prompt panel. This is a proposal, not a collected result.",
"treatment": "Add an approved prerequisites block to every listed page in each assigned treatment cluster; preserve all existing facts, access and URLs.",
"control": "Retain the saved page versions in each assigned control cluster.",
"unit": "source-cluster",
"sharedChanges": "none",
"surface": "fictional-surface-v1; no engine is called by this example",
"context": "Fictional fixed English context; register the actual surface, model when exposed, locale, account and session rules before a real pilot.",
"blockRationale": "Pair similar documentation tasks before assignment; these labels are teaching assumptions, not empirical proof of comparability.",
"units": [
{
"id": "U01",
"block": "integration",
"pages": [
"https://docs.example/crm",
"https://docs.example/crm/limits"
],
"prompts": [
{
"id": "P01a",
"text": "What must I configure before using the CRM integration?",
"candidateSources": [
"https://docs.example/crm",
"https://docs.example/crm/limits"
]
},
{
"id": "P01b",
"text": "Which permissions does the CRM setup require?",
"candidateSources": [
"https://docs.example/crm"
]
}
]
},
{
"id": "U02",
"block": "integration",
"pages": [
"https://docs.example/helpdesk"
],
"prompts": [
{
"id": "P02a",
"text": "What must I configure before using the helpdesk integration?",
"candidateSources": [
"https://docs.example/helpdesk"
]
},
{
"id": "P02b",
"text": "Which permissions does the helpdesk setup require?",
"candidateSources": [
"https://docs.example/helpdesk"
]
}
]
},
{
"id": "U03",
"block": "data-export",
"pages": [
"https://docs.example/warehouse"
],
"prompts": [
{
"id": "P03a",
"text": "What must I configure before using the data warehouse integration?",
"candidateSources": [
"https://docs.example/warehouse"
]
},
{
"id": "P03b",
"text": "Which permissions does the data warehouse setup require?",
"candidateSources": [
"https://docs.example/warehouse"
]
}
]
},
{
"id": "U04",
"block": "data-export",
"pages": [
"https://docs.example/analytics"
],
"prompts": [
{
"id": "P04a",
"text": "What must I configure before using the analytics platform integration?",
"candidateSources": [
"https://docs.example/analytics"
]
},
{
"id": "P04b",
"text": "Which permissions does the analytics platform setup require?",
"candidateSources": [
"https://docs.example/analytics"
]
}
]
},
{
"id": "U05",
"block": "identity",
"pages": [
"https://docs.example/saml"
],
"prompts": [
{
"id": "P05a",
"text": "What must I configure before using the SAML single sign-on integration?",
"candidateSources": [
"https://docs.example/saml"
]
},
{
"id": "P05b",
"text": "Which permissions does the SAML single sign-on setup require?",
"candidateSources": [
"https://docs.example/saml"
]
}
]
},
{
"id": "U06",
"block": "identity",
"pages": [
"https://docs.example/scim"
],
"prompts": [
{
"id": "P06a",
"text": "What must I configure before using the SCIM provisioning integration?",
"candidateSources": [
"https://docs.example/scim"
]
},
{
"id": "P06b",
"text": "Which permissions does the SCIM provisioning setup require?",
"candidateSources": [
"https://docs.example/scim"
]
}
]
},
{
"id": "U07",
"block": "automation",
"pages": [
"https://docs.example/webhooks"
],
"prompts": [
{
"id": "P07a",
"text": "What must I configure before using the webhooks integration?",
"candidateSources": [
"https://docs.example/webhooks"
]
},
{
"id": "P07b",
"text": "Which permissions does the webhooks setup require?",
"candidateSources": [
"https://docs.example/webhooks"
]
}
]
},
{
"id": "U08",
"block": "automation",
"pages": [
"https://docs.example/scheduler"
],
"prompts": [
{
"id": "P08a",
"text": "What must I configure before using the scheduled jobs integration?",
"candidateSources": [
"https://docs.example/scheduler"
]
},
{
"id": "P08b",
"text": "Which permissions does the scheduled jobs setup require?",
"candidateSources": [
"https://docs.example/scheduler"
]
}
]
}
],
"sentinels": [
{
"id": "S01",
"text": "Which platform is best for managing these integrations?",
"role": "monitoring-only"
}
],
"windows": {
"baseline": [
-14,
-1
],
"deploymentDay": 0,
"followup": [
14,
27
]
},
"repeatsPerPromptPerWindow": 3,
"outcome": "An inspectable answer contains a resolved link to one of the registered owned documentation URLs; retain the raw answer and full source list.",
"missingHandling": "pause-primary-contrast",
"analysisUnit": "source-cluster",
"weighting": "equal-unit",
"plannedContrast": "Average the prompt/repeat values within each cluster and window. Compare changes within each assigned block, then average the block contrasts. No values or effects are calculated in this design-only artifact.",
"powerAnalysis": "not-completed; eight clusters are a teaching pilot, not a justified confirmatory sample",
"interferenceReview": "declared-source-overlap-only; actual provider retrieval and wider spillover remain unverified",
"stopRule": "Pause interpretation if a control page changes, cross-group exposure appears, a required observation is missing, or the registered surface/protocol changes. Report the reason and retain all records.",
"registration": {
"status": "local-draft",
"reference": null
}
}
What receives the change?
Suppose an agency wants to add approved setup prerequisites to selected SaaS documentation. The proposed intervention changes page content. A prompt is an observation probe: repeating it ten times does not create ten independently assigned page changes. Register the pages that receive the edit and keep the measurement unit separate from the assignment unit.
Where several pages share an intervention or likely retrieval exposure, keep them together as one source cluster. The example puts a CRM setup page and its limits page in the same unit. Its two prompts travel with that unit. Do not put one in treatment and call the other an untouched control simply because the questions use different words.
Microsoft's design guidance discusses choosing a randomization unit and leakage through shared components. Applying that principle here requires an operator's source map and a review of shared templates, navigation, brand-wide changes and cross-page references. The helper checks declared overlap; unknown retrieval paths remain unknown.
Which prompts can serve as controls?
A control prompt belongs to a source cluster whose registered pages retain their saved versions. Its role depends on exposure, not on receiving a control label. Before assignment, list candidate sources for each question and investigate prompts that plausibly reach multiple proposed units. Combine the connected units, remove the question with a recorded reason, or choose a different design.
Keep a broad brand question as a separate context check if it can draw from the whole site. The example registers one such sentinel as monitoring-only and excludes it from the assigned-cell schedule. Arrange its separate context observations; do not silently pool it with the treatment/control contrast.
Google explains that AI features can fan out into related searches and return different answers and links. A preflight map cannot guarantee what a live provider will retrieve. If later evidence shows cross-group exposure, preserve the observation and pause the intended interpretation. A domain-wide change may leave no credible within-site control.
How should you pair and assign the units?
Choose pairing criteria before looking at post-change outcomes. Candidate criteria include document template, buyer task, source availability and baseline behavior. In a real project, retain the records supporting those comparisons. Similar labels alone do not establish exchangeability, adequate precision or a credible counterfactual.
The teaching plan uses four blocks with two source clusters each. One member of each block receives treatment and the other remains control. There are 2 to the fourth power, or sixteen, permitted allocations. NIST's blocked-design guidance supports accounting for selected nuisance factors; the exact source-cluster construction here is an original adaptation.
The helper can draw one allocation using system randomness and print the plan checksum, time, selected index and assignment. Save that output with the plan before collection. Do not repeat the draw until preferred pages land in treatment. If an allocation must satisfy additional restrictions, predeclare and implement that restricted assignment mechanism before drawing.
| Teaching block | Two source units | Rule |
|---|---|---|
| Integration | CRM setup + limits; helpdesk setup | One cluster treated, one control |
| Data export | Warehouse setup; analytics setup | One cluster treated, one control |
| Identity | SAML setup; SCIM setup | One cluster treated, one control |
| Automation | Webhooks setup; scheduled jobs | One cluster treated, one control |
What belongs in the frozen protocol?
Specify one supported edit, the retained control version, a falsifiable hypothesis, all unit and prompt IDs, candidate source URLs, assignment record, and the named provider surface. Register wording, context, repeat slots, answer/source retention, URL ownership, and the primary outcome before the follow-up window. Different APIs, consumer apps and search features need separate records.
The example proposes a baseline before day zero and a follow-up after an illustrative waiting interval. These relative windows demonstrate ordering; they are not a recommended fourteen-day crawl or citation lag. Choose actual windows from the task, collection feasibility and observed exposure. Record a changed deployment boundary before interpreting the results.
The plan also records stopping rules, missingness, weighting and the intended cluster-level contrast. Its local-draft label is deliberate. The Center for Open Science describes preregistration as specifying the plan in advance and submitting it to a registry. Copying a worksheet or generating a checksum alone does not complete that process. Disclose later deviations and distinguish exploratory work.
What did the runnable design checks establish?
On September 13, 2026, the Python exercise enumerated all sixteen allocations. Every block had one treatment and one control, each of the eight clusters appeared in treatment exactly eight times, and all repeated observations stayed with their assigned cluster. Reordering the input units or prompts preserved the allocation mapping.
Sixteen assigned prompts, three repeat slots and two windows produce ninety-six planned cells per allocation. Every cell retains status planned and value null. The retained receipt also includes one actual random draw, which selected index 15. That is an executed allocation-helper result, not a launched experiment or a favorable answer observation.
Ten modified plans were rejected, including a source split between units, a prompt crossing source clusters, a shared-template edit, missing runs treated as zero, and repeated prompt runs misidentified as assignment units. These checks test whether the coded protocol is internally consistent. They do not establish the source map's real-world completeness, statistical power or a treatment effect.
| Executed check | Recorded result |
|---|---|
| All permissible paired allocations | 16 checked; each block balanced |
| Assignment units / primary planned cells | 8 source clusters / 96 cells |
| Incorrect plan variants | 10 rejected with explicit reasons |
| Input-order change | Allocation mapping preserved |
| Provider answers / estimated effect | 0 collected / none calculated |
How do you run the helper?
From the repository checkout, run python3 test/fixtures/study_design/design.py to check the included plan. Add --exercise to reproduce the exhaustive allocation and rejection checks. Add --draw to create a fresh proposed assignment; save the exact output before collecting any pilot outcomes. Python's standard library is sufficient and the helper makes no network requests.
Use --plan path/to/your-plan.json to validate or draw from a plan with the same schema. The helper supports paired blocks and a pilot-feasibility purpose; it is not a general experimental-design package. The included fictional surface must be replaced with a real, precisely documented collection context. The exercise mode is specific to the bundled teaching plan.
A passing check means the declared inputs satisfy this protocol. It does not certify a production rollout or permission to change a customer's pages. Assign an operator, retain the before versions, verify the deployed treatment and untouched controls, and inspect source exposure before treating the pilot as interpretable.
How will the observations be analyzed?
Keep all assigned units in the record. The proposed contrast first averages repeated prompt values within each cluster and window, then compares the changes within each treatment/control block and averages those block contrasts. That keeps a frequently sampled cluster from acquiring extra assignment weight. The helper does not calculate this contrast because no answers have been collected.
The conservative example pauses the primary contrast if required observations are missing. Retain failures and unavailable runs with null outcomes, and report coverage separately. Do not turn them into non-citations or drop inconvenient assigned units after seeing results. More elaborate missing-data strategies require an explicit pre-analysis method and sensitivity assessment.
Eight clusters are a small teaching pilot, not a justified confirmatory sample. Repeating prompts can characterize within-unit variability but cannot manufacture independent assignments. A confirmatory design needs a suitable power or sensitivity analysis, an estimator and uncertainty method that respect assignment, and a credible interference assessment. A raw treatment/control difference alone is not a causal result.
When should you use ordinary retesting instead?
If there is only one eligible page, a shared sitewide intervention, unreliable source separation or too little capacity to assess the design, use the existing fixed-prompt retest as a descriptive workflow. A stable before/after record can still support an operational decision. It should describe what appeared after the change without claiming the change caused it.
Keep later citation observations, identified referral visits and paid conversions as separate outcomes. SearchPilot's published GEO testing approach emphasizes page groups and traffic; this guide does not substitute a small prompt panel for that business-outcome evaluation. Choose the measure that answers the registered question and preserve no-change and adverse findings.
After the team approves its observation protocol, RankEcho's Proof Loop supports recurring same-prompt observations around a recorded shipment. Preview that workflow and retain the external design packet alongside it. Cluster assignment, preregistration, power calculations and causal estimation are not presented as native RankEcho features.
Frequently asked questions
Only if the group definition follows credible treatment exposure. Questions that can retrieve the same edited source are not isolated by a label. Assign the source unit and keep its prompts together.
No. It reports executed assignment and protocol checks on a fictional prospective plan. No provider answers were collected and no visibility or causal effect was calculated.
The eight-unit plan is a teaching pilot. Adequacy for inference depends on the estimand, assignment, variation, dependence and effect size of interest; this helper supplies no power result.
No. Inspect the actual assignment and exposure, interference, missingness, concurrent changes and analysis assumptions. A control label alone cannot remove those problems.
Sources reviewed
Material technical claims below were checked against primary provider documentation. The sources support the documented control or signal, not a guarantee of indexing, ranking, an AI impression, or a citation.
5 claim-level source records
| Claim reviewed | Official source | Review record |
|---|---|---|
| Microsoft's experimentation guidance treats the randomization unit and leakage through shared components as design decisions. | Microsoft Research: pre-experiment design | Checked 2026-09-13 · Primary source checked September 13, 2026 · Applied to an original source-cluster planning protocol. The checker rejects declared shared sources and cross-cluster prompts; it cannot discover all real provider interference. · Confidence: High |
| NIST describes randomized blocks as a way to account for selected nuisance factors within a designed experiment. | NIST: randomized block designs | Checked 2026-09-13 · Primary source checked September 13, 2026 · The teaching plan pairs eight source clusters into four blocks. All sixteen paired allocations were executed and checked for one treatment and one control per block. The block labels are assumptions, not measured comparability. · Confidence: High |
| The Center for Open Science distinguishes a plan registered before a study from analyses chosen after seeing results. | Center for Open Science: preregistration | Checked 2026-09-13 · Primary source checked September 13, 2026 · The copied plan remains a local draft. A recorded assignment and a file checksum are useful records but are not a public preregistration, a power calculation or an endorsement. · Confidence: High |
| Google describes query fan-out, differences between AI search surfaces, and variation in returned responses and links. | Google: AI features and websites | Checked 2026-09-13 · Primary source checked September 13, 2026 · This makes source overlap and named-surface collection relevant concerns. The local exercise does not call Google or measure a provider's actual retrieval graph. · Confidence: High |
| SearchPilot's GEO testing guide describes control and variant page groups and a traffic-oriented evaluation task. | SearchPilot: GEO testing | Checked 2026-09-13 · Primary source checked September 13, 2026 · First-party methodology and product positioning, not independent proof of this protocol. RankEcho's guide covers a narrower prospective prompt/source assignment task and does not claim native GEO A/B testing. · Confidence: High |
