An experiment is a measured before and after: mark a change, and we report whether it actually worked — the difference between baseline and treatment, with a confidence interval.
Experiments and proof is the half of MentionBeat that answers "did any of this work?" with a number instead of a story. You mark a change, it measures a baseline before and a treatment run after, and reports the difference with a confidence interval — including the times the honest answer is that nothing moved. It is the difference between a team that learns which GEO tactics pay and one that repeats last quarter's guesses with more conviction. Four steps, and the discipline is entirely in the first one: declaring what you are about to change before you change it.
Name the change and the prompts it should affect. Declaring the target first is what stops a result being rationalised afterwards.
A run against the current state, sampled repeatedly, so the starting point has an interval rather than a single value.
Publish the page, send the correction, fix the crawler block. The intervention date is recorded automatically.
The same suite, the same engines, the same sampling — after the change has had time to be crawled and retrieved.
Same prompts, same engines, same repeat count. Changing the suite mid-experiment invalidates the comparison, so the tool holds it fixed and tells you if a prompt was retired underneath you.
A single line you can put in a board deck, with everything underneath it available when someone pushes back.
The last line is deliberate. An experiment that hides its confounders is a marketing asset, not a measurement.
This is a before-and-after with repeated sampling, not a randomised controlled trial. You cannot show half of ChatGPT a different version of your website, and any tool claiming a true control group is describing something it has not built.
What the design does give you: a measured baseline instead of a remembered one, a known intervention date, an interval wide enough to absorb ordinary drift, and a null result reported as a null result.
Not every fix deserves an experiment. These are the ones where teams most often assume an effect that is not there.
| Change | Typical time to show | Why it is worth isolating |
|---|---|---|
| Shipping a comparison page family | 1–3 weekly cycles | The most common GEO investment, and the effect varies hugely by category |
| Unblocking an AI crawler | 1–2 cycles | Should be dramatic on grounded engines and irrelevant on parametric ones — a clean test of the diagnosis |
| Adding structured data sitewide | 2–4 cycles | Widely assumed to work; worth knowing whether it does for you |
| Correcting a third-party source | 2–6 cycles | Effort is high and the effect is invisible without measurement |
| Rewriting an opening to be answer-first | 1–2 cycles | Cheap to do, easy to over-attribute across a whole site |
Two clocks, again: page and retrieval fixes can land inside a cycle or two, while entity and reputation work compounds over quarters. Every experiment states which one it is on before you start waiting.
Get a baseline this week so the next thing you ship has something to be measured against.