Skip to content
Why GEO Why us Pricing
Get started free
HomeProductExperiments & proof

Did any of this actually work?

An experiment is a measured before and after: mark a change, and we report whether it actually worked — the difference between baseline and treatment, with a confidence interval.

How an experiment is built

Experiments and proof is the half of MentionBeat that answers "did any of this work?" with a number instead of a story. You mark a change, it measures a baseline before and a treatment run after, and reports the difference with a confidence interval — including the times the honest answer is that nothing moved. It is the difference between a team that learns which GEO tactics pay and one that repeats last quarter's guesses with more conviction. Four steps, and the discipline is entirely in the first one: declaring what you are about to change before you change it.

Declare

Name the change and the prompts it should affect. Declaring the target first is what stops a result being rationalised afterwards.

Baseline

A run against the current state, sampled repeatedly, so the starting point has an interval rather than a single value.

Ship

Publish the page, send the correction, fix the crawler block. The intervention date is recorded automatically.

Treatment

The same suite, the same engines, the same sampling — after the change has had time to be crawled and retrieved.

Same prompts, same engines, same repeat count. Changing the suite mid-experiment invalidates the comparison, so the tool holds it fixed and tells you if a prompt was retired underneath you.

What a result looks like

A single line you can put in a board deck, with everything underneath it available when someone pushes back.

  • The delta — mention rate +9 points, runs #3 → #5
  • The interval — and whether the move clears it
  • The verdict — moved beyond noise, or no measurable change
  • The scope — which prompts and engines carried the change
  • The evidence — the actual answers, before and after, side by side
  • The caveats — what else changed in the window that could explain it

The last line is deliberate. An experiment that hides its confounders is a marketing asset, not a measurement.

Honest about the design

This is a before-and-after with repeated sampling, not a randomised controlled trial. You cannot show half of ChatGPT a different version of your website, and any tool claiming a true control group is describing something it has not built.

What the design does give you: a measured baseline instead of a remembered one, a known intervention date, an interval wide enough to absorb ordinary drift, and a null result reported as a null result.

The statistics, in detail →

Which changes are worth measuring

Not every fix deserves an experiment. These are the ones where teams most often assume an effect that is not there.

ChangeTypical time to showWhy it is worth isolating
Shipping a comparison page family1–3 weekly cyclesThe most common GEO investment, and the effect varies hugely by category
Unblocking an AI crawler1–2 cyclesShould be dramatic on grounded engines and irrelevant on parametric ones — a clean test of the diagnosis
Adding structured data sitewide2–4 cyclesWidely assumed to work; worth knowing whether it does for you
Correcting a third-party source2–6 cyclesEffort is high and the effect is invisible without measurement
Rewriting an opening to be answer-first1–2 cyclesCheap to do, easy to over-attribute across a whole site

Two clocks, again: page and retrieval fixes can land inside a cycle or two, while entity and reputation work compounds over quarters. Every experiment states which one it is on before you start waiting.

Questions about proof

Is this a real A/B test?
It is a before-and-after with repeated sampling, not a randomised split — you cannot show half of ChatGPT one version of your site. That limit is stated on every result. What the design does buy you is a measured baseline, a known intervention date and an interval wide enough to be honest about ambient drift.
How long should an experiment run?
Long enough for the change to be crawled and retrieved, which is typically one to three weekly cycles for page-level work. Entity and reputation changes take longer and the tool says so rather than declaring a null result early.
What if the result is nothing?
Then it says so: no measurable change beyond the interval. That is a real finding and often a cheap one — it stops a team repeating an expensive tactic for another quarter on the strength of a coincidence.
Can I attribute revenue to this?
Not directly, and we will not pretend otherwise. The buying decision happens inside an answer, before any click your analytics could count. What you can attribute is the visibility change itself, and connect it to pipeline the way you already connect other upper-funnel work.

Keep going

Stop guessing whether it worked.

Get a baseline this week so the next thing you ship has something to be measured against.