Skip to content
Why GEO Why us Pricing
Get started free
HomeEngine Weather Report

The Engine Weather Report

The Engine Weather Report is a recurring measurement of how consistently each AI answer engine responds: how often it contradicts itself between repeats of one question, how far its answers swing from run to run, and how much the sources it cites turn over. It runs on MentionBeat's own public prompt suite, so no customer's data is in it, and every figure is published with the number of runs behind it.

Measured

Three things, per engine, every run

  • Answer flip — disagreement between repeats of one question
  • Rate swing — movement between consecutive runs
  • Source churn — turnover in the domains it cites

A component an engine cannot produce is left out, never counted as zero.

The current reading

Lower is steadier. An engine listed without a score has not been measured over enough runs yet — the count is shown so you can see exactly how thin the evidence is.

The first edition has not been published yet. This page only ever shows measured runs. Two of the three components compare one run to the next, so the suite has to build a few weeks of history before a reading means anything — and we would rather show you nothing than fill the gap with a plausible number.

Once published, each edition carries its run count, the suite it was measured on and the date, so any figure here can be challenged.

What the three components measure

The composite is the mean of whichever of these can be computed for that engine. We publish the parts as well as the score, because a reader who disagrees with how we combine them should be able to ignore our index and use the columns.

ComponentWhat it capturesWhy it matters to you
Answer flipHow often repeats of the same question, inside one run, disagreed about whether a brand was named.High flip means one query tells you almost nothing — you need repeats before the number is real.
Rate swingHow far the engine's mention rate moved between consecutive runs.An engine that swings a lot will manufacture "wins" and "drops" that are nothing but noise.
Cited-source churnHow much the set of domains the engine cited turned over between runs.High churn means the sources feeding your answers are being reshuffled, so a citation win may not hold.

How to use this

This is not a league table of which engine is "best". A volatile engine can be the one your buyers use most — the reading tells you how much sampling it takes before you can trust what you measured there.

  • Calm — a single run is a reasonable read
  • Variable — trust the interval, not the point estimate
  • Stormy — add repeats before drawing a conclusion

Want this for your own prompts rather than ours? Tracking reports the same three components per engine on your own suite, and the free page checker is the fastest way to see where you stand first.

Why we publish it

Because it is the clearest evidence for the thing we argue about most: a single screenshot of an AI answer is not a measurement, and a tracker reporting one unqualified number is reporting noise with confidence.

This report only exists because we sample repeatedly and compute uncertainty properly. If that were decoration, there would be nothing to publish.

Why we build it this way →

Fair questions

What does a high volatility score mean?
That the engine gave materially different answers to the same questions across repeats and across runs, and changed the sources it cited. It does not mean the engine is bad — it means a single measurement of it is a weak reading, so you should sample it more and trust the interval rather than the point estimate.
Whose data is this measured on?
Our own. The report runs on a fixed public prompt suite that MentionBeat operates, not on customer measurements. We do not publish customer data, even anonymised.
Why is an engine sometimes listed with no score?
Because two of the three components compare one run to the next, so an engine needs several runs of history before a volatility figure means anything. Until it has them we publish the run count and say the score is withheld, rather than scoring it on thin evidence.

Measure your own weather.

The same three components, computed on your prompts and your engines, with a confidence interval on every rate.