The Engine Weather Report is a recurring measurement of how consistently each AI answer engine responds: how often it contradicts itself between repeats of one question, how far its answers swing from run to run, and how much the sources it cites turn over. It runs on MentionBeat's own public prompt suite, so no customer's data is in it, and every figure is published with the number of runs behind it.
Three things, per engine, every run
A component an engine cannot produce is left out, never counted as zero.
Lower is steadier. An engine listed without a score has not been measured over enough runs yet — the count is shown so you can see exactly how thin the evidence is.
The first edition has not been published yet. This page only ever shows measured runs. Two of the three components compare one run to the next, so the suite has to build a few weeks of history before a reading means anything — and we would rather show you nothing than fill the gap with a plausible number.
Once published, each edition carries its run count, the suite it was measured on and the date, so any figure here can be challenged.
The composite is the mean of whichever of these can be computed for that engine. We publish the parts as well as the score, because a reader who disagrees with how we combine them should be able to ignore our index and use the columns.
| Component | What it captures | Why it matters to you |
|---|---|---|
| Answer flip | How often repeats of the same question, inside one run, disagreed about whether a brand was named. | High flip means one query tells you almost nothing — you need repeats before the number is real. |
| Rate swing | How far the engine's mention rate moved between consecutive runs. | An engine that swings a lot will manufacture "wins" and "drops" that are nothing but noise. |
| Cited-source churn | How much the set of domains the engine cited turned over between runs. | High churn means the sources feeding your answers are being reshuffled, so a citation win may not hold. |
This is not a league table of which engine is "best". A volatile engine can be the one your buyers use most — the reading tells you how much sampling it takes before you can trust what you measured there.
Want this for your own prompts rather than ours? Tracking reports the same three components per engine on your own suite, and the free page checker is the fastest way to see where you stand first.
Because it is the clearest evidence for the thing we argue about most: a single screenshot of an AI answer is not a measurement, and a tracker reporting one unqualified number is reporting noise with confidence.
This report only exists because we sample repeatedly and compute uncertainty properly. If that were decoration, there would be nothing to publish.
The same three components, computed on your prompts and your engines, with a confidence interval on every rate.