Skip to content
Why GEO Why us Pricing
Get started free
HomeProductVisibility tracking

How often does an assistant name you?

Visibility tracking is the measurement layer that answers how often an assistant names you: your buyers’ real questions, run across every engine on a schedule, sampled enough to give a rate with an interval.

What a run actually does

Visibility tracking is the measurement half of MentionBeat: a suite of the questions your buyers actually ask, run across every answer engine on a schedule, each prompt sampled several times so the result is a rate with a confidence interval rather than a screenshot. It answers the only question that matters at the start — when someone asks an assistant about your category, how often does your name come out, and is this week's change real? A run is a defined experiment, not a crawl. You control every part of it, and the console prices it before it starts.

The prompt suite

MentionBeat reads your site, derives a product brief, and proposes prompts from it — grouped by buyer intent and journey stage, and tagged to an ICP. You approve, reword or retire each one. Reworded prompts are retired and re-added rather than edited in place, so a trendline never silently changes what it was measuring.

The engine mix

Each prompt is put to the engines you enable. Web-grounded engines search live and cost more; parametric ones answer from training alone. Both are worth tracking and they routinely disagree — the split is reported, never averaged into one misleading figure.

Repeat sampling

Every prompt is asked several times per engine. One answer is an anecdote; a sample is a measurement. The repeats are what make a week-on-week move interpretable instead of a coin flip.

Parsing and scoring

Each answer is parsed for whether you were named, where you appeared, whether the mention was a recommendation, and how you were characterised. Rivals in the same answer are scored the same way, which is what produces share of voice.

The metrics, and what each is for

Four rates and one roll-up. They move independently, and the differences between them are usually the interesting part.

MetricWhat it countsWhat a drop usually means
Mention rateShare of sampled answers that name you at allYou are missing from the retrieval set, or the category framing moved
Recommendation rateShare of answers that name you as a suggested option, not just in passingYou are known but not preferred — usually a comparison or proof gap
Share of voiceYour mentions as a proportion of all brands mentionedA rival got louder, even if your own rate held
Sentiment when mentionedHow you are characterised in the answers that do name youA negative source is being retrieved and repeated
Visibility IndexA 0–100 weighted roll-up of the above — published arithmetic, not a black boxThe one-line trend for a status update — always decomposable

Every rate is reported with a 95% confidence interval built from bootstrap resampling, clustered by prompt. A change is flagged as movement only when it clears that interval.

Engine weather: which engines to trust this week

Not every engine is equally steady, and the unsteady ones make a single run misleading. Each engine gets a volatility read from three things we already measure: how often repeats of the same question disagree, how far its mention rate swings between runs, and how much its cited sources churn.

Use it to decide where extra repeats buy you precision, and which engine's one-off number to treat with caution before you quote it in a meeting.

  • Calm — a single run is a reasonable read
  • Variable — trust the interval, not the point estimate
  • Stormy — add repeats before drawing a conclusion

Each score names the components it stands on. Where a component cannot be computed — a memory-only engine cites nothing, so it has no source churn — it is left out rather than counted as zero.

Slice it the way your team argues about it

A single headline rate hides the disagreement that makes the number actionable. Every metric can be broken down by:

  • Engine — ChatGPT may recommend you where Gemini never names you
  • ICP and persona — strong with enterprise buyers, absent for self-serve
  • Journey stage — visible at comparison, invisible at discovery
  • Market and language — a suite scoped per country, reported per country
  • Grounded vs parametric — whether the model knows you or has to look you up
  • Run history — every past run kept, so a trendline is real history, not a redraw

Drift alerts

Runs are scheduled on an interval you set — fortnightly by default, because page and retrieval fixes take days to land and that is the shortest gap where a change is usually real rather than noise. When a tracked rate moves beyond its interval, or a rival overtakes you on a prompt you used to win, you get told rather than discovering it a month later in a dashboard.

Alerts name the prompt, the engine and the size of the move, so the notification is the start of the investigation instead of a nudge to go and look.

See the answers behind a move →

Is 31% good?

On its own, no share-of-voice number answers that. So every measured project is placed against the others we measure — same suite construction, same judging, same repeat sampling — and you are told your percentile rather than left to guess whether your rate is strong or embarrassing.

  • Like for like — the cohort is projects measured the same way, not a survey of self-reported numbers
  • Aggregates only — quartiles and medians cross the boundary; no other customer's numbers, brands or prompts ever do
  • A floor on the cohort — below a minimum number of peers there is no comparison, because a “percentile” against three projects is noise wearing a statistic's clothes
  • Per metric — mention rate, share of voice and the Visibility Index are each ranked separately

Goals that read from the measurement

Set a target on a metric — mention rate on a topic, share of voice, the index — with a date. Progress is not something you update in a spreadsheet: it reads from your runs, and a goal is only marked achieved by a measured run that clears it.

Because the benchmark knows what the cohort looks like, a goal can also be proposed rather than guessed: a target that would move you from the middle of the field to the top quartile is a more useful number than a round one someone picked in a planning meeting.

Proving a change caused the move →

Questions about measurement

How many prompts do I need?
Enough to cover the decisions you care about, which is usually 20 to 60 for a single product. A suite that small is fine because the statistical power comes from repeat sampling, not from prompt count — asking 500 prompts once tells you less than asking 30 prompts five times each.
Why does the same prompt give different answers?
Because generative engines are stochastic and, when grounded, depend on what the live web returned that second. That variance is real and it is the reason a single screenshot is not measurement. Sampling each prompt repeatedly turns the variance into an interval you can reason about.
Can I track more than one market or language?
Yes. A suite can be scoped by market and language, and results are reported per market so a strong position in one country does not hide a weak one in another.
What is the Visibility Index?
A single 0–100 roll-up of the underlying rates — mention, recommendation and share of voice, weighted by prompt importance. It exists to give one trendline for a status update; every component stays visible underneath so the number can always be taken apart.

Keep going

Get your baseline this week.

1,500 free credits is enough for a real first suite across every engine. No card, no site changes.