MentionBeat is a measurement platform for AI answers: it asks real engines what your buyers ask, repeatedly, and reports every result with a confidence interval. This page is the whole method, including the limitations most vendors leave out.
Last updated 6 August 2026 · No account required · Every claim here corresponds to shipped, tested code.
The Visibility Index, in full
0.4 × share of voice
+ 0.4 × mention rate
+ 0.2 × recommendation rate
Unbranded prompts only, with a confidence interval from the same bootstrap as its parts. Emitted only when every component exists — never quietly reweighted around a missing one.
Every response comes from the real engine's own API, called with your tracked prompt at a realistic sampling temperature. We do not infer AI answers from Google rankings, and we never silently substitute one model for another.
| Engine | How we query it | Web-grounded? |
|---|---|---|
| ChatGPT | OpenAI API | Optional (OpenAI web search) |
| Claude | Anthropic API | Optional (Anthropic web-search tool) |
| Gemini | Google Gemini API | Optional (Google Search grounding) |
| Perplexity | Perplexity Sonar API | Always (natively grounded) |
| Grok | xAI API | Optional (Live Search over web + X) |
| DeepSeek | DeepSeek API | No — parametric only, and we refuse to label it otherwise |
| Google AI Overviews | SERP capture (SerpApi) | Always |
| Google AI Mode | SERP capture (SerpApi) | Always |
| Microsoft Copilot · Meta AI | Not yet measured. No API or SERP surface exists; a browser-capture adapter is on the roadmap. We would rather show a gap than simulated data. | |
LLMs are non-deterministic. A single query is an anecdote, not a measurement.
Headline visibility uses only unbranded prompts — the questions a real buyer asks before they know your name.
Brand-bearing prompts are still tracked, but reported separately and excluded from headline share of voice. Asking an engine “tell me about Acme” and counting the reply as visibility is grading your own exam.
Responses are scored by an LLM judge that extracts brand mentions, position, sentiment, recommendation strength, cited domains and factual claims. Because judges are also LLMs, we treat the judge itself as an instrument to calibrate.
Our composite score is not a proprietary mystery. It is 0.4 × share of voice + 0.4 × mention rate + 0.2 × recommendation rate, computed on unbranded prompts only, carrying a confidence interval from the same cluster bootstrap as its parts. It is emitted only when every component exists — never silently reweighted around a missing one.
Those weights are chosen for explainability, and they carry that provenance in the product. As measured before/after experiments accumulate alongside AI-referred traffic, we fit outcome-calibrated weights against them — and any weight change would be announced and versioned, never slipped into your trend.
Not every question is asked equally often, so prompts can carry a demand weight and the headline metrics can be reported weighted as well as flat. Weighted and unweighted figures come from the same bootstrap, so the two are directly comparable.
Where a weight is a modelled estimate it is labelled as one, in the product and here. An estimated weight is relative, on a 0–100 scale within your own prompt set — never a claim about absolute search volume. It is built from a log-scaled demand proxy (keyword and Trends volume, optionally calibrated against your own Search Console impressions), multiplied by mild intent-class priors: chat usage skews toward informational and comparison questions and away from purely transactional ones, so those classes are nudged up and down by no more than 25%. The priors are priors, not measurements, which is exactly why they are kept small.
A prompt we have no demand signal for keeps a uniform weight and is flagged as unweighted, rather than being quietly assigned a number.
Some questions only the whole measured field can answer, so these lenses aggregate across projects — under hard floors, stated with every result.
Per-engine volatility: within-run answer flips, run-to-run rate swings, and cited-source churn. Each score names exactly which components it stands on, and a component we cannot compute stays absent rather than counting as zero.
Where your number sits among projects measured the same way. Withheld, with the reason shown, until a cohort holds at least five peer projects — a percentile against two peers is both meaningless and a privacy leak.
Which third-party domains shape AI answers across the field. A domain appears only when at least three unrelated projects' answers cite it, and only ever as shares, so no row is traceable to any one customer's data.
What a kind of content change — an FAQ block added, schema markup fixed, an entity listing claimed — has measurably moved, pooled across comparable experiments. Confounded experiments are excluded, and no number is shown below three measured experiments of that kind.
We don't only count mentions — we fact-check what engines say about you against your product brief: a provenance-tracked store of approved facts, each carrying its source and human-review status.
Content changes are evaluated as baseline → treatment experiments: per-metric lift with confidence intervals, significant only when those intervals separate.
A delta is only as trustworthy as the comparison behind it. If the questions changed, the answer changed for a reason that has nothing to do with your content — so we check that first and say so, rather than letting a setup change take the credit.
Where a prompt suite did change between two runs, we also recompute both numbers over only the prompts they share, giving you a like-for-like delta beside the raw one.
Any of these between baseline and treatment marks the result confounded instead of reporting it as a win:
The authoring studio can preview engine behaviour through a persona simulation when no API key is configured. Every such run carries its label from the moment it is created to the moment it is read.
Modelled data never appears in measurement reporting as though it were live. If a vendor key is missing you see a gap or a label, never an imitation of the real thing.
The same rule governs this website. A number we publish either came from a real run and carries its interval, or it is cited and dated to someone else.
Answers from a real engine API call. The only kind that reaches measurement reporting.
A persona simulation used for authoring preview. Never counted, never charted as visibility.
A view combining both. Labelled as such, so the blend is never mistaken for a measurement.
LLM output is non-deterministic, so a tracker reporting a single unqualified number is reporting noise with confidence. These six requirements follow from that one fact. They are not our preferences — they are what honest measurement of a non-deterministic system requires, and each takes minutes to verify in a demo.
| Requirement | Why it matters | MentionBeat |
|---|---|---|
| Repetition and confidence intervals | One sample per prompt makes every week-over-week “change” indistinguishable from noise. | Multiple repetitions per (prompt × engine); 95% CI on every headline metric; deltas flagged only when intervals separate. |
| Uncertainty computed at the prompt level | Repetitions of one prompt are correlated; treating them as independent samples understates uncertainty — the standard mistake. | Cluster bootstrap of 2,000 resamples over prompts, not responses; Wilson intervals on small slices. |
| A calibrated judge | If an LLM scores the answers, that judge has its own error rate. Uncalibrated, it can shift a headline number by more than the change you are trying to detect. | Consensus voting, validation against human gold labels (Cohen's κ, sensitivity and specificity), Rogan–Gladen bias correction on mention rates. |
| Grounded and parametric reported separately | A web-grounded answer and a from-memory answer measure different channels; averaging them hides the difference that matters. | Every response is classified, and the two are never averaged together. |
| No brand-name leakage in headline metrics | Asking “tell me about [Brand]” and counting the reply measures nothing. | Headline metrics come only from unbranded category prompts; brand-bearing prompts are a separately reported diagnostic slice. |
| Immutable raw answers | If only scores are stored, the methodology can never be audited or improved retroactively — the dashboard becomes the only evidence of itself. | Every raw response is stored, and every metric is recomputable from the raw log without re-spending a single API call. |
Evaluating any AI-visibility tracker, this one included? Ask how it handles these six. The answers separate measurement from screenshots faster than any feature list.
A methodology that lists no limitations is marketing. These are ours, stated plainly.
We query engine APIs; consumer web interfaces can route to different configurations. A browser-capture adapter with a published API-versus-UI reconciliation report is planned. Until then, our results measure the engines' API-served answers.
We have no clickstream data of real user prompts. Prompt sets are generated and human-curated, and demand weights are estimates unless you supply real volumes.
Not IP-of-origin. Region-parameterised querying is in development; SERP engines already accept country and language parameters, chat APIs do not.
AI Overviews and AI Mode are read through a SERP provider's rendering. The absence of an AI answer for a query is itself recorded as a data point.
Repetition and intervals bound it; they do not make an LLM deterministic. Treat every point estimate as the centre of an interval.
The free tier runs a real suite across every engine, and you have already read the method. Want the statistical detail? The metrics module implements everything above — ask and we will show you the code.