- Measure before you change anything. Without a baseline you cannot tell an improvement from a bad week, and these engines vary run to run.
- Content shape has the strongest evidence. The original GEO research measured visibility gains of up to ~40% from adding citations, quotations and statistics to source pages.1
- Corroboration is slower but compounds. What independent sources say about you feeds the answers more than your own marketing copy does.
- Retrievability is a gate, not a lever. It is worth nothing until it is broken — and then it caps everything else at zero.
- Judge results by intervals, not point estimates. 24% to 29% is very often noise wearing a suit.
Most advice on this question skips the hard part. It lists tactics without saying which ones the evidence supports, how long each takes to show up, or how you would ever know it worked. That is comfortable to write and useless to act on, because these four levers differ by roughly an order of magnitude in both effort and effect.
This page is the lever catalogue: what each one is, what it is worth, what it costs, and the test that separates a real improvement from a good week. If you want the argument for why the goal is being named rather than ranked, that is a different piece — how to rank in AI search covers the mental model and the loop. This page assumes you have accepted it and want to know what to actually change.
First, be precise about which number you're improving
Brand visibility in AI search is also known as AI visibility, LLM visibility or AI share of voice, depending on who is selling you something. Whatever it is called, it is three different measurements that people use interchangeably, and they move independently:
- Mention rate — of the answers to your category's buyer questions, how many name you at all.
- Share of voice — of every brand named across those answers, what fraction is you.
- Recommendation rate — how often you are actively endorsed rather than listed in passing.
These come apart in practice. A brand can raise its mention rate while its share of voice falls, because the whole category got more crowded. Pick the one your team will be judged on before you start, or you will end up arguing about which number to report after the fact. The metric definitions go through each with a worked example.
One distinction decides how fast anything works. When an engine answers by retrieving live pages, changing those pages can show up within days to weeks. When it answers from model memory, nothing you publish today changes it until the model is retrained — months, and not on your schedule. Measure the two separately. Averaging them hides the half you can actually move.
Get a baseline before you touch anything
This is the step teams skip and then regret, because without it every later result is unfalsifiable. These engines are non-deterministic: ask the same question five times and you can get five different brand line-ups. So a single “before” reading is not a baseline — it is one draw from a distribution.
A usable baseline is a fixed prompt suite of buyer questions that never name your brand, run repeatedly across the engines you care about, with the rates written down alongside their confidence intervals. The sampling method has the full recipe, including a spreadsheet version that costs nothing but an afternoon. Freeze that suite. Changing the questions between waves means you are measuring your question set, not your visibility.
The four levers, ranked
Ranked by strength of published evidence, not by how often you will see them recommended:
| Lever | What the evidence says | Effort | Time to show up |
|---|---|---|---|
| 1. Source content shape | Strongest. Adding citations, quotations and statistics to source pages measured up to ~40% visibility gain in the original GEO study1 | Low–medium — editing pages you already own | Days to weeks, once re-crawled |
| 2. Third-party corroboration | Strong but indirect. Engines lean on independent sources; you influence rather than control them | High — outreach, relationships, time | Weeks to months |
| 3. Technical retrievability | Binary. No measured “gain” when it works; total loss when it does not | Low, one-off — then monitor | Immediate once fixed |
| 4. Correcting what's wrong | Protective. Stops a wrong answer costing you the deal; rarely raises the rate by itself | Medium — depends on other publishers | Weeks to a training cycle |
Lever 1 — make your pages worth quoting
This is the highest-yield change available to most teams, and the reason is mechanical: a generative engine composing an answer needs material it can lift and attribute. A page of confident adjectives gives it nothing to work with. A page with a specific number, a dated source and a quotable sentence gives it three things.
The original GEO research tested exactly this — rewriting source content to add citations, quotations and statistics — and measured visibility improvements of up to roughly 40% in generative engine responses.1 That is the single strongest published finding in this field, and it is about the pages you already control.
In practice: put the direct answer in the first paragraph under each heading rather than building to it; replace vague claims with figures you can source; add a dated, attributed quote where you have one; and make the comparison table you have been avoiding, because comparison tables are unusually liftable. The formats that win citations goes deeper, and anatomy of a quotable product page works through a single page end to end.
Lever 2 — earn corroboration you don't own
Engines weigh what independent sources say about you more heavily than what you say about yourself, which is uncomfortable but not surprising — it is the same instinct a careful buyer has. This makes third-party presence a genuine lever, and the slowest one on the list.
The practical version is unglamorous: be in the roundups and comparison articles your category's buyers read; be accurately represented on the review platforms they check; be present in the communities where the question gets asked, without astroturfing them. Why third-party mentions drive recommendations makes the full case, and where LLMs actually learn about your brand covers which sources carry disproportionate weight.
Two honest caveats. You are influencing other people's publishing decisions, so timelines are not yours to set. And this lever is the easiest to do badly — manufactured mentions are detectable, and the downside is worse than the upside.
Lever 3 — check you are retrievable at all
This lever behaves differently from the others: there is no gain from doing it well, only a catastrophe from doing it badly. If an AI crawler cannot fetch your page, or fetches it and finds an empty shell because the content renders client-side, then every hour spent on levers 1 and 2 is capped at zero for the grounded channel.
It takes an afternoon to rule out. Confirm the AI crawlers you want are not blocked in robots.txt or by your CDN or WAF — managing AI crawlers covers which agents matter and the traps, and the free access checker runs live probes that catch the CDN-level blocks a robots.txt read alone will miss. Then confirm your content survives without JavaScript; rendering and why AI crawlers miss content explains the failure mode.
Lever 4 — fix what they get wrong about you
Being visible for the wrong facts is its own problem: a retired plan quoted as current, a price that changed last year, a capability you never had. This lever rarely raises your mention rate — it protects the value of the mentions you already get.
The honest boundary here matters. Nobody can edit a model's memory. What you can do is change what it retrieves, and correct the third-party sources the wrong claim came from. Detecting and fixing brand hallucinations covers tracing an error back to its source, which is the part most teams miss.
What doesn't work, or isn't proven
- Keyword stuffing for engines. There is no keyword to rank for; the engine composes prose. Density does nothing.
- Publishing an
llms.txtand stopping. Adoption is thin and the evidence is weak — what the evidence actually says. It costs an hour, so ship one if you like, but do not count it as a lever. - Paying for a mention. There is no submission portal and no ad slot inside the answer for brand facts. Anyone selling one is selling something else.
- Volume for its own sake. Forty thin pages perform worse than four quotable ones, because the unit that gets lifted is the passage, not the page count.
MentionBeat runs your category's buyer questions across 6 AI surfaces, repeatedly, and reports mention rate, share of voice and accuracy with 95% confidence intervals — so the next change you make has something to be measured against.
Get a free visibility reportHow to tell whether it worked
Re-run the frozen suite with the same number of repeats, and compare intervals rather than point estimates. If your mention rate went from 24% to 29% but the two intervals still overlap heavily, you have not shown an improvement — you have shown a number that moved, which is not the same claim.
Three rules keep this honest. Change one lever at a time where you can, or accept that you will not know which one paid. Give a content change at least one re-crawl cycle before judging it. And keep the raw answers, not just the scores, so a result can be re-examined later without re-running everything. Why one ChatGPT query is not a measurement covers the sample-size maths.
If it did not move, that is information too. The most common reasons, in order: the change was cosmetic rather than substantive; the pages have not been re-crawled; the answers in question are parametric rather than grounded, so nothing published this quarter could have affected them; or the sample is too small to detect a change of the size you achieved.
Frequently asked questions
Four, in descending order of evidence: reshape your source content so it is quotable (citations, quotations, statistics — up to ~40% measured gain1), earn third-party corroboration, make sure you are technically retrievable, and correct what the engines get factually wrong about you. Everything else is either unproven or a rebrand of one of these four.
Days to weeks for web-grounded answers, once the changed page is re-crawled. Potentially never, on your timeline, for answers drawn from model memory — those only change when the model is retrained. Measure the two channels separately so you can see the half that is actually responding.
No — and neither can anyone else, including us. Nobody controls these engines and there is no submission portal for brand facts. What is achievable is raising the rate at which you are named across a set of buyer questions, and being able to prove the rate moved rather than asserting it.
Compare confidence intervals, not point estimates. Take a baseline, re-measure with the same frozen prompt suite and the same repeat count, and treat the change as real only when the intervals stop overlapping. Anything less is a number that moved, which is a weaker claim than an improvement.
Sources & further reading
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., Deshpande, A. — "GEO: Generative Engine Optimization", KDD 2024 / arXiv:2311.09735 — the source of the ~40% figure, measured on the GEO-Bench benchmark.
- Google Search Central — "AI features and your website" — how Google frames answer surfaces, crawling and site-owner controls.


