The verdict column does not rank the metrics by precision or by cost, and several of the four survivors are expensive and noisy. It ranks them by whether a movement implies an action: a fall in mention rate on one engine sends you to that engine’s citation list, while a fall in a blended score sends you nowhere in particular. The six underlying quantities are defined in the sibling article on what visibility actually means.
Which GEO metrics hold up, and which ones fail?
Every metric in this category has a failure mode, and knowing it is what separates a number you can act on from one you have to defend. The blended cross-engine score is the default dashboard output and it destroys the only decision it supports.
On this page
Four GEO metrics hold up: mention rate and citation rate reported per engine with an interval, share of voice against a competitor set you name and freeze, and framing accuracy read by a person on a sample. Four fail: a single blended cross-engine score, a per-prompt week-over-week arrow, any traffic or revenue figure derived from citations, and sentiment scoring. The blended score fails hardest, because a SIGIR 2026 study of 11,500 queries measured Jaccard similarities of 0.11 to 0.18 between the sources Google organic results, AI Overviews and Gemini retrieve.1
Key takeaways- A metric holds up only if a movement in it changes what somebody does next.
- Report per engine and print the interval: a cross-engine average hides the two engines that actually moved.
- Never convert citations into traffic or revenue. Pew Research Center found users clicked a source inside an AI summary in about 1% of visits.8
Which metrics survive scrutiny, and which fail?
A metric survives scrutiny when a movement in it changes what somebody does next; four of the eight below clear that bar and four do not.
The last column is the useful one: it names the observation that would prove a given reading wrong.
| Metric | Verdict | What would make a reading of it wrong |
|---|---|---|
| Mention rate, per engine, with an interval | Holds up | The interval overlaps last period’s, or the panel version changed between them |
| Citation rate, per engine, with an interval | Holds up | It moved opposite to mention rate and only one of the two was reported |
| Share of voice, fixed competitor set | Holds up | A brand entered or left the counted set inside the comparison window |
| Framing accuracy, human-read sample | Holds up | Two readers applying the rubric disagree on more than a fifth of answers |
| Position in answer | Use with care | Read on a single answer, or compared across engines that cite different source counts |
| One blended cross-engine visibility score | Fails | Always: the components move independently, so the average has no referent |
| Per-prompt score with a week-over-week arrow | Fails | The movement is smaller than the interval on a handful of runs of one prompt |
| Traffic or revenue derived from citations | Fails | Always: the coefficient linking the two has never been measured |
Why does a blended cross-engine score break its own decision?
A blended score breaks its own decision because the only action a visibility metric supports is “work on this engine next”, and averaging across engines deletes exactly that information. It is the default output of almost every dashboard in this category.
Start with how little the engines share. A SIGIR 2026 study of 11,500 queries measured URL-level Jaccard similarity of 0.11–0.18 between Google organic results, AI Overviews and Gemini.1 Read that concretely: if two surfaces each cite ten URLs for one question, a Jaccard of 0.15 means about three appear on both lists. Those three surfaces are built by one company, on one index, inside one product family; two unrelated assistants overlap less. The 2026 critical survey draws the conclusion in a sentence worth quoting exactly: these results “refute the notion of a global GEO ranking. Visibility is indexed by engine and surface.”2
Now watch what averaging does to a real week. Suppose your mention rate rises 8 points on one engine, falls 6 on another and is flat on three. The blended score moves about half a point and the report says “stable”, while two engines have just done something worth investigating. The failure is not imprecision; it is that the aggregate is a well-behaved average of quantities that are not measuring the same thing. The mechanism behind the disagreement is worked through in the article on whether AI engines agree.
There is one legitimate use for a cross-engine roll-up, and it is not a score. Count the engines on which you cleared a threshold: named in at least a quarter of answers on three engines out of five is a defensible summary sentence, because it is a count of independently defensible facts rather than an average over incommensurable ones.
Why is framing accuracy the metric nobody buys?
Framing accuracy is the metric nobody buys because it cannot be automated and it makes the report look worse, and the metric everybody needs because every other number treats a confidently wrong description of your product as a win. It is also the only one of the eight with an independently established base rate, and that base rate is bad.
Three independent measurements, three different methods, the same direction. A Tow Center for Digital Journalism audit published 6 March 2025 ran 1,600 queries (twenty publishers, ten articles each, eight engines) and found more than 60% returned incorrect answers, with the best-performing engine still wrong 37% of the time.3 The same centre’s November 2024 study of 200 quotes found 153 responses partially or entirely incorrect, with uncertainty signalled seven times.4 Independent work found 51.5% of answer sentences fully supported by their attached citations,5 and a 2026 preprint classified roughly 11% of 98,020 atomic claims as insufficiently supported.6 That last figure is lower because it counts claims rather than responses, and a response is wrong if any claim in it is wrong. Quote whichever you like, with its unit attached.
Measuring it on your own product is cheaper than those studies make it sound. Sample twenty answers a month, weighted towards the classes where an error costs a deal: comparison and objection prompts. Score each against a fixed rubric with a few factual axes: pricing model, what the product does, what it does not do, integrations, and any safety or compliance claim. Record a binary per axis rather than a sentiment, and keep the wrong sentence verbatim so it can be shown to whoever owns the page that produced it. Two readers should agree on four answers in five; if they do not, the rubric is too vague rather than the engines too erratic.
The output is not a percentage for a slide but a list of specific false sentences and the pages that most plausibly caused them. The article on citation accuracy covers what to do with that list.
What makes share of voice fragile?
Share of voice is fragile because it is the only surviving metric whose denominator is a judgement call, so it can be moved several points without anything happening in the world. Adding two minor competitors to the counted set lowers your share; quietly dropping them raises it. The fix is not statistical: name the competitor set, freeze it for the quarter, version it, and print the version next to the number.
A second fragility is structural and cannot be fixed by discipline, only disclosed. Citation-based share of voice has an engine-dependent denominator, because engines attach very different numbers of sources to an answer. A vendor index published 26 June 2026, built on 126 million US AI-search prompts collected between January and April 2026, reported one assistant citing an average of about 15 sources per response against another citing about 3.7 The direction is corroborated by academic measurement of citations per response, so treat the ratio as real and the exact means as indicative. Being one of three cited sources and one of fifteen are different events, so a citation-based share is not comparable across engines and a mention-based share is the safer default.
Which numbers should never reach a dashboard at all?
Four numbers should never reach a dashboard, and each fails for its own reason rather than for being imprecise. The first is any traffic or revenue figure derived from citations. The coefficient that would license the conversion has never been measured: the field’s own 2026 survey rates the claim that citation scores predict clicks, conversions or revenue at very low confidence,2 and Pew Research Center’s July 2025 browsing study found users clicked a source inside an AI summary in about 1% of visits.8 Multiplying a citation count by an invented click rate produces a number with an unknown error term and a decimal point, not an estimate.
The second is a per-prompt score with a week-over-week arrow: a handful of runs of one prompt carries an interval far wider than any movement worth reporting, so the arrow is drawn through noise, and the arithmetic is in the sibling article on statistical power. The third is sentiment scoring, which collapses the signal you need: an answer can be warm in tone while stating your pricing model incorrectly, and the error is what costs the deal.
The fourth is any composite index whose formula is not printed next to it. If a number is a weighted blend of presence, position and prominence, the weights are the finding, and a reader who cannot see them cannot check whether last quarter’s figure was computed the same way. An index with a published formula is a metric; an index without one is a claim.
How do you decide whether a new metric is worth adding?
Put a candidate metric through four questions, and add it only if it answers all four. Does it have a denominator you control and can state? If the denominator is the engine’s traffic, prompt volume or user base, you cannot state it and the metric is a guess wearing a percent sign. Can you compute an interval around it? If the observation is not a countable event over a countable set of runs, there is no interval and no way to tell a change from a fluctuation.
Does a decision change when it moves? Write down, before you collect anything, the specific action a two-standard-error move would trigger. If you cannot name one, the metric is decoration on a page where attention is finite. Can it be falsified? Name the observation that would show a reading was wrong; the last column of the table above exists to force this. A metric with no falsifying observation describes a mood rather than measuring anything.
Run your existing dashboard through the same four questions: in most programmes two or three panels fail on question three alone.
The verdicts on this page are arguments, not measurements. No study has compared these eight metrics head to head for predictive validity, because that needs an outcome to predict, and the outcome everyone wants, revenue, is the one the field’s own survey rates at very low confidence. So “holds up” here means “has a stateable denominator, a computable interval and a decision attached”, not “has been shown to lead sales”. Framing accuracy is the weakest-defined survivor: there is no standard rubric, no published inter-rater agreement for this task, and the base rates quoted above come from studies of news publishers rather than of products.
Everything above is doable in a spreadsheet: fix the panel, run it, count named-in-prose and URL-attached separately per engine, compute a binomial interval on each, and read twenty answers a month against a rubric you wrote. Bavior runs a fixed prompt panel across five engines on a schedule, reports per engine rather than as a blended score, and records every cited URL including the ones that are not yours. What it does not do is forecast traffic or revenue from citations, because the coefficient does not exist. It does not score framing accuracy either, because judging whether a sentence about your product is true is a person reading answers against your rubric, not a classifier. The free visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).
- Grossman et al., SIGIR 2026; 11,500 queries; URL-level Jaccard 0.11–0.18 across Google organic, AI Overviews and Gemini: arxiv.org/abs/2604.27790
- “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035; confidence table, and “these results refute the notion of a global GEO ranking. Visibility is indexed by engine and surface” (survey preprint): arxiv.org/abs/2607.14035
- Jaźwińska & Chandrasekar, Tow Center for Digital Journalism, “AI Search Has a Citation Problem”, 6 Mar 2025; 1,600 queries, 20 publishers, 8 engines; >60% incorrect: cjr.org
- Tow Center for Digital Journalism, “How ChatGPT Search (Mis)represents Publisher Content”, Nov 2024; 200 quotes from 20 publishers; 153 of 200 responses partially or entirely incorrect: cjr.org
- Liu, Zhang, Liang, “Evaluating Verifiability in Generative Search Engines”, 2023; 51.5% of sentences fully supported by their citations: arxiv.org/abs/2304.09848
- Xu, Iqbal & Montgomery, 2026; ~11% of 98,020 atomic claims classified as insufficiently supported (preprint): arxiv.org/abs/2605.14021
- AI visibility index, published 26 Jun 2026: 126 million US AI-search prompts collected January–April 2026; one assistant averaged ~15 cited sources per response against ~3 for another; 36 brands visible across every platform covered. Vendor-published; described rather than linked, per this curriculum’s sourcing rule.
- Pew Research Center, 22 Jul 2025; users clicked a source inside an AI summary in about 1% of visits: pewresearch.org
Frequently asked questions.
Is there a benchmark AI visibility score I should aim for?
No, and any figure offered as one is comparing panels that are not comparable. Your rate is a function of the prompts you chose: a team measuring narrow bottom-of-funnel questions will show a lower number than a team measuring broad category questions about the same product, and neither number says anything about the other. The comparisons that mean something are your own panel against itself at a fixed version, and share of voice against a named competitor set inside that same panel on the same day. Both are internal, which is exactly why they are defensible.
My dashboard shows one AI visibility score. Is it useless?
It is not useless as a tripwire and it is useless as a diagnosis, which is the distinction worth holding. A single blended number can tell you that something moved enough to look at. It cannot tell you which engine moved, and since a SIGIR 2026 study measured URL-level Jaccard of 0.11–0.18 between three surfaces built by one company, the components are close to independent. Keep the composite if it is already on the page, print the per-engine split underneath it, and make sure every action anyone takes is triggered by the split rather than by the composite.
How do I measure framing accuracy without hiring anyone?
Read twenty answers a month against a fixed rubric and score each factual axis as a binary. Pick the axes that cost you deals: pricing model, what the product does, what it does not do, integrations, and any compliance or safety claim. Record the wrong sentence verbatim rather than a score, because the sentence is what you can act on. Weight the sample towards comparison and objection prompts, where an error is most expensive. Two readers should agree on at least four answers in five; when they do not, the rubric is too vague, not the engines too erratic.
Why can't I convert citations into an estimated traffic number?
Because the coefficient you would multiply by has never been measured, and the one relevant public figure says it is close to zero. Pew Research Center's July 2025 browsing study found users clicked a source inside an AI summary in about 1% of visits, and the field's own 2026 survey rates the claim that citation scores predict clicks, conversions or revenue at very low confidence. You can still state a modelling assumption out loud, "if 1% of readers click and the prompt is asked N times", provided you present it as an assumption with N unknown, not as a measurement. The moment it appears as a chart it will be quoted as one.
The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.
Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com