Home/Learn GEO/Six things
Measuring AI visibility · Definitions

What are the six things people call AI visibility?

Six separate measurements share one label, they move independently, and most reporting shows two of them. The gap between mention and citation is the one that changes what you do on Monday.

On this page
Share this
Share on X Share on LinkedIn
The short answer

AI visibility is six different measurements wearing one label: mention rate, citation rate, share of voice, position in the answer, framing accuracy, and source attribution. They move independently, so a single blended visibility score has averaged away the one thing you could have acted on. That they come apart is measurable rather than theoretical: an evaluation of generative search engines found that on average only 51.5% of generated sentences were fully supported by the citations attached to them,7 so being cited and being described correctly are plainly not the same quantity.

Key takeaways
  • The six are not variants of one number. Mention rate and citation rate are shares of answers, share of voice is a ratio between brands, position is an ordinal, framing accuracy is a human judgement, and source attribution is a list of URLs rather than a score.
  • Mention against citation is the gap that changes what you do. A link with no name in the prose is a sentence problem on a page you own; a name with no link comes from training data, which on-page work will not move inside a release cycle.
  • Source attribution is the only one that names an action, because it records every competitor and third-party URL an answer cited, not just yours. All six are indexed per engine, so five engines means thirty numbers.
Definitions

What are the six quantities called AI visibility?

Six separate quantities travel under the name AI visibility, and a typical dashboard reports two of them. Each definition below stands on its own, because these six get swapped for one another constantly.

01

Mention rate

Mention rate is the share of answers whose prose names your brand, whether or not a link is attached. It is the quantity a reader actually experiences, because most readers never click a source.

02

Citation rate

Citation rate is the share of answers that attach one of your URLs as a source, whether or not the prose names you. It is the only one of the six your server logs can partially corroborate.

03

Share of voice

Share of voice is your mentions divided by all brand mentions in the same answer set. It is a ratio, so it moves when a competitor moves even though nothing about you changed.

04

Position in answer

Position in answer is where in the response you appear. The founding academic metric in this field weighted attributed text by a function that decays exponentially with position.1

05

Framing accuracy

Framing accuracy is whether the sentence written about you is true. It is the only one of the six that can be negative: a confidently wrong description scores as presence on all five of the others.

06

Source attribution

Source attribution is the full list of URLs an answer cited, including every competitor’s and every third party’s. It is the only one of the six that tells you what to do next.

Two properties make the six non-interchangeable, and both are documented rather than asserted. The first is that they answer different questions about different stages of the pipeline: the 2026 critical survey of generative engine optimization ends on the instruction to “produce a relevant, comprehensive, verifiable, clearly structured, and technically retrievable page; then measure retrieval, citation, and fidelity separately.”2 Retrieval, citation and fidelity are three measurements in that sentence, not one.

The second is that every one of the six is indexed per engine. A SIGIR 2026 study of 11,500 queries measured URL-level Jaccard similarity of 0.11–0.18 between Google organic results, AI Overviews and Gemini, three surfaces built by one company on one index.3 Six quantities across five engines is thirty numbers, not one; the measuring AI visibility stage sets out the reporting shape for them.

The important gap

Why do mention rate and citation rate move apart?

Mention and citation come out of different halves of the system, so the direction of the gap between them is a diagnosis rather than a curiosity. A citation is produced when a retrieved document is attached to a sentence; a mention is produced when the model writes your name into the prose. Two decisions, at different points, on different inputs.

A link with no name means retrieval worked and identification failed. Your page was fetched, judged relevant and attached, and then the passage the engine used did not say who was speaking. This is the most fixable failure in the whole field, because it is a sentence-construction problem on a page you own: the quoted passage says “the tool costs $99 a month” when it needed to say “Bavior costs $99 a month”. Name the subject inside the sentence that carries the fact, not two paragraphs above it.

A name with no link means the model is drawing on training data rather than on what it just retrieved. That is a much slower problem. The description came from the corpus the model was trained on, which you cannot inspect, cannot edit, and cannot expect to change inside one release cycle. On-page work does not move it; what moves it, slowly, is what third parties write about you, which is the subject of the off-site GEO stage.

The gap is not small. A vendor study published in June 2026, covering 115 prompts and 3,981 domain appearances across four assistants in fourteen countries, reported that roughly 62% of source appearances were links whose brand the answer text never named, and found the ratio inverted between engines: one assistant named brands in 83.7% of answers while citing them in 21.4%, another named them in 20.7% while citing them in 87%.4 Treat those percentages as directional: a 115-prompt sample cannot support two decimal places, and no academic replication exists. What survives is the qualitative claim, that the two rates are not proxies for each other and on some engines are nearly inverse.

The denominator

What does share of voice actually depend on?

Share of voice depends entirely on a denominator you chose, which makes it the easiest of the six to move without doing any work. Add two small competitors to the counted set and your share falls; drop them and it rises. Neither movement is a fact about the engines. The fix is procedural rather than statistical: name the competitor set in the report, freeze it for the quarter, and record the version alongside every figure so two reports can be compared at all.

There is a second denominator problem that is specific to citation-based share of voice, and it is not fixable by discipline. Engines attach different numbers of sources to an answer. In a 602-prompt study that collected 21,143 search-layer citations across three platforms, mean citations per response were 6.88 for ChatGPT, 12.06 for Google’s AI surfaces and 16.35 for Perplexity.5 Being one of seven cited sources is a different event from being one of sixteen. A citation-based share of voice is therefore not comparable across engines even before any question of preference arises, which is one more reason the per-engine split in the measuring AI visibility stage is not a stylistic choice.

Mention-based share of voice is the more robust of the two, because prose length varies far less between engines than source count does. Report it that way, and say which one you picked.

Position

Does where you appear inside the answer matter?

Position inside the answer is worth recording and is the weakest-evidenced of the six as a business metric. It exists because the paper that started this field built its headline metric on it: the KDD 2024 GEO paper scored a source with Position-Adjusted Word Count, which weights the words attributed to a source by a function that decays exponentially with where those words sit in the response.1 That is a well-specified measurement, validated against other properties of the generated answer rather than against clicks, conversions or revenue.

Two rules make it usable anyway. Read it only in aggregate across a panel, because a single answer’s ordering is one draw from a distribution that re-rolls every run. And read it as an ordinal rather than a distance: first-of-three and first-of-sixteen are not the same position.

Fidelity

Why is framing accuracy counted separately?

Framing accuracy is counted separately because every other metric treats a false sentence about you as a success. If an engine writes that your product has no free tier when it does, and links your pricing page while doing it, your mention rate, citation rate, share of voice and position all improve. The only number that moves the right way is the one nobody collects.

The base rate justifies the effort. A Tow Center for Digital Journalism audit published on 6 March 2025 ran 1,600 queries, twenty publishers with ten articles each across eight engines, and found more than 60% returned incorrect answers, with the best-performing engine still wrong 37% of the time.6 Independent work on generative search found only 51.5% of sentences in answers were fully supported by the citations attached to them.7 Neither study measured brands, and both measured the property that matters here: the attached link does not make the sentence true. The sibling article on which metrics hold up covers how to sample this cheaply.

The sixth quantity

Why log everybody’s cited URLs and not just yours?

Logging every cited URL, not only the ones pointing at you, is what converts a visibility number into a list of things to do this week. Your own rate tells you where you stand; the full citation list tells you what is standing there instead of you.

Three actions come out of it directly. Third-party pages that recur across your shortlist prompts, such as a comparison roundup, a forum thread or a directory listing, are the off-site targets, ranked by frequency in your own panel rather than by anybody’s authority score. Your own pages split into the ones that get cited and the ones that never do, which turns a content backlog into an ordered one. And pages carrying stale facts about you, an old price or a discontinued integration, become a correction list, which is a different job from a content job.

This only works if the log is raw. Store, per run: the engine and mode, the model version if it is exposed, the prompt identifier and its class, the timestamp, the full answer text, every cited URL in order, and a flag for whether your brand name appeared in the prose. All six quantities are computable from that record and none of them from a summary written afterwards. The 2026 critical survey’s minimum-study checklist asks for the same discipline, including the instruction most teams break: keep the runs where nothing happened in the denominator.2

The honest limit of this article

Two of the six quantities have no published evidence linking them to anything a business cares about. Position in the answer comes from a metric validated against other properties of the generated text, never against clicks or revenue. Share of voice has no external validation at all: it is a ratio anyone can define, and no study has tested whether moving it moves anything else. The field’s own 2026 survey rates the claim that citation scores predict clicks, conversions or revenue at very low confidence, and that rating covers all six. The 62% link-without-name figure quoted above rests on a 115-prompt vendor sample with no replication. Use these six because they are the decomposition that makes your decisions legible, not because any of them has been shown to be a leading indicator of sales.

Where a product fits, and where it does not

You can collect all six by hand and the method above is the whole method: fix a prompt list, run it on a schedule, paste each answer into a spreadsheet with its cited URLs, and mark whether your name appeared in the prose. A hundred runs a month is an afternoon of work and it produces a defensible number. Bavior runs a fixed prompt panel across five engines on a schedule, records every cited URL including the ones that are not yours, and reports per engine rather than as a blended score. What it does not do is forecast traffic or revenue from any of these six, because the field’s own survey rates that link at very low confidence, and it does not judge framing accuracy for you, because deciding whether a sentence about your product is true is a human reading answers, not a classifier. The free visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).

Sources, all checked 29 Aug 2026
  1. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, Deshpande, “GEO: Generative Engine Optimization”, KDD ’24; Position-Adjusted Word Count weights attributed text by an exponentially decaying function of position: arxiv.org/abs/2311.09735
  2. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035; confidence table, minimum-study checklist, and the instruction to measure retrieval, citation and fidelity separately (survey preprint): arxiv.org/abs/2607.14035
  3. Grossman et al., SIGIR 2026; 11,500 queries; URL-level Jaccard 0.11–0.18 across Google organic, AI Overviews and Gemini: arxiv.org/abs/2604.27790
  4. Ghost-citation study, published June 2026: 115 prompts, 3,981 domain appearances, four assistants, fourteen countries; ~62% of source appearances were links whose brand the answer never named; one assistant 83.7% mention / 21.4% citation, another 20.7% / 87%. Vendor-published; described rather than linked, per this curriculum’s sourcing rule.
  5. Zhang, He, Yao, “From Citation Selection to Citation Absorption: A Measurement Framework for GEO Across AI Search Platforms”, arXiv:2604.25707 (preprint); 602 controlled prompts, 21,143 search-layer citations; mean citations per response 6.88 / 12.06 / 16.35: arxiv.org/abs/2604.25707
  6. Jaźwińska & Chandrasekar, Tow Center for Digital Journalism, “AI Search Has a Citation Problem”, 6 Mar 2025; 1,600 queries across eight engines: cjr.org
  7. Liu, Zhang, Liang, “Evaluating Verifiability in Generative Search Engines”, 2023; 51.5% of sentences fully supported by their citations: arxiv.org/abs/2304.09848
FAQ

Frequently asked questions.

Is a mention better than a citation, or the other way round?

They answer different questions, so the useful move is to report both and read the gap. A mention is what a reader sees, and most readers never click a source, so a mention with no link still reaches the person. A citation is the machine-verifiable half: it means your page was retrieved and attached, which is the part your own work can influence directly. If you are cited but not named, fix the sentence on the page so the fact and the brand name sit in the same sentence. If you are named but not cited, the description is coming from training data, and only what third parties publish about you will move it.

How many of the six should a small team actually track?

Four, and the fourth is the one usually skipped. Track mention rate and citation rate per engine because they diverge and the direction is a diagnosis; track mention-based share of voice against a competitor set you name and freeze; and log every cited URL in the answer, including your competitors' and every third party's, because that list is the only part of the data that names an action. Position in the answer is worth recording but not worth deciding on. Framing accuracy is worth sampling by hand once a month rather than measuring every run, because it needs a person to read the sentence.

Can I add the six numbers together into one visibility score?

No, because they have different units and different denominators, and the sum destroys the only decisions they support. Mention rate and citation rate are shares of answers; share of voice is a ratio between brands; position is an ordinal; framing accuracy is a proportion of statements judged by a human. Averaging them produces a number that cannot tell you which engine to work on or which of the six failed. On top of that, each is measured per engine, and a SIGIR 2026 study found URL-level Jaccard of 0.11–0.18 between three surfaces built by one company, so a cross-engine blend is averaging things that barely overlap.

What should I log on every run so I can compute all six later?

Seven fields, logged raw: the engine and mode, the model version if it is exposed, the prompt identifier and its class, the run timestamp, the full answer text, every cited URL in the order it appeared, and a flag for whether your brand name appeared in the prose. All six quantities are computable from that record and none of them is computable from a summary written afterwards. Keep the runs where nothing happened, because the 2026 critical survey's minimum-study checklist is explicit that null outcomes stay in the denominator, and deleting them is the single most common way a panel starts flattering itself.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Six numbers, not one.
Start with the two that disagree.

Start free trial