Home/Learn GEO/What cannot be measured
Measuring AI visibility · Limits

What are the limits of AI visibility measurement?

Four questions in this field have no instrument, and pretending otherwise is the main way GEO reporting loses its credibility. The useful move is to bound each one out loud rather than estimate it quietly.

On this page
Share this
Share on X Share on LinkedIn
The short answer

Four things in AI visibility have no instrument: how often a prompt is asked, the revenue a citation produced, what a model says about you when it is not searching, and the answer any one buyer actually saw. Each is missing the observation it would need rather than waiting for a better tool, so bound it out loud instead of estimating it quietly. On the second, Pew Research Center’s March 2025 browsing panel found users clicked a link inside a Google AI summary in just 1% of visits to a page that carried one, so most of these touches leave no event for analytics to record.2

Key takeaways
  • Report coverage, not volume: no engine publishes prompt frequency, so a mention rate describes the questions you chose and is never a market share.
  • Do not convert citations into revenue. The field’s own 2026 survey grades “citation scores predict clicks, conversions, or revenue” at very low confidence, its lowest.3
  • Bound each gap in three sentences: name the estimand, name the instrument you actually have, and name the observation that would show the reading was wrong.
The missing denominator

Why can nobody tell you how often your prompt is asked?

No engine publishes prompt frequency, and nothing reconstructs it, so every rate you report has a numerator you measured and a denominator you invented. Why the number does not exist, and what a vendor’s prompt-volume figure is actually made of, is worked through in the prompt research stage. What matters for a report is the consequence: your mention rate is a property of the questions you chose, not of the market.

That has one concrete effect on what a report may claim. “We are named in 34% of answers” is a defensible sentence about your panel. “We are named in 34% of AI answers about our category” is not, because the second version implies a denominator nobody has. Write the panel into the sentence and the claim becomes true.

The bound available to you is coverage rather than volume. Take the demand you can observe: question-shaped queries in your own search console, first questions from discovery calls, tickets opened in the first week of a trial. State what share of those your panel represents and you have a real fraction with a real denominator, the strongest coverage claim available here. One published signal helps you weight the panel: a 55,393-query study collected between 13 March and 21 April 2026 measured a Google AI Overview activation rate of 13.7% overall against 64.7% for queries phrased as questions.1 Question-shaped prompts are where AI answers happen, so a panel weighted towards them is measuring where the surface actually exists.

Attribution

Can a citation be attributed to revenue?

No, and the reason is mechanical rather than statistical: in the ordinary case the reader never clicks, so there is no event for any analytics system to record. Pew Research Center’s March 2025 browsing panel put clicks on a source inside an AI summary at about 1% of visits, and the magnitude of what a citation is worth as an impression is covered in the GEO versus SEO stage.2 The field’s own 2026 survey rates the claim that citation scores predict clicks, conversions or revenue at very low confidence, its lowest grade.3

Three mechanisms break the chain, and knowing which is operating tells you which workaround is worth trying. The first is the missing click: a reader satisfied by the answer never becomes a session. The second is the missing referrer: chat clients frequently send none at all, so a visit that did happen arrives as direct traffic. The third is the indistinguishable agent: a Tow Center for Digital Journalism analysis published 30 October 2025 reports that to a website, an agentic browser’s agent “is indistinguishable from a person using a standard Chrome browser”, and that such browsers retrieved subscriber-only material the same vendors’ standard interfaces could not.4 That class of traffic is growing quickly: Cloudflare’s network-wide 2025 review reported user-triggered crawling growing more than fifteenfold across the year.5 The share of your AI-mediated touches that are invisible in analytics is therefore rising, not falling.

No workaround recovers the missing coefficient, and be suspicious of anything presented as one. State a modelling assumption as a sentence instead, with its inputs named and its unknowns marked unknown, and never render it as a chart. A chart is quoted as a measurement by the second meeting.

The other answer

Can you measure what a model says without retrieval?

You can sample it, and you cannot track it over time, because the instrument changes underneath you. When a model answers without searching, the description of your product comes from training data you cannot inspect, cannot edit and cannot influence on a schedule. Sampling that today is easy: ask the question with search disabled and read what comes back. The problem is the second measurement.

A time series needs a fixed instrument, and a model is not one. Each release is a new corpus, a new cut-off and new post-training, so a change between two no-search measurements may be a change in your reputation or may be a change in the model, and nothing in the data distinguishes them. Even the version label often does not help: one product name can route to different checkpoints without announcement. The 2026 critical survey rates “commercial engines differ from one another and vary over time” at high confidence, which is exactly what a no-retrieval time series would need to be free of.3

So treat the no-search run as a periodic snapshot with the model version stamped on it, read qualitatively. It answers a genuinely useful question: what does this system believe about us when it is not reading anything. It is also the earliest warning you get that a wrong fact about your product has settled into the corpus. Report it as quoted text next to its date and version, never as a trend line.

One run, one session

Can you measure the answer a real buyer actually saw?

No, and this is the limit most often mistaken for noise in the data. A panel runs a clean-room session: fresh context, no signed-in account, a fixed locale, search left on, one run per prompt. A buyer holds none of that still. Anthropic documents that Claude “saves memory as a set of individual topics as you chat”, carries that into later conversations, and lets the user switch it off, so two people asking the same question are not querying the same system.7 Nothing you run reproduces their state.

The other half is non-determinism, and unlike personalisation it can at least be quantified. The March 2026 statistical treatment of AI visibility reports that “citation rankings are unstable across samples, not only among top-ranked domains but throughout the frequently cited domain set”, and that bootstrap intervals place many apparent differences between domains inside “the noise floor of the measurement process”.6 The 2026 critical survey records the same pattern from the audit side: commercial audits show “low source overlap, substantial run-to-run variability, and persistent fidelity gaps”.3

The working rule that follows is to stop treating one run as an observation. Repeat each prompt, report a rate with an interval rather than a point, and call a week-on-week move real only once it clears that interval; the survey’s own checklist asks studies for multiple paraphrases and repetitions for this reason.3 How many repeats that takes is arithmetic, worked out in the statistical-power article. What repetition never buys is the buyer’s own answer, which stays unobservable, so keep your session settings written beside the number rather than in an appendix.

Proxies

Which proxies are worth keeping, and how should they be labelled?

Three proxies are worth keeping, and each is only safe once it carries an explicit label saying what it cannot support.

01

Self-reported attribution

A free-text “how did you hear about us” field at signup is the only place a buyer can tell you an assistant was involved. Label: self-selected, unprompted, and systematically under-reported by people who do not remember the touch.

02

Crawler fetches of high-intent pages

Search-side and user-triggered fetches of your pricing and comparison pages are a real demand signal, covered in the crawler-log article. Label: leading indicator of interest, not an outcome, and not attributable to a person.

03

Branded and direct traffic

A rise in branded search or direct visits alongside a rise in mention rate is weak corroboration. Label: confounded with every other marketing activity in the same window, and useless without a holdout period.

The label is the whole discipline here. A proxy stripped of its caveat becomes a measurement in the retelling, so write the caveat into the chart title rather than a footnote, and the number stays honest when the slide is forwarded without you.

The method

How do you bound a question you cannot measure?

Bounding a question means answering it with a stated range and a stated assumption instead of a point estimate, and it takes three sentences. Name the estimand: write down precisely what you wish you knew, as in “how many buyers formed an opinion of us inside an assistant last quarter”, because the vague version is what lets a proxy quietly stand in for it. Name the instrument: say what you actually have and what it measures, so that a 200-observation panel measures the share of your chosen questions that named you, on five engines, in a date range. Name the falsifier: say what observation would show the reading was wrong, which for a bound usually means naming the assumption whose failure breaks it.

Written out, a bounded claim reads like this: “On our 60-prompt panel, four of five engines named us in 28–36% of answers in July, at 200 observations per engine. We do not know how often those questions are asked, so this is not a market share. If our buyers ask questions materially different from our panel, this number does not describe them.” That paragraph is shorter than most executive summaries and it cannot be misquoted, which is the point.

The same discipline is what the research literature asks of itself. The 2026 critical survey’s recommendation is to build a retrievable page and then “measure retrieval, citation, and fidelity separately”, and its minimum-study checklist requires the product, mode, model, date, locale, account and whether search was enabled to be recorded on every reading, and null outcomes to be retained rather than dropped.3 None of this stops you measuring. It stops the measurement being asked to carry a claim it was never built for.

The honest limit of this article

The four limits described here are structural today, not laws of nature, and two of them could close. If an engine ever published prompt frequency, or if assistants began passing a consistent referrer, the first two sections of this page would be obsolete within a quarter, and there is no way to know in advance whether that will happen. The other two are harder: a no-retrieval time series needs a fixed instrument and a retrained model is not one, while the answer one person saw is observable by nobody but them. Be careful, too, with the strongest number quoted here. The roughly 1% click figure comes from one behavioural panel of 900 US adults in March 2025, measuring Google AI summaries specifically; it is the best independent evidence available and it is one study, on one surface, in one country, in one month.

Where a product fits, and where it does not

Bounding costs nothing but discipline: write the estimand, the instrument and the falsifier into the report, add a free-text attribution field at signup, and stop converting citations into sessions. Bavior runs a fixed prompt panel across five engines on a schedule, records every cited URL including the ones that are not yours, and reports per engine rather than as a blended score. What it does not do is any of the four things on this page: it cannot tell you how often a question is asked, because no engine publishes that; it does not forecast traffic or revenue from citations, because the field’s own survey rates that link at very low confidence; it cannot tell you what a model believes about you when it is not searching, beyond letting you read a snapshot for yourself; and it cannot show you the answer one buyer saw, because that session was theirs. The free visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).

Sources, all checked 30 Aug 2026
  1. Xu, Iqbal & Montgomery, 2026; 55,393 queries collected 13 March – 21 April 2026; “overall AIO activation is 13.7%, rising to 64.7% for question-form queries” (preprint): arxiv.org/abs/2605.14021
  2. Pew Research Center, 22 Jul 2025; browsing panel of 900 US adults, 68,879 searches, March 2025; on clicking a link inside an AI summary, “this occurred in just 1% of all visits to pages with such a summary”: pewresearch.org
  3. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035; confidence table and minimum-study checklist (survey preprint): arxiv.org/abs/2607.14035
  4. Chandrasekar & Jaźwińska, Tow Center for Digital Journalism, “How AI browsers sneak past blockers and paywalls”, 30 Oct 2025; an agentic browser is “indistinguishable from a person using a standard Chrome browser”: cjr.org
  5. Cloudflare Radar 2025 Year in Review, data 1 Jan – 2 Dec 2025; “AI ‘user action’ crawling increased by over 15x in 2025”: blog.cloudflare.com
  6. Sielinski, “Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement”, Mar 2026, rev. 26 Aug 2026, arXiv:2603.08924 (preprint): arxiv.org/abs/2603.08924
  7. Anthropic, “Use Claude’s chat search and memory to build on previous context” (memory saved across chats, user-controllable; first-party): support.claude.com/en/articles/11817273
FAQ

Frequently asked questions.

How do I know how many people are asking AI about my category?

You cannot, because no engine publishes prompt frequency and no proxy reconstructs it. What you can measure honestly is coverage rather than volume: take the questions you have actually observed buyers asking, meaning question-shaped queries in your own search console, first questions from discovery calls and tickets opened during trials, then state what share of those your panel contains. "This panel covers 40 of the 55 distinct questions we have observed" is a true sentence with a real denominator. "We appear in 34% of AI answers in our category" is not, because nobody has that denominator.

Can I put a revenue number on AI visibility for my board?

Not as a measurement, and you can state a bounded claim instead. The chain from citation to revenue breaks in three places. The reader usually does not click, and Pew Research Center's March 2025 browsing panel put clicks on a source inside an AI summary at about 1%. Chat clients often send no referrer, so the visits that do happen arrive as direct traffic. And agentic browsers are indistinguishable from a person in server logs. Write a sentence that names the assumption and marks the unknowns as unknown, and keep it out of chart form, because a chart gets quoted as a measurement.

Should I track what a model says about me when search is off?

Sample it periodically, stamp it with the model version, and read it as text rather than plotting it. A no-search answer comes from training data you cannot inspect or edit, which makes it the earliest warning that a wrong fact about your product has settled into the corpus, and that is genuinely worth knowing. What it cannot support is a time series, because each model release is a new corpus and a new cut-off, so a change between two measurements may be a change in the model rather than in your reputation, and nothing in the data separates the two.

What is the difference between bounding a question and estimating it?

An estimate offers a point value and hides its assumptions; a bound offers a range and names them. Bounding takes three sentences: state the estimand, meaning precisely what you wish you knew; state the instrument, meaning what you actually have and what it measures, including the panel, the engines and the date range; and state the falsifier, meaning the observation that would show the reading was wrong. "On our 60-prompt panel, four of five engines named us in 28–36% of answers in July, at 200 observations per engine; we do not know how often those questions are asked, so this is not a market share" cannot be misquoted, which is the whole benefit.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Bound what you cannot measure.
Then the rest can be trusted.

Start free trial