Home/Learn GEO/A defensible report
Measuring AI visibility · Reporting

What does a defensible AI visibility report contain?

Every figure carries its observation count, its date range and its panel version, or it cannot be compared with anything. A holdout set is what separates your work from the drift of the platform underneath it.

On this page
Share this
Share on X Share on LinkedIn
The short answer

A defensible AI visibility report gives mention rate and citation rate per engine, each with an interval and each stamped with its observation count, date range and panel version. It also carries a holdout set of prompts nobody acted on, so platform drift stays separable from your own work. Per-engine reporting is not fussiness: a SIGIR 2026 study of 11,500 queries measured URL-level Jaccard similarity of 0.11 to 0.18 between Google organic results, AI Overviews and Gemini, so one blended score describes no engine that exists.5

Key takeaways
  • Four stamps travel with every figure: engine, observation count, date range and panel version. A number missing one cannot be compared with last month’s.
  • A frozen holdout of ten to fifteen prompts, mirroring the panel’s class mix, separates your effect from engine drift.
  • Keep off the page: cross-engine averages, arrows smaller than your minimum detectable effect, and revenue inferred from citations.
Provenance

What must travel with every figure on the page?

Four things travel with every number, or the number cannot be compared with anything: the engine it came from, the observation count behind it, the date range it covers, and the version of the prompt panel that produced it. A figure missing one of them is not a measurement: a reader cannot tell whether this month’s value and last month’s came from the same instrument.

The panel version is the one teams skip, and the one that quietly invalidates a year of history. Adding six prompts changes the denominator of every rate, so a rise from 29% to 34% across two versions may be an artefact of which questions were asked. Use a two-part version, printed beside every figure. A minor bump is a wording fix that does not change what a prompt asks; comparisons across it need only a footnote. A major bump is a prompt added, removed or re-scoped; there you re-baseline rather than compare. Keep the prompt list in version control so anyone can diff two versions.

One scheduling rule follows: never change the panel in the same period as the content or off-site work you want to measure. Do that and the two causes are confounded beyond recovery. Change the panel in a quiet period, take a fresh baseline, then act.

The interval belongs on the figure too. The 2026 critical survey’s minimum checklist for a GEO study asks for product, mode, model, date and locale, whether search was enabled, repeated runs in multiple time windows, retrieval and generation kept apart, and null outcomes retained in the denominator.1 A commercial report meeting that research checklist is simply one an outsider could check.

The control

What is a holdout set, and why does a report need one?

A holdout set is a slice of your prompt panel that you deliberately never act on: no content written for it, no off-site work aimed at it, no removal when it looks bad. It exists so that engine-wide drift shows up separately from your work. Without one, a rise and a platform update are indistinguishable in a dashboard, as are a fall and a bad month.

Build it once, at panel creation. Ten to fifteen prompts is enough, and they must mirror the class mix of the rest of the panel: classes drift at different rates, so an all-definitions holdout will not track a change that hit comparison questions. Freeze it in the same version scheme as everything else, and mark it so a new team member cannot accidentally optimise for it.

Read it as a difference in differences. Your effect is roughly the change in the worked set minus the change in the holdout. If both fall eight points, the engine moved and you did nothing wrong. If the worked set rises nine and the holdout is flat, you have evidence. That subtraction is what turns a before-and-after number into a claim about your own work.

Two honest limits. A holdout of a dozen prompts carries a wide interval of its own, so read its movement as a direction rather than a precise correction: a noisy number minus a noisy number is not a clean one, and the combined comparison needs more data than either side alone (the arithmetic is in the article on statistical power). And it controls only for changes that hit both sets equally: it catches a model swap, not a competitor campaign aimed at the questions you were working on.

The reason to pay for all this is that platform drift is large. The 2026 critical survey rates “commercial engines differ from one another and vary over time” at high confidence.1 A vendor time series makes the size concrete: between 1.9 million AI Overview citations sampled in July 2025 and 4 million cited URLs across 863,000 keyword result pages in March 2026, the share of cited pages also ranking in the organic top ten fell from roughly 76% to roughly 38%.6 A team that ran a before-and-after across that window, with no holdout, would have measured the platform and written it up as their own result.

The page itself

What belongs on the page, and what stays off?

Five things belong on the page and five are actively harmful; the line is whether an item survives questioning by somebody who did not run the panel.

On the page

  • Mention rate and citation rate per engine, each with an interval and an observation count
  • The worked set beside the holdout set, so drift and effect are separable at a glance
  • A split by prompt class, because a gain in one class and a loss in another cancel in the total
  • The cited-URL leaderboard for the whole panel, including every competitor and third party
  • Two or three verbatim answers, at least one of which described you wrongly

Off the page

  • Any single number that averages across engines
  • Arrows on movements smaller than your minimum detectable effect
  • Screenshots offered as evidence rather than as illustration
  • Estimated sessions or pipeline attributed to citations
  • Runs quietly dropped because you were absent from them

Two entries need their reasons stated. The verbatim wrong answer belongs on the page because presence metrics score a confidently false description of your product as a win, and the base rate is not low: a Tow Center for Digital Journalism audit of 1,600 queries across eight engines, published 6 March 2025, found they “provided incorrect answers to more than 60 percent of queries”.2 One quoted sentence does more to get a fix funded than any chart.

The revenue line stays off because the coefficient does not exist. The field’s own survey rates the claim that citation scores predict clicks, conversions or revenue at very low confidence,1 and Pew Research Center’s July 2025 browsing study found users clicked a link inside an AI summary in about 1% of visits.3 If leadership needs a commercial framing, state the modelling assumption out loud and label it as one; the moment it becomes a chart it will be quoted as a measurement.

Per engine

Why can a report not average across engines?

Because the engines are not looking at the same web. A SIGIR 2026 study compared the sources returned for 11,500 queries by Google organic search, AI Overviews and Gemini Flash 2.5, and measured URL-level Jaccard similarity of 0.18 between AI Overviews and the organic page, 0.16 between organic and Gemini, and 0.11 between the two AI surfaces: “regardless of query subset and which pair of search engines is compared, the retrieved lists are dissimilar, despite all three being developed by Google.”5 Three surfaces from one vendor already disagree.

A blended score then hides what a reader needs. Sit at 41% on one engine and 9% on another, and the average of 25% is true of neither; it moves when the weighting changes, not when your visibility does. The 2026 critical survey arrives at the same place from the audit literature, reporting “low source overlap, substantial run-to-run variability, and persistent fidelity gaps” across commercial engines.1

Report a small table instead, one row per engine with its own count, interval and panel version. If leadership wants one tracked number, make it a single named engine and say why that one matters commercially.

Cadence

How often should the report actually be published?

Publish the numbers monthly and run the panel weekly, because those are two different jobs. A weekly panel carrying around 200 observations per engine can only distinguish a move of roughly 13 to 14 points from noise, while a monthly aggregate of the same runs brings that down to about 6.6 points, and almost every real change lands between those two figures.

The weekly read is qualitative and genuinely useful. Read the answers rather than the rates, and look for three things: a competitor or third-party URL newly appearing in the citation list, a false statement about your product that was not there last week, and a prompt that stopped returning an AI answer at all. Each is an event rather than a rate, so none needs statistical power. Log them and act; save the arithmetic for the monthly page.

Quarterly, re-read the panel against the business: prompts age, and a panel that is never retired slowly measures a question nobody asks. That is a major bump and a fresh baseline, so schedule it.

Null results

What should the report say when nothing moved?

It should say “no detectable change” and print the size of change that would have been detectable. “Mention rate 31% on this engine, 200 observations, no detectable change from 29% last month; the smallest change this panel could have detected is about 13 points” tells a reader exactly what was and was not learned. “Up 2 points” tells them something false.

Two habits protect the null result from being tidied away. Keep every run in the denominator, including the ones where you did not appear: the 2026 critical survey’s checklist requires it, and quietly dropping empty runs is the most common way a panel starts flattering the team that owns it.1 And report an interval rather than a point estimate. The March 2026 statistical treatment of AI-visibility measurement found citation counts power-law shaped and rankings “unstable across samples, not only among top-ranked domains but throughout the frequently cited domain set”, and used bootstrap resampling to report that uncertainty.4 Where a rate is very low or very high the normal approximation degrades, so an exact or bootstrap interval is the honest instrument; that caveat is ordinary statistics, not a finding of the paper.

Expect the null more often than not: the NeurIPS 2025 C-SEO Bench found that “most current C-SEO methods are not only largely ineffective but also frequently have a negative impact on document ranking”.7 Even so, a flat month is rarely a month with nothing to report. The citation list will have changed when your rate did not, and that is where the actions are: whose pages replaced whose, which of your own pages never appear, and which answers now say something untrue about you. A report whose value depends on the rate going up will eventually be written to make the rate go up.

The honest limit of this article

The holdout design here has not been validated for this setting, and cannot easily be. A difference-in-differences comparison assumes the two sets would have drifted in parallel had you done nothing, and no published study has tested whether AI answer engines drift in parallel across prompt classes; the one relevant measurement, of citation counts over nine consecutive days, found instability throughout the frequently cited set rather than a common trend. So a holdout tells you reliably that something platform-wide happened, not by how much: show the two sets side by side rather than publishing the subtraction as one adjusted figure. Everything else here is a discipline rather than a finding: provenance stamps make a report checkable, which is not the same as making it right.

Where a product fits, and where it does not

This report is a spreadsheet and a decision to be honest in it. Version the prompt list, mark a dozen prompts as the holdout and never touch them, stamp every figure with engine, count, dates and version, and write “no detectable change” when that happened. Bavior runs a fixed prompt panel across five engines on a schedule, keeps every run including the empty ones, records every cited URL including the ones that are not yours, and reports per engine. What it does not do is decide which prompts are your holdout, enforce that nobody optimises for them, or forecast traffic or revenue from citations. The free visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 30 Aug 2026).

Sources, all checked 30 Aug 2026
  1. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035; Tables 5 and 6 (survey preprint): arxiv.org/abs/2607.14035
  2. Jaźwińska & Chandrasekar, Tow Center for Digital Journalism, “AI Search Has a Citation Problem”, 6 Mar 2025; 1,600 queries across eight engines; >60% incorrect: cjr.org
  3. Pew Research Center, 22 Jul 2025; users clicked a source inside an AI summary in about 1% of visits: pewresearch.org
  4. Sielinski, “Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement”, Mar 2026, arXiv:2603.08924 (preprint): arxiv.org/abs/2603.08924
  5. Grossman et al., “How Generative AI Disrupts Search”, SIGIR 2026; 11,500 queries; URL-level Jaccard 0.11–0.18 across Google organic, AI Overviews and Gemini: arxiv.org/abs/2604.27790
  6. AI Overview citation time series: 1.9 million citations sampled July 2025, compared with 4 million cited URLs across 863,000 keyword result pages in March 2026; the share of cited pages also ranking in the organic top ten fell from ~76% to ~38%. Vendor-published; described, not linked, per this curriculum’s sourcing rule.
  7. Puerto et al., “C-SEO Bench: Does Conversational SEO Work?”, NeurIPS 2025 Datasets & Benchmarks Track, arXiv:2506.11097: arxiv.org/abs/2506.11097
FAQ

Frequently asked questions.

How big should a holdout set of prompts be?

Ten to fifteen prompts, chosen to mirror the class mix of the rest of the panel. The mix matters more than the count: if a third of your working prompts are comparison questions, a third of the holdout should be, because prompt classes drift at different rates and a holdout made only of definition questions will not track a change that hit comparison questions. Freeze it at panel creation, version it with everything else, and mark it so clearly that a new team member cannot accidentally write content for it. Its own interval will be wide, so read its movement as a direction rather than as a precise correction.

Why does the panel need a version number?

Because adding or removing a prompt changes the denominator of every rate on the page, so two figures produced by different panel versions are not comparable even though they look identical. Use a two-part version: a minor bump for a wording fix that does not change what a prompt asks, across which comparisons are fine with a footnote, and a major bump for a prompt added, removed or re-scoped, across which you re-baseline instead of comparing. Print the version next to every figure and keep the prompt list in version control so anyone can diff two versions. Never change the panel in the same period you ship the work you are trying to measure.

Should an AI visibility report include estimated traffic?

No, because the coefficient you would multiply by has never been measured. The field's own 2026 survey rates the claim that citation scores predict clicks, conversions or revenue at very low confidence, and Pew Research Center's July 2025 browsing study found users clicked a source inside an AI summary in about 1% of visits. If leadership needs a commercial framing, write it as a sentence that names the assumption ("if roughly 1% of readers click and this question is asked often") rather than as a chart, because a chart will be quoted as a measurement by the second meeting.

What do I write in the report when the numbers did not move?

Write "no detectable change" and print the smallest change the panel could have detected next to it. "Mention rate 31%, 200 observations, no detectable change from 29% last month; smallest detectable change about 13 points" tells a reader exactly what was and was not learned, while "up 2 points" tells them something false. Then report what did change: the citation list moves even in months when your rate does not, and that is where the actions are, in whose pages replaced whose, which of your own pages never appear, and which answers now say something untrue about you.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Stamp every figure with its provenance.
Then it can be defended.

Start free trial