Home/Learn GEO/How much you own
Off-site GEO · Sizing the problem

How much of an AI answer do you actually own?

Before deciding what off-site work is worth, size the part on-site work can never reach. The published averages are not your share, and the arithmetic between them is where GEO plans go wrong.

On this page
Share this
Share on X Share on LinkedIn
The short answer

A minority of it on a typical commercial question, and the number worth planning against is one you measure on your own prompt set rather than one you read in a study. Pages you can edit today are one bucket of three, next to third-party pages about you and pages unrelated to you. The only public academic taxonomy of AI citations puts its largest single category, labelled Official, at 34.22% of citations on ChatGPT, 46.35% on Google’s AI surfaces and 44.07% on Perplexity, which means most citations on every engine measured are not official, and that share covers every entity the answer names rather than you alone.1

Key takeaways
  • “Largest single category” and “most of the answer” are different claims. Official leads on all three engines measured and peaks at 46.35%, so 54% to 66% of citations are something else.
  • The category is shared. A Google AI answer in that dataset cites 12.06 sources on average and the official slice splits across every entity named, so one brand’s slice is a fraction of a fraction.
  • Your reachable share is smaller again: an owned page must be fetchable and relevant to a sub-query before its prose counts.
  • Measure your own share from your own citation log. At 95% confidence it takes roughly 200 logged answers before a seven-point move beats noise.
Definition

What counts as a page you own?

Ownership needs a narrow operational meaning here: a cited URL sits on a domain whose text you can change today without asking anybody. That splits every citation into three buckets. Owned is your site, your documentation, your changelog, your pricing page. Influenced is third-party pages that mention you and that you move only by giving somebody a reason to edit them: roundups, review directories, community threads, news coverage. Neither is the rest, which on most questions is the bulk of the answer.

Two things get miscounted. A page on a domain you rent rather than own, a marketplace listing or a profile on somebody else’s platform, behaves like an influenced page: the operator can change the template, the canonical or the robots rules without telling you. And a brand mention with no link is not a citation at all. It is a separate quantity, kept apart in six things called visibility, and your owned share of citations and of the prose are different numbers with different fixes.

The denominator

Why is the largest citation category not your share?

A category share and a brand share have different denominators, and nobody does the division out loud.

Start with the table. An April 2026 preprint with an open dataset analysed 602 controlled prompts and 21,143 search-layer citations across ChatGPT, Google and Perplexity. Its source-type composition reads Official 34.22%, News 31.17% and Vertical 22.13% on ChatGPT; Official 46.35% on Google; Official 44.07% on Perplexity.1 Official is the largest single category on all three, and a minority of the whole on all three, because even the highest figure leaves 53.65% of citations going somewhere that is not official.

Both halves have to be stated together, because either alone points the wrong way. Quote the first and you conclude AI answers are mostly brand sites and your own writing decides the outcome. Quote the second and you conclude official sources barely feature, which the same table refutes. What survives: official sources are the strongest single entry route into an answer while still leaving most of it to other people.

Then divide. The same dataset puts mean citations per response at 6.88 on ChatGPT, 12.06 on Google and 16.35 on Perplexity, medians 6, 12 and 17.1 Apply the official share and a median Google answer carries five or six official-site citations, spread across whichever vendors and institutions the question required. On a shortlist naming six products your site is one of six at best, so the honest expectation for one brand is a low single-digit percentage of an answer.

Read the label as printed. The column is Official with no definition given, and the paper’s external-validity section states the source-type taxonomy contains unknown and noisy values.1 Treat the ordering as informative and the decimals as not; what sources get cited covers the rest of that taxonomy.

Reach

How much of the answer can on-site work reach?

Only the owned bucket, and only the part of it retrieved for the question asked. Two audits size the gap between ranking and retrieval. A 4,706-query audit of Google AI Overviews in Findings of ACL 2026, data collected September 2025 in the US and Germany, reports that “on average 53% (27%) of domains that AIO consults are not contained in top-10 (top-100) Organic search results”.2 A 55,393-query preprint collected 13 March to 21 April 2026 found 29.8% of AI Overview reference domains absent from the corresponding first page.3 One counts domains consulted, the other domains referenced, so quote both windows rather than a point from either.

That cuts both ways. A page ranking nowhere for the visible question can still be retrieved for one sub-question of it, which is how a documentation page enters an answer its homepage could never reach. It also means your ranking is a weaker lever over the answer than over the ten blue links beneath it, even though Google states its “generative AI features on Google Search are rooted in our core Search ranking and quality systems” (updated 10 Jul 2026).8 Retrievability sits in front of all of it, the subject of retrievability as a ceiling.

How much movement exists inside the owned bucket has been measured twice, and neither answer is large. C-SEO Bench, a NeurIPS 2025 benchmark of ten methods across six domains, found the best content method moved retail citation rank 0.36 places against 2.77 for moving a source to position one in the context, and that “out of 54 cases, we uncover only three where the ranking improvements are statistically significant”.4 A KDD 2026 arena study of ten strategies found editing body text alone scored 0.53, 0.84 and 0.47 on its three visibility measures against a 0.58, 1.00 and 0.50 baseline, 9%, 16% and 6% below changing nothing.5 The repair order is in the technical checklist.

Method

How do you measure your own owned share?

Count it from your own citation log, because every published average is an average over a query mix that is not yours. The procedure is a spreadsheet: run a fixed prompt set on a schedule, record every cited URL in every answer rather than only yours, and tag each one owned, influenced or neither. Then keep two numbers apart. Owned citation share is owned URLs over all cited URLs: how much of the answer you supply. Owned answer coverage is the share of answers carrying at least one owned URL: how often you are present at all. The second is always larger, and reporting it under the first name is how this gets inflated.

Logging everybody’s URLs is what makes the number worth having, because the same log doubles as the list of third-party pages currently deciding your category. Ten fields to log has the record shape. The field people skip is the full cited URL rather than the domain, and an outdated comparison page is not the same problem as a company homepage.

A team that ran forty prompts a week while logging only whether their own domain appeared spent a quarter reporting a flat line. Re-running it with every cited URL captured showed two moving parts underneath: owned share had roughly doubled off a tiny base, while one roundup cited in about a third of the answers still carried a price they no longer charged. The action was an email to a reviewer, which no work on their own site would have surfaced.

Precision

How many answers before the share means anything?

Owned citation share is a proportion, so its uncertainty depends on sample size and nothing else. This site quotes the 95% Wilson interval, whose half-width at the worst case of p = 0.5 is 1.96 divided by twice the square root of (n plus 3.84). Five answers gives roughly ±33 points, thirty gives ±17, two hundred gives ±7. An owned share moving from 4% to 6% across thirty answers has not moved, whatever the relative-change column says. Statistical power has the full arithmetic.

Two further sources of movement have nothing to do with you. Repeated runs of one prompt at temperature zero changed 9% to 28% of decisions in a study reported by the 2026 critical survey, so run-to-run variance survives with the sampling knob off.6 A tracking study of four engines over 45 days, in the same survey, found daily source-level Jaccard similarity of 0.34 to 0.42, meaning a large fraction of cited sources turns over overnight; it ran on a small Swiss query universe, so read the direction rather than the coefficient.6

What follows is a reporting cadence rather than a dashboard. Fix the prompt set and the run count, print the interval beside the point estimate, and compare quarters instead of weeks. A number this small and this noisy still earns its place, because its job is to size a decision.

Consequences

What does the split change about where the next hour goes?

It changes the ceiling on each kind of hour rather than the value of any of them. On-site work operates on the owned bucket, whose effect sizes are modest and whose retrieval precondition is close to binary. Off-site work operates on the influenced bucket, where most of the citations in a commercial answer live. The order that falls out is unglamorous: make the owned pages fetchable first, because a zero there cancels everything downstream; then find your actual owned share; then spend the remaining hours on the third-party pages your log says are deciding the answer.

Two things this measurement is not. It is not a revenue number: the 2026 critical survey rates the evidence that citation scores predict clicks, conversions or revenue Very low,6 and Pew’s panel of 900 US adults across 68,879 searches in March 2025 found a click inside an AI summary “occurred in just 1% of all visits”.7 And it is not a ranking. The only defensible outcome language here is mention share, citation share and share of voice, each indexed to a named engine, prompt set and date.

Raising the owned share is not automatically the goal. On a question whose honest answer compares five vendors, an answer built mostly from one vendor’s pages would be worse, and no engine is going to build it. The goal that survives the evidence is that the sources an engine does use describe you correctly, which is what the rest of the off-site stage is about.

The honest limit of this article

The 34.22% to 46.35% figures come from a single preprint whose 602 prompts were designed rather than sampled from real user behaviour, whose Official column is never defined, and whose own external-validity section says the source-type taxonomy contains unknown and noisy values. More importantly, no published study reports how owned citation share is distributed across individual brands. The arithmetic here divides a category average by a citation count instead of measuring per-brand shares, so it establishes an order of magnitude and not a value. What is solid is the pair of retrieval-gap audits, one of them peer-reviewed, and the instruction to measure your own share, which holds whatever these figures turn out to be.

Where a product fits, and where it does not

Everything above is a fixed prompt list and a spreadsheet, and it costs nothing. Bavior does not raise your owned share, does not write or edit pages on your domain, does not crawl your site, and cannot make an unfetchable page fetchable; the third-party pages in your log are not ours to change either. What it does is the collection step: it runs a fixed prompt set across five engines on a schedule and records every cited URL, yours and everybody else’s, which turns owned share into a number rather than an impression. Where a cited source is a live discussion thread it drafts a reply on an account you control, which you approve before anything posts. Results come back as mention share, citation share and share of voice per engine, never as a rank. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 30 Aug 2026).

Sources, all checked 30 Aug 2026
  1. Zhang, He, Yao, “From Citation Selection to Citation Absorption”, 28 Apr 2026, arXiv:2604.25707 (preprint, open dataset); source-type and citations-per-response tables, §10.2 “The source-type taxonomy contains unknown and noisy values”: arxiv.org/abs/2604.25707
  2. Kirsten et al., “Characterizing Web Search in The Age of Generative AI”, Findings of ACL 2026; 4,706 queries, September 2025, US and Germany; “on average 53% (27%) of domains that AIO consults are not contained in top-10 (top-100) Organic search results”: aclanthology.org/2026.findings-acl.526
  3. Xu, Iqbal & Montgomery, “Measuring Google AI Overviews”, 2026 (preprint); 55,393 queries, 13 March to 21 April 2026; “29.8% of AIO reference domains do not appear anywhere on the corresponding first page”: arxiv.org/abs/2605.14021
  4. Puerto et al., “C-SEO Bench: Does Conversational SEO Work?”, NeurIPS 2025 Datasets & Benchmarks Track, arXiv:2506.11097; Table 3 retail, §6.2: arxiv.org/abs/2506.11097
  5. Kim et al., “SAGEO Arena”, KDD 2026, arXiv:2602.12187; Table 2, “Body Text only”: arxiv.org/abs/2602.12187
  6. “A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (survey preprint); Table 5, the temperature-zero decision-flip range, the 45-day Jaccard tracking study: arxiv.org/abs/2607.14035
  7. Pew Research Center, 22 Jul 2025; 900 US adults, 68,879 searches, March 2025; a click inside an AI summary “occurred in just 1% of all visits”: pewresearch.org
  8. Google Search Central, “Google’s Guide to Optimizing for Generative AI Features on Google Search”, updated 10 Jul 2026 (first-party); “rooted in our core Search ranking and quality systems”: developers.google.com/search/docs/fundamentals/ai-optimization-guide
FAQ

Frequently asked questions.

What share of an AI answer comes from pages a brand controls?

A low single-digit percentage on most commercial questions, though no study measures it per brand. The largest source category in the one public academic taxonomy is labelled Official, at 34.22% to 46.35% of citations depending on the engine, but that covers every organisation the answer names. Divide it by the six to seventeen sources a typical answer cites and one brand's slice is small. Measure your own instead of assuming this one.

Does "official sources are the largest category" mean AI answers are mostly brand sites?

No, and stating only that half is misleading. Official is the largest single category on all three engines measured, and it still peaks at 46.35%, so between 54% and 66% of citations are news, vertical publications, discussion threads or something else. Both facts come from the same table. The useful reading is that official sources are the strongest single entry route into an answer while leaving most of it to other people.

How many logged answers do I need before my owned share means anything?

Around 200 if you want to see a seven-point change. Owned citation share is a proportion, and the 95% Wilson interval at the worst case is 1.96 divided by twice the square root of (n plus 3.84): five answers is about plus or minus 33 points, thirty is 17, two hundred is 7. Engine variance adds to that, so compare quarters rather than weeks.

If most of the answer is other people's pages, is on-site work still worth doing?

Yes, with a smaller expected return than most plans assume. On-site work is the only thing that moves your owned bucket, and being fetchable and relevant is a precondition for everything else. What it cannot do is reach the majority of citations in a commercial answer. Fix retrievability first because a zero there cancels the rest, then treat off-site work as the part addressing the other bucket.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Most of the answer is somebody else’s.
Find out how much of it is yours.

Start free trial