Home/Learn GEO/Retrievability ceiling
Technical GEO · Planning

What does a retrievability ceiling cost you in wasted work?

Every hour you spend on phrasing, structure and authority is multiplied by a number you have probably never measured. When it is zero, so is the return, and nothing in your analytics says so.

On this page
Share this
Share on X Share on LinkedIn
The short answer

It costs you the whole of it, because the stages of an AI answer multiply rather than add. The 2026 critical survey writes the identity out and states the consequence in one line: a high conditional probability of citation “does not compensate for a low probability of retrieval”.1 One published measurement found 27.1% of URLs already cited in AI answers could not be scraped at all, being inaccessible, removed or non-textual.1

Key takeaways
  • The purchasable multiplier is small: the one controlled benchmark measured the best content method lifting citation rank 0.36 places in retail against 2.77 for position one in context, with three of 54 cases significant.
  • The evidence shows retrieval failure is common in the population counted. It does not give your site’s rate; no study reports the distribution across sites.
  • Three checks in cost order name your regime before you commit a budget, and the first takes minutes and needs no tool.
  • Read zeros, not rates. No successful search-side fetch in 30 days is a finding; a count falling from 12 to 7 is noise.
  • Content work can lower the ceiling: in the one end-to-end test, optimizing body text alone degraded visibility at every stage, cutting final citation about 6%.
The arithmetic

Why does one retrieval failure multiply your budget by zero?

Because the identity behind a citation is a product of probabilities, not a sum of contributions. The 2026 critical survey sets it out as the probability that the engine ran a search at all, times the probability that your page was retrieved given that it did, times the probability that a retrieved page was cited.1 Almost every hour a team spends goes into that last term, where phrasing, passage structure, evidence and authority live. The first two scale that hour rather than sitting beside it, and a term of zero is not recoverable by improving anything to its right.

What makes the model sharp is the size of the factors. The retrieval term is close to binary per page: fetchable and indexed for a given question, or not. The content term has been measured. C-SEO Bench, a NeurIPS 2025 benchmark covering ten methods across six domains, more than 1.9k queries and 16k documents, reports the strongest content method in retail improving citation rank by 0.36 ±1.47 places, against 2.77 ±2.31 for the same document placed first in the model’s context.2 Its verdict on significance is blunter than any effect size: “Out of 54 cases, we uncover only three where the ranking improvements are statistically significant.”2

Put the two together and a planning rule falls out. Content work buys a small and noisy multiplier on a term you may not have: spent where the retrieval term is zero it returns zero, and spent where that term is unknown it returns an unknown. Which pages are which is the cheapest question in this field, and it is almost never asked first. How a page drops out at retrieval belongs to the retrieval stage article; this one is about what the resulting ceiling does to a plan.

The evidence

How big is the ceiling, and what has been measured?

One number has been published. It measures a different population from the one you care about, and no study reports how retrievability is distributed across sites.

The published quantity is 27.1%. Allaham and Diakopoulos, reported through the 2026 critical survey, found that share of URLs cited in AI answers “were not scraped because they were inaccessible, removed, or non-textual”.14 Read the denominator first: it counts URLs an engine had already judged worth citing, and whether a third-party scraper could fetch them rather than whether the engine could. The failure rate among pages that never got cited is therefore necessarily higher, and nobody has measured it.

So the figure establishes that retrieval failure is common rather than an edge case, at roughly a quarter in the one population counted. It does not establish your own rate. There is no published distribution of retrievability by site, so a sentence like “about a quarter of our pages are unreachable” has nothing behind it, and a plan built on 27.1% is built on a borrowed denominator.

The ceiling is not one number per site in any case, because retrieval runs per sub-question rather than per page. A 4,706-query audit of Google AI Overviews in Findings of ACL 2026, data collected September 2025, found 53% of the domains an AI Overview consults absent from the organic top 10 and 27% absent from the top 100.7 Keep the verb straight: that study measures domains consulted rather than cited. Retrievability is a property of a page-and-question pair, so it has to be measured on your own URLs against your own prompt list.

Triage

How do you tell you are in the ceiling regime before you spend?

Three checks, in cost order, on the twenty URLs that would answer a buying question. Each returns a yes or no, and the pattern across them names your regime.

01

Is the answer text in the initial HTML?

Fetch each URL with a plain command-line request and confirm the text is present with no JavaScript executed. Minutes of work, no tool needed.

02

Did a search-side crawler succeed?

Filter your access log by search-side token and look for one successful fetch of each URL in 30 days. Method: the crawler-log article; which tokens count is the next one.

03

Does the page ever appear as a source?

Run a fixed prompt list across the engines you care about and record which URLs each answer cites. Only this one costs money to run continuously.

Read the pattern, not the individual answers. A no at check one or two puts you in the ceiling regime: the retrieval term on those pages is zero or unknown, and content spend against them returns near zero however good the work is. Two yeses and a no at check three puts you above the ceiling and below the bar, the only regime where the C-SEO Bench multiplier is an honest expectation. Three yeses means you are measuring rather than repairing.

One statistical caution, because ignoring it turns the triage into noise: at the volumes a single site produces, read zeros and not rates. A page with no successful search-side fetch in 30 days is worth acting on; a page whose fetch count fell from 12 to 7 is not. A 95% Wilson interval on 30 observations has a half width of about 17 points at worst, so a rate from 30 tries cannot separate 40% from 60%. The statistical power article has the arithmetic.

Budget

What do you buy first, and what can wait?

The ordering comes from the cost shapes, not from taste.

Line itemCost shapeWhat the evidence supportsWhen to buy
Fetch and index-eligibility check, 20 URLsOne-off, hoursRequirement stated first-party by Google5First
Repairing dead URLs, gates and blocked tokensOne-off, mostlySame first-party requirement5First
New pages for sub-questions you missRecurringRelevance and context position rated high confidence1After the gate
Rewriting phrasing and passage structureRecurring0.36 rank places in retail; 3 of 54 significant2After the gate
An answer-side measurement panelRecurring, smallNothing else names your regime1Alongside both

Google is explicit about the gate, which makes the first line item easy to write into a plan. To be shown as a supporting link in AI Overviews or AI Mode, a page “must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements”, and the same page adds “There are no additional technical requirements” (updated 10 Dec 2025).5 A separate page states that the generative features “are rooted in our core Search ranking and quality systems” (updated 10 Jul 2026).6 Two pages, two dates, one budget consequence: there is no second index to buy your way into.

The asymmetry in the second column is what you defend to a finance team. Rows one and two are bounded and remove a known cause of zero. Rows three and four are unbounded, recurring, and their measured effect is a fraction of a rank place with a standard deviation several times its own size.2 Any expected-value calculation that takes the multiplication seriously puts the bounded work first.

Backfire

Can content work push your own ceiling down?

Yes, and there is one end-to-end measurement of it. SAGEO Arena, accepted at KDD 2026, puts retrieval and reranking back around the generation step instead of holding a candidate set fixed, then runs ten strategies over 171,003 documents and 2,700 queries. Its finding is flat: “optimizing body text alone consistently degrades visibility across all stages.”3 Against baseline, average top-20 presence at retrieval fell about 9%, top-10 presence after reranking about 16%, and final citation about 6%.13

The mechanism matters more than the percentages, because it names which rewrites are risky. The strategies that lost most at retrieval replaced common expressions with domain-specific or uncommon vocabulary, which the paper attributes to lexical mismatch with user queries.3 A rewrite that makes a page sound more expert can match fewer of the queries that used to retrieve it.

The survey states the same result in the notation above: if a transformation raises the conditional probability of citation but lowers the probability of retrieval, “the total effect may be negative”.1 That converts into a test-design rule you can hold an agency to. A content method evaluated inside a fixed-context benchmark measures the last term with the first two held constant, which is not the quantity you are buying. Re-run checks one and two on any page you rewrite. A rewrite followed by the page falling out of retrieval has not underperformed; it has lowered the ceiling over everything you do next.

Above the ceiling

What sits above the ceiling that no retrieval fix reaches?

Two further terms, and neither is yours to move. The first is whether the engine ran a search at all. In the configuration Schulte and colleagues studied, 57.8% of ChatGPT repetitions did not activate web search.1 That figure comes from a small Swiss query universe and should not be carried into your market as a rate, but the direction is the point: on a run with no search, a perfectly retrievable page is not retrieved, and nothing on your site changes that.

The second is the commercial term, the weakest link in the chain. The survey’s confidence table rates the claim that “Citation scores predict clicks, conversions, or revenue” as very low, resting on “one suggestive quasi-experiment and a few industry claims” with causality not established.1 Downstream of the citation, Pew Research Center’s panel of 900 US adults across 68,879 searches in March 2025 found clicking a source inside a Google AI summary “occurred in just 1% of all visits”.8

So budget a retrievability fix as what the evidence supports: removing a known cause of zero, at a bounded one-off cost, with no promise about what fills the space. The sentence that survives a board meeting is that you were structurally excluded from a surface and now you are not. Any pipeline number attached to it will not survive.

The honest limit of this article

The multiplication at the centre of this argument is a modelling choice made in a survey preprint, not a mechanism any engine has published. The 27.1% figure is a preprint result reported second-hand, it measures third-party scrapability of already-cited URLs, and it is not your rate. The C-SEO Bench effect sizes come from a benchmark with fixed candidate sets, and the SAGEO percentages from a single reproduction. What this article claims with confidence is the ordering, which survives large errors in any of those sizes.

Where a product fits, and where it does not

Two of the three checks above are free and already yours: a command-line fetch of your twenty highest-intent URLs, and a filter over your own access log. No tool can do the repairs they surface. Bavior touches only the third check, on the answer side: it runs a fixed prompt set across five engines on a schedule and records which sources each answer cited, which is what separates the ceiling regime from the one above it. It does not fetch your pages, read your logs or edit your robots.txt, and cannot make a blocked page retrievable, so a ceiling problem stays yours whatever you buy. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 30 Aug 2026).

Sources, all checked 30 Aug 2026
  1. “A Critical Survey of Generative Engine Optimization (2023–2026)”, dated 15 Jul 2026, arXiv:2607.14035; the citation identity, §7.4, the 27.1% and 57.8% figures, Table 5 (preprint): arxiv.org/abs/2607.14035
  2. Puerto, Gubri, Green, Oh, Yun, “C-SEO Bench: Does Conversational SEO Work?”, NeurIPS 2025 Datasets & Benchmarks Track, arXiv:2506.11097; Table 3 retail and §6.2: arxiv.org/abs/2506.11097
  3. Kim et al., “SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization”, accepted at KDD 2026, arXiv:2602.12187; Table 2 and its body-text analysis: arxiv.org/abs/2602.12187
  4. Allaham & Diakopoulos, 2026; 27.1% of cited URLs not scraped, being inaccessible, removed or non-textual (preprint, reported in note 1)
  5. Google Search Central, “AI Features and Your Website”, updated 10 Dec 2025; snippet eligibility, “There are no additional technical requirements” (first-party): developers.google.com/search/docs/appearance/ai-features
  6. Google Search Central, “Google’s Guide to Optimizing for Generative AI Features on Google Search”, updated 10 Jul 2026; the “rooted in our core Search ranking and quality systems” sentence (first-party): developers.google.com/search/docs/fundamentals/ai-optimization-guide
  7. Kirsten et al., Findings of ACL 2026; 4,706 queries, data September 2025; “on average 53% (27%) of domains that AIO consults are not contained in top-10 (top-100) Organic search results”: aclanthology.org/2026.findings-acl.526
  8. Pew Research Center, 22 Jul 2025; 900 US adults, 68,879 searches, March 2025; clicking a link inside an AI summary “occurred in just 1% of all visits to pages with such a summary”: pewresearch.org
FAQ

Frequently asked questions.

If retrievability is the ceiling, should I stop writing content?

No, you should sequence it. The multiplication says content work is scaled by retrievability rather than replaced by it, so the same work returns almost nothing on unreachable pages and returns its measured effect on reachable ones. Run the three checks first, fix whatever they surface, then spend the content budget on the pages that passed. The order changes the return, not the activity.

How do I put a number on my own retrievability ceiling?

You measure it, because no study reports it for you. Take the twenty URLs that would answer a buying question, check each for answer text in the initial HTML and for one successful search-side crawler fetch in the last 30 days, and count the failures. That count is your ceiling on those pages. The published 27.1% describes a different population and is not a site-level prior.

Does fixing retrievability increase revenue?

Nothing published establishes that. The 2026 critical survey rates the claim that citation scores predict clicks, conversions or revenue as very low confidence, resting on one suggestive quasi-experiment with causality not established. Pew's March 2025 browsing panel found a click on a source inside a Google AI summary in just 1% of visits. Budget the fix as removing a known cause of zero, not as buying a return.

Can a rewrite make a page less retrievable than it was?

Yes, and it has been measured once end to end. SAGEO Arena, accepted at KDD 2026, found that optimizing body text alone consistently degraded visibility at every stage, with final citation down about 6%. The strategies that lost most at retrieval swapped common expressions for domain-specific or uncommon vocabulary, creating a lexical mismatch with the queries people actually type.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Retrievability caps every dollar spent after it.
Find out where yours sits.

Start free trial