Ranking remains the strongest single predictor of being cited, and it has stopped being sufficient. A 4,706-query audit of Google AI Overviews in Findings of ACL 2026, on September 2025 data, found 53% of the domains it consults absent from the organic top 10 and 27% from the top 100.11 A 55,393-query preprint collected between 13 March and 21 April 2026, roughly twelve times the sample and six months later, found 29.8% of AI Overview reference domains absent from the corresponding first page.13 Quote that as a range, roughly 30% to 53% with both collection windows named, rather than as a point from either study; and keep the verb straight, because Kirsten measures the domains the system consults, which is not identical to the domains it cites.
An 11,500-query SIGIR 2026 study measured URL-level Jaccard similarity of 0.11–0.18 between Google organic, AI Overviews and Gemini.12 Fan-out is the mechanism: retrieval runs per sub-query, so pages that rank nowhere for the visible question can rank precisely for one of its parts.
Be sceptical of any single overlap percentage, including the ones above, because the number depends entirely on a denominator that is rarely stated. “What share of AI answers contain at least one top-ten URL” and “what share of individual citations come from the top ten” are different questions, and published figures across 2025 and 2026 range from roughly a sixth to over three quarters depending on which was asked, when, and on what sample. A number without its denominator and its date is not a measurement. The defensible synthesis is directional: top-ten ranking earns the high-frequency citations, and a large and growing share of the long tail comes from pages that do not rank on page one at all.
The honest limit of this article
The two most-quoted numbers here are the weakest links. The 27.1% unscrapable figure is from a preprint and measures URLs already cited, not the general web. The JavaScript conclusion rests on a single log study from December 2024 whose execution-detection method was published separately and is one-sided: a missing render-time beacon cannot separate a crawler that does not render from one whose beacon was blocked. It is unreplicated twenty months on and no vendor documentation either way, and the products have changed a lot since. Neither claim is settled. The parts of this article that are solid are the first-party robots.txt token splits and Google’s own eligibility statement, because those come from the companies that operate the systems.
Where a product fits, and where it does not
The entire retrieval audit is free and takes about an hour: fetch each important URL with a plain command-line request and confirm the answer text is in the HTML with no JavaScript; check your robots.txt against all four token splits above; confirm the pages are indexed and snippet-eligible; and crawl your own site for 404s and redirect chains, which AI crawlers follow far worse than Googlebot does. No tool is needed for any of that, and no tool can do the fixing. Bavior works on the stage after this one: it runs a fixed prompt set across five engines on a schedule and records which sources each answer cited, which tells you whether a retrieval fix changed who shows up, and where a cited source is a live discussion thread it drafts a reply on an account you control, which you approve before anything posts. It does not crawl your site, does not edit your robots.txt, and cannot make a blocked page retrievable. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).