Home/Learn GEO/Stage 4
How AI search works · Stage 4

How do selection and reranking decide which passages an AI answer uses?

This is the stage that quietly kills most content work, because the thing being judged is not the page you published. It is a chunk of roughly a hundred words that a reranker can lift out of it.

On this page
Share this
Share on X Share on LinkedIn
The short answer

Selection and reranking is the stage where an AI engine merges the candidate sets returned for every fan-out sub-query, removes duplicates, reorders what is left by how well each passage answers the specific question, and cuts the list until it fits the model’s context budget. The unit being scored is a passage, not a page, which is why a great page with its answer in paragraph nine competes as a mediocre passage. In a vendor dataset of 15,699,298 Google AI Mode citations published July 2026, 47.7% resolved to a scroll-to-text highlight of one specific passage rather than a plain page link, with a median highlighted passage of 117 words.1

Key takeaways
  • Four operations run in order: merge the candidate sets, collapse near-duplicates, rerank by how well each passage answers the sub-question, then cut to the context budget. Only what survives the cut reaches the writing stage.
  • Domain authority is not what is scored here, so a clear answer on a weaker site can beat a buried answer on a stronger one. Authority did its work earlier, at retrieval.
  • Position inside the model’s context is the strongest documented lever, and you do not set it directly: in a NeurIPS 2025 benchmark on GPT-4o-mini, improving a source’s context position moved it 2.77 rank places in retail, against 0.36 for the best content rewrite.5
  • Every documented control is subtractive. Google names nosnippet, data-nosnippet, max-snippet and noindex; no engine lets you nominate a preferred passage.6
Definition

What happens at selection and reranking?

Selection and reranking turns several candidate lists into one short, ordered list of evidence that fits inside a fixed context budget. Four operations happen in sequence. The candidate sets returned for each fan-out sub-query are merged; near-duplicates and syndicated copies are collapsed so the same claim does not occupy three slots; what remains is reordered by a model scoring how well each passage answers the specific question rather than how authoritative its domain is; and the tail is cut until the evidence fits. Google describes the same shape in first-party language: grounding retrieves pages “by relying on our core Search ranking systems”, after which “our systems then review the specific information from those retrieved pages to generate a more reliable and helpful response”.9 That review is stage four.

That a model reads the passages, rather than a formula scoring a link graph, is not a guess about closed systems. It is the published technique: an EMNLP 2023 study found properly instructed language models “can deliver competitive, even superior results to state-of-the-art supervised methods on popular IR benchmarks” as re-ranking agents.10 A reranker is a reader, and it reads chunks.

The cut is the part people underestimate. Everything upstream produced candidates in the hundreds or thousands; what reaches the writing stage is a handful. In a 602-prompt academic sample published April 2026 the mean number of citations on a finished answer was 6.88 for ChatGPT, 12.06 for Google’s AI surfaces and 16.35 for Perplexity.3 Almost all of the pool was discarded here, and nothing discarded can be recovered by better prose further down. The how AI search works stage places this stage inside the full pipeline.

The unit of competition

Why is the unit a passage and not a page?

The unit is a passage because 47.7% of Google AI Mode citations resolve to one specific highlighted stretch of text rather than to the page as a whole, and the median highlighted stretch is 117 words.

From a year-long vendor dataset of 15,699,298 Google AI Mode citations across 148 industries, published July 2026.1

47.7%of citations were passage highlightsnot plain page links
117words, median highlighted passageroughly one tight paragraph
85%of passages were self-containedreadable with no surrounding page
48%of repeatedly cited passages opened with a questionagainst 22% of one-off passages

An engine quoting a scroll-to-text highlight is telling you, in the URL itself, exactly which words it selected, and that is the closest thing to a public window into stage four that exists. The same dataset found 80% of extracted passages put the answer in the first sentence, that 80.9% of passages were cited only once while about 2,300 were reused 61 times or more, and that the single most recycled passage was cited 661 times.1 A page is not what competes. A paragraph is, and a good one can be reused hundreds of times.

Two caveats travel with those numbers. They are descriptive, not causal. And there is a base-rate problem the study does not solve: it reports what share of extracted passages are answer-first, without reporting what share of all web paragraphs are, so “80% of extracted passages are answer-first” is not proof that being answer-first causes extraction. Treat it as corroboration of a plausible mechanism rather than a measured lift.

Dilution

Why does a buried answer lose?

A buried answer loses because the reranker scores the chunk, and the eight paragraphs surrounding your answer are not evidence for it, they are noise around it. A 3,000-word page whose answer sits in paragraph nine presents one relevant passage embedded in a lot of material about something adjacent. A shorter page whose second sentence answers the question presents a passage that is almost entirely on-target. Page-level authority does not rescue the first case, because authority is not what is scored here.

Position within the page appears to matter as well, though the evidence is vendor-published and self-declared correlational. A June 2026 study of 10,000 AI Overview responses and 100,000 citation placements, matched to source locations via scroll-to-text fragments, found 38% of citations were pulled from the first 100 words of a page, up from 20% in a smaller comparison set a year earlier; its authors state plainly that the data is correlational and that they “can show where AI models cite from, not definitively state why”.2 The practical reading is not “put everything at the top”. It is that preamble, brand throat-clearing and “in this article we will cover” occupy the most valuable real estate on the page.

Position in the model’s context, once a passage is selected, matters too, and the support is stronger than for anything about page shape. The 2026 critical survey grades position in context Strong, conditional on the document already having been retrieved, and query-document relevance Strong in controlled settings.4 The NeurIPS 2025 C-SEO Bench puts numbers on it: on GPT-4o-mini in retail, the best traditional intervention, improving the source’s position in the LLM context, moved the target 2.77 rank places on average, against 0.36 for the strongest of the ten content-rewriting methods it tested. Its own wording carries no multiple: making the target document “the first one in the LLM context window leads to far greater citation ranking gains in the LLM response than any C-SEO method”.5 Why has its own literature: performance “is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts”.8 You do not set that position. You influence it by being the passage that most exactly answers the sub-question.

Signals

What decides which passage wins?

Relevance to the query decides more than any formatting feature measured: model-judged relevance correlates with citation influence at r = 0.4322 across platforms, ahead of every structural signal in the same dataset.

Descriptive associations from 602 controlled prompts, 21,143 search-layer citations and 18,151 fetched pages, published April 2026. Top quartile against bottom quartile on mean citation influence.3

Page featureAssociation with citation influenceHow to read it
Model-judged relevance to the queryStrongest cross-platform signal, r = 0.4322Answer the actual question, not the topic
Answer-to-citation embedding similarityr = 0.3561Say it in the words the answer would use
Contains numbers or statistics0.1171 vs 0.0725, or +61.6%Real, dated, attributed numbers only
Contains definition markers0.1252 vs 0.0795, or +57.3%Write “X is …” as a flat sentence
Heading countTop quartile 10.59 vs 0.85 headingsConfounded; good pages have more headings
List densityTop quartile 8.94× the bottomSame confound; do it for readers
Q&A page format0.0947 vs 0.1005, or −5.7%Slightly negative; possibly thin FAQ pages

The authors’ own caveat governs how far this can be pushed: they cannot isolate whether headings cause absorption, since well-made pages both get cited and happen to carry more headings, and the negative Q&A result may be confounded with thin FAQ pages.3 The signals also differ by engine in the same dataset: Google’s strongest were embedding similarity plus definition markers, Perplexity’s were relevance plus heading count and length. The critical survey’s instruction on this whole family of findings is the right one to carry, and it grades document structure only moderate and heterogeneous: “Structure should be evaluated stage by stage, not treated as a universal talisman.”4

Direct control

Can you control which passage an engine uses?

Only negatively, and only on the surfaces that document a control. Google’s guidance names four: “to limit the information shown from your pages in Search, use nosnippet, data-nosnippet, max-snippet, or noindex controls”, and its AI features run off the same snippet eligibility.6 That gives you a way to exclude a specific block of text from being quoted, or to cap how much of a page can be shown, both of which reduce your exposure rather than steer it.

There is no positive equivalent. No engine offers a way to mark a paragraph as the one you would like quoted, and no markup, file or directive nominates one. The only positive influence is compositional: make the passage that answers the sub-question the clearest and most exactly on-target thing on the page, and give it a heading that states the question. That is not a trick, which is why it keeps working.

Length

Does a longer page win at this stage?

Length is close to irrelevant to citation position, and the two large datasets that touch the question disagree because they measure different outcomes. A December 2025 vendor study of 560,346 AI Overviews, 1,677,876 cited URLs and 174,048 pages with extractable content measured the Spearman correlation between word count and citation position at 0.04; 53.4% of citations went to pages under 1,000 words.7 The April 2026 academic dataset, measuring mean influence rather than position, found influence rising with length past 3,000 words.3

Different outcome variables, different samples, no reconciliation available. The defensible position is that length is neutral as a lever, and that what is worth optimising is how many distinct sub-questions the page genuinely answers and how cleanly each answer sits in its own passage. Adding words to hit a target is the one version of this that is definitely wrong.

The honest limit of this article

No engine publishes how it chunks a page, so the widely repeated claim that AI systems split content into 200 to 400 token passages is an assertion about the internals of closed retrieval stacks, not a fact: useful as a mental model, worthless as a specification, and anyone giving you a token count for your own page is guessing. The passage statistics here come from one descriptive vendor dataset with no base rate, and the feature associations come from a preprint whose authors reserve confirmatory analysis for future work and state they cannot separate cause from correlation. Stage four is the best-instrumented stage after retrieval, and it is still measured entirely from the outside.

Where a product fits, and where it does not

You can read stage four directly, for free, with no software: run your buyer questions on Google AI Mode, click the citations, and watch where the scroll-to-text highlight lands on each cited page. That habit teaches more about passage selection than any guide, including this one. Bavior does the repetitive half: it runs a fixed prompt set across five engines on a schedule and records which sources each answer cited, so you can see whether a rewrite changed anything over dozens of runs instead of one, and where a cited source is a live discussion thread it drafts a reply on an account you control, which you approve before anything posts. It does not rewrite your pages, does not score your passages, and cannot make a reranker prefer you. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).

Sources, all checked 29 Aug 2026
  1. Passage-extraction dataset, published Jul 2026: 15,699,298 Google AI Mode citations across 2.7 million pages; 47.7% scroll-to-text highlights, median 117 words, ~85% self-contained, 80% answer-first, 48% of repeatedly cited passages question-led against 22% of one-off ones. Vendor-published, described not linked.
  2. Citation-placement study, published Jun 2026: 10,000 AI Overview responses, 100,000 placements matched via scroll-to-text fragments; 38% of citations from the first 100 words, against 20% in a 2025 comparison set; authors state the data is correlational. Vendor-described.
  3. Zhang, He, Yao, “From Citation Selection to Citation Absorption”, 28 Apr 2026, arXiv:2604.25707 (preprint, open dataset). arxiv.org/abs/2604.25707
  4. Martinez, “Optimizing Visibility in Generative Engines: A Critical Survey of GEO (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (survey preprint; Table 4 grades the levers). arxiv.org/abs/2607.14035
  5. Puerto, Gubri, Green, Oh, Yun, “C-SEO Bench: Does Conversational SEO Work?”, NeurIPS 2025 Datasets & Benchmarks Track, arXiv:2506.11097. arxiv.org/abs/2506.11097
  6. Google Search Central, “AI features and your website” (snippet controls; first-party, updated 10 Dec 2025). developers.google.com/search/docs/appearance/ai-features
  7. Content-length study, published Dec 2025: 560,346 AI Overviews, 1,677,876 cited URLs, 174,048 pages with extractable content; word count against citation position, Spearman 0.04. Vendor-described.
  8. Liu et al., “Lost in the Middle: How Language Models Use Long Contexts”, TACL 2024, arXiv:2307.03172. arxiv.org/abs/2307.03172
  9. Google Search Central, “Optimizing your website for generative AI features on Google Search” (first-party, updated 10 Jul 2026). developers.google.com/search/docs/fundamentals/ai-optimization-guide
  10. Sun et al., “Is ChatGPT Good at Search? Investigating LLMs as Re-Ranking Agents”, EMNLP 2023, arXiv:2304.09542. arxiv.org/abs/2304.09542
FAQ

Frequently asked questions.

How long should a passage be to get quoted by AI search?

The median highlighted passage in a July 2026 vendor dataset of 15.7 million Google AI Mode citations was 117 words, roughly one tight paragraph, and about 85% of highlighted passages were understandable without the surrounding page. That is a description of what gets extracted, not a target to hit, and no engine publishes a chunk size. The useful version of the finding is structural rather than numeric: write each answer so it is complete inside one paragraph, because that is the shape of the thing an engine lifts.

Can I tell an AI engine which paragraph to quote?

Not positively: no engine offers a way to nominate a preferred passage, and no markup, file or directive does it. The controls that exist are all subtractive: Google's documentation names nosnippet, data-nosnippet, max-snippet and noindex as the ways "to limit the information shown from your pages in Search", and its AI features run off the same snippet eligibility. So you can exclude a block from being quoted or cap how much of the page can be shown, but the only positive influence available is making the passage that answers the question the clearest and most self-contained thing on the page.

Does page authority decide which passage gets selected?

Not at this stage: reranking scores how well a passage answers the specific sub-question, which is why a clear answer on a weaker page can beat a buried answer on a stronger one. Authority matters earlier, at retrieval, where ranking governs whether the page becomes a candidate at all. The distinction has a practical consequence: if a competitor with a weaker site keeps getting quoted, the problem is usually not their domain strength but the shape of their paragraph, and you can check that in a minute by clicking the citation and seeing which words were highlighted.

Do headings and lists make a page more likely to be cited?

The association is real and the causation is unproven. In a 602-prompt academic dataset published April 2026, pages in the top quartile of citation influence carried a median 10.59 headings against 0.85 in the bottom quartile, and 8.94 times the list density, but the authors state directly that they cannot isolate whether headings cause absorption, since well-made pages both get cited and happen to carry more headings. The 2026 critical survey rates document structure as "moderate and heterogeneous" and advises testing without assuming the direction of effect. Do it because it makes the page usable; promise nothing for it.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

The page is not what competes.
The paragraph is.

Start free trial