Home/Learn GEO/The pipeline
How AI search works · Deep dive

What are the five stages of the AI search pipeline?

A tactic can win the stage it was designed for and still lose the answer, because the stages compose. This is the whole pipeline end to end, and a table that sorts every common tactic into the stage it acts on.

On this page
Share this
Share on X Share on LinkedIn
The short answer

The AI search pipeline has five stages: interpretation, fan-out, retrieval, selection and reranking, and synthesis with attribution. They compose rather than add, so a tactic that wins the stage it was designed for can still lose the answer, and you own an input to only the last three. In the one published benchmark that ran all five stages over 171,003 documents and 2,700 queries, the average body-only rewrite cut top-20 retrieval presence from 0.58 to 0.53 and the final citation rate from 0.50 to 0.47.3

Key takeaways
  • Stages one and two read nothing you publish, so any service sold as improving how the engine reads a question is mislabelled work on a later stage.
  • Stage three is a gate, not a ranking factor. Google requires that a page be “indexed and eligible to be shown in Google Search with a snippet” before it can be a supporting link at all.1
  • Almost everything with strong evidence behind it acts at stage three; almost everything sold as GEO acts at stage five, which the 2026 critical survey rates low confidence.2
End to end

What are the five stages of the AI search pipeline?

The five stages are interpretation, fan-out, retrieval, selection and reranking, and synthesis with attribution, and they run in that order. Read each as what it consumes and what it emits: the output of one stage is the entire universe the next stage gets to work with.

1

Interpretation

Consumes your message plus the conversation so far. Emits an intent, a set of resolved entities, constraints such as locale and recency, and a decision about whether to search at all.

2

Fan-out

Consumes that intent. Emits several search queries across subtopics and data sources. One buying question silently becomes searches about cost, setup, alternatives and risk.

3

Retrieval

Consumes each generated query. Emits a candidate set per query from an index or a live fetch. Pages that are unindexed, blocked, dead or unfetchable never enter this set.

4

Selection and reranking

Consumes the merged candidate sets. Emits an ordered shortlist of passages that fits a context budget. The unit is a passage, not a page, and most candidates are dropped here.

5

Synthesis and attribution

Consumes the shortlist. Emits the answer text plus links attached to some of its claims. Wording and quotability act here, on material that already survived stage four.

The composition

Your visibility is the product of all five, not the sum. When one 2026 benchmark restored the retrieval and reranking stages, the average body-only rewrite lost about 9% of its top-20 presence upstream and gained nothing downstream.3

Leverage

Which stages do you own any input to?

You own an input to stages three, four and five, and nothing at all at stages one and two. Interpretation runs before any document is touched, on the user’s words, the conversation history and the engine’s own priors alone. Fan-out is generated from that intent by the engine’s own model of the subtopics. Neither stage reads your site, so any tactic sold as improving them is either mislabelled or empty.

At stage three your inputs are index membership, fetchability and relevance. Google states the eligibility rule for its own surfaces directly: a page “must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements”, and adds that no special files or schema.org markup are needed to appear.1 Its optimization guide names the mechanism behind that rule, describing retrieval-augmented generation as improving answers “by relying on our core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index”.6 At stage four your input is passage structure, meaning whether the answer to a specific sub-question sits in one liftable chunk. At stage five your input is wording.

There is also a large part of the pipeline you own no input to, and it is not a stage: it is everybody else’s pages. Stages three and four run over an index built from the whole web, so on most commercial questions most retrieved candidates are third-party reviews, roundups and discussion threads. One 2026 vendor study of four brands, covering 47,097 citations across three engines, put the share of cited material a brand owns at roughly 2%; four brands is illustrative rather than a constant, but the order of magnitude is the point.9

The sorting table

Which stage does each class of GEO tactic act on?

Every common tactic acts on exactly one stage, and which stage decides how much weight the claim deserves before you evaluate the tactic itself. The evidence ratings below are Google’s own eligibility documentation and the 2026 critical survey’s labels, not ours.12

StageTactic classes that act thereEvidenceHow to treat a claim aimed here
1 · InterpretationNothing you publishNot applicableReject. A claim to change how the engine reads a question is not advice.
2 · Fan-outCovering the sub-questions of a topicMechanism first-party documented; effect sizes vendor-onlyDo it, and do not quantify it.
3 · RetrievalBeing indexed, fetchable and server-rendered; ranking for the query; keyword stuffingStrongest: a first-party eligibility requirement; stuffing null or negativeDo first. This is the ceiling on everything after it.
4 · Selection and rerankingAnswer-first self-contained passages; real, dated, attributed numbers; headings, lists and tablesStrong once retrieved; structure moderate and heterogeneousDo second. Test structure rather than assuming its direction.
5 · SynthesisAuthoritative tone, citation-shaped phrasing, a fixed recipe copied across domainsWeak and unstable; poor generalizationDo last, and only if it costs nothing upstream.
No stageExtra AI text files and special markupContradicted by the engine’s own documentation6Reject. There is nothing here to act on.

Read the table top to bottom and the pattern is clear: almost everything with strong evidence behind it sits at stage three, and almost everything sold as “GEO” sits at stage five. That is not a coincidence, because stage five is the only stage a copywriter can act on without touching infrastructure. The survey rates “a document already placed in the context can causally alter its rank, citation or use” as high confidence while noting that this evidence “does not address organic retrieval”, and rates “a white-hat GEO intervention durably improves organic discoverability across multiple engines” as low.2 Those two rows are the shape of the field’s evidence base, and why the how AI search works stage sorts tactics by stage first. The last row has a first-party verdict behind it: Google’s guide lists “LLMS.txt files and other ‘special’ markup” among the things you can ignore, because “Google Search itself doesn’t use them”.6 The last column is the practical use: when someone recommends a tactic and cannot say which stage it acts on, that inability is itself the answer.

Composition

Why can a stage-five tactic lose at stage three?

A stage-five tactic loses at stage three whenever a rewrite that makes a passage more quotable simultaneously makes the page less relevant to the query that would have retrieved it, because the pipeline multiplies the two effects rather than adding them. A KDD 2026 benchmark built an end-to-end environment of 171,003 documents and 2,700 queries with retrieval and reranking reinstated, and tested ten body-only strategies. Averaged across them, top-20 presence fell from 0.58 to 0.53, post-rerank top-10 presence from 1.00 to 0.84, and the citation rate from 0.50 to 0.47; the worst single strategy took top-20 presence to 0.37 and the citation rate to 0.39.3 The survey’s reading is exact: a rewrite “may therefore perform well once injected while making the document less retrievable or less competitive upstream.”2

Two other benchmarks land in the same place from different directions. A NeurIPS 2025 datasets-and-benchmarks paper ran “a total of ten methods” over a corpus it describes as “more than 1.9k queries and 16k documents”, and concluded that these methods are “not only largely ineffective but can actually have the opposite expected effect by decreasing document ranking”: “Out of 54 cases, we uncover only three where the ranking improvements are statistically significant”, and none at all for question answering.4 On position it stays qualitative and prints no multiple: making a document the first one in the context window “leads to far greater citation ranking gains in the LLM response” than any of the content methods it tested.4 The 2024 KDD paper that founded the field reported much larger gains, but on a custom metric inside a purpose-built two-step engine, for a page already sitting in the five-document context; the survey calls the popular generalisation of that number “a relative maximum on one metric under a specific configuration”.52 Both results are honest; only one of them ever ran a retrieval step. Where GEO advice goes wrong takes both benchmarks apart in full.

Sequencing

In what order should you do the work?

Work the stages in pipeline order, because each one gates the next: retrievability, then coverage, then passage structure, then wording. Retrievability first because it is a gate rather than a ranking factor. A 2026 preprint that audited 26,266 URLs cited by generative search engines could not fetch 27.1% of them at all: they pointed at PDFs, images or other non-text formats, or the content was inaccessible or removed.7 That is the failure rate among pages good enough to have been cited already. Coverage second, because fan-out means a page answering four sub-questions is eligible at four retrieval points and a page answering one is eligible at one.

Passage structure third. Selection operates on chunks: a vendor dataset of 15,699,298 Google AI Mode citations published in July 2026 resolved 47.7% of them to scroll-to-text highlights of specific passages rather than plain page links, median highlighted passage 117 words, about 85% of them self-contained.8 Wording last, and lightly, because it is the stage with the weakest evidence and the only one where a change can cost you upstream.

Diagnosis

How do you tell which stage is failing?

Measure the stages separately, because a single blended visibility number cannot tell you which one failed, and the fix for each is different work by a different person. The survey states the sequence plainly: “produce a relevant, comprehensive, verifiable, clearly structured, and technically retrievable page; then measure retrieval, citation, and fidelity separately.”2

The benchmark that restored the full pipeline reports three numbers per strategy instead of one: presence in the top 20 after retrieval, presence in the top 10 after reranking, and the citation rate in the finished answer. The three move independently. Adding structural information to the same documents raised average top-20 presence from 0.58 to 0.71 while the citation rate barely moved, going from 0.50 to 0.52.3 A single blended score would have shown a small positive and hidden both halves of that result.

The founder version needs no instrumentation and takes an afternoon. Stage three: fetch the page with a plain request, confirm the answer text is in the HTML, then confirm it is indexed and snippet-eligible.1 Stages two and three: search the sub-questions your topic fans out into and see whether anything of yours ranks. Stage four: read the passages engines quoted from whoever they cited instead. Stage five: compare that wording with yours. Four checks, four different failures, and only the last is a writing problem.

The honest limit of this article

Five stages is an analytic split, not an architecture diagram. No engine publishes its pipeline, and some products clearly do not run the stages once each: deep-research and agentic modes visibly loop fan-out, retrieval and selection several times before writing anything. The boundaries are drawn where the published evidence changes character, not where a system diagram would put them. Treat the model as a way to sort claims, not as a specification.

Where a product fits, and where it does not

Bavior only helps with the part of that diagnosis that repeats. It runs a fixed prompt set across five engines on a schedule and records which sources each answer cited, so stage-three and stage-four outcomes read as a distribution instead of a screenshot, and where a cited source is a live discussion thread it drafts a reply on an account you control, which you approve before anything posts. It does not crawl your site, it does not fix retrievability for you, and it cannot promise a citation: the four manual checks above stay yours either way. The free AI visibility check and the free GEO audit need no paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).

Sources, all checked 30 Aug 2026
  1. Google Search Central, “AI features and your website” (fan-out, eligibility, no special markup; first-party; updated 10 Dec 2025): developers.google.com/search/docs/appearance/ai-features
  2. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (survey preprint): arxiv.org/abs/2607.14035
  3. Kim et al., “SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization”, KDD 2026, arXiv:2602.12187v2 (revised 7 Aug 2026); 171,003 documents, 2,700 queries; stage-level figures from Table 2: arxiv.org/abs/2602.12187
  4. Puerto et al., “C-SEO Bench: Does Conversational SEO Work?”, NeurIPS 2025 Datasets & Benchmarks Track, arXiv:2506.11097; ten methods over “more than 1.9k queries and 16k documents”: arxiv.org/abs/2506.11097
  5. Aggarwal et al., “GEO: Generative Engine Optimization”, KDD 2024, arXiv:2311.09735: arxiv.org/abs/2311.09735
  6. Google Search Central, AI optimization guide (RAG on the Search index; LLMS.txt unused; first-party; updated 10 Jul 2026): developers.google.com/search/docs/fundamentals/ai-optimization-guide
  7. Allaham & Diakopoulos, “Synthetic Sources?: Auditing Generative Search Engine Citations”, 22 May 2026, arXiv:2605.23684 (preprint); 26,266 cited URLs, 72.9% scraped: arxiv.org/abs/2605.23684
  8. Passage-extraction dataset, Jul 2026: 15,699,298 Google AI Mode citations; 47.7% scroll-to-text highlights, median passage 117 words, ~85% self-contained. Vendor-published; described, not linked.
  9. Content-recency study, Jul 2026: 7,683 pages, 47,097 citations, three engines, four brands; brands owned roughly 2% of the pages cited about them. Vendor-published; described, not linked.
FAQ

Frequently asked questions.

Which stage of the AI search pipeline should I fix first?

Fix retrieval first, because it is a gate rather than a ranking factor: a page that cannot be indexed or fetched scores zero at every later stage regardless of how well it is written. Confirm the page is indexed, returns its text to a plain HTTP fetch with no JavaScript, is not blocked for the engine's search crawler, and is genuinely relevant to a query someone would issue. Only then work on passage structure, and only then on wording, which is the stage with the weakest evidence and the only one where a change can cost you upstream.

Is the five-stage model how the engines are actually built?

No engine publishes its pipeline, so the five-stage model is inferred rather than documented. It is assembled from first-party statements about individual steps, since Google describes query fan-out and states its indexing requirement in its own documentation, from academic benchmarks that rebuild retrieval and reranking in the open, and from black-box audits. It predicts observed behaviour well enough to sort claims by mechanism. It is not a specification, and deep-research modes visibly loop the middle stages several times rather than running them once each.

Can a page rewrite make my AI visibility worse?

Yes, and it has been measured. A KDD 2026 benchmark that reinstated retrieval and reranking over 171,003 documents and 2,700 queries tested ten body-only rewriting strategies; averaged across them, top-20 presence fell from 0.58 to 0.53, post-rerank top-10 presence from 1.00 to 0.84, and the citation rate from 0.50 to 0.47, and the worst single strategy took top-20 presence to 0.37. The mechanism is that a rewrite tuned for quotability can dilute the topical relevance that got the page retrieved in the first place. A separate NeurIPS 2025 benchmark reported that such methods are "not only largely ineffective but can actually have the opposite expected effect by decreasing document ranking".

Which stages of AI search can I not influence at all?

Stages one and two, interpretation and fan-out, take no input from your site. Interpretation runs before any document is touched, using only the user's words, the conversation history and the engine's own priors. Fan-out is generated from that intent by the engine's own model of the topic's subtopics. Neither reads your pages, so nothing you publish can move them, and any service sold as improving them is either mislabelled work on a later stage or empty. The useful response to that constraint is to choose your measurement prompts deliberately, since the prompt is the one stage-one input you control.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Five stages, one failing.
Find out which one is yours.

Start free trial