You own an input to stages three, four and five, and nothing at all at stages one and two. Interpretation runs before any document is touched, on the user’s words, the conversation history and the engine’s own priors alone. Fan-out is generated from that intent by the engine’s own model of the subtopics. Neither stage reads your site, so any tactic sold as improving them is either mislabelled or empty.
At stage three your inputs are index membership, fetchability and relevance. Google states the eligibility rule for its own surfaces directly: a page “must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements”, and adds that no special files or schema.org markup are needed to appear.1 Its optimization guide names the mechanism behind that rule, describing retrieval-augmented generation as improving answers “by relying on our core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index”.6 At stage four your input is passage structure, meaning whether the answer to a specific sub-question sits in one liftable chunk. At stage five your input is wording.
There is also a large part of the pipeline you own no input to, and it is not a stage: it is everybody else’s pages. Stages three and four run over an index built from the whole web, so on most commercial questions most retrieved candidates are third-party reviews, roundups and discussion threads. One 2026 vendor study of four brands, covering 47,097 citations across three engines, put the share of cited material a brand owns at roughly 2%; four brands is illustrative rather than a constant, but the order of magnitude is the point.9