Home/Learn GEO/Fan-out expansion
Prompt research · Source 3 of 5

How do you expand a head term into fan-out sub-questions?

Along six fixed axes, four to six questions per head term, each labelled as reconstructed rather than attested. Google documents that its AI surfaces issue multiple related searches across subtopics, which makes these literal retrieval targets and makes discipline about invention the whole job.

On this page
Share this
Share on X Share on LinkedIn
The short answer

Fan-out expansion is writing, for each head term in your category, the four to six sub-questions a buyer asks next (cost, setup time, alternatives, risk, fit and proof) and entering each one as its own panel prompt. It ranks third among prompt sources because nobody is attested to have asked these: you wrote them, so label every entry as reconstructed. The reason to write them anyway is that sub-questions reach where the head term does not: an audit of 4,706 queries found 53% of the domains an AI Overview consults absent from the organic top 10.2

Key takeaways
  • Six axes cover a buying decision: cost, setup time, alternatives, risk, fit and proof. Write four to six per head term, in your buyers’ words, and stop there.
  • Admit an expansion only if somebody asked it, an engine suggested it, or the axis is universal, and record which test it passed.
  • Sub-questions are retrieval targets, not brainstorming: Google documents that AI Overviews and AI Mode may run a query fan-out across subtopics and data sources.1
Definition

What is fan-out expansion in prompt research?

Fan-out expansion is the deliberate reconstruction of the sub-questions a buyer asks around a head term, so that a prompt panel covers a topic rather than a phrase. You take a two-word category term that carries no question at all, and write the four to six questions a reasonable person asks between first hearing of the category and paying for something in it. Each becomes a separate panel entry with its own presence rate.

It shares a name with something else and the two must not be confused. The engine’s own fan-out is an internal step: Google documents that AI Overviews and AI Mode “may use a ‘query fan-out’ technique”, which it glosses as “issuing multiple related searches across subtopics and data sources” to develop one response, and nobody outside Google can see the queries it generates.1 The mechanism is How AI search works’s query fan-out page. Yours is a reconstruction used as a sampling frame: a way of choosing which questions to measure, not a claim about the engine’s internals.

The six axes

Which four to six sub-questions should you write?

Six axes, in the order a buyer walks them: cost, setup time, alternatives, risk, fit and proof. Write four to six per head term, not all six every time.

AxisThe question underneath itWorked expansion of “reddit marketing tools”
CostWhat does this class of thing cost, and what am I comparing it to?What does it cost to run Reddit marketing for a small SaaS?
Setup timeHow long before it does anything?How long does it take before Reddit posts bring any traffic?
AlternativesWhat else solves this, including doing nothing?What are the alternatives to marketing on Reddit for early-stage products?
RiskWhat goes wrong, and how badly?Can promoting your own product on Reddit get your account banned?
FitDoes it work for someone in my situation?Is Reddit marketing worth it for a two-person team?
ProofHas it worked for anyone like me?Do B2B companies actually get customers from Reddit?

The axes are stable across categories and the wording never is, which is the useful asymmetry here. Every buyer of a considered purchase asks some version of all six, so the axes catch the question you would otherwise have forgotten. The phrasing has to come from your own buyers, because the words they use are the retrieval targets and the words you would use are not.

Four to six per head term is a working limit rather than a finding. Below four you have not covered the decision; above six you are generating questions to fill a template, and every extra reconstructed entry dilutes the attested share of your panel. Twelve genuinely different sub-questions usually means the head term is two categories wearing one name, and splitting it produces a better panel than expanding it.

The grounding rule

How do you stop expansion from becoming invention?

Admit an expansion only if it passes one of three tests, and record which one. Somebody asked it: the question, or something close to it, appears in a call transcript, a ticket or a community thread, in which case it is not an expansion any more and should be filed with its evidence. The engine suggested it: the follow-up questions an assistant offers under an answer, and the related-searches strip on a results page, are the platform’s own guess at the next question, a weaker warrant than a human but a real one. The axis is universal: cost and risk apply to every considered purchase. An expansion passing none of the three is a hypothesis; tag it, and expect it to be the first thing you retire.

Generating the list with a language model fails systematically rather than randomly. A model asked for example questions samples its training distribution, which is dominated by how your category is already written about in public. It returns the median phrasing of questions that are already well covered, and under-produces the questions nobody has published on, which are exactly the questions where you are invisible. Used to rewrite a list you grounded in calls it is a reasonable tool; used to originate the list, it builds a panel that measures the consensus of your category’s existing content.

Keep the reconstructed share of the panel visible in your reporting. A panel that is 25% expansion is a research instrument with a labelled assumption; one that is 70% expansion is a survey of your own imagination, and it will still produce a confident percentage every week. The source ranking sets the composition caps.

Why it works

Why are these retrieval targets rather than guesses?

Because retrieval runs per sub-query, so a sub-question is a separate entry point rather than a rephrasing of the head term. Google documents the mechanism: its AI surfaces may issue multiple related searches across subtopics and data sources for one response.1 A page that answers one of those searches can be retrieved for it while ranking nowhere for the visible question.

The gap between what ranks and what gets cited is measured, and should be quoted as a dated range rather than a point. A 4,706-query audit of Google AI Overviews published in Findings of ACL 2026, on data collected in September 2025, reports that “on average 53% (27%) of domains that AIO consults are not contained in top-10 (top-100) Organic search results”; a 55,393-query preprint collected between 13 March and 21 April 2026 measured 29.8% of AI Overview reference domains absent from the corresponding first page.23 Roughly 30% to 53%, across two windows six months apart, with different verbs: one counts domains consulted, the other domains referenced. A SIGIR 2026 study of 11,500 queries measured URL-level Jaccard similarity of 0.11 to 0.18 between Google organic results, AI Overviews and Gemini.4 Fan-out is the most plausible mechanism for a gap that size, and plausible is the honest word: no study instruments the fan-out step.

A vendor study published in December 2025, over 10,000 keywords and 173,902 URLs, reported a Spearman correlation of 0.77 between how many reconstructed fan-out queries a page ranks for and its odds of AI Overview citation.5 Read it carefully: correlational, vendor-published, and a reconstruction of the kind this article describes rather than the engine’s real queries. Corroboration for the approach, not evidence about the mechanism.

The boundary

How is this different from writing one page that covers the fan-out?

One is a measurement instrument and the other is a content decision, and conflating them corrupts both. A panel entry is a question you intend to measure; a page section is a question you intend to answer. Most teams collapse the two by writing panel entries only for questions they have already covered, which guarantees a flattering number and removes the panel’s ability to tell them anything. The panel should contain questions you have no page for.

The reverse mistake is turning every panel entry into a section. A 2026 end-to-end benchmark over 171,003 documents and 2,700 queries found that optimising body text alone, averaged across ten strategies, reduced top-20 presence by about 9%, post-rerank top-10 presence by 16% and citation by about 6%.6 Padding a page with sub-question headings you cannot answer specifically is what that looks like from the inside: the page gets longer, its topical signal gets thinner, and it competes worse upstream than the extra entry points win downstream.

The rule that keeps both honest is to let the panel lead by a quarter. Measure the sub-questions first, find the ones where an engine answers with somebody else’s page, and write only those. That sequence gives you the before-and-after that makes any later claim defensible, and the 2026 critical survey is blunt about how little else is: its confidence table rates “a white-hat GEO intervention durably improves organic discoverability across multiple engines” at low.7

Ranking still applies

Do the pages you write for these sub-questions still need to rank?

Mostly yes, and Google says so about its own surfaces. Its guide to optimising for generative AI features, last updated 10 Jul 2026, states that “the best practices for SEO continue to be relevant because our generative AI features on Google Search are rooted in our core Search ranking and quality systems”.8 A separate page, last updated 10 Dec 2025, is blunter about the gate: to be shown as a supporting link in AI Overviews or AI Mode a page “must be indexed and eligible to be shown in Google Search with a snippet”, and “there are no additional technical requirements”.1 An expansion that turns into a page therefore inherits every indexing problem you already had; no amount of sub-question coverage routes around a noindex.

The qualifier is the gap in the previous section. If half the domains an overview consults sit outside the ranked ten, ranking for a sub-question is a strong prior rather than the whole route, and the sources filling the rest were retrieved for a query nobody typed. That is why a panel measures sub-questions rather than positions: a rank check tells you where you stand on the question a person asked, and a panel entry tells you whether anything of yours reached the searches an engine issued underneath it. Run both, and treat the disagreement as the finding rather than as a bug in one instrument.

The honest limit of this article

Nobody can check this source’s work. Fan-out queries are internal to the engines, no engine publishes them, and the only published correlation between fan-out coverage and citation uses a reconstruction rather than the real thing, so the justification for this source rests on one documented sentence from one engine plus an unexplained gap between ranking sets and citation sets. The six axes are a convention we find useful, not a measured taxonomy; a different practitioner would name five or eight. And the failure mode here is invisible by construction: a well-formed sub-question nobody has ever asked looks exactly like a good panel entry, produces a presence rate every week, and tells you nothing.

Where a product fits, and where it does not

Expansion costs nothing and takes an afternoon: list your head terms, walk each down the six axes, write the question in your buyers’ words, mark which grounding test it passed, and stop at six. Then search each one yourself and note who is cited, which tells you where you stand before measuring anything formally. Bavior cannot do the expansion for you: it does not reconstruct fan-out queries and does not sell a list of them, because that would be inference presented as telemetry. What it does is the repetition afterwards, running your fixed panel across five engines on a schedule and recording which sources each answer cited, so you can see whether coverage work changed who gets retrieved; where a cited source is a live discussion thread it drafts a reply on an account you control, which you approve before anything posts. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).

Sources, all checked 30 Aug 2026
  1. Google Search Central, “AI Features and Your Website” (page last updated 10 Dec 2025): the query fan-out sentence, and the requirement that a supporting link be indexed and snippet-eligible (first-party): developers.google.com/search/docs/appearance/ai-features
  2. Kirsten et al., Findings of ACL 2026; 4,706 queries issued in September 2025 from the US and Germany; 53% of consulted domains absent from the organic top 10, 27% from the top 100: aclanthology.org/2026.findings-acl.526
  3. Xu, Iqbal & Montgomery, 2026; 55,393 trending queries collected 13 March to 21 April 2026; 29.8% of AI Overview reference domains absent from the corresponding first page (preprint): arxiv.org/abs/2605.14021
  4. Grossman et al., SIGIR 2026; a public benchmark of 11,500 queries; Jaccard similarities between 0.11 and 0.18 across Google organic results, AI Overviews and Gemini: arxiv.org/abs/2604.27790
  5. Vendor fan-out study published 6 December 2025: 10,000 keywords, 33,000 reconstructed fan-out queries, 173,902 URLs; Spearman 0.77, and pages ranking for the head term plus a fan-out 161% more likely to be cited. Correlational and vendor-published; described rather than linked, per this curriculum’s sourcing rule.
  6. Kim et al., “SAGEO Arena”, KDD 2026 (arXiv:2602.12187v2, 7 Aug 2026); 171,003 documents, 2,700 queries; Table 2 body-text-only averages across ten strategies: hit rate at 20 down 9%, at rerank 16%, citation 6%: arxiv.org/abs/2602.12187
  7. Martinez, “A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035; Table 5 rates durable cross-engine white-hat improvement at low confidence: arxiv.org/abs/2607.14035
  8. Google Search Central, “Optimizing your website for generative AI features” (page last updated 10 Jul 2026): SEO best practices stay relevant because Google’s generative AI features are rooted in its core Search ranking and quality systems (first-party): developers.google.com/search/docs/fundamentals/ai-optimization-guide
FAQ

Frequently asked questions.

Can I see the actual fan-out queries an engine generated?

No engine exposes them, and any list presented as your fan-out queries is a reconstruction. Google documents that its AI surfaces may issue multiple related searches across subtopics and data sources, and it publishes nothing about which searches were issued for a given question; Search Console reports AI-feature traffic blended into the Web search type with no per-query breakdown. A good reconstruction is still useful, being a defensible guess at the questions a buyer asks next, but it is inference rather than telemetry, and anything sold as the real thing should be priced accordingly.

How many head terms should I expand?

Enough that expansion supplies about a quarter of the panel, which for a 40-prompt panel is roughly two to three head terms. The constraint is not how many head terms exist but how much reconstructed material a panel can carry before it stops measuring your buyers: no single source should exceed about 40% of the entries, and expansion is the one source where nobody is attested to have asked the question. Two head terms expanded properly along the six axes beats six head terms expanded into a template you filled in to reach a number.

Is it wrong to include a sub-question I have no page for?

It is the point. A panel entry is a question you intend to measure, not a question you have already answered, and the entries where an engine answers with somebody else's page are the only ones that produce a decision. A panel built only from questions your site already covers guarantees a high presence rate and tells you nothing you did not know. Let the panel lead your content by about a quarter: measure the sub-questions first, find the gaps, and write only the pages the measurement asked for.

Should I write one page per sub-question?

Usually not. Cover the cluster on one page, and only for the sub-questions you can answer with something specific. A 2026 benchmark over 171,003 documents and 2,700 queries found body-text-only optimisation cutting average top-20 presence by about 9% and final citation by about 6%, which is what dilution looks like when it is measured: a section that says nothing thins the page's topical signal without buying a real entry point. Add a sub-question heading when you have a specific, sourced answer under it, and leave the rest to the panel until you do.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

One head term is one question.
Your buyer asks six.

Start free trial