Home/Learn GEO/Third-party pages
Off-site GEO · Mechanism

Why do third-party pages dominate shortlist questions?

Not because independent writers are trusted more than you are. A question asking for several named options is answered better by a document containing several, and yours contains one.

On this page
Share this
Share on X Share on LinkedIn
The short answer

A shortlist question asks for a set of named options with reasons attached, and a vendor page is a document about one option, so the roundup is a closer match for the query that was actually issued before anyone judges whose prose is better. The asymmetry is structural rather than editorial, which is why it does not respond to rewriting the page you own. In the most detailed public breakdown of AI citations, covering 21,143 search-layer citations, citations doing comparison work inside an answer averaged 0.1524 on the study’s influence proxy against 0.0529 for reference-only citations, the widest gap in its semantic-role table.1

Key takeaways
  • The cause is query-document match, not credibility: to build a set out of single-vendor pages an engine must find each candidate, decide who belongs, and invent the comparison. One roundup supplies all three.
  • Definitional and how-to questions behave differently because the evidence they need is about one subject, which is the form your own documentation is best placed to publish.
  • The column reading 34.22% to 46.35% in that study is labelled Official and covers every organisation’s own site in an answer, not how much of one answer a single brand owns.
  • No published study reports the first-party share of citations by question class, and the percentages that circulate for it carry no method.
  • Buying a slot on a shortlist page is a different transaction from earning one, and where a platform or community requires disclosure of a commercial relationship, disclosure is required.
The structural case

Why is a shortlist question a bad match for your own page?

Because the two documents are answering different questions. “Best project tracker for a five-person agency” is a request for a set: several named candidates, each with a reason attached and a boundary where it stops being the right pick. A vendor page is a document about one entity, and its job is to end the comparison rather than conduct it. Retrieval scores documents against the query as issued, so a page that already has the shape of the answer wins the match on structure, before anything about quality or trust is weighed.

Read the failure from the engine’s side and it stops looking like a verdict on your writing. To build a shortlist out of single-vendor pages, a system has to find each candidate separately, decide which candidates belong in the set at all, and construct the comparison itself from documents each written to win it. One roundup hands over all three jobs pre-solved, and nothing in that requires the roundup to be independent, well written, or even current.

The one systematic review of this literature grades the same boundary: it rates “query–document relevance and context position are major determinants” at high confidence, and a white-hat intervention durably improving discoverability across engines at low.3 Relevance to the issued query is the part of this system with strong evidence behind it, and relevance is exactly where a single-vendor page is structurally short.

Before the answer

What happens to that question before anything is written?

It gets taken apart and searched several times. Google documents that AI Overviews and AI Mode “may use a ‘query fan-out’ technique”, “issuing multiple related searches across subtopics and data sources”, and the same page names what it is built for: nuanced questions that would previously have taken multiple searches, “from exploring a new concept, to comparing options, and beyond”.4 A shortlist question is therefore rarely one query by the time retrieval sees it, and several of the sub-queries it becomes are themselves comparative.

The pool assembled from those sub-queries is wider than the first page of results. A 4,706-query audit of Google AI Overviews published in Findings of ACL 2026, on data collected in September 2025, reports that on average 53% of the domains an overview consults are not in the organic top 10, and 27% are not in the top 100.6 An 11,500-query SIGIR 2026 study measured URL-level Jaccard similarity of 0.11–0.18 between Google organic, AI Overviews and Gemini, Jaccard being the intersection over the union of the two URL sets.7 Neither number is a citation share, and both point the same way here: the pool being assembled is not your category’s results page.

What sits underneath is still ordinary search: Google states that its “generative AI features on Google Search are rooted in our core Search ranking and quality systems”.5 The system is doing what search has always done, against a query you did not get to write.

The evidence

What does the citation data actually support?

Less than the argument above needs, and more than nothing. The most detailed public breakdown is a descriptive preprint on 602 controlled prompts, 21,143 search-layer citations and 18,151 fetched pages, scored on a constructed influence proxy.1

What the citation was doing in the answerCitationsMean influence
definition1,6630.1531
comparison1,7190.1524
evidence6,2160.1235
example1,4680.1047
background5,5820.0801
reference1,2980.0529

Read the left column before the right one. These labels describe what a citation was doing inside one answer, not what kind of page it came from, so “comparison” here is a job, not a genre. The page-side table runs in the same direction: pages containing comparison content averaged 0.1389 against 0.0894 for pages without, which the authors print as a relative difference of +55.28%.1 Their own summary of the question-type figure is that “comparison questions have the highest reported average influence among question types”.

What none of this establishes is the sentence this article opened with. The study reports no split of citations into first-party and third-party by question class. Its influence score is an observational proxy built from repeated reference, early position, paragraph coverage and text overlap, not a measurement of what the model relied on. And the authors set an explicit ceiling on their own claims, treating direct counts and descriptive contrasts as findings while reserving causal optimization prescriptions for intervention experiments that have not been run. The tables are consistent with the structural argument, but they were not built to test it.

The contrast

Why do definitional and how-to questions not behave this way?

Because the evidence those questions need is about one subject, and on one subject you are the closest available document. Ask what a category is, or how to configure something, and the answer wants a definition and a procedure scoped to a single entity. In the same dataset, pages carrying definition markers averaged +57.33% higher mean influence than pages without them, and pages carrying how-to content +41.20%.1 Those are the two forms your own documentation is best placed to supply, and that is the whole asymmetry: third parties are not preferred, the evidence each question class needs is simply published by different people.

This is also where the most-quoted figure in the file gets misread. That study’s source-type table reports official sources at 34.22% of citations on ChatGPT, 46.35% on Google’s AI surfaces and 44.07% on Perplexity, the largest of three categories alongside news and vertical.1 The column is labelled Official, and it counts every organisation’s own site appearing in an answer: manufacturers, institutions, agencies, and every vendor named, not one of them. It is also reported per platform, never per question class. So it cannot be read either as your share of an answer, which is a separate sizing question, or as proof that definitional questions pull official pages while shortlist questions pull roundups. That per-class split is the number nobody has published, and the class taxonomy flags the same gap.

Your own roundup

Should you publish your own comparison page?

It clears the structural bar in this article and fails a different one, so publish it for a different reason than citations. A genuine multi-product comparison on your domain is set-shaped, which is exactly what the query wants, so it is eligible in a way your product page is not. What it does not clear is selection: it is one candidate among many for the same sub-query, it carries an obvious interest, and the outcome does not depend on you.

The intervention evidence is discouraging about the usual response, which is to keep rewriting. A NeurIPS 2025 benchmark tested ten conversational-SEO methods across more than 1.9k queries and 16k documents and reports that “out of 54 cases, we uncover only three where the ranking improvements are statistically significant”.2 Changing the words on a page you already own is the class of intervention that keeps not working. Changing which pages exist about your category is a different class, and the honest version of it is slow.

So the defensible reason to publish a comparison is supply, not placement. Roundup authors omit products whose pricing, limits and constraints they cannot verify, and an engine writing a comparison needs the same facts from somewhere. A page that gives every other product a real “best for” a reader could act on, states a genuine limitation of your own, and recommends a free option where a free option actually wins is usable by both. One where everything but your product is bad is usable by neither.

The line

Where is the line between earning a place and buying one?

The conclusion this mechanism invites is “so get onto third-party pages”, and that sentence quietly contains two different transactions. Being included because a reviewer judged you belonged is one thing. Paying for the slot is another: a purchase of position in a document whose value to a reader, and to a system reading on a reader’s behalf, comes entirely from the belief that the ranking reflects a judgement. Where a platform, publisher or community requires disclosure of a commercial relationship, disclosure is required, and that obligation does not weaken because part of the audience is now machines.

The same line rules out the other shortcut. Creating the third-party page yourself, under a byline that does not disclose the relationship, is your own page wearing someone else’s name; it fails the first time a reader checks, and on community platforms it is removed by moderators rather than by an algorithm you could out-run. The boundary article sets this out in full, and this page does not restate it.

What is left is unglamorous and durable. Be verifiable, so a reviewer can confirm a claim instead of dropping you. Correct the record where a widely cited roundup has your price or your limits wrong, which is the fastest-returning work here and the one nobody schedules. Earn coverage with something real, such as original data or a tool people use. None of it is fast, and all of it survives the next retrieval change.

The honest limit of this article

The structural argument here is an explanation, not a measurement. The one number that would test it directly, the first-party versus third-party share of citations broken out by question class, has not been published by anyone; a figure near 80% for the third-party share of shortlist citations circulates widely with no method attached, and nothing above rests on it. The influence tables come from a descriptive preprint on a single dataset of 602 prompts collected in one window, using a proxy its authors are careful to call a proxy, and length and structure there may stand in for editorial quality rather than cause anything. Treat any percentage attached to this mechanism as unverified until someone measures it.

Where a product fits, and where it does not

The diagnostic here is free and takes an afternoon: run your ten most commercial shortlist questions on each engine you care about, record every cited URL, and sort them into pages you own, pages that describe you, and pages that do not mention you. That list, usually about fifteen URLs, is the concrete version of everything above. Bavior runs a fixed prompt set across five engines on a schedule and records which sources each answer cited, so the list stays current instead of being rebuilt by hand; where a cited source is a live discussion thread it drafts a reply on an account you control, which you approve before anything posts. It does not place you on anyone’s roundup, does not contact publishers, and cannot buy, negotiate or arrange an inclusion. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 30 Aug 2026).

Sources, all checked 30 Aug 2026
  1. Zhang Kai, He Xinyue, Yao Jingang, “From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms”, 28 Apr 2026 (v2, 29 Apr 2026), arXiv:2604.25707 (descriptive preprint, public dataset); 602 controlled prompts, 21,143 valid search-layer citations, 23,745 citation-level feature records, 18,151 fetched pages; semantic-role, evidence-genre and source-type tables: arxiv.org/abs/2604.25707
  2. Puerto, Gubri, Green, Oh, Yun, “C-SEO Bench: Does Conversational SEO Work?”, NeurIPS 2025 Datasets & Benchmarks Track, arXiv:2506.11097; ten methods, more than 1.9k queries and 16k documents; “out of 54 cases, we uncover only three where the ranking improvements are statistically significant”: arxiv.org/abs/2506.11097
  3. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (survey preprint); the confidence table rating relevance and context position high and durable cross-engine gains low: arxiv.org/abs/2607.14035
  4. Google Search Central, “AI features and your website”, last updated 10 Dec 2025 (query fan-out, comparing options, snippet eligibility; first-party): developers.google.com/search/docs/appearance/ai-features
  5. Google Search Central, “Google’s Guide to Optimizing for Generative AI Features on Google Search”, last updated 10 Jul 2026 (the “rooted in our core Search ranking and quality systems” sentence; first-party): developers.google.com/search/docs/fundamentals/ai-optimization-guide
  6. Kirsten et al., Findings of ACL 2026; 4,706-query audit of Google AI Overviews, data collected September 2025; “on average 53% (27%) of domains that AIO consults are not contained in top-10 (top-100) Organic search results”: aclanthology.org/2026.findings-acl.526
  7. Grossman et al., SIGIR 2026; representative sample of 11,500 queries; URL-level Jaccard 0.11–0.18 across Google organic, AI Overviews and Gemini: arxiv.org/abs/2604.27790
FAQ

Frequently asked questions.

Does this mean my own site does not matter for shortlist questions?

It matters differently. Your pages are the supply of verifiable facts that roundup authors and engines both draw on, so pricing, limits and constraints that cannot be confirmed anywhere get you omitted from comparisons you never see. What your pages will not do is win the match against a set-shaped query on their own, because a document about one product is answering a narrower question than the one that was asked.

Official sources are 34% to 46% of AI citations. Is that not proof my own pages win?

No, because of what that column counts. It is a source-type label covering every organisation's own site appearing in an answer, including institutions and every vendor named, so it is split across all of them rather than held by one. It is also reported per platform, not per question class, so it says nothing about whether a shortlist question behaves like a definitional one.

Should I pay to be listed in a roundup that AI answers already cite?

Treat it as a different transaction from being included on the merits, and price it that way. A paid slot buys position in a document whose worth to readers rests on the ranking reflecting a judgement, and where the platform or community requires a commercial relationship to be disclosed, disclosure is required regardless of who is reading. The boundary article covers the full case, including undisclosed authorship.

Why does a definitional question pull my documentation when a shortlist question does not?

Because the evidence each one needs is published by different people. A definition or a procedure is scoped to a single subject, and on a single subject your own documentation is the closest available document. A shortlist needs contrastive evidence across several named options, which no single-vendor page contains, so the closest available document is somebody else's.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Most of the answer is written by other people.
Find out which pages, in your category.

Start free trial