Home/Learn GEO/Relevance and ranking
GEO versus SEO · Topic

Why do relevance and ranking dominate AI answer engines?

Because the largest controlled benchmark in the field tested 54 combinations of rewriting method and domain, found 3 significantly positive and none in question answering, and measured plain context position as worth several times the best rewrite.

On this page
Share this
Share on X Share on LinkedIn
The short answer

Relevance and context position dominate because they are the only two levers that act before the model writes anything, and the only two the 2026 critical survey grades as strongly supported.2 C-SEO Bench, presented at NeurIPS 2025 over more than 1.9k queries and 16.3k documents, benchmarked ten rewriting methods across six domains and two tasks: only 3 of its 54 method and domain cases came out significantly positive, and none of those were in question answering.1

Key takeaways
  • Two levers are graded strong by the survey and nothing else is: query and document relevance, and position inside the retrieved context.2
  • Position beats prose. In the benchmark's retail domain, moving a source to context position 1 improved its citation rank by 2.77 places on average, against 0.36 for the best rewriting method.1
  • Rewriting gains decay as rivals copy them; the benchmark calls the dynamic “congested and zero-sum”. Genuine relevance is judged against the query, not against the other documents.1
  • One content lever survives at moderate-to-strong support: verifiable, dated, attributed evidence. Fixed formatting recipes generalise poorly, and keyword stuffing scored below doing nothing.2
  • Do not restructure a page that already ranks first. The measured effect on top-ranked sources was a 20% to 30% loss of visibility.3
The benchmark

What did C-SEO Bench actually test?

C-SEO Bench tested ten conversational-SEO rewriting methods across six domains and two tasks, on a scale large enough to detect an effect if one existed.

1.9k+queriesacross 6 domains
16.3kdocumentsQA and product recommendation
54method-domain casesten methods, six domains
3significantly positivezero in question answering

C-SEO Bench is a benchmark for whether rewriting a document to please a generative engine actually works, and its answer is mostly no. It was presented at the NeurIPS 2025 Datasets and Benchmarks Track, covers two tasks, question answering and product recommendation, over six domains, and its code and data are public.1 It is also the result the GEO versus SEO stage is built around.

The abstract states the result without hedging: “most current C-SEO methods are not only largely ineffective but also frequently have a negative impact on document ranking, which is opposite to what is expected. Instead, traditional SEO strategies, those aiming to improve the ranking of the source in the LLM context, are significantly more effective.” The paper's own summary is that making the target document first in the context window “leads to far greater citation ranking gains in the LLM response than any C-SEO method”.

Two details make that harder to dismiss than a single experiment usually is. The first is breadth. The paper reports 54 method and domain cases, enough that scattered wins would show up, and its finding is that “out of 54 cases, we uncover only three where the ranking improvements are statistically significant”, with no method effective at all for the question answering task. The second is the direction of the failures. Several transformations did not merely fail to help, they reduced the document's rank: adding statistics lowered rankings in 19 of the 24 settings where it was measured.1 A tactic that is neutral is cheap; a tactic that is negative is a tax you pay for reading the wrong blog post.

The mechanism

Why is context position the same thing as ranking?

Context position is what ranking converts into, because the retrieval layer's output ordering becomes the order of documents inside the model's prompt. An answer engine retrieves a candidate set, keeps the top handful, concatenates them into a context window and generates from that window. Whatever decided the ordering of the candidate set has therefore decided the ordering of the context. Google's own documentation describes the mechanism in exactly those terms: retrieval-augmented generation is used “to improve the quality, accuracy, and freshness of AI responses by relying on our core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index”.5

That is why ranking is the only lever that pays twice. It decides whether you are in the candidate set at all, and then it decides where you sit inside it. Google states the entry condition plainly: to be eligible for its generative features, “a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements”.5 The 2026 critical survey grades query and document relevance as strong in controlled settings, and position in context as strong conditional on the document already being retrieved, and nothing else strong.2 Every content tactic in the literature operates on the second stage only, on a document that has already won the first.

One caveat belongs in the same breath, because it is routinely dropped. C-SEO Bench did not manipulate a live search ranking; it operationalised “traditional SEO” as placing the source at context position 1 in its own retrieval setup. The finding is therefore precise about what it measured, which is that being first in the context is worth far more than being better written. It remains an inference, not a measurement, to say that a specific SERP improvement buys a specific citation gain. That study has not been published.

Adoption decay

What happens to a rewriting gain once everyone adopts it?

A rewriting gain shrinks as more documents in the same candidate set adopt the same rewrite, and C-SEO Bench measured that decay directly. The benchmark varied the number of documents in a candidate set that had been optimised with the same method and found that “as we increase the number of C-SEO adopters, the overall gains decrease, depicting a congested and zero-sum nature of the problem”.1 The 2026 critical survey lists competitive erosion of individual gains at moderate confidence, one of three moderate rows in its table of the field's principal claims.2

This is the difference between a stock and a flow, and it should decide where recurring budget goes. Genuine relevance is a stock: a page that is the best answer to a question is not made less relevant by a competitor also becoming relevant, because relevance is judged against the query, not against the field. Formatting tactics are a flow: they work by making your document stand out inside a candidate set, which means their value falls exactly as fast as other documents adopt them.

The 2024 KDD paper shows the same redistribution from the other end. Its per-rank table has the three best-performing tactics reducing visibility for an already first-ranked source while roughly doubling it for the fifth-ranked source: adding citations moved rank-1 visibility by -30.3% and rank-5 visibility by +115.1%, with adding quotations at -22.9% and adding statistics at -20.6%.3 These techniques do not create answer share; they move it from stronger sources to weaker ones. If you are the strong source, the same technique is a cost.

The other half

Does this mean content work is worthless?

No, because one category of content work has moderate-to-strong support, and it is the category that also makes a page better. The critical survey grades extractable evidence, meaning figures, definitions and comparisons, as moderate to strong, conditional on being truthful, attributed and intent-matched. Its stated criterion is worth quoting exactly: “The criterion is therefore not to ‘add numbers,’ but to provide relevant, verifiable, dated, and properly attributed evidence.”2

What is not supported is the recipe version. Fixed formatting templates are graded “poor generalization”; authoritative tone is weak and unstable; keyword stuffing is null or negative across multiple benchmarks, and in the KDD experiment scored 17.7 against a do-nothing baseline of 19.3, the only tactic to finish below doing nothing.3

And there is a composition risk that only shows up when you measure the whole pipeline. SAGEO Arena reinstated retrieval and reranking over 171,003 documents and 2,700 queries and found body-only optimisation reducing average top-20 presence by about 9%, post-rerank top-10 presence by 16%, and final citation by about 6%.4 The survey's reading is the practical rule: a rewrite “may therefore perform well once injected while making the document less retrievable or less competitive upstream”. Never buy quotability with topical relevance.

Valuation

What is a top context position actually worth?

It is worth a larger share of the answers, and the published evidence stops there. The critical survey's table of confidence in the field's principal claims puts “citation scores predict clicks, conversions, or revenue” in the lowest band it uses, very low, on the basis of “one suggestive quasi-experiment and a few industry claims”, with causality not established.2 The chain from rank to context position to citation is measurable. The chain from a citation to a visit, and from a visit to a customer, is not measured at all yet.

Independent behavioural data suggests the last link is thin on its own terms. Pew Research Center tracked 68,879 Google searches in March 2025 and found that users who encountered an AI summary clicked a traditional search result on 8% of visits, against 15% of visits where no summary appeared. Clicks on a link inside the summary itself happened on just 1% of visits.6 Winning the top context slot buys a mention in front of a reader who mostly will not click through.

Two practical consequences follow. Value ranking work by the share of answers that name you rather than by forecast traffic, because share of voice is the variable anyone can actually measure today. And treat priced-up outcome claims with suspicion: the same survey table has a bottom row for the number the industry still repeats. Against “GEO increases visibility by 40%” it records the verdict “rejected as a general claim”, on the grounds that the figure is a relative maximum on one metric under one configuration.2 Rank first because it moves the only variable measured cleanly, not because someone has priced the outcome for you.

Application

What order should you actually work in?

1. Be genuinely the best answer to the question

Relevance is the only lever graded strong that you can move on your own side. Pick the exact question, answer it completely, and do not dilute the page with adjacent topics to catch more queries, because that dilution is what the retrieval benchmark penalised.

2. Earn the rank that buys context position

Everything ordinary SEO does to improve position is also the highest-leverage GEO work available, because Google states its AI surfaces rely on the core Search ranking systems. Links, internal structure and indexation belong here, not in a separate AI workstream.

3. Add verifiable, dated, attributed evidence

Real numbers with real sources and real dates, in the body text. This is the one content lever with moderate-to-strong support, and the survey is explicit that fabricated statistics can raise reuse while destroying the reason anyone should reuse you.

4. Leave your already-first pages alone

If a page already ranks first for its query, the measured effect of aggressive answer-shaped rewriting is a 20% to 30% loss in answer share. Edit those pages for accuracy and freshness; do not restructure them to chase a citation you may already be getting.

The honest limit of this page

The position-versus-rewriting gap is a ratio we derived, not a figure the paper quotes: C-SEO Bench reports an average rank improvement of 2.77 for context position 1 against 0.36 for the best rewriting method, in one domain on one model, and describes the gap in words as “far greater” rather than as a multiple. “3 of 54 significantly positive” is a statistical bar rather than an effect size, so it tells you how often a method beat noise, not how large the wins were. C-SEO Bench, like almost every study in the field, supplies its own candidate set rather than querying a live commercial engine, so it measures ordering effects inside retrieval it controls. Nobody has published a controlled experiment that moves a real page's rank on a live engine and measures the change in its citation rate. Until that exists, “ranking dominates” is a strong inference from convergent evidence, not a demonstrated coefficient.

Where a product fits, and where it does not

Nothing on this page needs a tool: read the benchmark, check whether your target pages rank for the questions they answer, add dated and attributed evidence to the ones that do not yet carry any, and stop rewriting the ones already sitting first. The only part software helps with is knowing whether any of it changed the answers, and only because a single check cannot tell you: Bavior runs a fixed prompt set across five engines on a schedule and records which sources each answer cited, so a shift shows up as a change in a distribution rather than a change in a screenshot. It does not improve rankings, and no measurement product does. Start with the free visibility check; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 30 Aug 2026).

Sources, all checked 30 Aug 2026
  1. Puerto, Gubri, Green, Oh, Yun, “C-SEO Bench: Does Conversational SEO Work?”, NeurIPS 2025 Datasets & Benchmarks Track; 2 tasks, 6 domains, more than 1.9k queries and 16.3k documents, ten methods, 54 method and domain cases, code and data public: arxiv.org/abs/2506.11097
  2. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (preprint); 45 studies reviewed, Table 4 on lever support and Table 5 on claim confidence: arxiv.org/abs/2607.14035
  3. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, Deshpande, “GEO: Generative Engine Optimization”, KDD ’24; Table 2 per-rank visibility changes and the keyword-stuffing result: arxiv.org/abs/2311.09735
  4. Kim, Jeong, Kim, Lee, Lee, “SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization”, KDD 2026 (arXiv:2602.12187v2, 7 Aug 2026); 171,003 documents, 2,700 queries, retrieval and reranking reinstated: arxiv.org/abs/2602.12187
  5. Google Search Central, “Optimizing your website for generative AI features on Google Search”, updated 10 Jul 2026 (first-party): developers.google.com/search/docs/fundamentals/ai-optimization-guide
  6. Pew Research Center, “Google users are less likely to click on links when an AI summary appears in the results”, 22 Jul 2025; 68,879 searches tracked in Mar 2025, 8% versus 15% click rates: pewresearch.org
FAQ

Frequently asked questions.

What is C-SEO Bench and what did it find?

C-SEO Bench is a benchmark presented at the NeurIPS 2025 Datasets and Benchmarks Track that tests whether rewriting a document for a generative engine works. It covers two tasks, six domains, more than 1.9k queries and 16.3k documents, and ten conversational-SEO methods. The paper reports 54 method and domain cases, of which only three showed statistically significant ranking improvements, and none of those were in question answering. Its abstract states that most such methods "are not only largely ineffective but also frequently have a negative impact on document ranking", and that traditional ranking work is "significantly more effective".

Is context position really worth several times more than rewriting?

In the retail domain of that benchmark, yes, and the qualifier matters. C-SEO Bench reports an average citation-rank improvement of 2.77 places for moving a source to context position 1, against 0.36 for the best-performing conversational-SEO method, a ratio of roughly 7.7 to 1 on one model in one domain. The paper itself does not print a multiple; its wording is that context position 1 gives "far greater citation ranking gains in the LLM response than any C-SEO method". The benchmark also supplies its own candidate documents rather than querying a live commercial engine, so the honest reading is that being first in the context is worth far more than being better written, not that a specific SERP improvement buys a specific citation gain.

Why do GEO tactics stop working over time?

Because they compete for share inside a fixed candidate set, so their value falls as adoption rises. C-SEO Bench varied how many documents in a set had been optimised with the same method and found that "as we increase the number of C-SEO adopters, the overall gains decrease, depicting a congested and zero-sum nature of the problem". The 2026 critical survey grades competitive erosion of individual gains at moderate confidence. Genuine relevance behaves differently: it is judged against the query, not against the other documents, so a competitor improving does not make you less relevant.

Should I rewrite a page that already ranks first?

Not for citation purposes, because the measured effect is negative. The KDD 2024 GEO paper's per-rank table shows the three best-performing rewriting tactics reducing visibility for an already first-ranked source by 20% to 30%, while roughly doubling it for the fifth-ranked source. The mechanism is redistribution: the tactics make lower-ranked sources more quotable and take answer share from the incumbent. Edit a first-ranking page for accuracy and freshness, and put restructuring effort into pages ranking third to tenth instead.

If ranking dominates, is any content tactic worth doing?

One category is, and the survey grades it moderate to strong: extractable evidence, meaning figures, definitions and comparisons, conditional on being truthful, attributed and intent-matched. Its criterion is explicit: "The criterion is therefore not to 'add numbers,' but to provide relevant, verifiable, dated, and properly attributed evidence." What fails is the recipe version. Fixed formatting templates are graded poor-generalization, authoritative tone weak and unstable, and keyword stuffing scored below the do-nothing baseline in the KDD experiment at 17.7 against 19.3.

Does a citation in an AI answer bring traffic?

Not reliably, and nobody has demonstrated that it does. The 2026 critical survey puts "citation scores predict clicks, conversions, or revenue" in its very-low confidence band, resting on one suggestive quasi-experiment and a few industry claims, with causality not established. Pew Research Center's March 2025 tracking of 68,879 Google searches found users clicked a traditional result on 8% of visits where an AI summary appeared, against 15% where none did, and clicked a link inside the summary on 1% of visits. Treat a citation as share of voice in front of a reader, not as a click you have already earned.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Rank first. Rewrite second.
Measure both as a distribution.

Start free trial