Home/Learn GEO/The diagnostic
How AI search works · Diagnostic

Why does so much GEO advice go wrong?

Four questions sort almost any claim in this field into “supported”, “measured somewhere that skipped the hard part”, and “no mechanism”. Here they are, with the results they were built from.

On this page
Share this
Share on X Share on LinkedIn
The short answer

Most GEO advice goes wrong because it acts on the last stage of the pipeline and was measured in a setting where the first four stages had been switched off. Sort any claim by asking which stage it acts on, and whether retrieval was running when somebody measured it. A KDD 2026 benchmark that reinstated retrieval and reranking over 171,003 documents found body-only optimisation cut average top-20 presence by 9% and final citation by 6%.1

Key takeaways
  • Ask which of the five pipeline stages a tactic acts on. Interpretation and fan-out take no input from your site at all, so a tactic aimed there has no mechanism to work through.
  • Ask whether retrieval and reranking were running when the effect was measured. A study that places the document in the model’s context first cannot say whether the page would have been retrieved.
  • Ask who the tactic is for. In the founding paper’s own table the three strongest rewrites cost an already first-ranked source 20 to 30% of its visibility while roughly doubling a fifth-ranked one.2
The root cause

Why does so much GEO advice go wrong?

Because the stage that is easiest to act on is also the stage that is easiest to measure in isolation, and the two facts compound into an industry. Rewriting a page is something one person can do in an afternoon without touching infrastructure, so that is where the advice concentrated. Measuring a rewrite is easy if you hand the model a fixed set of documents and vary only the wording, so that is where the early experiments concentrated. The result is a large body of genuine findings about what happens to a document that is already in the model’s context, generalised into claims about what happens to a page on the open web.

The 2026 critical survey of the field draws the line in its own confidence table. That a document already placed in the context can causally alter its rank, citation or use is rated high confidence, annotated with the note that this evidence “does not address organic retrieval”. That a white-hat intervention durably improves organic discoverability across multiple engines is rated low. That citation scores predict clicks, conversions or revenue is rated very low.3 Almost every disappointment in this field lives in the gap between the first row and the other two. The how AI search works stage is the pipeline this diagnostic sorts against.

The diagnostic

How do you sort a GEO claim by the stage it acts on?

Sort a GEO claim by asking which of the five pipeline stages it acts on, whether anything you own is an input at that stage, what evidence exists for the effect, and what the effect size is.

Ask those four in order. Most claims stop at the first or second.

01

Which stage does it act on?

If the honest answer is interpretation or fan-out, there is no mechanism: those stages run before any document is touched and take no input from your site. If the answer is retrieval, the evidence is strongest. If it is synthesis, the evidence is weakest and the tactic can cost you upstream.

02

Were the earlier stages running when it was measured?

If the study placed the document in the model’s context and then varied the wording, the result is real and cannot speak to whether the page would have been retrieved. That single design choice separates most of this field’s optimistic results from its pessimistic ones.

03

What is the denominator, and the date?

A percentage in this field is unreadable without both. Per-answer and per-citation shares differ by tens of points on the same data, and these systems change month to month: one engine’s citation rate for a single large forum moved roughly 50 points inside one month in 2025.

04

Does it still work once everyone does it?

A NeurIPS 2025 benchmark found gains declining as adoption rises, describing the problem as “congested and zero-sum”.4 Anything that works because few people do it has an expiry date, and the advice you are reading is what brings it forward.

The reversal

What did the end-to-end benchmark actually show?

The end-to-end benchmark showed that the same rewrites that win once a document is in context can lose the document its place in the context. A KDD 2026 benchmark paper built an environment of 171,003 documents and 2,700 queries with retrieval and reranking reinstated, rather than fixing the candidate set in advance as earlier benchmarks did, and measured body-only optimisation reducing average top-20 presence by about 9%, top-10 presence after reranking by 16%, and final citation by about 6%. Applying an automated optimiser to the body alone produced larger losses.1

The 2026 critical survey’s reading dissolves the apparent contradiction with the KDD 2024 results rather than picking a side: a rewrite “may therefore perform well once injected while making the document less retrievable or less competitive upstream.”3 Both sets of numbers are honest measurements of different things, and the composition is what you actually experience.

Be precise about what this does not show. It is a single benchmark, accepted at KDD 2026 and built by one team. It tests body-only optimisation, not every possible intervention. And a controlled corpus is not the live web. What it establishes is narrower and still decisive: a content tactic evaluated only after retrieval has not been evaluated, and at least once, when someone checked, the total came out negative.

The rank-1 penalty

Do GEO tactics hurt you if you already rank first?

Yes. In the KDD 2024 results the three best-performing rewrite tactics cut visibility for an already first-ranked source by 20–30% while roughly doubling it for the fifth-ranked one, so the same tactic is a gain for a challenger and a cost for a leader.

The table shows relative change in visibility by the source’s original Google rank, from the KDD 2024 paper’s own results table. Almost nobody quotes this table.2

Rewrite tacticIf the source already ranked firstIf it ranked fifth
Cite sources−30.3%+115.1%
Quotation addition−22.9%+99.7%
Statistics addition−20.6%+97.9%
Authoritative tone−6.0%+6.1%
Fluency optimisation−2.0%+2.2%

These tactics redistribute share of the answer toward lower-ranked sources rather than creating visibility from nothing: a levelling effect, not a universal lift, and it changes who should adopt them. If you are the challenger at position five, the upside in this measurement is large. If you already own the top result for a query, aggressive rewriting is a downside risk on your best asset, and the same paper’s figures say so.

The practical rule is unglamorous: test on pages you are not already winning with, keep changes small enough to reverse, and measure presence and citation separately over repeated runs rather than reading a single before-and-after screenshot as a result.

The audit

Which popular claims fail this test outright?

Four of the most repeated GEO claims fail the stage test outright: a headline percentage lift, a single-platform share of all AI citations, schema markup as a citation lever, and llms.txt as a visibility file. Each is stated in full, with its correction, below.

Fails the stage test

  • “GEO lifts visibility by 40%” is rejected by the 2026 critical survey as a general claim: it is a relative maximum on one metric, under one configuration, for a page already in context
  • “Schema markup increases AI citations” fails: the only matched-control test, 1,885 pages against 4,000 controls, measured −4.6% on AI Overviews, and Google states no special schema.org data is needed
  • “Publish an AI text file so engines can find you” fails: Google states Search itself does not use such files, and a 137,210-domain study found 97% received zero requests
  • “One forum is 40.1% of AI citations” traces to a chart of per-response presence rates whose published values sum to roughly 208%; they were never shares
  • “X% of AI citations come from the organic top 10” is unreadable without a denominator and a date; published figures span roughly a sixth to over three quarters
  • “No AI crawler renders JavaScript” is single-sourced to one December 2024 log study, never replicated, and false for Google, Applebot and agentic browsers
  • “Optimise the page for the AI’s keywords” fails: keyword stuffing is the only tactic in the KDD 2024 results that scored below doing nothing at all

Survives the stage test

  • Be indexed and snippet-eligible: the engine’s own stated requirement, stage 3
  • Return your text without JavaScript: free, and removes an engine-specific risk at stage 3
  • Check the search-bot token, not the training token: every vendor documents the split
  • Be genuinely relevant to a real sub-question: the strongest lever in the survey
  • Answer that sub-question in the first sentence, self-contained: stage 4
  • Put real, dated, attributed numbers and flat definitions in the text: stages 4 and 5
  • Report per engine, over repeated runs, and separate retrieval from citation from accuracy
What is left

What survives the test?

What survives is a short, boring list that the 2026 critical survey states in one sentence: “produce a relevant, comprehensive, verifiable, clearly structured, and technically retrievable page; then measure retrieval, citation, and fidelity separately.”3 Every clause of that does work. Relevant is stage three. Comprehensive is stage two, because fan-out retrieves per sub-question. Verifiable and clearly structured are stage four. Technically retrievable is the gate in front of all of it. And measuring the three outcomes separately is what stops you from optimising the wrong one for a quarter.

Google’s own documentation says something compatible and even shorter, which is worth holding onto when a vendor tells you otherwise: to appear as a supporting link in its AI features, “a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements”, and “you don’t need to create new machine readable files, AI text files, or markup to appear in these features.”5 Everything sold above that line should be able to name the stage it acts on and the study that measured it there.

The honest limit of this article

The diagnostic sorts claims by mechanism, not by whether they are true. A claim can name the right stage and still be wrong, and a claim can fail the test and still describe something real that nobody has measured properly yet, because absence of evidence in a field this young is weak evidence of absence. Two of the debunkings above rest on single studies, which is exactly the sin they are debunking: the schema result is one matched-control test on pages that were already heavily cited, and the JavaScript conclusion is one log study from December 2024. These systems also change monthly, so a null result with a date on it is a null result at that date. Re-check anything here that is more than six months old before you act on it.

Where a product fits, and where it does not

The diagnostic itself is the deliverable, and it is free: take the last five pieces of GEO advice you were given, write the stage number next to each one, and delete the ones that cannot name a stage. Then run the survivors as small reversible experiments on pages you are not already winning with, measuring over repeated runs rather than a single screenshot. Bavior helps only with that last measurement problem: it runs a fixed prompt set across five engines on a schedule and records which sources each answer cited, so a change shows up as a shift in a distribution rather than as an anecdote, and where a cited source is a live discussion thread it drafts a reply on an account you control, which you approve before anything posts. It does not audit GEO advice for you, does not rewrite pages, and cannot promise a citation or a ranking. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).

Sources: all checked 29 Aug 2026
  1. Kim et al., “SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization”, KDD 2026, arXiv:2602.12187. arxiv.org/abs/2602.12187
  2. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, Deshpande, “GEO: Generative Engine Optimization”, KDD 2024, arXiv:2311.09735 (relative change by original rank, and the keyword-stuffing result). arxiv.org/abs/2311.09735
  3. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (confidence table, and the rejection of the 40% generalisation). arxiv.org/abs/2607.14035
  4. Puerto, Gubri, Green, Oh, Yun, “C-SEO Bench: Does Conversational SEO Work?”, NeurIPS 2025 Datasets & Benchmarks Track, arXiv:2506.11097. arxiv.org/abs/2506.11097
  5. Google Search Central (first-party). “AI features and your website”, last updated 10 Dec 2025, source of both quotations here: developers.google.com/search/docs/appearance/ai-features. “Optimizing your website for generative AI features on Google Search”, last updated 10 Jul 2026: developers.google.com/search/docs/fundamentals/ai-optimization-guide
  6. Structured-data matched-control study, May 2026: 1,885 pages that added JSON-LD between Aug 2025 and Mar 2026, matched against 4,000 control pages, citations measured 30 days before and after; AI Overviews −4.6% (significant), other surfaces not significant. Vendor-published; described rather than linked, per this curriculum’s sourcing rule.
  7. AI text file study, Jun 2026: 137,210 domains; 97% of such files received zero requests and no bots probed for files that did not exist. Vendor-published; described rather than linked, per this curriculum’s sourcing rule.
  8. Most-cited-domains study, Nov 2025: 230,000 prompts, 100 million-plus citations, Jul–Oct 2025; one engine’s citation rate for a single large forum moved roughly 50 points within one month, and the widely quoted 40.1% figure traces to per-response presence rates that sum to roughly 208%. Vendor-published; described rather than linked, per this curriculum’s sourcing rule.
FAQ

Frequently asked questions.

Is "GEO increases visibility by 40%" true?

The 2026 critical survey of the field rejects it as a general claim, stating that the figure is "a relative maximum on one metric under a specific configuration". The underlying result is a rise from 19.3 to 27.2 on Position-Adjusted Word Count, a measure of what share of an answer's words are credited to your source, inside a purpose-built engine that fetched the top five Google results and had GPT-3.5 write over them, with the optimised page already among those five. It is a real measurement of one stage. It is not a forecast of what a rewrite does to a page on the open web.

Does adding schema markup get me cited more by AI search?

The only matched-control test available says no, and the engine's own documentation agrees. That study tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 against 4,000 control pages that never added it, measuring citations 30 days before and after, and found AI Overview citations down 4.6% with other surfaces showing no significant change. Google states directly that "there's also no special schema.org structured data that you need to add" to appear in its AI features. Ship schema for rich results and machine-readability hygiene; do not treat it as an AI-citation lever.

If I already rank first, should I apply GEO rewrites to that page?

Be careful. The founding paper's own results table shows these tactics measurably reducing the share of the answer credited to a source that already ranked first: cite sources −30.3%, quotation addition −22.9%, statistics addition −20.6%, against roughly a doubling for a source that ranked fifth. The tactics redistribute answer share toward lower-ranked sources, so they are a levelling effect rather than a universal lift. Test on pages you are not already winning with, keep changes small enough to reverse, and measure over repeated runs rather than a single before-and-after check.

How do I tell whether a GEO study is worth acting on?

Ask four questions in order. Which pipeline stage does the claim act on? If it is interpretation or fan-out, there is no mechanism, because those stages take no input from your site. Were retrieval and reranking running when it was measured, or was the document placed in the model's context in advance. What is the denominator and the date, since per-answer and per-citation shares differ by tens of points and these systems change monthly. And does the effect survive widespread adoption? A NeurIPS 2025 benchmark found gains declining as more actors adopt the same tactic, calling the problem "congested and zero-sum".

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Name the stage, or drop the tactic.
Then measure what actually moved.

Start free trial