Home/Learn GEO/Do the tactics work
GEO versus SEO · Evidence

Do GEO tactics actually work?

Three benchmarks have tested the question, each restoring a stage the previous one skipped. Read in order they tell a single story, and it is not the story the tactic lists tell.

On this page
Share this
Share on X Share on LinkedIn
The short answer

Some do, conditionally, and by much less than the tactic lists suggest. The 2024 KDD benchmark that named the field tested nine rewriting methods inside a synthetic engine, on a document already in the model's context: its best method moved a custom visibility metric from 19.3 to 27.2, while keyword stuffing scored 17.7, below doing nothing.1 The field's own 2026 survey rejects “GEO increases visibility by 40%” as a general claim, calling the figure “a relative maximum on one metric under a specific configuration.”4

Key takeaways
  • The founding benchmark never tested retrieval: the optimised page was already one of five documents handed to the model, so what it measured was attribution inside a fixed context.1
  • Both benchmarks that put retrieval back found the effect shrink or reverse: body-only rewriting cut post-rerank top-10 presence by 16% and final citation by about 6% across 171,003 documents.3
  • Almost all the gain went to sources that started low: citing sources, quotations and statistics cut the rank-1 source's share by 20 to 30%, so a page that already wins the query is the one with something to lose.
The founding study

What did the founding GEO benchmark actually measure?

It measured how much of a generated answer's text gets attributed to a document that was already inside the model's context, not whether the document was retrieved, and not clicks, rankings or traffic.

Method testedPosition-Adjusted Word CountAgainst a 19.3 baseline
No optimization (baseline)19.3n/a
Keyword stuffing17.7Worse than doing nothing
Unique words20.5Marginal
Authoritative tone21.3Small
Easy-to-understand22.0Small
Technical terms22.7Small
Cite sources24.6Fourth best
Fluency optimization24.7Third best
Statistics addition25.2Second best
Quotation addition27.2Best of the nine

Almost nobody who quotes the result knows what the metric is. Position-Adjusted Word Count is the share of the generated answer's words credited to your source, discounted by an exponentially decaying function of where your citation appears. It measures how much of the text is yours, not whether anyone clicked.1 The benchmark is large, 10,000 queries from nine public sources across 25 domains, which makes the design problem in the next section a flaw rather than a limit of scale.

The caveats

Why does the headline percentage not mean what it is used to mean?

The headline number is the relative distance from 19.3 to 27.2 on one custom metric in one artificial setup, and the field's own 2026 survey rejects it: it lists “GEO increases visibility by 40%” under claims rejected as general, on the grounds that the figure is “a relative maximum on one metric under a specific configuration.”4 Three caveats travel with it and are almost never repeated alongside it.

The first is the engine. The main experiment did not query ChatGPT, Google or Bing: the authors built a synthetic engine that fetched the top five Google results and had a 2023-generation model answer from those five sources.1 The second follows: the optimised page was already one of those five, so retrieval was never part of the test. The survey grades the finding accordingly, rating it high confidence that a document already placed in the context can causally alter its rank, citation or use, while noting in the same row that this evidence “does not address organic retrieval.”4 The third sits in the same results table: keyword stuffing finished below the do-nothing baseline, the only method in the set to do so, and it has since been graded null or negative across multiple benchmarks.4

Two details from the paper itself sharpen the picture. Its live-engine check ran on 200 examples only, by the authors' own footnote, and there the best rewrite gained 21% while keyword stuffing lost 9% against baseline.1 And the paper's illustrative output for its “cite sources” method attributes a claim to a research organisation that does not appear to exist, which makes the tactic as operationalised “add citation-shaped text” rather than “add true citations”. The survey's warning is the one to keep: a fabricated statistic may increase reuse while degrading epistemic quality.4 Engines already misattribute at scale, with a Tow Center audit of 1,600 queries across eight engines finding more than 60% returned incorrect answers, so a citation-shaped invention is a liability that compounds.6

Redistribution

Who actually gains when one of these tactics works?

Almost all of the gain goes to sources that started low: citing sources, adding quotations and adding statistics roughly doubled visibility for the fifth-ranked source while cutting the first-ranked source's share by 20–30%.

Relative change in visibility by the source's original Google rank, from the same paper's second table.

MethodRank 1Rank 2Rank 3Rank 4Rank 5
Authoritative tone−6.0%+4.1%−0.6%+12.6%+6.1%
Fluency optimization−2.0%+5.2%+3.6%−4.4%+2.2%
Cite sources−30.3%+2.5%+20.4%+15.5%+115.1%
Quotation addition−22.9%−7.0%+3.5%+25.1%+99.7%
Statistics addition−20.6%−3.9%+8.1%+10.0%+97.9%

Read that as redistribution rather than creation. These rewrites do not enlarge the amount of answer text available to be attributed; they change who inside a fixed candidate set gets attributed, and the mechanism moves share from the incumbent to the challengers. Which makes the tactic's value a function of where you already stand: at rank 1, aggressive answer-shaped rewriting is a measurable downside risk rather than an upside bet.

The same paper's third table adds the other half of the qualification: the best method depends on the domain, citation-adding on factual and legal-government queries, quotation-adding on explanation and history, statistics-adding on debate and opinion.1 No single ordering held across domains, the property the 2026 survey later graded when it rated fixed formatting recipes as generalising poorly.4

The reversal

What happened when retrieval and reranking were put back?

Both benchmarks that restored a stage the founding paper had skipped found the effect shrinking or reversing. C-SEO Bench, presented at the NeurIPS 2025 Datasets and Benchmarks Track, tested ten conversational-SEO methods across six domains on more than 1.9k queries, and reported that most such methods are “not only largely ineffective but also frequently have a negative impact on document ranking”: 3 of 54 cases significantly positive, none in question answering, with moving a source to context position one worth 2.77 rank places in retail against 0.36 for the best rewrite.2 The sibling article on relevance and ranking goes through that benchmark in detail.

SAGEO Arena is the one that closes the arc, because it reinstated retrieval and reranking rather than supplying a fixed context. Over 171,003 documents and 2,700 queries it found body-only optimisation reducing average top-20 presence by about 9%, post-rerank top-10 presence by 16%, and final citation by about 6%.3 The survey's reading is the sentence to keep: a rewrite may perform well once injected while making the document less retrievable or less competitive upstream.4

Put the three results side by side and the contradiction dissolves. The founding paper measured the last stage of the pipeline in isolation and found the rewrites help there; the later work measured all the stages composed and found the same rewrites cost more upstream than they gain downstream. Which is why the practical rule is a constraint rather than a tactic: never optimise the body in a way that dilutes the page's topical relevance to the question it is trying to be retrieved for.

The synthesis

Which tactics survive all three studies?

Two survive with real support, genuine relevance and truthful extractable evidence, while fixed recipes, authoritative tone and keyword stuffing do not.

Grades below are the 2026 critical survey's own, not ours.

LeverSurvey's gradeWhat the benchmarks add
Query–document relevanceStrong, in controlled settingsThe only lever that beat every rewrite tested
Position in contextStrong, once retrieved2.77 rank places in retail, best rewrite 0.36
Extractable evidence: figures, definitions, comparisonsModerate to strongMust be truthful, dated, attributed, intent-matched
Recency, prices, explicit datesModerateA 252,000-trial factorial found effects for prices and recent dates
Document structureModerate and heterogeneousTest without assuming the direction of effect
Fluency and simplificationWeak to moderate24.7 against a 19.3 baseline, in the synthetic rig only
Authoritative toneWeak and unstable21.3, and may conflict with credibility
Fixed formatting recipesPoor generalization3 of 54 cases significantly positive
Keyword stuffingNull or negative17.7 against a 19.3 baseline; the field's most replicated negative

One entry in that table is worth reading twice. Extractable evidence is the only content lever graded moderate-to-strong, and its conditions are binding rather than decorative: the criterion is not to add numbers but to provide relevant, verifiable, dated and properly attributed evidence.4 The only academic taxonomy of AI citations published so far points the same way descriptively, associating pages containing numbers with 61.6% higher mean influence and definition markers with 57.3% higher, while cautioning that it cannot separate cause from correlation.7

One popular tactic is absent from the table because the evidence points the wrong way. “Schema markup improves AI citations” is contradicted by the only matched-control test published: a vendor study of 1,885 pages that added JSON-LD, matched against 4,000 controls over 30-day windows, found AI Overview citations down 4.6% and significant, with no significant change on the two other platforms tested.11 Google's own documentation agrees that no special structured data is needed to appear in its AI features.5 Ship schema for rich results and machine-readability; claim nothing further for it.

Measurement

How would you test one of these tactics on your own site?

Not with one before-and-after screenshot. Three findings govern your own test: answers drift, surfaces disagree, and gains decay as competitors adopt the same method.

Start by fixing what you ask. An edit's effect has to be read against the engine's own variance, and a 2026 statistical treatment found many apparent differences sitting inside the measurement's noise floor,10 which makes one question asked once before and once after uninformative. Fix a prompt set, repeat it on a schedule, and record which sources each answer cited, so what you compare is a distribution rather than a screenshot.

Then fix where you ask. A SIGIR 2026 study of 11,500 queries measured URL-level Jaccard similarity of 0.11–0.18 between three answer surfaces belonging to one company,9 which is the survey's basis for stating that visibility is “indexed by engine and surface.”4 A result on one surface is not a result on the others, and a tactic that worked somewhere worked on one engine, in one domain, in one month.

Finally, expect the answer to decay. C-SEO Bench found gains shrinking as more documents in the same candidate set adopted the same method,2 so anything that tests positive this quarter is being measured against competitors who have not adopted it yet. Two things are worth holding fixed throughout: change one thing at a time, and keep the page's topical relevance intact.

The honest limit of this article

Nobody has run a controlled experiment on a live commercial engine from end to end. Every result above comes from a rig the researchers built: the founding engine was synthetic and ran on a 2023-generation model, and both later benchmarks supply their own retrieval rather than querying a production system. So “ranking dominates rewriting” is a strong inference from convergent evidence, not a coefficient measured in the wild, and the survey grading all of it is a preprint, though all three benchmarks are peer-reviewed. Treat the negative results as sturdier than the positive ones: they are the ones that replicated.

Where a product fits, and where it does not

Everything actionable on this page is free and takes an afternoon of editing: make the page genuinely the best answer to the question it targets, put real dated attributed numbers and plain definitions in the body, leave your already-first-ranking pages alone, and delete any keyword-stuffed paragraph you find. No software is involved in any of that. What software changes is only the measurement above: Bavior runs a fixed prompt set across five engines on a schedule and records which sources each answer cited. It does not improve your rankings, does not rewrite your pages, and cannot promise any tactic here will work on your domain. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 30 Aug 2026).

Sources, all checked 30 Aug 2026
  1. Aggarwal et al., “GEO: Generative Engine Optimization”, KDD ’24 (30th ACM SIGKDD, Aug 2024); GEO-bench, 10,000 queries, 25 domains. arxiv.org/abs/2311.09735
  2. Puerto et al., “C-SEO Bench: Does Conversational SEO Work?”, NeurIPS 2025 Datasets & Benchmarks Track; ten methods, more than 1.9k queries and 16k documents, 54 cases. arxiv.org/abs/2506.11097
  3. Kim et al., “SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization”, KDD 2026 (arXiv:2602.12187v2, 7 Aug 2026); 171,003 documents, 2,700 queries. arxiv.org/abs/2602.12187
  4. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035; confidence table and lever-support table. arxiv.org/abs/2607.14035
  5. Google Search Central, “AI features and your website” (no special structured data required; first-party). developers.google.com/search/docs/appearance/ai-features
  6. Jaźwińska & Chandrasekar, “AI Search Has a Citation Problem”, Tow Center for Digital Journalism, Columbia, 6 Mar 2025; 1,600 queries across eight engines. cjr.org
  7. Zhang, He & Yao, “From Citation Selection to Citation Absorption”, arXiv:2604.25707 (preprint, descriptive statistics only); 602 controlled prompts, 21,143 search-layer citations. arxiv.org/abs/2604.25707
  8. Vishwakarma et al., 2026; 252,000-trial factorial finding effects for explicit prices and recent dates but weak effects for formatting changes alone (reported in the critical survey at note 4)
  9. Grossman et al., SIGIR 2026; 11,500 queries; URL-level Jaccard 0.11–0.18 across Google organic, AI Overviews and Gemini. arxiv.org/abs/2604.27790
  10. Sielinski, “Quantifying Uncertainty in AI Visibility”, Mar 2026, arXiv:2603.08924 (preprint). arxiv.org/abs/2603.08924
  11. Vendor matched-control study of 1,885 pages that added JSON-LD between August 2025 and March 2026, against 4,000 control pages, citations measured 30 days before and after, published May 2026; AI Overview citations −4.6% (significant), the two other platforms non-significant. Described rather than linked, per this curriculum's sourcing rule.
FAQ

Frequently asked questions.

Is the famous "40% lift" from the GEO paper real?

It is a real number from a real paper and it does not generalise, which is why the 2026 critical survey lists it among claims rejected as general. The figure is the relative distance from a 19.3 baseline to 27.2 on Position-Adjusted Word Count, a custom metric of how much of an answer's text is credited to your source, measured inside a synthetic engine built from five Google results plus a 2023-generation model, with the optimised document already in the context. Quote it only with those conditions attached, or not at all.

Should I add statistics and quotations to my pages?

Add real ones, with real sources and real dates, and expect a moderate rather than a dramatic effect. Extractable evidence is the only content lever the 2026 survey grades moderate-to-strong, and its conditions are binding: the criterion is not to add numbers but to provide relevant, verifiable, dated and properly attributed evidence. The failure mode this tactic invites is exactly the one the founding paper's own worked example fell into, inventing a citation-shaped attribution to an organisation that does not exist. Engines cannot check it and readers eventually can.

My page already ranks first. Should I still rewrite it for AI answers?

The measured effect on an already-first-ranked source is negative, so treat aggressive answer-shaped rewriting as a downside risk there. The founding paper's per-rank table shows citing sources, quotations and statistics reducing visibility for the rank-1 source by 20–30% while roughly doubling it for the rank-5 source; the mechanism is redistribution inside a fixed candidate set rather than growth. Edit those pages for accuracy, freshness and clarity, and leave the structural gymnastics for pages that are not already winning the query.

If most tactics do not work, what is left to do?

Three things, in order: be genuinely the best answer to the question, earn the ranking that puts you high in the retrieved context, and put truthful dated attributed evidence in the body. Those are the only levers the 2026 survey grades above weak, and the first two are ordinary search work rather than anything new. What is left after that is off-site: on commercial questions much of what gets cited is third-party coverage and community threads rather than any vendor's own pages, which is a different budget line and a different team.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

The benchmarks are public.
Read them before you buy a tactic list.

Start free trial