Home/Learn GEO/Stage 5
How AI search works · Stage 5

What happens during synthesis and attribution, the last stage of AI search?

This is the only stage where wording changes anything, which is why almost all GEO advice is aimed here, and why almost all of it is measured in a setting that skips the four stages before it.

On this page
Share this
Share on X Share on LinkedIn
The short answer

Synthesis and attribution is the stage where a model writes the answer over the passages selected at stage four, then attaches source links to some of its claims. Attaching a link is not the same as checking a sentence against the page behind it: a 2023 evaluation of generative search engines found that on average, a mere 51.5% of generated sentences are fully supported by citations.5

Key takeaways
  • A citation is a link attached to a claim; a mention is your brand named in the answer text. They come apart constantly, so keep two columns rather than one visibility score.
  • Being cited is not being described correctly. A 2026 audit of Google AI Overviews found 11.0% of 98,020 atomic claims unsupported by the pages cited beside them, with omission the larger failure mode.8
  • Wording acts here and only here, on material that already survived retrieval, ranking and selection. Every clean measurement of the stage put the page in the model’s context first, so none of it describes the open web.
  • Write each sentence so it stays true with no neighbours: subject, qualifier, figure and date inside the sentence itself.
  • A fabricated statistic scores the same as a real one on every published metric, because nothing in the pipeline checks. That is an argument for being the source that survives checking.
Definition

What happens at synthesis and attribution?

The model receives the shortlist of selected passages, writes an answer over them, and then attaches links from claims back to sources. Those are two separable operations and it matters that they are separable, because the writing does not have to be a faithful summary of the sources for a link to be attached to it. Attribution in these systems is a post-hoc association between a generated sentence and a document that was in context, not a verified derivation of the sentence from the document.

That is the whole basis of the accuracy problem below. It also explains the odd asymmetry that trips up brand teams: the engine can name a competitor in prose while linking to your page, or link your page while naming nobody, because the two outputs are produced by different mechanisms with no requirement that they agree. The four stages that decide what reaches this one are in the how AI search works stage.

Two outcomes

What is the difference between a mention and a citation?

A citation is a link the engine attaches to a claim; a mention is your brand name appearing in the answer’s prose. They come apart constantly, and only one of them survives being read aloud. A June 2026 vendor study covering 115 prompts, 3,981 domain appearances, four platforms and 14 countries found roughly 62% of source appearances were links whose brand the answer never named, and reported two engines with nearly inverted profiles: one naming brands in about 84% of appearances while linking in about 21%, the other naming in about 21% while linking in about 87%.6 The sample is small; the direction is the finding, and the decimals are not.

The practical consequence is that “AI visibility” needs two columns, not one. A link with no mention delivers referral traffic and nothing to a reader who never clicks. A mention with no link delivers recall and no traffic. Which one you need depends on whether you are trying to be discovered or to be remembered, and a report that collapses them into a single score has thrown away the distinction that determines what to do next.

Fidelity

Does being cited mean being described accurately?

No, and the measured error rate is high enough to change how you write. The Tow Center for Digital Journalism at Columbia tested 200 quotes from 20 publishers against ChatGPT Search in November 2024 and found 153 of 200 responses partially or entirely incorrect, with uncertainty signalled only 7 times.3 A follow-up in March 2025 ran 1,600 queries across eight engines, 20 publishers by 10 articles by 8 assistants, and found more than 60% of those queries answered incorrectly, with the worst assistant wrong 94% of the time, producing 154 citations that resolved to error pages across its 200 prompts. Licensing agreements conferred no accuracy benefit, and publishers who had blocked the crawlers were still cited.4

Independent work points the same way from other angles. A 2023 evaluation of generative search engines found that on average, a mere 51.5% of generated sentences are fully supported by citations, and only 74.5% of citations support their associated sentence.5 A 2026 audit decomposed 7,491 Google AI Overviews into 98,020 atomic claims and checked each against the pages cited beside it: 11.0% came back unsupported, of which 4.1% were directly contradicted by or in conflict with the cited content and a further 7.0% were not addressed in the cited source text at all.8 Omission, not contradiction, is the larger failure: the common case is not an engine arguing with your page but an engine crediting it with something it never said. A separate 2026 study measured credible-source shares of 71.4% to 86.3% depending on the assistant and the topic.8 Different methods, different denominators, one conclusion: the link being attached to a sentence tells you almost nothing about whether the sentence is true of the source.

So optimising to “get cited” is not the same as optimising to be represented correctly, and the second is the one that protects you. A sentence that is only true in the presence of its neighbours will be lifted without them.

The evidence

Why does wording only act at this stage?

Wording acts only at synthesis because every earlier stage has already decided which passages are in the room; rewriting changes how much of the answer a passage wins, never whether it was retrieved.

The table below gives the KDD 2024 results, in the paper’s own metric. Baseline “no optimization” scores 19.3 on Position-Adjusted Word Count, the share of the answer’s words credited to your source, discounted if you are cited late.1

Rewrite tacticPAWC in a purpose-built engineWhat it is evidence of
No optimization19.3The baseline
Keyword stuffing17.7Worse than doing nothing
Cite sources24.6Downstream effect only
Statistics addition25.2Downstream effect only
Quotation addition27.2Best of nine, downstream effect only

Every number in that table was produced with the optimised page already inside the engine’s context. The experiment fetched the top five Google results for a query and had GPT-3.5 write an answer over those five documents, with the page under test among them.1 That design is exactly what makes it a clean measurement of stage five, and exactly what makes it silent about stages two through four. The 2026 critical survey rates the causal claim it supports as high confidence while noting that this evidence “does not address organic retrieval”, and separately rejects the popular generalisation of the paper’s headline as “a relative maximum on one metric under a specific configuration”.2

The composition is where it gets expensive. A 2026 KDD benchmark that reinstated retrieval and reranking over 171,003 documents and 2,700 queries measured body-only optimisation reducing average top-20 presence by about 9%, top-10 presence after reranking by 16%, and final citation by about 6%.7 A stage-five gain bought with a stage-three loss is a loss. Wording is the last lever to pull and the one to pull most gently.

The dark twin

What failure mode does this stage invite?

Fabricated attribution, meaning text that looks like a citation with no real source behind it, is the specific temptation stage five creates, and the founding paper of this field demonstrates it by accident. Its illustrative output for the “cite sources” tactic credits a survey conducted by an organisation that does not appear to exist, and scores that example at a 132.4% relative improvement.1 The tactic as operationalised there is “add citation-shaped text”, not “add true citations”, and the metric cannot tell the difference because no part of the pipeline checks.

The 2026 survey names the problem in one sentence: “adding a fabricated statistic may increase reuse while degrading epistemic quality. The criterion is therefore not to ‘add numbers,’ but to provide relevant, verifiable, dated, and properly attributed evidence.”2 It also draws the boundary in a form you can apply without a lawyer. Its tests are whether the commercial intent is disclosed and whether competitors are represented without fabricated disparagement, and on that reading reorganising your paragraphs or adding a verified primary source passes, while a planted testimonial or an instruction written to make a model favour you does not.2

The asymmetry is what should decide this for you. A real number with a named, dated, linked source earns the same extraction benefit as an invented one and carries none of the liability, and in a field where more than 60% of queries in a 1,600-query test across eight engines already returned incorrect answers, being the source that survives checking is a durable position rather than a tactic.4

Application

So what should you actually write?

Write every sentence you would like quoted so that it is still true with no surrounding context. Bake the subject, the qualifier and the date into the sentence itself: not “this rose 30% last year” but “on average, 53% of the domains a Google AI Overview consults do not appear in the organic top 10, measured across 4,706 queries in a study published in Findings of ACL 2026”.9 The second version survives being lifted; the first becomes a claim you did not make, attributed to you, in an answer you cannot edit.

The survey’s own summary of what can reasonably be recommended is shorter than most GEO checklists: “produce a relevant, comprehensive, verifiable, clearly structured, and technically retrievable page; then measure retrieval, citation, and fidelity separately.”2 Separately is the operative word, because a single blended visibility number cannot tell you which of the three moved. The same survey asks that audits pair citation rates with a matrix covering tone, attribution accuracy and factual support, which is the three-column habit this stage rewards.

Then keep the stage-five work proportionate. Its evidence base is the narrowest in the pipeline, its effects are measured downstream-only, and the survey rates authoritative tone as weak and unstable, while rating extractable evidence, meaning real figures, definitions and comparisons, as moderate to strong, conditional on being truthful, attributed and intent-matched.2 That is the whole of the defensible advice for this stage. Everything else sold under the heading of writing for AI is either a restatement of it or is not supported.

The honest limit of this article

Every measured stage-five result in this article comes from a setting where the document was already in the model’s context, because that is the only way anyone has found to isolate the stage. None of these numbers tells you what happens to a page on the open web, and the one benchmark that reinstated the earlier stages found the composition negative. The metric itself is also not the outcome you care about: it counts share of the answer’s words, and the 2026 critical survey rates the claim that citation scores predict clicks, conversions or revenue at very low confidence. The mention-versus-citation split rests on a 115-prompt vendor sample, which is enough to show the gap exists and not enough to size it.

Where a product fits, and where it does not

Auditing this stage needs no software, only patience: ask your buyer questions on each engine, then for every answer record three separate things, namely whether your page was linked, whether your brand was named in the prose, and whether what the answer said about you was actually true. The third column is the one nobody keeps, and it is the one that catches a misdescription before a customer does. Bavior automates the collection, not the judgement: a fixed prompt set across five engines on a schedule, with the cited sources recorded per run, and where a cited source is a live discussion thread it drafts a reply on an account you control, which you approve, edit or reject before anything posts. It does not write your pages, does not verify what an answer said about you, and cannot correct an engine that describes you wrongly. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).

Sources, all checked 30 Aug 2026
  1. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, Deshpande, “GEO: Generative Engine Optimization”, KDD 2024, arXiv:2311.09735: arxiv.org/abs/2311.09735
  2. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (survey preprint): arxiv.org/abs/2607.14035
  3. Tow Center for Digital Journalism, Columbia, “How ChatGPT Search (Mis)represents Publisher Content”, Nov 2024: cjr.org
  4. Jaźwińska & Chandrasekar, Tow Center, “AI Search Has a Citation Problem”, 6 Mar 2025; 1,600 queries across eight engines: cjr.org
  5. Liu, Zhang, Liang, “Evaluating Verifiability in Generative Search Engines”, 2023; 51.5% of generated sentences fully supported by citations, 74.5% of citations supporting their sentence: arxiv.org/abs/2304.09848
  6. Ghost-citation study, Jun 2026: 115 prompts, 3,981 domain appearances, 4 platforms, 14 countries; ~62% of citations were links whose brand was never named in the answer text. Vendor-published; described rather than linked, per this curriculum’s sourcing rule.
  7. Kim et al., “SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization”, KDD 2026 (arXiv:2602.12187v2, 7 Aug 2026); 2,700 queries over 171,003 documents: arxiv.org/abs/2602.12187
  8. Xu, Iqbal, Montgomery, “Measuring Google AI Overviews”, 2026 (preprint); 11.0% of 98,020 atomic claims unsupported by the cited pages: arxiv.org/abs/2605.14021. The credible-source shares of 71.4–86.3% are Vykopal et al., 2026, as reported in the critical survey at note 2.
  9. Kirsten, Große Perdekamp, Wu, Upadhyay, Gummadi, Zafar, “Characterizing Web Search in The Age of Generative AI”, Findings of ACL 2026, pp. 10827–10848; 4,706 queries, 53% of AI Overview domains outside the organic top 10: aclanthology.org/2026.findings-acl.526
FAQ

Frequently asked questions.

Why does an AI answer link my site but never say my brand name?

Because the answer text and the attached links are produced by different operations, with no requirement that they agree. A June 2026 vendor study of 115 prompts and 3,981 domain appearances across four platforms found roughly 62% of source appearances were links whose brand the answer never named in prose, and that engines differ enormously: one named brands in about 84% of appearances while linking in about 21%, another was almost the reverse. Track links and named mentions as two separate columns, because they deliver different things: one brings traffic, the other brings recall.

If an AI engine cites my page, does that mean it described me correctly?

No: attachment is not verification, and the measured error rates are high. The Tow Center at Columbia found 153 of 200 ChatGPT Search responses partially or entirely incorrect in November 2024, and in a March 2025 follow-up of 1,600 queries across eight engines more than 60% returned incorrect answers, with one assistant wrong 94% of the time. A 2023 evaluation found only 51.5% of sentences in generative-search answers were fully supported by their attached citations. The defence is writing each sentence so it stays true when quoted with no surrounding context.

Does adding statistics and quotes to my page increase AI citations?

In the one setting where it has been measured cleanly, yes, but that setting skipped retrieval entirely. The KDD 2024 paper scored quotation addition at 27.2 and statistics addition at 25.2 on its Position-Adjusted Word Count metric against a 19.3 baseline, with the optimised page already inside the engine's five-document context. A 2026 benchmark that reinstated retrieval and reranking found body-only optimisation cutting top-20 presence by about 9% and final citation by about 6%. Real, dated, attributed figures are worth adding; a rewrite that dilutes the page's topical relevance is not.

Is it worth writing in an authoritative tone for AI search?

The 2026 critical survey rates authoritative tone as weak and unstable support, and warns it may conflict with credibility, which makes it one of the lowest-rated levers in the field. What the same survey rates moderate to strong is extractable evidence: real figures, explicit definitions and clear comparisons, conditional on being truthful, attributed and matched to the query's intent. In practice that means a flat sentence with a named source and a date beats a confident sentence with neither, and it also survives the accuracy checks that confident phrasing does not.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Cited is not the same as described correctly.
Check both.

Start free trial