Home/Learn GEO/Do engines agree
How AI search works · Topic

Do AI engines agree on who to cite? Three academic audits say mostly not.

Two assistants asked the same question return largely different sources, and so do two surfaces built by the same company on the same index. That single fact decides how AI visibility has to be measured and reported.

On this page
Share this
Share on X Share on LinkedIn
The short answer

AI engines mostly do not agree on which sources to cite, so a citation on one engine is weak evidence about any other. The starkest measurement comes from three surfaces built by one company on one index: an 11,500-query study presented at SIGIR 2026 reports average URL-level Jaccard similarity below 0.2 across Google organic results, AI Overviews and Gemini.1

Key takeaways
  • Jaccard is the intersection over the union, not the share of your cited URLs. At 0.18, only 18% of the sources either surface returned were returned by both.
  • The divergence starts before content quality is judged: engines crawl under different tokens, expand the question into different sub-queries, and cite different numbers and types of source.
  • One blended AI visibility score hides the only thing that is actionable. Report one row per engine per surface, with its date and locale, and do no arithmetic across rows.
The measurements

How much do two AI engines actually overlap?

Two AI engines overlap far less than most people assume: two surfaces built by one company on one index reach a URL-level Jaccard of only 0.11 to 0.18, and two unrelated assistants shared 26% of their domains in the one audit that measured it.

Four independent academic measurements, each with a different method, denominator and collection window. Read the last two columns before the number.

StudySample and collection windowFindingWhat it measures
Grossman et al., SIGIR 202611,500 queriesURL Jaccard 0.11–0.18Set overlap between Google organic, AI Overviews and Gemini
Kirsten et al., Findings of ACL 20264,706 queries, collected Sep 202553% absent from organic top 10; 27% absent from top 100Per domain AI Overviews consults, against the same query's organic results
Xu, Iqbal & Montgomery, 2026 preprint55,393 queries, collected 13 Mar – 21 Apr 202629.8% absent from the organic first pagePer AI Overview reference domain, against the same query's first page
Li & Sinnamon, 2024672 responses, 355 unique domains26% cited by bothDomain overlap between two assistants

The 2026 critical survey draws the conclusion in a single sentence worth quoting exactly: these results “refute the notion of a global GEO ranking. Visibility is indexed by engine and surface.”4 There is no such object as your AI search position. There are five or six of them, only loosely correlated. Between two unrelated assistants the gap is wider still: the 2024 audit of Bing Chat and Perplexity found 74% of the domains it saw were cited by exactly one of the two.3

Reading the number

What does a Jaccard of 0.15 actually mean?

Jaccard similarity is the size of the intersection divided by the size of the union, which is not the share of your cited URLs, though it is routinely repeated as if it were. If AI Overviews cites ten URLs for a query and Gemini cites ten, an overlap of 0.15 means roughly three URLs appear on both lists and about seventeen distinct URLs exist across the two. Grossman et al. put the strongest pair in their own words: “on average, only 18% of the sources returned by either the AIO or traditional SERP will be retrieved by both search engines.”1

The three pairs are not equally bad, and the abstract does not tell you that. It reports only a “<0.2 average Jaccard similarity”; the pairwise values sit in Table 2, where AI Overviews against organic results is 0.18, organic against Gemini is 0.16, and AI Overviews against Gemini is 0.11.1 The pair most people treat as interchangeable, the two generative surfaces, is the pair that agrees least.

Repetition does not rescue the picture either. The critical survey summarises Schulte et al. as observing daily source-level Jaccard scores of approximately 0.34 to 0.42 across four engines over 45 days, with similar levels for repetitions inside 24 hours, and uses that spread to argue for seven to eight repetitions per prompt.11 Read the caveat with the number: the survey says it derives from a small universe of Swiss queries. Even so, the tightest comparison published does not reach one half.

Mechanism

Why do engines reading the same web disagree this much?

Engines disagree because four things differ before content quality is ever considered: how they crawl, how they expand the question, how many sources they cite, and which kinds they prefer.

Those four differences compound, and none can be fixed once for every engine.

01

They do not read the same web

Each engine crawls under its own tokens, documented separately by each vendor: OpenAI splits GPTBot from OAI-SearchBot, Anthropic splits ClaudeBot from Claude-SearchBot, Google splits Google-Extended from Googlebot.5 A robots.txt rule that blocks one leaves the others untouched, so the candidate pools differ before ranking begins.

02

They expand the question differently

Google documents that AI Overviews and AI Mode may use a query fan-out technique, issuing multiple related searches across subtopics and data sources to develop a response.6 Sub-queries are generated per engine, so two engines retrieve for different questions before either of them ranks anything.

03

They have different citation budgets

On 602 controlled prompts, mean citations per response were 6.88 for ChatGPT, 12.06 for Google's AI surfaces and 16.35 for Perplexity.7 An engine that cites 16 sources reaches further down its candidate list than one that cites 7, so the marginal source differs by construction.

04

They prefer different kinds of source

In the same dataset, sources the paper labels official were 34.22% of ChatGPT's citations against 46.35% of Google's, while news was 31.17% for ChatGPT against 18.99% for Google.7 A page that is the right type for one engine is the wrong type for another.

Those four differences multiply rather than average out, which is why the observed overlap is so low. It also explains a pattern most teams read as a mystery: a page can be cited constantly by one engine and never appear in another, with nothing wrong with the page. Before diagnosing content, check whether the engine that ignores you can fetch you at all. That is the subject of Technical GEO, and it is a different investigation from writing better answers.

The denominator problem

Does ranking on Google still predict being cited?

Ranking remains the strongest single predictor of being cited, and it is no longer sufficient on its own. You cannot express that as one overlap percentage, because the published figures answer two different questions. “What share of AI answers contain at least one page that also ranks?” and “what share of individual citations come from pages that also rank?” have wildly different answers on identical data. A commercial study of 362,000 US queries in October 2025 reported both: 94% of AI Overviews included at least one URL from the organic top 20, while 56% of individual citations came from the top 20.8 Same data, two numbers, both correct, thirty-eight points apart.

The academic figures sit in the same place once you read their denominators, and together they form a range with two dates on it rather than a point. Kirsten et al.'s 4,706-query audit, collected in September 2025 in the US and Germany, counts per domain consulted: “on average 53% (27%) of domains that AIO consults are not contained in top-10 (top-100) Organic search results.”2 Xu, Iqbal and Montgomery, collecting 55,393 trending queries between 13 March and 21 April 2026, report that 29.8% of AI Overview reference domains do not appear anywhere on the corresponding first page.10 Quote that as a range, roughly 30% to 53% with both windows named, and keep the verbs apart: Kirsten measures the domains the system consults, which is not identical to the domains it cites, and neither says whether an AI Overview usually contains something that ranks; it usually does.

The figures move fast: one provider published a per-citation overlap of 54.5% for September 2025 and roughly 17% for February 2026, same tool, same nine industries, no explanation for the gap.9 Treat any overlap number older than about six months as expired, and one with no date as fiction.

The honest limit of this article

Every overlap figure here is a snapshot of a moving system, and no academic study covers all the engines a founder cares about. The SIGIR work compares three Google surfaces, the ACL audit is Google-only, and the two-assistant audit measured Bing Chat and Perplexity as they were in 2024. Nobody has measured ChatGPT, Perplexity, Google, Copilot and Claude with one method in one window, so the general claim, that the engines disagree a lot, is well supported while no specific pairwise number transfers to your prompt set. Measure your own set rather than importing these.

Reporting

Why is a single blended AI visibility score meaningless?

Averaging across engines destroys exactly the information a decision needs, because the average of two loosely correlated series is not a summary of either. A brand cited in 40% of runs on one engine and 0% on another has the same blended score as one cited in 20% on both, and the two situations demand opposite work. The first is an access problem on one engine; the second is a weak-but-broad position everywhere.

Weighting makes it worse rather than better, because a citation is not one unit across engines. In the 602-prompt dataset, mean influence per cited page was 0.2713 for ChatGPT against 0.0584 for Google and 0.0646 for Perplexity, roughly a four-fold difference in how much of the answer one citation shapes.7 Adding a ChatGPT citation to a Perplexity citation adds two different units, and no single scalar fixes that.

A blended number has one narrow use: a coarse quarter-to-quarter trend line, where you only care whether the whole picture is moving. It cannot support prioritisation, competitive comparison, or a verdict on whether a specific change worked, which is what most people reach for it to do.

Method

How should you report AI visibility instead?

One row per engine per surface, and no arithmetic across rows. This runs in a spreadsheet.

Make the engine and surface part of every number

Not “AI visibility 18%” but “cited in 7 of 40 runs, Google AI Mode, en-US, August 2026”. AI Overviews and AI Mode are different surfaces and have to be separate rows, since one company's own surfaces overlap at Jaccard 0.11–0.18.

Track the cited URLs, not only your own presence

The set of sources each engine keeps citing is the map of who you are competing with there, and it differs per engine. On most commercial questions most of that list belongs to someone else.

Expect a fix to land unevenly, and say so up front

A change that raises citation rate on one engine may do nothing on another for months, or never. A per-engine target committed before shipping stops that being read as failure.

Pick the one or two engines your buyers use

Divergence means covering five engines equally costs five times as much for the same confidence. Sampling two properly beats sampling five badly, and the choice should follow where your traffic already comes from.

The reporting discipline above is the whole method, and it needs no software. The how AI search works stage supplies the pipeline stages that explain the divergence; the measuring AI visibility stage turns per-engine rows into a full measurement plan; the sibling article on why AI answers change covers the run-to-run variance that sits inside each row.

Where a product fits, and where it does not

You can run the per-engine table above by hand: fixed prompts, one tab per engine, a tally of runs and cited URLs, no averaging. Bavior automates that chore, running a fixed prompt set against five engines on a schedule and recording which sources each answer cited, per engine rather than blended. It does not produce a single visibility score, cannot tell you why two engines disagree about you, and has no influence over which engine your buyers use. The free AI visibility check and free GEO audit run without a paid plan; paid tracking starts from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).

Sources, all checked 30 Aug 2026
  1. Grossman et al., “How Generative AI Disrupts Search”, SIGIR 2026, 11,500 queries; abstract “<0.2 average Jaccard similarity”, Table 2 the pairwise 0.18, 0.16 and 0.11: arxiv.org/abs/2604.27790
  2. Kirsten et al., Findings of ACL 2026, 4,706-query audit of Google AI Overviews, data collected September 2025 in the US and Germany; “on average 53% (27%) of domains that AIO consults are not contained in top-10 (top-100) Organic search results”: aclanthology.org/2026.findings-acl.526
  3. Li & Sinnamon, “Generative AI Search Engines as Arbiters of Public Knowledge”, Proceedings of ASIST 61(1), 2024; 672 responses across Bing Chat and Perplexity, 355 unique domains, 26% cited by both: arxiv.org/abs/2405.14034
  4. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (survey preprint): arxiv.org/abs/2607.14035
  5. First-party crawler documentation on the training-token and search-token split: OpenAI developers.openai.com/api/docs/bots, Anthropic support.claude.com/en/articles/8896518, Perplexity docs.perplexity.ai/guides/bots
  6. Google Search Central, “AI features and your website” (query fan-out, first-party documentation, last updated 10 Dec 2025): developers.google.com/search/docs/appearance/ai-features
  7. Zhang, He, Yao, “From Citation Selection to Citation Absorption: A Measurement Framework for GEO Across AI Search Platforms”, 28 Apr 2026, arXiv:2604.25707 (preprint, open dataset of 602 prompts): arxiv.org/abs/2604.25707
  8. 362,000-query US desktop study, October 2025, reporting both 94% of AI Overviews containing an organic top-20 URL and 56% of individual citations coming from the top 20. Vendor-published; described, not linked, per this curriculum's sourcing rule.
  9. Per-citation organic-overlap figures of 54.5% (September 2025) and roughly 17% (February 2026), from one provider's same tool across the same nine industries, unreconciled. Vendor-published; described, not linked, per this curriculum's sourcing rule.
  10. Xu, Iqbal & Montgomery, 2026, 55,393 trending queries collected 13 March to 21 April 2026; “29.8% of AIO reference domains do not appear anywhere on the corresponding first page” (preprint): arxiv.org/abs/2605.14021
  11. Schulte et al., 2026, daily source-level Jaccard of approximately 0.34 to 0.42 across four engines over 45 days, on a small Swiss query universe; quoted as summarised in the critical survey at note 4.
FAQ

Frequently asked questions.

If I get cited by ChatGPT, will Perplexity cite me too?

Usually not, on the evidence available. An audit of 672 responses across Bing Chat and Perplexity found 355 unique cited domains, of which only 26% appeared on both. Even two surfaces built by one company on one index overlap at URL-level Jaccard 0.11 to 0.18, per an 11,500-query study presented at SIGIR 2026. Treat a citation on one engine as evidence about that engine only, and check the others separately rather than assuming they follow.

What is the real overlap between AI citations and organic top-10 rankings?

There is no single correct number, because the published figures answer two different questions. Measured per AI answer, meaning does this answer contain at least one page that also ranks, one 362,000-query commercial study found 94% of AI Overviews included a top-20 URL. Measured per citation, meaning what share of individual cited pages also rank, the same sample gave 56% from the top 20, and a 4,706-query academic audit found 53% of the domains AI Overviews consults absent from the organic top 10 on September 2025 data. A 55,393-query preprint collected in March and April 2026 measured 29.8% on the comparable first-page test, so quote the range rather than a point. Always ask which denominator and which collection window a number uses before repeating it.

Should I optimise for one AI engine or all of them?

Pick the one or two your buyers actually use, because divergence makes broad coverage expensive rather than thorough. Covering five engines to the same confidence costs roughly five times the sampling, and a fix that works on one surface frequently does nothing on another: one company's own organic, AI Overviews and Gemini surfaces overlap at Jaccard 0.11 to 0.18. The exception is technical retrievability, which is the one layer that pays off on every engine at once, since being indexable, fast and server-rendered helps everywhere.

Is one blended AI visibility score ever useful?

Only as a coarse quarter-over-quarter trend line, and never for a decision. A brand cited in 40% of runs on one engine and 0% on another scores the same blended number as one cited in 20% on both, while the two situations need opposite work. Weighting does not rescue it either: mean influence per cited page was 0.2713 for ChatGPT against 0.0584 for Google and 0.0646 for Perplexity in a 602-prompt academic dataset, so citations on different engines are not the same unit.

Why does an engine cite me on Google but never in ChatGPT?

Most often because the two engines are not reading the same version of the web before ranking anything. Each vendor documents separate crawler tokens, with OpenAI splitting GPTBot from OAI-SearchBot, Anthropic splitting ClaudeBot from Claude-SearchBot and Google splitting Google-Extended from Googlebot, so a robots.txt rule, a WAF rule or a rendering requirement can remove you from one engine's candidate pool while leaving another untouched. Check fetchability per engine before rewriting the page.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

There is no single AI position.
Start reporting one row per engine.

Start free trial