Home/Learn GEO/What engines quote
Content for AI citation · The unit

What does an AI engine actually quote, a page or a passage?

Everything after this in Content for AI citation depends on the answer. If the thing an engine takes is a span of text rather than a document, then a span of text is the object you are writing.

On this page
Share this
Share on X Share on LinkedIn
The short answer

A passage. An engine lifts a span of text out of your page and sets it next to spans lifted from other sites, so the object competing for a place in the answer is the span, not the document that happens to contain it. The academic work on this relationship is measured sentence by sentence for exactly that reason: a 2023 evaluation of four generative search engines found that on average, a mere 51.5% of generated sentences are fully supported by their citations.2

Key takeaways
  • The addressable unit is a passage of roughly a hundred words, judged on its own merits. A 2,500-word article is about twenty separate entries in that competition, two or three of which do all the work.
  • Quotability is a property of a span, not of a page. The subject, the qualifier, the figure and the date all have to sit inside the span, because everything around it is gone by the time the engine reads it.
  • The most detailed published proxy for how deeply a source shapes an answer is built out of paragraph coverage and n-gram overlap, which is to say it measures passages rather than pages.1
  • In that same dataset, citations used as direct factual support average 0.1241 mean influence against 0.0775 for background-only use, on 5,411 and 5,100 observations.1
  • Google tells you to ignore “chunking” as a tactic.6 Both things hold: write in self-contained units because readers and extractors both need them, and ship no machine format nobody asked for.
The unit

Does an engine cite your page or quote a passage from it?

It quotes a passage and attaches your URL to it. The two look like one thing in the rendered answer, which is why the distinction gets lost, but they are separable and fail separately. The link is an address; the passage is the content. Whether the passage reads correctly on its own is decided before your domain, your authority or your word count enters the picture at all.

The clearest description of the unit comes from an industry collection published in July 2026: a year of AI Mode citations, 15,699,298 of them, resolving to 4.6 million unique highlighted passages across 2.7 million pages. In that dataset 47.7% of citations were scroll-to-text highlights rather than plain links, the median highlighted passage ran 117 words, 80% put the answer in the first sentence, and roughly 85% were fully self-contained. Repeatedly quoted passages opened with an explicit question 48% of the time against 22% for one-off passages.

Two caveats belong with those figures and they are load-bearing. The study reports no base rate, so “80% of quoted passages answer first” corroborates a plausible mechanism without proving one; nobody has published the share for all web paragraphs. And it is vendor research, so this curriculum describes it precisely enough to find rather than linking it, and rests the argument on the work in the next section.

Evidence

What evidence says the passage is the unit?

Look at how the independent researchers had to build their instruments. A 2026 preprint on 602 controlled prompts, 21,143 search-layer citations and 18,151 fetched pages needed a measure of how much a cited page actually shaped an answer, and the score it constructed weights repeated reference at 0.20, early appearance at 0.15, paragraph coverage at 0.20, TF-IDF cosine similarity at 0.25 and n-gram overlap at 0.20.1 Three of those five components compare stretches of the answer against stretches of the source. You cannot build that instrument at page granularity, because at page granularity there is nothing to compare.

The same paper states the mechanism as a hypothesis worth quoting exactly: a page “becomes valuable to a generative engine when it can be decomposed into reusable, semantically aligned information units”, and it notes that evidence genres such as definitions, statistics, comparisons and step sequences “create reusable support units, while Q&A formatting is only a surface wrapper.”1 Its authors call the work descriptive and reserve confirmatory analysis for later, so treat the framing as a well-specified hypothesis rather than a finding.

The independent accuracy literature points the same way by measuring at the same grain. The 2023 verifiability evaluation scored sentences against citations and found 51.5% of sentences fully supported and 74.5% of citations supporting their sentence.2 A 2026 audit of Google AI Overviews decomposed answers into 98,020 atomic claims and found 11.0% unsupported by the pages cited beside them.3 Whether those failures matter for a brand is a separate question, and synthesis and attribution answers it. What they establish here is narrower and more useful: every serious attempt to study the source-to-answer link has had to work below the page.

Properties

What makes a passage survive being lifted out of its page?

A passage survives extraction when a stranger can read it with the page deleted and lose nothing. Five properties do most of that work, and all five are about what sits inside the span rather than around it.

01

The subject is named inside the span

“It”, “the product” and “this approach” resolve against a heading the extractor does not take. Name the thing again in the first sentence even when repetition reads slightly heavy on the page.

02

The qualifier travels with the claim

A limit stated two paragraphs earlier is not part of the passage. “Free on annual plans above ten seats” has to be in the same span as the word free, or the span is quotable and wrong.

03

Every number carries its denominator and date

A percentage with no sample and no collection window cannot be checked by anyone downstream, and the 2026 critical survey is explicit that the criterion is not to add numbers but to supply relevant, verifiable, dated and properly attributed evidence.5

04

Nothing points forward or backward

Openers that reference their neighbours are structurally unquotable, because the neighbour is gone. The house list of banned openers exists for this reason alone, and a passage failing it is invisible rather than penalised.

05

One paragraph carries one claim

A paragraph holding three claims can be selected for the one you care least about and quoted with the other two attached. Splitting them gives each claim its own chance and its own failure mode.

None of this is a ranking factor

These properties decide whether a span is usable once it has been retrieved and selected. They do nothing about getting it there, which is settled earlier and is why crawler logs come before this stage.

The test

How do you test a passage before you publish it?

Copy the paragraph into an empty document and read it cold. Three questions settle it. Can a reader name what the paragraph is about without scrolling anywhere? Is every qualifier that would change the meaning present in the text in front of them? Would the paragraph still be true in six months, or does it need a date it does not have? A paragraph that fails any of the three is not bad writing; it is writing that only works in place, which is a different defect and an invisible one, because nothing in your analytics reports a passage that was never eligible.

A B2B documentation team that ran this over its pricing and integration pages found the pattern that usually shows up: the qualifiers had drifted into a shared paragraph at the top of each page, so twelve individually accurate sections were each quotable and each incomplete. Moving one clause back into each section changed no facts and added about forty words per page. That kind of edit is cheap, and the team could not attribute any later change in citations to it, because a single site with no control group cannot separate an edit from everything else moving in the same window.

The failure worth naming is the opposite one. A team that read the same evidence and chopped a long guide into thirty question-headed fragments lost the connective reasoning that made the guide worth citing, and ended up with thirty spans that each restated a definition available in twenty other places. The paper behind most of the passage evidence anticipates that outcome: Q&A format was the one genre in its table associated with lower mean influence, at 0.0947 against 0.1005, which its authors read as a surface wrapper doing none of the work.1

Usage style

Which passages get used as fact rather than as background?

The same dataset labels how each citation was used inside the answer. Direct factual support outscores background-only use by about a third, on nineteen thousand labelled citations.1

How the passage was usedCitationsMean influenceAgainst direct support
Direct factual support5,4110.1241Reference point
Synthesised across sources3,9670.0964−22%
Paraphrased5,3050.0955−23%
Background only5,1000.0775−38%

The percentages in the last column are our arithmetic on the published means, and the ordering is the part worth keeping. A passage that supplies a checkable fact gets used as one; a passage that supplies atmosphere gets used as atmosphere, and atmosphere is where a brand ends up when its pages describe rather than state. Writing a specific number with its date and its denominator is not a formatting trick here, it is the difference between the top row and the bottom one.

A second cut of the same records scores what a passage is for rather than how it was used, and which sources get cited covers that table along with the source-type mix. The two cuts agree, which is mild evidence and not independent evidence, since they come from one collection window on three platforms.

Consequences

What does this change about how you write a page?

The page stops being the thing you optimise and becomes the container you assemble. A 2,500-word article is roughly twenty independent entries at the selection stage, most of which will never be eligible for anything; the useful question at the outline stage is which four or five spans you actually want quoted, and whether each of them answers a question somebody asks in those words. That reframing changes outlines more than it changes sentences.

It licenses less than it appears to. Deliberately re-shaping text to win selection is the intervention that keeps testing null: a NeurIPS 2025 benchmark of ten conversational-SEO methods reported that out of 54 cases, only three showed statistically significant ranking improvements.4 Question-shaped headings have two independent lines of support, since a 55,393-query study measured Google AI Overview activation at 64.7% on question-phrased queries against 9.5% on non-question ones,3 and the passage collection found question-led spans recycled about twice as often. Everything else in this article is a hypothesis with an argument attached, and the levers with evidence grades them one by one.

The honest limit of this article

Google contradicts a careless reading of everything above, in its own guide, updated 10 July 2026: “For Google Search, you can ignore tactics like ‘chunking’ content, creating unnecessary AI text files (like llms.txt), or pursuing inauthentic mentions.”6 It rejects chunking as a machine-readability format, and it does not say passages are not the unit; the same guide describes query fan-out and requires a page to be indexed and snippet-eligible.7 The defensible position is that self-contained sections help a reader and an extractor for the same reason, and that no published experiment shows shaping them on purpose reliably moves citations. The passage statistics above are vendor-published; the influence figures come from one preprint whose authors call it descriptive statistics; none of it is causal.

Where a product fits, and where it does not

Nothing in this article needs a tool. The passage test is a paragraph, an empty document and three questions, and the rewriting is yours either way; no software can decide which four spans of your page are worth being quoted for. Bavior works on the measurement side of the same loop: it runs a fixed prompt set across five engines on a schedule and records which sources each answer cited, so a before-and-after around a rewrite has numbers on both ends instead of one. It does not read your drafts, does not score passages, does not edit pages, and cannot tell you whether a span was the reason an engine cited you. The free AI visibility check and free GEO audit run without a paid plan; paid tracking starts from $99/mo billed monthly, or $79.17/mo billed annually (as of 30 Aug 2026).

Sources, all checked 30 Aug 2026
  1. Zhang, He & Yao, “From Citation Selection to Citation Absorption”, arXiv:2604.25707 (preprint; the authors call it descriptive statistics); 602 controlled prompts, 21,143 search-layer citations, 18,151 fetched pages; influence score §5.2, usage style §8.3, evidence-container hypothesis §9: arxiv.org/abs/2604.25707
  2. Liu, Zhang & Liang, “Evaluating Verifiability in Generative Search Engines”, arXiv:2304.09848; “a mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentence”: arxiv.org/abs/2304.09848
  3. Xu, Iqbal & Montgomery, 2026 (preprint); 55,393 trending queries collected 13 March to 21 April 2026; 11.0% of 98,020 atomic claims unsupported; AI Overview activation 64.7% on question-form queries against 9.5% otherwise: arxiv.org/abs/2605.14021
  4. Puerto, Gubri, Green, Oh, Yun, “C-SEO Bench: Does Conversational SEO Work?”, NeurIPS 2025 Datasets & Benchmarks Track, arXiv:2506.11097; ten methods, more than 1.9k queries and 16k documents; §6.2: “Out of 54 cases, we uncover only three where the ranking improvements are statistically significant”: arxiv.org/abs/2506.11097
  5. “Optimizing Visibility in Generative Engines: A Critical Survey”, 15 Jul 2026, arXiv:2607.14035 (survey preprint): arxiv.org/abs/2607.14035
  6. Google Search Central, “Google’s Guide to Optimizing for Generative AI Features on Google Search”, last updated 10 Jul 2026 (first-party; the chunking recap): developers.google.com/search/docs/fundamentals/ai-optimization-guide
  7. Google Search Central, “AI features and your website”, last updated 10 Dec 2025 (first-party; query fan-out and the indexed-and-snippet-eligible requirement): developers.google.com/search/docs/appearance/ai-features
  8. Industry passage-level collection, July 2026: a year of AI Mode citations, 15,699,298 of them, resolving to 4.6 million unique highlighted passages across 2.7 million pages; median passage 117 words. Vendor-published, so described rather than linked.
FAQ

Frequently asked questions.

Does an AI engine read my whole page before citing it?

Something in the pipeline has your page, but the thing that reaches the answer is a span of text from it. The most detailed published proxy for how much a source shapes an answer is built from paragraph coverage, TF-IDF cosine similarity and n-gram overlap, all of which compare parts of the answer against parts of the source. Write so that any single section can be lifted out and still be accurate, because that is the object being judged.

How long should a quotable passage be?

The one published passage-level collection reports a median highlighted passage of 117 words, which is a description of what engines happened to take rather than a target to write to. Treat it as a sanity check: a section that runs to six hundred words probably contains several claims that would each be better off with their own heading, and a section of thirty words probably cannot carry its own subject, qualifier and date.

Should I convert my pages into FAQ format to get quoted?

No, and the evidence points mildly the other way. In the 2026 citation-absorption dataset, Q&A page format was associated with slightly lower mean influence, 0.0947 against 0.1005, which its authors read as a surface wrapper rather than a source of extractable evidence. Question-shaped headings are well supported and are a different intervention. Ask a real question in the heading, then answer it in the first sentence of a normal, self-contained section.

If passages are the unit, does Google not say to ignore chunking?

It does, and both statements hold at once. Google's guide, updated 10 July 2026, lists chunking among tactics you can ignore, and it is rejecting a machine-readability format rather than describing what a generated answer contains. The same guide describes query fan-out and requires a page to be indexed and snippet-eligible. Write self-contained sections because a reader and an extractor need the same thing, not because an engine asked for a format.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Engines quote passages,
not pages. Write the passage.

Start free trial