Home/Learn GEO/Citation accuracy
How AI search works · Topic

AI citation accuracy: being cited is not the same as being described correctly.

The independent research on this is unusually clear and unusually bad. It changes what you should write, because the sentence an engine lifts out of your page will be read without the paragraph that made it true.

On this page
Share this
Share on X Share on LinkedIn
The short answer

Being cited is not the same as being described correctly, and every independent measurement of that gap has found it wide. Treat a citation as evidence that an engine judged your page relevant, never as evidence that the sentence beside your link is true. The Tow Center for Digital Journalism at Columbia put 200 quotes from 20 publications to ChatGPT Search in November 2024 and got partially or entirely incorrect responses 153 times, while the system acknowledged an inability to answer only 7 times.1

Key takeaways
  • Neither half of the legal toolkit worked. Licensing deals bought no accuracy, and publishers who had blocked crawlers were cited anyway.2
  • A citation is not a credibility signal either: credible-source shares ran 71.4% to 86.3% by assistant and topic.8
  • The one lever you control is the sentence. Put the subject, the qualifier and the date inside it, so it stays true with the page deleted.
The measurements

How often does an AI answer misrepresent the source it cites?

Often enough that a citation is not evidence of accurate description: more than 60% of answers were wrong across eight engines in the Tow Center’s March 2025 test.

Four independent measurements, three different methods, all pointing the same way. None of them is vendor research.

153 / 200

Tow Center, November 2024

200 direct quotes drawn from 20 publications, each fed to ChatGPT Search with a request to identify the source. 153 responses were partially or entirely incorrect, and the system acknowledged an inability to answer just 7 times.1

> 60%

Tow Center, March 2025

1,600 queries built from 20 publishers, 10 articles each, across 8 engines. More than 60% returned incorrect answers. The best performer was still wrong 37% of the time; the worst was wrong 94% of the time, and out of the 200 prompts it was given, 154 of its citations led to error pages.2

51.5%

Sentence-level support, 2023

Only 51.5% of sentences in generative-search answers were fully supported by the citations attached to them, in the first systematic verifiability evaluation of these systems.3

11.0%

Atomic claims, 2026

A 40-day study of 55,393 trending queries decomposed Google AI Overview responses into 98,020 atomic claims and found 11.0% of them unsupported by the pages cited beside them.4

The four numbers are not measuring the same thing, and the gap between them is informative rather than contradictory. The Tow Center studies test attribution: given a passage, does the engine name the right publisher and link the right article. The sentence-level and claim-level studies test entailment: is the statement actually supported by the source hanging off it. Attribution fails most of the time and entailment fails on a large minority, which fits the mechanism, because the model writes fluently from material it read and then attaches sources afterwards.

Failure modes

What do the errors actually look like?

Four distinct failure modes appear in the independent work, and they need different responses from you. The first is confident misattribution: the answer names a source that did not say the thing. The Tow Center’s November 2024 test found ChatGPT Search hedged only 7 times across 200 attempts, so the wrong answers arrived in the same tone as the right ones.1

The second is the broken citation. Out of the 200 prompts the March 2025 study gave Grok 3, 154 of its citations led to error pages: links that look like evidence and go nowhere.2 The third is partial support: the claim is roughly right, the source is real, and the specific number or qualifier in the sentence is not in the source. This is the one that hurts a brand most, because it is invisible to a reader and produces a plausible-sounding claim you never made about your own pricing, limits or capabilities.

The fourth is the composite. When an engine cannot fetch the primary source, it assembles an answer from whatever else discusses it. The Tow Center observed exactly this in October 2025: with the publisher blocking the crawler, the agent produced a composite summary drawing on tweets about the article, syndicated versions, citations in other outlets and related coverage.5 The output looks like a description of you. It is a description of what other people said about you.

The source pool

Does a citation mean the source is credible?

No, and the field’s 2026 critical survey gives that finding a section heading of its own: citation implies neither credibility nor support.7

Credibility and support are measured separately, and both come out mixed. A peer-reviewed EACL 2026 study of chat assistants observed credible-source shares of 71.4% to 86.3%, depending on the assistant and the topic, with more misinformation sources in some GPT configurations than in Perplexity or Qwen.8 Read the other way round, between roughly 14% and 29% of the sources behind an answer fell outside the study’s credible set.

The pool is also getting harder to read. An audit of ChatGPT, Copilot, Gemini and Perplexity across 712 real questions about politics, health and the environment found evidence of AI-generated sources on all four engines, at roughly 16% of cited sources.9 That changes the target. You are not competing to be mentioned most often inside a vetted library; you are trying to be the clearest correct first-party statement of a fact, in a pool that already contains synthetic pages describing your category.

What does not work

Does a licensing deal or a crawler block protect you?

Neither one did in the only independent test of the question. The Tow Center’s March 2025 study reports that licensing agreements produced no accuracy benefit in the tests it ran that February, and that publishers who had blocked crawlers were cited anyway: Perplexity’s free version correctly identified all ten excerpts from paywalled National Geographic articles, although that publisher disallows Perplexity’s crawlers and has no formal relationship with the company.2 The contract does not buy fidelity, and the robots.txt rule does not buy absence.

The October 2025 follow-up showed why blocking backfires specifically. Agentic browsers, which are Chromium sessions driving a real browser on a user’s behalf, retrieved the full text of a 9,000-word subscriber-only article that the standard ChatGPT and Perplexity interfaces refused, because the publisher had blocked those companies’ crawlers.5 Blocking changes which sources describe you; it does not remove you from the answer, and it usually makes the description worse.

The defensible posture is the opposite of a wall: be the easiest correct source to fetch and quote, and keep the facts an engine will be asked about in dated, server-rendered text on your own domain.

The honest limit of this article

All the strong accuracy evidence here comes from news-publisher test sets built between November 2024 and March 2025, and news attribution is a harder task than describing a SaaS product’s pricing page. No independent study has measured how often AI answers misstate a company’s own product facts, so the error rates above should be read as evidence that misrepresentation is common and unpredictable, not as a rate that applies to your brand. The writing advice that follows is also correlational: it comes from what extracted passages look like, not from a controlled test showing that writing that way causes correct quotation.

Craft

How do you write a sentence that survives being quoted alone?

Bake the subject, the qualifier and the date into the sentence itself, so that the sentence is still true with the page removed. A year-long dataset of 15.7 million AI Mode citations resolved to 4.6 million highlighted passages with a median length of 117 words, of which roughly 85% were fully self-contained and 80% put the answer in the first sentence.6

Survives extraction

  • “Bavior’s entry plan is $99 per month billed monthly, as of 29 Aug 2026.” Subject, number, billing basis and date all inside one sentence.
  • “A query fan-out is the expansion of one question into several searches across subtopics, documented by Google for AI Overviews and AI Mode.” A flat definition with its attribution attached.
  • “In a 4,706-query audit published in Findings of ACL 2026, 53% of the domains Google AI Overviews consulted were absent from the organic top 10.” Sample, venue, figure and scope in one line.

Gets misquoted

  • “This means it costs $99.” The referent, the billing period and the date are all in the previous paragraph, which will not travel with it.
  • “As we saw above, most citations come from outside the top 10.” Structurally unquotable: no subject, no source, no scope.
  • “Studies show a 40% improvement.” No named study, no metric, no conditions, and the field’s 2026 survey files this figure as rejected for general use, a relative maximum on one metric under one configuration.7

The rule generalises to one test you can apply to any sentence in five seconds: delete every other sentence on the page and read it again. If it is still true, still attributable and still about the right subject, it survives extraction. If it needs the paragraph above it, an engine will eventually quote it without that paragraph, and the resulting claim is one you did not make but will be seen to have made.

Two cautions belong with this advice. The first is that citation-shaped writing is not the same as truthful writing: the 2026 critical survey states that adding a fabricated statistic may increase reuse while degrading epistemic quality, and that the criterion is relevant, verifiable, dated and properly attributed evidence rather than the presence of numbers.7 The second is that none of this helps a page that never gets retrieved.

Response

What can you do when an engine describes you wrongly?

Less than you would like, and the useful actions are indirect. There is no correction endpoint at any major engine, no ticket queue for a wrong sentence about your product, and no published turnaround. What you can do is change the material the answer is assembled from, then wait for a recrawl.

In practice that is four moves. Put the disputed fact in plain server-rendered text on your own domain, in a sentence that carries its own subject and date. Make sure that page is fetchable by the engine that got it wrong, checking the search token rather than the training token; every vendor documents them separately. Look at what the wrong answer actually cited: if the error came from a third-party page or a stale discussion thread, the fix lives there and not on your site. And record the error with its date, engine and prompt, because a single wrong answer is a draw from a distribution and you need repeats before you know whether it is systematic.

Set expectations honestly. This is slow, partial and unguaranteed, and the same evidence that shows misrepresentation is common also shows that neither contracts nor blocking fixed it. The realistic goal is to make the correct version the easiest one to quote, which is the work the how AI search works stage recommends everywhere else in the pipeline.

Where a product fits, and where it does not

Everything above is doable unaided: write self-contained sentences, keep facts in server-rendered text, and log wrong answers in a spreadsheet. Bavior helps only with the logging half. It runs a fixed prompt set across five engines on a schedule and records which sources each answer cited, which is how you find out whether a wrong description repeats or was a one-off. It cannot correct an engine, cannot get a bad answer retracted, does not check whether the answer text is factually right, and no product can. The free AI visibility check and free GEO audit run without a paid plan; paid tracking starts from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).

Sources: all checked 30 Aug 2026
  1. Tow Center for Digital Journalism, Columbia, “How ChatGPT Search (Mis)represents Publisher Content”, 27 Nov 2024; 200 quotes, 20 publications; ChatGPT “returned partially or entirely incorrect responses on a hundred and fifty-three occasions, though it only acknowledged an inability to accurately respond to a query seven times” · cjr.org
  2. Jaźwińska & Chandrasekar, Tow Center, “AI Search Has a Citation Problem”, 6 Mar 2025; 1,600 queries, 20 publishers, 8 engines; more than 60% incorrect; “Out of the 200 prompts we tested for Grok 3, 154 citations led to error pages” · cjr.org
  3. Liu, Zhang & Liang, “Evaluating Verifiability in Generative Search Engines”, 2023; “on average, a mere 51.5% of generated sentences are fully supported by citations” · arxiv.org/abs/2304.09848
  4. Xu, Iqbal & Montgomery, “Measuring Google AI Overviews”, 13 May 2026 (preprint); 55,393 trending queries over 40 days; “decomposing responses into 98,020 atomic claims, 11.0% are unsupported by the cited pages” · arxiv.org/abs/2605.14021
  5. Chandrasekar & Jaźwińska, Tow Center / CJR, “How AI browsers sneak past blockers and paywalls”, 30 Oct 2025 · cjr.org
  6. Year-long dataset of 15,699,298 AI Mode citations across 148 industries, resolving to 4.6 million highlighted passages on 2.7 million pages, July 2026; median passage 117 words, roughly 85% self-contained, 80% answering in the first sentence. Vendor-published; described rather than linked, per this curriculum’s sourcing rule.
  7. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (survey preprint); section 8.3 is headed “Citation Implies Neither Credibility nor Support”, and its confidence table rejects the 40% visibility figure as a general claim, calling it “a relative maximum on one metric under a specific configuration” · arxiv.org/abs/2607.14035
  8. Vykopal, Pikuliak, Ostermann & Šimko, “Assessing Web Search Credibility and Response Groundedness in Chat Assistants”, EACL 2026, pp. 2539–2560; credible-source shares of 71.4–86.3% depending on the assistant and topic · aclanthology.org/2026.eacl-long.115
  9. Allaham & Diakopoulos, “Synthetic Sources? Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources”, 22 May 2026 (preprint); 4 engines, 712 queries; “evidence of AI-generated sources being cited across all four generative search engines (~16% of cited sources)” · arxiv.org/abs/2605.23684
FAQ

Frequently asked questions.

If an AI engine cites my page, does the answer describe my page correctly?

Not reliably. The Tow Center tested 200 quotes from 20 publications against ChatGPT Search in November 2024 and found 153 responses partially or entirely incorrect, with the system acknowledging an inability to answer only 7 times, and a March 2025 follow-up across 1,600 queries and eight engines found more than 60% returned incorrect answers. Separately, only 51.5% of sentences in generative-search answers were fully supported by the citations attached to them. Treat a citation as evidence of relevance, never as evidence of accurate description.

Will blocking AI crawlers stop engines from describing my product wrongly?

No, and it usually makes the description worse. The Tow Center's March 2025 study found publishers who had blocked crawlers were still cited, and licensing agreements conferred no accuracy benefit either. Its October 2025 follow-up showed what happens when a block does hold: the engine assembled a composite answer from tweets, syndicated copies and third-party coverage instead. Blocking changes which sources describe you rather than whether you are described, and agentic browsers are ordinary Chromium sessions that a robots.txt rule does not reach.

How should I write so an AI answer quotes me correctly?

Write every sentence you would like quoted so it stays true with the whole page deleted. Put the subject, the qualifier and the date inside the sentence: "our entry plan is $99 per month billed monthly, as of 29 Aug 2026" rather than "this means it costs $99". A year-long dataset of 15.7 million AI Mode citations resolved to 4.6 million highlighted passages with a median length of 117 words, roughly 85% of them fully self-contained and 80% answering in the first sentence. The pattern is correlational, but the failure it prevents is not hypothetical.

Does being cited mean the engine picked a credible source?

No. A peer-reviewed EACL 2026 study of chat assistants measured credible-source shares of 71.4% to 86.3% depending on the assistant and the topic, which means between roughly 14% and 29% of the sources behind an answer fell outside its credible set. A separate 2026 audit of ChatGPT, Copilot, Gemini and Perplexity over 712 real questions found evidence of AI-generated sources on all four engines, at roughly 16% of cited sources. Credibility and accurate support are separate properties, and a citation guarantees neither.

Can I get an AI engine to correct a wrong answer about my company?

There is no correction endpoint at any major engine, so the practical route is indirect and slow. Put the disputed fact in plain server-rendered text on your own domain in a self-contained, dated sentence; confirm the engine that got it wrong can actually fetch that page, checking its search crawler token rather than its training token; look at what the wrong answer cited, because the error often lives on a third-party page rather than yours; and log the error with engine, prompt and date so you can tell a systematic problem from a single unlucky draw.

Which AI engine is most accurate at citing sources?

All eight engines tested were bad, and the ranking is dated enough that it should not drive a decision. In the Tow Center's March 2025 study of 1,600 queries, the best performer still returned incorrect answers 37% of the time, ChatGPT incorrectly identified 134 articles while signalling low confidence just 15 times out of 200 responses, and the worst performer was wrong 94% of the time, with 154 of its citations across 200 prompts leading to error pages. Engine behaviour moves within months, so treat the finding as "misrepresentation is common everywhere" rather than as a league table.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Write the sentence you want quoted.
It will be quoted without the paragraph.

Start free trial