How do you appear in AI search results, and what actually moves the answer?
Appearing is a supply problem before it is a ranking problem. Three things decide it: whether an engine can fetch you, whether the document you published matches the question that was asked, and whether anything you shipped this month changed either one.
AI visibility12 min readUpdated 12 Sep 20268 sources, all primary
You appear in AI search results by clearing a mechanical eligibility gate, then publishing the shape of evidence the question needs in the places an engine already pulls from. Google states the gate plainly: a page must be indexed and eligible to be shown in Google Search with a snippet, and there are “no additional technical requirements”.2 What sits behind the gate is not random either: across 602 controlled prompts and 21,143 search-layer citations, official, news and vertical sources carried 79.12% to 87.52% of all citations, depending on the platform.1
Key takeaways
Eligibility is mechanical and free: indexed, fetchable by the search-side crawlers, no stray snippet directive.
Engines run separate crawlers for separate jobs, so blocking the training crawler does not remove you from a search index.
Match the document to the question. Citations doing comparison work averaged 0.1524 on one study’s influence proxy against 0.0529 for reference-only citations.
No platform citation share is stable enough to plan around: the literature reports low source overlap and substantial run-to-run variability.
Monitoring tells you where you stand. Every lever below is a publishing or outreach action, so a dashboard alone moves nothing.
The mechanism
Where do AI search results actually come from?
From an index, a fan-out and a fetch, and only the first looks like classic SEO. Google documents that its AI features “may use a query fan-out technique”, issuing multiple related searches across subtopics and data sources, so the question a person typed is rarely the query retrieval sees.2 The features underneath are “rooted in our core Search ranking and quality systems”.3
The crawler layer is where the two systems visibly diverge, and the split is documented by the engines rather than inferred. OpenAI runs GPTBot for training data, OAI-SearchBot to surface websites in ChatGPT’s search features, and ChatGPT-User for pages a person asked ChatGPT to open, where “robots.txt rules may not apply”.4 Perplexity documents the same shape, and its user-triggered fetcher “generally ignores robots.txt rules”.5
Crawler token
Documented purpose
robots.txt
GPTBot
Content that may train foundation models
Applies, a licensing choice
OAI-SearchBot
Surfacing websites in ChatGPT search
Applies, blocking removes those answers
ChatGPT-User
Opening a page a person asked for
May not apply, user-initiated
PerplexityBot
Search and linking, not model training
Applies
Perplexity-User
A user action inside Perplexity
Generally ignored
Google-Extended
AI training and grounding elsewhere at Google
Separate control, not Search eligibility
One robots.txt line can remove you from a surface you meant to keep. Blocking the training crawler costs nothing in search. Blocking the search-side crawler is a visibility decision, and teams make the second by accident while intending the first. Which token does what is set out in which crawler tokens matter.
The supply side
Which three source types do AI engines pull from?
The most detailed public breakdown labels them official, news and vertical, and those three carry most of every platform’s citations in that dataset.1
87.52%of ChatGPT citationsthe three types combined
46.35%Google’s official sharelargest single bucket
31.17%ChatGPT’s news shareits second bucket
21,143search-layer citations602 prompts, three platforms
Source type
ChatGPT
Google
Perplexity
Official
34.22%
46.35%
44.07%
News
31.17%
18.99%
16.07%
Vertical
22.13%
22.00%
18.99%
Combined
87.52%
87.34%
79.12%
01
Pages you own
The largest single bucket on every platform measured, and the only one you can publish into this afternoon. Also the most under-supplied: price, limits and migration path are often published nowhere a machine can read them.
02
Coverage you earn
News was 31.17% of ChatGPT citations in that dataset, so editorial coverage is a supply line rather than a vanity metric. The unit of work is something a writer can verify: original data, a benchmark, a free tool.
03
Documents that compare you
Roundups, category explainers and community threads describe several named options at once, the shape a shortlist question needs. The structural reason is argued in why third-party pages dominate.
Read the labels before the percentages. These are the study’s own categories, collected in one window, and the table does not break community platforms out as a row, so “vertical” is not a synonym for “forums”.1 What survives that caution is the ordering: your own site is the largest bucket everywhere, earned coverage is next, and documents where somebody else compares you carry the rest. Three production lines, three lead times, and a program running one of them is capped by that bucket.
The leverage question
Is Reddit still the highest-leverage place to appear?
For one narrow class of questions, yes, and the reason is document shape rather than domain authority. A question of the form “what do people actually use for X” asks for several named options with lived reasons attached. A thread where twenty people answer that from experience is comparison-shaped evidence about a category. Your product page is reference-shaped evidence about one option.
The citation tables run the same way: comparison citations numbered 1,719 and averaged 0.1524 on the influence proxy, second only to definition at 0.1531 and roughly three times the 0.0529 of reference-only citations. At page level, documents containing comparison content averaged 0.1389 against 0.0894, printed as +55.28%. The authors describe high-influence pages as “richer in extractable evidence such as definitions, numerical facts, comparisons, and procedural steps”.1
That makes the unit of work a thread and a question, not a domain, so find both before posting anywhere. In Bavior, open reddit.com from Sources: Cited pages lists the specific threads answers cited, and Top prompts lists the buying questions that led engines to cite them. Join the threads that answer questions you care about, and skip the rest of the site.
Sample data from a demo workspace. Thread titles are illustrative, not real Reddit posts. The red box marks Top prompts, the questions behind the citations; the counts come from one demo prompt set and are not a platform share.
Where the leverage stops is competition, and that has been measured too. C-SEO Bench found that “as we increase the number of C-SEO adopters, the overall gains decrease”, calling the problem congested and zero-sum.8 Community platforms have that property plus a moderator: what works when three companies do it gets a rule written about it when three hundred do. Disclose the commercial relationship wherever a platform requires it, and see how community threads get selected for the rest.
About the percentage you have seen quoted
A figure near 40% for Reddit’s share of ChatGPT citations circulates constantly and no peer-reviewed study supports it. The measurements that exist are commercial panels using different denominators, and the ones we checked on 12 Sep 2026 disagreed with each other by more than the effect anyone is trying to measure. The literature says why that is expected: engines vary substantially in source diversity and stability,6 and the critical survey records low source overlap and substantial run-to-run variability across commercial audits.7 A platform share is a dated snapshot of one panel, and no plan should rest on it.
Worth doing on community platforms
Answer from operating experience, in threads that already rank for the question you care about
Publish the missing number, price, limit or migration path, so the thread becomes quotable
Disclose the commercial relationship wherever the subreddit or platform requires it
Not worth doing
Buying a mention in a thread you never took part in
Pasting one comparison paragraph across many subreddits, the congestion effect C-SEO Bench measured
Setting a target for one platform’s citation share, which you neither control nor can reproduce
The other lines
What else moves the answer besides community threads?
Eligibility comes first because it is binary, and the trap sits in the same documentation: the snippet controls, nosnippet, data-nosnippet and max-snippet, govern what AI features may show from your page too.2 A directive added years ago to protect a paywall can be the reason you never appear.
Structured data is the next thing teams over-invest in, and Google is unusually blunt: “Structured data isn’t required for generative AI search, and there’s no special schema.org markup you need to add”.3 Markup earns its place as parsing hygiene for the features that consume it. It is not a lever on whether an answer quotes you, so budget it that way.
The rewrite meant to win a citation is the most common failure here, and it has been measured twice. C-SEO Bench reports most conversational-SEO methods are “not only largely ineffective but also frequently have a negative impact on document ranking”.8 The survey of 45 studies lands in the same place: relevance and context position are the most reproducible levers, and citation-oriented rewrites can impair retrieval.7 A page rewritten into a quotable answer that then falls out of the retrieval pool has lost the only round that mattered.
Fix eligibility before anything else
Indexed, fetchable by each search-side token, no stray snippet directive. Full ordering in the technical GEO checklist.
Publish evidence, not adjectives
Definitions, numbers, comparisons and procedures scored highest. A paragraph that survives being lifted out of the page beats a section that only reads well in place.
Earn mentions a stranger can verify
Original data, a public benchmark or a free tool gives a writer something to cite, and an engine a second document that names you.
Correct pages that already describe you
A roundup carrying your old price keeps feeding that error into answers. Correcting it returns fastest. Start from where your category is cited.
To run that last step, open Sources in Bavior and read the label on each row of Sources to act on, because the same cited source can call for opposite work. Content gap means answers cite it but it never mentions you, a supply job for the lines above; Reputation means it mentions you in coverage that skews negative, and that row is where correcting the pages already describing you starts.
Sample data from a demo workspace. Red boxes: the two action labels.
Monitoring against execution
Why do most AI visibility tools only solve half the problem?
Because measuring and supplying are different jobs, and only one turns cleanly into a chart. A monitoring product runs a prompt set, records which sources each answer cited, and plots the trend. That work is necessary rather than decorative: the critical survey’s summary of commercial audits is low source overlap, substantial run-to-run variability and persistent fidelity gaps,7 and the ACL comparison of five generative systems against Google organic search reports “substantial variation among engines in their reliance on internal v.s. external knowledge, source diversity, and stability”.6 Without repeated runs you cannot separate change from noise.
What a chart cannot do is publish. Every lever on this page resolves to a document somebody writes, a directive somebody removes, or a thread somebody answers. The survey sets the ceiling in one sentence: already-retrieved content can causally alter its citation or use, but no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability.7 Work on the documents, measure over repeated runs, expect the measurement to be noisy.
One question separates the two halves of this category: after a week with a tool, what changed outside the dashboard? If the answer is a prompt list and an alert digest, the tool solved reporting, which is genuinely hard and worth paying for. It is simply not the half that gets you cited.
The 30 day plan
What does a 30 day AI visibility plan look like?
Four weeks, one job each, and a check that proves it landed. No budget required, no citation promised.
Arrivals in the logs from every token you allow, no nosnippet on key pages.
Days 8 to 14
Baseline: 20 to 30 real buying questions, three runs each, every cited URL logged.
One sheet, 60 to 90 runs, sorted into pages you own, describe you, or ignore you.
Days 15 to 21
Supply the gaps the baseline found: the comparison, the price and limits, the procedure.
Three pages live, each with a passage that answers one logged question on its own.
Days 22 to 30
Off-site and re-measure: answer two threads, correct one page that has you wrong, rerun the set.
A diff of the cited URL lists, read as direction rather than proof.
The honest limit of this plan
Thirty days is long enough to fix eligibility, build a baseline and ship three documents. It is not long enough to prove causation, and no single technique has yet shown a durable cross-platform effect.7 Expect a clean mechanical audit, a repeatable measurement and a supply line that did not exist before. Do not expect a number that only moves upward: the source mix moves on the engine’s schedule, not yours.
Where a product fits, and where it does not
The first two weeks above are free and manual by design. What breaks is week five, when the set has to be rerun on a schedule and the threads worth answering appear faster than anyone reads them. Bavior reruns a fixed prompt set across engines on a schedule and records which sources each answer cited, and where a cited source is a live discussion it drafts a reply in your voice on an account you control. Nothing publishes without your approval, from an aged Bavior account or from your own. It cannot buy a placement, negotiate an inclusion, or make any engine cite you. Plans are from $99/mo, against $3,000 to $20,000/mo for an agency retainer.¹
Sources, all checked 12 Sep 2026
Zhang Kai, He Xinyue, Yao Jingang, “From Citation Selection to Citation Absorption”, 28 Apr 2026 (v2, 29 Apr 2026), arXiv:2604.25707 (descriptive preprint, public geo-citation-lab dataset); 602 prompts, 21,143 search-layer citations, 18,151 fetched pages; source-type, semantic-role and evidence-genre tables: arxiv.org/abs/2604.25707
Kirsten et al., “Characterizing Web Search in The Age of Generative AI”, Findings of ACL 2026, pp. 10827–10848; five generative systems from three providers against Google organic: aclanthology.org/2026.findings-acl.526
Olivier Martinez, “Optimizing Visibility in Generative Engines: A Critical Survey (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (survey preprint); 45 studies, Nov 2023 to Jul 2026: arxiv.org/abs/2607.14035
Puerto et al., “C-SEO Bench: Does Conversational SEO Work?”, NeurIPS 2025 Datasets & Benchmarks Track, arXiv:2506.11097 (v1 6 Jun 2025, v3 20 Oct 2025); ten methods, two tasks, three domains each: arxiv.org/abs/2506.11097
How long does it take to appear in AI search results?
Eligibility changes land as fast as recrawling does, often days. Being cited is slower and noisier, because answers vary between runs of the same question. Plan on a mechanical fix inside a week, a baseline inside two, and a readable direction after at least three repeated runs of a fixed prompt set.
Do I need schema markup to appear in AI search results?
No. Google's own guidance states that structured data is not required for generative AI search and that there is no special schema.org markup to add. Markup remains useful for the search features that consume it, so treat it as parsing hygiene with a small budget, not as a lever on citations.
Does blocking GPTBot remove me from ChatGPT search results?
No. OpenAI documents separate tokens for separate jobs: GPTBot crawls content that may train foundation models, while OAI-SearchBot is what surfaces sites in ChatGPT's search features. Blocking the first is a licensing choice. Blocking the second is what removes you from those answers.
Can any tool guarantee that an AI engine will cite my brand?
No, and a vendor claiming otherwise is describing something nobody has demonstrated. The published survey of this field reports that no reviewed technique shows a stable, cross-platform causal effect on discoverability. What is buyable is faster measurement and faster supply of the documents an answer might draw on.
Bavior Editorial
The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.
Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com
You cannot publish a dashboard into an AI answer. Find the pages and threads your category is cited from.