Independence deserves the sharpest treatment of the three, because the mechanism behind it is measured rather than assumed. A 2026 study of 21,143 search-layer citations from 602 controlled prompts across ChatGPT, Google and Perplexity found official sources to be the largest single category on every platform, from 34.22% of citations to 46.35%.1 Engines retrieve official pages readily, so a prompt written in your own words will tend to retrieve you, and a panel of such prompts reports a rising presence rate that is a property of the panel rather than of your visibility. That is the anti-pattern, and it has its own article.
Which prompt sources are worth building a panel from?
Five places produce sentences a real buyer actually said. They are not equally good, they fail in different directions, and the composition of the panel decides what your visibility number can honestly mean.
On this page
Real prompts come from five places, ranked here by signal quality: recorded calls and support tickets, your own question-shaped search queries, fan-out expansion of head terms, community threads in your category, and churn and lost-deal reasons. Rank a source on three properties: attestation that a real person asked it, attribution to who and when, and independence from your own marketing wording. Collecting the question-shaped ones is not a stylistic preference: across 55,393 trending queries measured over 40 days in 2026, Google produced an AI Overview for 64.7% of question-phrased queries against 9.5% of the rest.6
Key takeaways- Attestation is the property you cannot manufacture, so calls and support tickets outrank the three sources where nobody is on record.
- Compose on purpose: in a 40-prompt panel, 12 calls and tickets, 8 of your own queries, 10 fan-out, 6 community threads and 4 lost deals, with no source past 40%.
- Nobody publishes how often a prompt is asked, so a panel is a sampling frame you declare rather than a share of a known population.
- Report per engine: URL-level overlap between Google organic, AI Overviews and Gemini runs at 0.11 to 0.18 Jaccard, so averaging five engines erases the difference.2
What makes one prompt source better than another?
Three properties, and a source is better when it has more of them: attestation, attribution, and independence from your own vocabulary.
Attestation
A real person is on record having asked this. A recorded call has it absolutely; a sub-question you wrote by expanding a head term has it not at all, and that gap is the largest quality difference in prompt research.
Attribution
You know who asked and when. That lets you say a question came from your segment last quarter rather than from anyone on the internet at any time, and it is what makes a panel prunable.
Independence
The wording did not come from you. A question phrased in your own feature vocabulary measures whether engines echo your marketing, which is the one failure mode that produces a rising number and no information.
Which five sources are they, in order?
In descending order of signal quality, with the first two worth more than the other three combined.
| Source | What it attests | Typical yield | How it fails |
|---|---|---|---|
| Calls and support tickets | A named person asked this, on a date, in their own words | 1–3 distinct questions per call; more from first-week tickets | Only people who contacted you; the silent majority is invisible |
| Your question-shaped queries | Someone typed this and reached your page | Dozens, already attributed to URLs | Google demand only, and the rare-query tail is withheld |
| Fan-out expansion | Nobody, these are reconstructed sub-questions | 4–6 per head term, unlimited in principle | Plausible questions nobody asks; the failure is invisible |
| Community threads | A buyer wrote this in public, for other buyers | Large, and skewed to whatever ranks | Survivorship, and the most volatile citation series measured |
| Churn and lost-deal reasons | A buyer said this while walking away | Few, often under ten, and all high-stakes | Post-rationalised reasons; small numbers, heavy consequences |
The order is by attestation first and independence second, which is why fan-out sits third despite being the easiest source to work with. This page stops at the ranking, and each source has its own article: calls and support tickets for verbatim discipline, question-shaped queries for the search-console filter and its truncated tail, fan-out expansion for the rule that keeps expansion from becoming invention, community threads for why they are useful twice, and churn and lost-deal prompts for the deals you never hear about.
One property cuts across all five and is easy to miss: a panel is per engine, not global. A SIGIR 2026 study of 11,500 queries measured URL-level Jaccard similarity of 0.11 to 0.18 between Google organic results, AI Overviews and Gemini, and an audit of 1,008 generative-search responses found only 26% of 355 domains were cited by both Bing Chat and Perplexity.23 The 2026 critical survey puts it in one line: “Visibility is indexed by engine and surface.”4
How often is each prompt actually asked?
Nobody publishes it. No generative engine reports how many people asked a given question, so a prompt panel cannot be weighted by demand the way a keyword list can.
The absence is structural rather than temporary. The 2026 critical survey describes the position content creators are actually in: they do not generally observe the retrieved set, the reranking score, or the generator’s internal states, which makes this “a black-box optimization problem under incomplete information”.4 Both of the largest independent measurements had to construct their own query samples: a public benchmark of 11,500 queries in one, 55,393 trending queries over 40 days in the other.26
The prompt you send is not even the unit that gets retrieved. Google documents AI Overviews and AI Mode using a “query fan-out” technique, “issuing multiple related searches across subtopics and data sources”.7 One panel entry becomes several searches you never see, so a volume figure attached to your sentence would not describe the queries that ran.
Rank prompts by consequence instead, meaning what it costs you when an engine answers one badly, and publish the frame beside the number. The survey makes the same point about published rates: two AI Overview activation figures, 64.7% and 51.5%, are not contradictory because the query distributions differ, and “a rate reported without a description of the sample has limited transportability”.4 Why prompt volume does not exist takes the argument apart in full.
How many prompts should come from each source?
A defensible starting composition for a 40-prompt panel is 12 from calls and tickets, 8 from your own question-shaped queries, 10 from fan-out expansion, 6 from community threads and 4 from churn and lost deals. The numbers are a convention rather than a measurement, and the reasoning is the part worth keeping: the two attested sources supply half the panel because attestation cannot be manufactured; fan-out gets a quarter because it alone reaches questions nobody has yet asked you; and the last two are small because community threads over-represent whatever ranks and lost-deal reasons are genuinely rare.
Two caps matter more than the exact split. No single source above roughly 40% of the panel, because past that point the panel measures that source’s bias rather than your buyers. A panel that is 80% fan-out is measuring your own imagination, and one that is 80% community threads tracks one engine’s retrieval policy toward forums. Every prompt class represented, in the taxonomy the prompt research stage sets out, because a panel missing the objection class will look stable while the questions that cost you deals go unmeasured.
Composition also has to survive being changed. When you add or retire prompts, bump the panel version and run the old composition alongside the new for one period, so the discontinuity is a number you can state rather than an invisible corruption of the series.
Why is this qualitative research with a quantitative output?
Because the collection step is qualitative sampling and only the measurement step is quantitative, which puts the validity of the final number in the sampling frame rather than in the sample size. You gather utterances from a population you cannot enumerate, using sources you chose, then measure a proportion over what you gathered. Every property of a good qualitative study therefore applies: document how you sampled, say who is excluded, and report the frame beside the estimate. A prompt panel without a stated frame is a survey without a methods section.
The measurement half has its own discipline, and the 2026 critical survey’s minimum checklist for a GEO study states it: record the product, mode, model, date, locale, account and whether search was enabled; run closely spaced repetitions in multiple time windows; keep retrieval, reranking, context and generation separated; and retain null outcomes in the denominator.4 The last instruction is where teams quietly cheat, because dropping the runs where nothing happened turns a panel into a highlight reel while every individual number stays technically true. A 2026 statistical treatment adds why it matters: citation distributions follow a power law, and many apparent differences between domains fall within the noise floor of the measurement process.5
State the consequence in your own reporting. Two teams in the same category, sampling honestly, will build different panels and publish different visibility numbers, and neither will be wrong. Your panel measures your buyers’ questions, not the category’s, which is a good reason to publish yours and a better reason to distrust a market-wide league table.
What does a finished panel entry look like?
Seven fields. The first two are the ones teams skip and later wish they had: the verbatim sentence, and where it came from.
Verbatim and normalised text
Keep both. The verbatim sentence is the evidence; the normalised one is what you run. Storing only the normalised version destroys your ability to check later whether you edited the meaning.
Source, date and speaker segment
Which of the five sources, when it was collected, and what kind of buyer said it. This is what lets you retire a question whose vocabulary has aged out.
Class, panel version and variants
One class label per prompt, the panel version it entered in, and the near-duplicate phrasings you merged in with a count. The variant list tells you which phrasing is common.
Deduplication is where panels quietly lose their frame, so make the rule explicit before you start merging. Two entries are the same prompt when a competent salesperson would give the same answer to both, and different prompts when the answer differs at all. Merge under the most common phrasing, keep the variants with counts, and never merge across classes: “is it safe?” and “how do I set it up safely?” look adjacent and carry different stakes.
The ranking on this page is reasoned, not measured. Nobody has compared panels built from different sources against a ground truth of what buyers ask an engine, because that ground truth is exactly what the section above says is not published. The three criteria are defensible and the ordering follows from them, but a team whose sales calls are unrepresentative could reasonably invert the top two. The composition numbers are a convention with reasoning attached, not a finding; adjust them with evidence from your own funnel.
Every source on this page is inside your own company or free to read, and the whole method is a spreadsheet with seven columns plus a week of reading call transcripts. Do that first: the hand-built panel is what makes any later measurement mean something. Bavior cannot build it for you. It has no access to your call recordings, your tickets or your lost-deal notes, it does not choose your prompts, and it cannot tell you whether your sampling frame is representative. What it does is the repetition: a fixed panel run across five engines on a schedule, with every cited URL recorded so you can see whether a gap is your page or somebody else’s thread, and where a cited source is a live discussion thread it drafts a reply on an account you control, which you approve before anything posts. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).
- Zhang, He & Yao, “From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms”, Apr 2026, arXiv:2604.25707; 602 controlled prompts, 21,143 valid search-layer citations; official sources 34.22% of citations on ChatGPT to 46.35% on Google (preprint, descriptive statistics only): arxiv.org/abs/2604.25707
- Grossman et al., SIGIR 2026; a public benchmark of 11,500 user queries; URL-level Jaccard 0.11–0.18 across Google organic, AI Overviews and Gemini; AI Overviews on 51.5% of queries: arxiv.org/abs/2604.27790
- Li & Sinnamon, 2024; audit of 1,008 generative-search responses; 355 unique domains identified, 26% of which were cited by both Bing Chat and Perplexity (academic, reported in the critical survey at note 4)
- “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (minimum checklist for a GEO study; black-box optimization under incomplete information; visibility indexed by engine and surface): arxiv.org/abs/2607.14035
- Sielinski, “Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement”, Mar 2026, arXiv:2603.08924 (preprint): arxiv.org/abs/2603.08924
- Xu, Iqbal & Montgomery, 2026; 55,393 trending queries collected 13 March to 21 April 2026; AI Overview activation 13.7% overall, 64.7% for question-form queries against 9.5% for non-question queries (preprint): arxiv.org/abs/2605.14021
- Google Search Central, “AI features and your website”, updated 10 Dec 2025 (the “query fan-out” technique, issuing multiple related searches across subtopics and data sources; first-party): developers.google.com/search/docs/appearance/ai-features
- Community-thread citation study, November 2025: 248,000 unique cited Reddit URLs across 217,000 prompts; cited threads averaged roughly 900 days old and 80 words, and 80% had fewer than 20 upvotes. Vendor-published, described rather than linked per this curriculum’s sourcing rule.
- Most-cited-pages study, October 2025: of the top 1,000 pages one assistant cited in September 2025, 28.3% had no organic keyword visibility at all. Vendor-published, described rather than linked per this curriculum’s sourcing rule.
Frequently asked questions.
Can I build a panel from just one source?
You can, and the panel will then measure that source's bias rather than your buyers' questions. Each source fails in a different direction: calls only reach people who contacted you, search console only sees Google demand with its rare-query tail withheld, fan-out expansion attests nothing because you wrote it, community threads over-represent whatever already ranks, and lost-deal reasons are rare and post-rationalised. Keeping any single source under roughly 40% of the panel is what stops one of those failure modes from becoming the whole measurement.
How do I know when I have collected enough prompts?
Stop adding sources when new collection stops producing new distinct questions, and stop sizing the panel when the confidence interval is narrow enough to act on. Those are two separate stopping rules and both matter: the first is the qualitative saturation test, which you can watch by plotting new distinct questions per batch of calls, and the second is arithmetic, because 40 prompts run 5 times each per engine gives roughly a plus-or-minus 7-point band around a 50% presence rate. A panel can be statistically adequate and still miss the questions your buyers ask, which is why saturation comes first.
Is it cheating to use an AI model to generate prompt ideas?
Using a model to expand a list you have already grounded in real calls is reasonable; using it to originate the list is not, and the failure is systematic rather than random. A language model asked for example questions samples its training distribution, which is dominated by the way your category is already written about publicly, so it returns the median phrasing of well-covered questions and under-produces exactly the questions where nobody has published anything, which are the ones where you are invisible. The test is unchanged: if you cannot point to a human who said something close to that sentence, it belongs in the fan-out bucket and should be labelled as reconstructed.
Should the same panel run on every engine?
Run the same prompts on every engine and report the results separately, never pooled. The evidence that pooling destroys information is direct: a SIGIR 2026 study of 11,500 queries measured URL-level Jaccard similarity of 0.11 to 0.18 between Google organic results, AI Overviews and Gemini, and an audit of 1,008 responses found only 26% of domains shared between Bing Chat and Perplexity. A single blended "AI visibility" percentage averages five different systems with different retrieval behaviour, and it will move for reasons you cannot attribute to anything you did.
The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.
Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com