Home/Learn GEO/Calls and tickets
Prompt research · Source 1 of 5

Why is the first question of a discovery call your best prompt?

It was spoken by a named person, on a date, before they had learned your vocabulary. Nothing else in prompt research has all three properties, and this method is about not damaging them on the way into the panel.

On this page
Share this
Share on X Share on LinkedIn
The short answer

The first question a buyer asks on a discovery call is the highest-signal prompt available to you, because it is attested, attributed and independent: a real person asked it, you know who and when, and they asked it before adopting your vocabulary. Tickets opened in the first week of a trial are the same artefact from the other side of the sale. Both also arrive in the shape engines reward: across 55,393 trending queries measured in 2026, Google AI Overview activation ran at 64.7% for question-form queries against 9.5% for everything else.1

Key takeaways
  • Take the first buyer-initiated question from every call and every first-week trial ticket, and store it verbatim beside a normalised copy.
  • Normalise with five edits and no more: strip identifiers, drop disfluencies, resolve pronouns, split compound questions, keep everything else.
  • Never translate a question into product vocabulary. Official pages already take 34.22% to 46.35% of citations by platform, so the rewrite retrieves your own documentation and measures nothing.4
  • Stop when two consecutive batches of five calls each add under one new question, and strip every identifier before the panel leaves the building.
Why this source ranks first

Why is a spoken first question worth more than a search query?

Because a spoken first question is the only artefact in prompt research that carries a person, a date and an uncontaminated vocabulary at the same time. A search query has the date and a partial identity but arrives pre-compressed into search grammar; a community thread has the words but not the person you sell to. The first question on a call is the sentence a buyer used while still describing the problem in their own terms, which is exactly the register people type into an assistant: long, conversational, and phrased as a question rather than as keywords.

The question form is also the form that most reliably produces an AI answer. A study of 55,393 trending queries collected between 13 March and 21 April 2026 measured Google AI Overview activation at 64.7% for question-phrased queries against 9.5% for non-question queries.1 Call questions arrive already in that shape, which is why they need so little processing: a human being did the work of turning a keyword into a prompt without knowing they were doing it.

There is a third reason, and it is the one that changes what you do next. The same measurement found that 29.8% of the domains an AI Overview cites do not appear anywhere on the corresponding first page of results.1 A vendor study of one assistant’s most-cited pages points the same way, at a number nobody outside that vendor can check.2 Either way, a research method that starts from keyword data cannot see that population. Calls can, because they start from the person rather than from the index.

Extraction

How do you extract them without contaminating them?

Copy the first buyer-initiated question in each call, and every question in a first-week trial ticket, exactly as spoken. Do the deciding afterwards, not during collection.

01

Define the unit before you start

A qualifying question is interrogative, uttered by the buyer, and asked before your demo begins. Fixing that first stops the collection drifting toward whichever questions the collector found interesting.

02

Copy, do not paraphrase

Keep the filler, the wrong noun, the hedge. A buyer who says “the posting thing” has told you the category has no settled name in their head, and you delete that by tidying it up.

03

Collect first, judge later

Deciding what belongs in the panel while reading transcripts merges two jobs and biases both. Take every qualifying question into a raw list, then select in a separate pass.

Two sampling decisions do more damage than any transcription error. Sample calls, not accounts: pulling questions from your ten favourite customers produces the questions of people who already bought, which is the opposite of the population you want to be visible to. Take a period, the last quarter say, and work through it. Include the calls that went nowhere: a prospect who asked one sharp question and never replied has told you something a won deal cannot, and their question is over-represented among the people asking an engine instead of asking you. The critical survey’s checklist carries the same instruction on the measurement side, and it applies to collection too: null outcomes stay in the denominator.3

Normalisation

How do you turn a spoken sentence into a panel prompt?

Apply five edits and no others, keeping the verbatim original in the next column so any later reader can check what you changed. Remove identifiers: names, companies, invoice numbers, anything that identifies the speaker. Drop disfluencies: the “um”, the false start, the repeated word. Resolve pronouns to the noun the speaker had just used, not to the noun you would have used. Split compound questions into one entry each, because a two-part question produces an answer you cannot score. Keep everything else, including their word for your category when it is wrong, because that word is what a real person would type.

What you must not do is the sixth edit everyone wants to make: translating the question into your product vocabulary. “Can I do this without getting my account banned?” becomes “What are the account safety controls?”, and the entry stops measuring anything. The rewritten version retrieves your own documentation because official sources are the largest single category of cited source, at 34.22% to 46.35% depending on platform, in a study of 602 controlled prompts producing 21,143 search-layer citations.4 The original retrieves whatever the open web says about bans in your category, which is the thing you actually wanted to know.

Store both columns permanently. Six months later, when the panel produces a number somebody disputes, the verbatim column is the only evidence that the prompts are real questions rather than your team’s idea of them, and an unauditable panel gets discounted by exactly the person you built it to convince.

Deduplication

How do you dedupe variants without editing them?

Group by the answer a question wants, elect one wording to run, and log the rest with counts. It is a grouping decision, never a rewriting one.

Two entries are duplicates when the same answer would satisfy both, not when they share words. “Does this work if my team is two people?” and “is this overkill for a small team?” share almost no vocabulary and want one answer; “how much does it cost?” and “is there a free tier?” look adjacent and want two. Sort by intended answer, then elect the wording that appeared most often.

What you must not do is average a group into a hybrid sentence nobody said. That is the rewrite problem in a different costume: the merged wording is yours, so the entry stops being attested at the moment you thought you were tidying. Keep the elected wording verbatim and the rest as logged variants beneath it, each with a count.

Retire a wording from the run, never from the record. Paraphrase is a design factor rather than noise: the critical survey’s minimum checklist puts intents, languages, paraphrases and an inclusion rule under the sample a study has to specify.3 A single wording is also a weak estimator: repeated sampling across three generative platforms found citation rankings unstable across samples, with many apparent differences between domains falling inside the measurement’s own noise floor.6 Rotate a logged variant into the run now and then, and report the panel rather than any single prompt.

Saturation

How many calls and tickets do you need?

Stop when new calls stop producing new distinct questions, which in qualitative research usually happens far sooner than teams expect.

BatchWhat you are watchingTypical decision
Calls 1–6The main themes appear; most questions are newKeep going, nothing here is stable yet
Calls 7–12New distinct questions per call falls sharplyCount variants, not new questions
Calls 13–20Mostly rephrasings of what you already haveApproaching saturation; phrasing counts are now the output
Two consecutive batches of 5 adding under one new question eachThe curve has flattenedStop; resume when your category’s vocabulary shifts

Saturation is a borrowed concept with a real measurement behind it, and the borrowing is what makes the stopping rule defensible rather than arbitrary. The best-known experiment on the question, a 2006 study in Field Methods analysing sixty in-depth interviews, reported that “saturation occurred within the first twelve interviews, although basic elements for metathemes were present as early as six interviews”.5 That study was about a different subject in a different decade, so treat it as a prior rather than a target: the right order of magnitude is a dozen, not a hundred, and the stopping rule is an observed flattening rather than a fixed number.

Plot the curve yourself: three spreadsheet columns, being new distinct questions per call, cumulative distinct questions, and calls processed. Two consecutive batches of five that each add fewer than one new question means the source is exhausted for now. That is also when the useful output changes, because the counts of how many buyers used each phrasing become the more valuable half of the dataset.

Before it leaves

What has to be stripped before a prompt leaves your company?

Every identifier, because a prompt panel is not an internal document: it is a list of sentences you will type into five third-party systems, on a schedule, for as long as the panel lives. That is a different risk profile from a transcript sitting in your own storage, and it is why identifier-stripping comes first in the normalisation list. Customer names, employers, deal sizes, ticket numbers: none of it improves the prompt, and all of it leaves the building the first time the panel runs.

Two practical rules follow. Keep the mapping from panel entry back to source call in your own systems, never in the prompt text, so you retain attribution without exporting it. And check what each engine documents about retention and training use for the surface you are on, because those terms differ by product and by plan and they change. The survey’s checklist already asks you to record the product, mode, model, date and locale of every run, and that record is what tells you later which terms applied.3

The reward is a shareable panel. A de-identified list of forty buyer questions is an artefact you can hand to a writer, a founder or an agency without a review, and it is often a more useful research output than the visibility number it was collected to compute. The Prompt research overview covers what to do with them next, and the other four sources are ranked in where prompts come from.

The honest limit of this article

Calls and tickets are a biased sample and no amount of collection discipline fixes it. Everyone in this dataset contacted you, so they had already found you, already trusted you enough to book, and already knew enough vocabulary to describe the problem out loud. The people you most want, who asked an assistant, got an answer naming somebody else, and never came near you, contribute nothing here and cannot be recovered from any source in this curriculum. The saturation numbers are borrowed from another field, and the composition advice is convention with reasoning attached, not a measured optimum. Collect from calls because it is the best evidence available, not because it is representative.

Where a product fits, and where it does not

This is the one part of prompt research that needs no software: a transcript folder, a spreadsheet with a verbatim column and a normalised column, and an afternoon per quarter. Reading twenty transcripts in one sitting tells you which words your buyers do not have, and no dashboard produces that finding. Bavior has no access to your calls, your ticket queue or your CRM, cannot extract questions from them, and cannot judge whether the ones you picked are representative. It works after the panel exists: running it across five engines on a schedule, recording which sources each answer cited, and drafting replies for cited threads on accounts you control, which you approve before anything posts. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).

Sources, all checked 29 Aug 2026
  1. Xu, Iqbal & Montgomery, 2026; 55,393 trending queries, 13 March to 21 April 2026; AI Overview activation 64.7% for question-form against 9.5% for non-question queries; “29.8% of AIO reference domains do not appear anywhere on the corresponding first page” (preprint): arxiv.org/abs/2605.14021
  2. Most-cited-pages study, Oct 2025: of the top 1,000 pages one assistant cited in Sept 2025, 28.3% had no organic keyword visibility. Vendor-published, so described rather than linked.
  3. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (minimum checklist: product, mode, model, date, locale; null outcomes retained): arxiv.org/abs/2607.14035
  4. Zhang, He & Yao, “From Citation Selection to Citation Absorption”, Apr 2026, arXiv:2604.25707; 602 controlled prompts, 21,143 search-layer citations; official sources 34.22% to 46.35% by platform (preprint): arxiv.org/abs/2604.25707
  5. Guest, Bunce & Johnson, “How Many Interviews Are Enough?”, Field Methods 18(1), 2006, 59–82; sixty in-depth interviews, saturation within the first twelve, basic elements for metathemes by six (peer-reviewed, paywalled): doi.org/10.1177/1525822X05279903
  6. Sielinski, “Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement”, Mar 2026, arXiv:2603.08924 (preprint; repeated sampling across three platforms, rankings unstable between samples, so the panel is the reporting unit): arxiv.org/abs/2603.08924
  7. Google Search Central, “AI features and your website” (query fan-out across subtopics; AI-feature traffic inside the “Web” search type; first-party, updated 10 Dec 2025): developers.google.com/search/docs/appearance/ai-features
  8. Tow Center, Columbia: “How ChatGPT Search (Mis)represents Publisher Content”, Nov 2024; of 200 quotes tested, 153 responses partially or entirely incorrect: cjr.org
FAQ

Frequently asked questions.

What if we do not record our sales calls?

Write the first question down during the call instead, in the buyer's words, before you answer it. A one-line note taken live is a weaker record than a transcript but keeps the property that matters: the words are theirs and the date is real. Support tickets need no recording at all and are usually the larger source anyway, because a ticket opened in the first week of a trial is already written down, already timestamped and already phrased as a question. Start there while you decide whether call recording is worth setting up, and note in your panel which entries came from live notes rather than transcripts.

Should I clean up a buyer's question before putting it in the panel?

Remove identifiers and disfluencies, and change nothing else, especially not their word for your category. A buyer who calls it "the posting thing" has told you the category has no settled name in their head, and that phrasing is closer to what they would type into an assistant than your product page's noun is. Rewriting it into your own vocabulary swaps a question about the open web for a question your own documentation answers, and official sources are already the largest category of cited source in published measurements, so the rewritten prompt will return you and measure nothing.

How many call questions belong in a 40-prompt panel?

Around a dozen, which is roughly a third, and no single source should exceed about 40% of the panel. Calls earn the largest share because they are the only source that is attested, attributed and independent of your own vocabulary at the same time, but a panel built entirely from them inherits their bias: everyone in that dataset had already found you and already decided to talk. The rest comes from your own question-shaped queries, fan-out expansion, community threads and lost-deal reasons, in a composition you write down and version alongside the panel.

Is there a privacy risk in putting customer questions into a prompt panel?

Yes, and it is different from the risk of holding the transcript, because a panel is typed into third-party systems repeatedly and on a schedule. Strip names, employers, deal sizes and ticket numbers during normalisation rather than later, keep the mapping from panel entry back to the source call inside your own systems, and record for every run which product, mode and model you used. The critical survey's minimum checklist asks for that record anyway, and it is also what tells you afterwards whose terms applied to the text you sent.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Your buyers already wrote the panel.
It is sitting in your call transcripts.

Start free trial