Home/Learn GEO/Question-shaped queries
Prompt research · Source 2 of 5

Which of your search console queries are really prompts?

Two filters find them: a question word, or seven words and up. That subset is the part of your existing demand that most resembles a prompt, it is already attributed to pages you own, and the report deliberately withholds the tail where most of it lives.

On this page
Share this
Share on X Share on LinkedIn
The short answer

The queries in your search console that most resemble prompts are the ones led by a question word and the ones that run long, roughly seven words and up. Both filters have measurement behind them: across 55,393 trending queries collected between 13 March and 21 April 2026, Google AI Overviews fired on 64.7% of question-form queries against 9.5% of non-question queries, a 6.8 times difference.1

Key takeaways
  • Take the union of two filters: a leading interrogative from the study’s own fifteen-word list, and query length.
  • Length is not just a proxy for phrasing: among non-question queries alone, activation climbs from 9.9% at one word to 38.7% at six or more.1
  • The report omits rare queries and caps the table at 1,000 rows, and that withheld tail is where question-shaped demand lives.4
The filter

Which existing queries most resemble a prompt?

Queries led by a question word, and queries of seven words or more: those two sets, unioned, are the closest thing in your own data to a prompt. Fix the word list in writing so the filter is reproducible, and prefer a published list to one you invent. The length threshold is a proxy for the same behaviour, catching the queries that were typed as sentences without an interrogative at the front.

The phrasing filter carries the sharpest evidence in this field. Across 55,393 trending queries issued over 40 days in spring 2026, Google AI Overviews activated on 13.7% of queries overall, but on 64.7% of question-form queries against 9.5% of non-question ones: a 6.8 times difference, chi-square 10,002.2, p < 10⁻³⁰⁰.1 Effects that large are rare, and this one was measured on live search, not in a simulation.

Length earns its place as a second filter because it moves the number independently of phrasing. In the same study, activation ran 3.4% at two words, 20.0% at four, 33.9% at five and 58.1% at six words or more; restricted to non-question queries alone it still climbed from 9.9% at one word to 38.7% at six or more.1 A long query is doing something a keyword is not, whether or not it opens with how. The study bins at six words and up, so a seven-word cut sits inside the strongest bin rather than at its edge; a vendor dataset of 146 million result pages put AI Overviews on 46.4% of queries of seven words or more, by a method it has never published.2

Both filters describe when an AI answer appears, not how often anyone asks. Keep that distinction live, because it changes what the output is for: this source tells you which of your existing demand is shaped like a prompt, and it cannot tell you which prompts exist that you have never ranked for. Those come from calls and tickets.

The ladder

Which question words actually trigger an answer?

Not equally, and the spread is wide enough to change which rows you work first. In the same 55,393-query dataset, activation by leading interrogative ran from 84.3% for how and 73.4% for why, through 72.6% for when and 65.5% for where, down to 47.9% for who and 39.8% for did.1 The pattern is legible: interrogatives that open an explanation get an AI answer, while the ones pointing at a single fact more often get an ordinary result.

Two cautions before you sort on that table. Several cells are tiny, a dozen queries or fewer, and a 100% rate computed on twelve queries is not a finding, so read only the rows with hundreds of queries behind them. The floor is also high: even the weakest interrogative still activated four to five times more often than a non-question query, so a low-ranking question word beats a keyword.

Borrow the study’s operating definition instead of writing your own, because it makes your filter reproducible and your numbers comparable to published work. A query counted as question-form there if its leading whole word was one of fifteen interrogatives: who, what, where, when, why, how, which, is, are, was, can, do, does, did, has.1 Leading whole word, not anywhere in the string, which keeps a query like “reddit rules on what you can post” out of the set.

The mechanics

How do you pull them out of Search Console?

Open the Performance report, set the longest date range available, apply a custom regex filter to the query dimension, and export, because the browser table caps out long before your data does.

1

Filter on the query dimension

The query filter accepts a regular expression in RE2 syntax, so one expression alternating the question words does the job in a single pass.5 Run the length filter as a second, separate export rather than folding both into one pattern.

2

Export rather than read

The table view stops at 1,000 rows.4 Export to a spreadsheet or pull the data through the API, because the rows past the cap are disproportionately the long, question-shaped ones you came for.

3

Sort by pages, not by clicks

Group the exported queries by the page they were attributed to. A cluster of long questions landing on one URL is a prompt cluster with a page already attached, the most actionable shape this data comes in.

One property of the report matters before you interpret anything in it: traffic from AI surfaces is already mixed into these numbers. Google’s documentation states that sites appearing in AI features such as AI Overviews and AI Mode “are included in the overall search traffic in Search Console”, and that “they’re reported on in the Performance report, within the ‘Web’ search type”.3 There is no separate breakdown and no way to isolate which impressions came from an AI answer, so the export tells you which question-shaped queries reach you and not whether an AI answer was involved, which is precisely the gap a prompt panel exists to fill.

The truncation

What will Search Console never show you?

The rare queries, which are the ones you most need. Search Console’s own documentation states that on the query tab “anonymized (rare) results are omitted from the table” while still counting toward the chart totals, that the table “can display a maximum of 1,000 rows, so some rare or long-tail rows might be omitted”, and that to protect user privacy the report “omits some queries that are searched a very small number of times”.4 Every one of those limits bites hardest at the long tail, where a nineteen-word question lives. The report is honest about this; the mistake is treating what survives the filter as the population rather than as the visible part of it.

A second blind spot sits underneath the first. Search Console can only show you queries where you appeared in Google’s results, so it cannot show a question you have never ranked for, and a large share of what gets cited was not on that first page anyway. In the same 55,393-query study, 29.8% of the domains an AI Overview cited did not appear anywhere on the first page of results for the same query.1 That is a domain-level count on Google’s own results page, not a claim about your keyword visibility, and its direction is the point: a method anchored to your own search performance is blind to the AI sourcing that never ran through the ranking you can see.

The third limit is scope. Search Console is Google, and panels are per engine: a SIGIR 2026 study built on a benchmark of 11,500 queries measured Jaccard similarity of 0.11 to 0.18 between the sources returned by Google Search, AI Overviews and Gemini.6 Question-shaped Google demand is a reasonable proxy for what people ask other assistants, and it is a proxy. Use it to populate part of the panel and never to validate the panel, because the validation would be circular.

Conversion

How do you turn a query into a panel entry?

Only when the query already contains the question: restore the words search grammar dropped, and stop there.

Exported queryPanel entryWhy
how long does reddit account age matter for postingHow long does a Reddit account need to age before it can post?The question is already there; you restored articles and a verb
is it against the rules to promote your own productIs it against the rules to promote your own product on Reddit?Ambiguity resolved with the context the searcher already had
reddit marketing toolsDo not convert hereA head term carries no question; expanding it is fan-out, not conversion
bavior pricingExcludeBranded and navigational, so it measures whether engines echo you

The rule that keeps conversion honest is narrow on purpose: you may restore the grammar a searcher dropped, and you may not add an intent they did not express. Search grammar strips articles, auxiliaries and politeness because those cost keystrokes and buy nothing in a keyword index; putting them back reconstructs the sentence rather than inventing one. Adding a qualifier the query never had, a segment or a price band or a competitor, creates a different question and quietly moves the entry from the attested pile to the reconstructed one.

Head terms are the common trap, because they look like the most valuable rows in the export. A two-word category term carries no question at all, so converting it means writing the question yourself, which is a different source with a different quality grade. Send those rows to fan-out expansion, where the reconstruction is labelled as such and the axes are fixed in advance. Keeping the two piles separate is what lets you report how much of your panel is attested and how much you wrote.

Exclusions

Which queries should you leave out?

Three sets, and the first is the one that quietly inflates most panels. Branded and navigational queries: anyone searching your product name will retrieve your pages, so a panel entry built from one measures whether engines echo your own site rather than whether you are visible to a stranger, the anti-pattern set out in why your feature list is not a source. Settled questions: if a query returns the same answer from the same source on every engine for two consecutive periods, it has stopped discriminating and belongs on the retirement list rather than in the panel. Questions with no commercial consequence: a definition query you rank for and cannot lose is pleasant to win and tells you nothing.

What to keep despite the temptation to cut: queries where you rank badly. Those rows have low clicks, sit at the bottom of every sorted export, and are the entries most likely to reveal that an engine is answering a real buying question with somebody else’s page. The 2026 critical survey’s minimum checklist for a GEO study makes the general form of this point about measurement, requiring that denominators and null outcomes be retained, and it applies just as forcefully at selection time.7 A panel assembled from your best-performing queries will report a high number and change nothing about your business.

The honest limit of this article

This source is the most convenient and the second-best, and the gap between those two facts is where teams go wrong. Search Console shows one engine’s demand, only where you already appeared, with the rare-query tail withheld and the table capped. The seven-word threshold is a convention rather than a constant: the academic study bins at six words and up, and the vendor measurement that produced seven has never published its method. Nothing here can tell you what buyers ask an assistant when they never search afterwards, the population this whole stage exists to reach. Use this source for volume and attribution, calls and tickets for evidence.

Where a product fits, and where it does not

The entire method here is free, takes about an hour, and needs no tool at any step: one regex filter, one export, one grouping by page, and a written rule for converting a query into a panel entry. Bavior does not connect to your Search Console, does not read your queries, and cannot recover the anonymized tail Google withholds, because that data is not released to anyone. What it does sits downstream of the panel: running a fixed prompt set across five engines on a schedule and recording which sources each answer cited, so you can see whether the question-shaped queries you already rank for are the ones engines answer with your pages. Where a cited source is a live discussion thread it drafts a reply on an account you control, which you approve before anything posts. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).

Sources, all checked 30 Aug 2026
  1. Xu, Iqbal & Montgomery, “Measuring Google AI Overviews”, arXiv:2605.14021 (preprint); 55,393 trending queries, 13 Mar to 21 Apr 2026; 13.7% activation overall, 64.7% question-form against 9.5% non-question (6.8 times, chi-square 10,002.2); Tables 4 and 5; 29.8% of cited domains off the first page: arxiv.org/abs/2605.14021
  2. Question-query study, January 2026: 146 million result pages; AI Overviews on 46.4% of queries of seven words or more. Vendor-published, method not disclosed; described rather than linked per this curriculum’s sourcing rule.
  3. Google Search Central, “AI features and your website”, updated 10 Dec 2025 (AI-feature traffic reported inside the Web search type; first-party): developers.google.com/search/docs/appearance/ai-features
  4. Google Search Console Help, “Troubleshooting data discrepancies” (anonymized rare queries omitted, 1,000-row table limit; first-party): support.google.com/webmasters/answer/17010575
  5. Google Search Console Help, “Advanced filtering and comparison” (regex query filter, RE2 syntax; first-party): support.google.com/webmasters/answer/17011165
  6. Grossman et al., SIGIR 2026; benchmark of 11,500 queries; Jaccard 0.11 to 0.18 between sources returned by Google Search, AI Overviews and Gemini, arxiv.org/abs/2604.27790
  7. “Optimizing Visibility in Generative Engines”, the 2026 critical survey of GEO, arXiv:2607.14035, 15 Jul 2026, Table 6 minimum checklist (denominators and null outcomes retained): arxiv.org/abs/2607.14035
FAQ

Frequently asked questions.

Why seven words rather than some other threshold?

Seven is a convention rather than a constant, and the academic data bins at six. The 55,393-query study reports activation by query length in bins, and its top bin is six words or more, where AI Overviews fired on 58.1% of queries. A January 2026 vendor analysis of 146 million result pages put the figure at 46.4% for queries of seven words or more, by a method it has never published. Either cut separates typed sentences from typed keywords well enough to be useful. What matters more than the exact number is that you take the union with the question-word filter, because length moves activation even among queries that are not questions at all.

Can I see which queries triggered an AI Overview?

No, and Google's documentation says so directly. Sites appearing in AI features are included in the overall search traffic in Search Console and reported within the Web search type, which means AI-surface impressions arrive blended with ordinary search results and there is no breakdown that separates them. You can see that a query brought impressions; you cannot see whether an AI answer was on the page or whether you were inside it. That gap is the reason a prompt panel exists at all, because running the questions yourself is currently the only way to observe what an answer contains.

Should I add the queries where I rank badly?

Yes, and they are usually the most informative entries in the panel. Low-click, low-position rows sit at the bottom of every sorted export and get cut first, but a real buying question you rank poorly for is exactly where an engine is likely to answer with somebody else's page, and that is a finding you can act on. The critical survey's minimum checklist makes the same point about measurement, requiring that denominators and null outcomes be retained; a panel assembled only from queries you already win reports a flattering number and produces no decisions.

What do I do about the queries Search Console hides?

Supplement them from sources that do not depend on Google's reporting, because the withheld tail cannot be recovered from any tool. Google omits queries searched a very small number of times from the query report as anonymized queries, and nobody has access to them, so the substitute is your own site-search logs, which have no such cap, plus the first questions from discovery calls and first-week support tickets, which reach the population that never searched at all. Treat search console as one contributor to the panel rather than its backbone, and cap it at roughly a fifth of the entries.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Your own queries already know the shape.
Filter for the questions.

Start free trial