Queries led by a question word, and queries of seven words or more: those two sets, unioned, are the closest thing in your own data to a prompt. Fix the word list in writing so the filter is reproducible, and prefer a published list to one you invent. The length threshold is a proxy for the same behaviour, catching the queries that were typed as sentences without an interrogative at the front.
The phrasing filter carries the sharpest evidence in this field. Across 55,393 trending queries issued over 40 days in spring 2026, Google AI Overviews activated on 13.7% of queries overall, but on 64.7% of question-form queries against 9.5% of non-question ones: a 6.8 times difference, chi-square 10,002.2, p < 10⁻³⁰⁰.1 Effects that large are rare, and this one was measured on live search, not in a simulation.
Length earns its place as a second filter because it moves the number independently of phrasing. In the same study, activation ran 3.4% at two words, 20.0% at four, 33.9% at five and 58.1% at six words or more; restricted to non-question queries alone it still climbed from 9.9% at one word to 38.7% at six or more.1 A long query is doing something a keyword is not, whether or not it opens with how. The study bins at six words and up, so a seven-word cut sits inside the strongest bin rather than at its edge; a vendor dataset of 146 million result pages put AI Overviews on 46.4% of queries of seven words or more, by a method it has never published.2
Both filters describe when an AI answer appears, not how often anyone asks. Keep that distinction live, because it changes what the output is for: this source tells you which of your existing demand is shaped like a prompt, and it cannot tell you which prompts exist that you have never ranked for. Those come from calls and tickets.