Home/Learn GEO/Six classes
Prompt research · Taxonomy

Which six prompt classes should you track separately?

A panel that reports one blended number averages six populations with different base rates, different definitions of a win, and different places where the fix lives. It moves when the mix moves.

On this page
Share this
Share on X Share on LinkedIn
The short answer

Six classes are worth separating: category definition, shortlist, comparison, fit, objection and how-to. They differ in how often you are named, in what counts as a win, and in which sources an engine retrieves, so a blended presence rate moves when the panel's class mix changes rather than when your visibility does. Comparison and objection carry the highest stakes, because there a citation attached to a wrong description is worse than silence: a Tow Center audit of 1,600 queries across eight engines found more than 60% returned incorrect answers.3

Key takeaways
  • Category definition and how-to are answerable from your own pages, shortlist and objection are not, comparison and fit sit in between: that split says which team can act.
  • A win differs by class: present anywhere in the list on shortlist, described accurately on comparison, answered honestly on objection. One presence flag asks the right question of two classes in six.
  • Five easy category-definition prompts added to a 40-prompt panel lift the blend from 37.4% to 41.0% with nothing else changed, so freeze the mix and print it.
The composition effect

Why does averaging the six classes destroy the signal?

Averaging destroys the signal because the classes have very different base rates, so the blend is as much a property of the panel's composition as of your visibility. Take a 40-prompt panel: 12 category-definition prompts where you are named 70% of the time, 10 shortlist at 15%, 8 comparison at 30%, 4 fit at 25%, 3 objection at 10% and 3 how-to at 45%. The blend is 37.4%. Add five more category-definition prompts you already win, change nothing else, and it reports 41.0%: a 3.6-point rise produced by an editorial decision.

The second reason is that a win is not the same event in each class. On a shortlist prompt, presence anywhere in the list is the whole outcome; on a comparison prompt it is worth nothing if the sentence next to your name is wrong; on a category-definition prompt the win is the correct category rather than a neighbouring one. A single presence rate asks the right question of two classes and the wrong question of four.

The third reason is that the classes are answered from different source populations, so they respond to different work: category-definition and how-to answers lean on pages a company can write, shortlist and objection answers on pages it cannot. One number says something moved without saying which team could have moved it. The prompt research stage builds the panel; this article is the tag on each row.

The taxonomy

What are the six classes, and what counts as a win in each?

Each of the six carries its own definition of a win: named at all, present in the list, described accurately, matched to your real constraints, answered honestly, and cited as the method.

ClassWhat a buyer typesWhat counts as a winWhere the fix usually lives
Category definition“What is X?” · “What kind of tool does Y?”Named at all, and placed in the correct categoryYour own pages and encyclopedic entries
Shortlist“Best X for Y” · “Top tools for Z”Present in the list, any positionThird-party roundups you do not control
Comparison“X vs Y” · “alternatives to X”Framed accurately, not merely presentSplit: your pages plus third-party write-ups
Fit“Is X right for a five-person team?”Your real constraints stated correctlyYour own pages, if the constraints are written down
Objection“Is X safe?” · “Does X get you banned?”The honest answer, sourced to somebody credibleForum and community threads
How-to“How do I do Z?”Your method cited as the methodYour own documentation and guides

Two of the six are unusually answerable from your own domain. The only academic taxonomy of AI citations published so far, 602 controlled prompts producing 21,143 search-layer citations across three assistants, found official sources the largest category on every platform it measured, at 34.22%, 46.35% and 44.07% of citations.8 That is the ceiling on a category-definition or how-to prompt: those are the questions your own documentation may answer, and where the official share is highest.

Tag exactly one class per prompt, and split a question that spans two: “is X or Y safer for a small team” is comparison and objection and fit at once, the answer populations differ, and a combined prompt corrupts two class rates. The same study measured a per-engine gap in how many sources an answer carries at all, with means of 6.88, 12.06 and 16.35 citations per response.8 A shortlist prompt has more than twice as many slots on one engine as on another, so class rates are reported per engine and never pooled.

Misrepresentation

Why do comparison and objection carry the highest stakes?

Because on those two a citation attached to a wrong description costs more than no citation at all, and wrong descriptions are the normal case. The Tow Center ran 1,600 queries across eight engines in March 2025 and found more than 60% returned incorrect answers; the worst engine was wrong 94% of the time and produced 154 citations leading to error pages across its 200 prompts, and even the best was wrong 37% of the time.3 Its earlier test of 200 quotes found 153 responses partially or entirely incorrect, with uncertainty signalled only 7 times.2

Two measurements put the same problem at sentence level. Earlier work found only 51.5% of sentences in generative-search answers were fully supported by the citations attached to them,4 and a 2026 study classified roughly 11% of 98,020 atomic claims as insufficiently supported.5 An engine right about your category and wrong about your pricing tier has still produced an answer that loses the deal, and a presence flag records it as a win.

Three cheap consequences follow. Log the full answer text on comparison and objection prompts, not only the presence flag: what you need to read is the sentence, not the boolean. Run those two classes at a higher run count, because you are estimating two rates from the same runs and the second, whether the description is right, needs a hand-read sample. And on objection prompts the fix usually sits off your domain, in the community thread the honest answer is retrieved from, which is the off-site GEO stage's territory.

The flat line

Why does the shortlist class barely respond to on-page work?

Because most of what an engine cites on a “best X” question is somebody else's page, which puts the lever outside your domain and release cycle. Two vendor datasets point the same way. One examined the top 1,000 pages a major assistant cited during September 2025 and reported only about a third fell into categories a business could compete in.11 Another tracked 7,683 pages and 47,097 citations across four brands between March and June 2026 and put brand-owned pages at roughly 2% of what gets cited about them.12 Neither is peer-reviewed, and the four-brand figure is illustrative, not a constant.

Set that against the academic taxonomy above, where official sources were 34–46% of citations, and the two look contradictory. They are not: the official share is a property of the prompt class, not of the web. A definitional question pulls official pages; a “best tools for” question pulls roundups. Nobody has published an official-source share by prompt class, so treat the gap as an argument for splitting the classes, not a settled number.

The consequence is a diagnosis rather than a tactic. If your panel is mostly shortlist prompts and the line is flat, that flatness tells you where the answer is sourced, not that your content is bad: publishing your own roundup competes in the minority slice, being named accurately in other people's competes in the majority.

The noise floor

Why can a class rate move when nothing about you changed?

Because two of the three things a class rate depends on belong to the engine rather than to you. The first is whether an AI answer appears at all: a 55,393-query study of Google AI Overviews measured 13.7% activation overall, rising to 64.7% on question-form queries.5 The classes are not phrased alike: category-definition, fit and objection prompts are natural questions, shortlist prompts are usually noun phrases. A class can look flat because its prompts rarely trigger an answer, which is why the log records whether an answer appeared separately from whether you were in it.

The second is that the source population behind a class is neither the ranked list nor a stable set. Google documents that AI Overviews and AI Mode may fan one prompt out into several searches across subtopics,10 and the same audit found nearly 30% of cited domains never appeared in the first-page results.5 A benchmark of 11,500 queries reported URL-level Jaccard similarities of 0.11 to 0.18 between Google organic, AI Overviews and Gemini, and less consistency across repeat runs.6 Repeated sampling across three engines put many apparent differences between domains inside the noise floor.9 Read a class move against that floor before crediting your own work.

Operating rules

How do you tag and report by class without inflating the panel?

Tag each prompt with exactly one class when you add it, freeze that tag for the life of the prompt, and publish a rate per class per engine with the count beside it. Freezing matters: re-tagging mid-quarter silently changes the denominator of two classes at once, and the move in both is an artefact you will spend a week explaining.

Class-level reporting costs sample size. Ten prompts run five times each is 50 observations, a 95% band of roughly ±14 points at a 50% rate, wide enough that a class rate of 30% and one of 44% are the same measurement. Panel-level numbers from the same runs are far tighter, which is why the panel stays the headline and the class split stays diagnostic. The sizing arithmetic is in the companion article.

Three more rules keep the tag honest. Hold the per-class counts roughly constant between periods and record them in the panel version, so a reader sees the mix behind the number. Never rebalance the mix in a period where you also claim an improvement, because the arithmetic at the top of this page will produce one for you. And keep the class in the log rather than re-deriving it from the prompt text later: it is a judgement made once, by the person who knew why the prompt was added.

The honest limit of this article

The six-class split is a working convention, not a validated instrument. No published study segments AI-answer prompts this way, so there is no external estimate of how far the classes' base rates really diverge. The 37.4% to 41.0% example above is arithmetic from plausible inputs, not a measurement. Two reasonable teams could land on five classes or eight. What is not optional is that some split exists and that any published number carries its class mix, because a blended rate that hides the mix can be moved by editing the panel instead of by earning anything.

Where a product fits, and where it does not

Classification is free: add a class column to your prompt sheet, tag each row once, and report six rates per engine instead of one. No software can do the tagging, because the class depends on why the prompt was added, which lives in the head of whoever heard a buyer say it. Bavior runs a fixed prompt set across five engines on a schedule and records which sources each answer cited, which makes a per-class source population visible rather than guessed; where a cited source is a live discussion thread it drafts a reply on an account you control, for you to approve before it posts. It does not decide your taxonomy, does not read your sales calls, and cannot tell you whether the sentence describing your product is accurate: that is a human read, on a sample. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).

Sources, all checked 29 Aug 2026
  1. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (survey preprint): arxiv.org/abs/2607.14035
  2. Tow Center for Digital Journalism, Columbia, “How ChatGPT Search (Mis)represents Publisher Content”, Nov 2024; 200 quotes from 20 publishers, 153 responses partially or entirely incorrect: cjr.org
  3. Jaźwińska & Chandrasekar, “AI Search Has a Citation Problem”, Tow Center, 6 Mar 2025; 1,600 queries across eight engines, more than 60% incorrect: cjr.org
  4. Liu et al., 2023; 51.5% of sentences in generative-search answers fully supported by their citations (reported in the critical survey at note 1)
  5. Xu, Iqbal & Montgomery, 2026 (preprint); 55,393 trending queries, 13 March to 21 April 2026; 13.7% AI Overview activation overall, 64.7% on question-form queries; roughly 11% of 98,020 atomic claims insufficiently supported: arxiv.org/abs/2605.14021
  6. Grossman et al., SIGIR 2026; representative sample of 11,500 queries; URL-level Jaccard 0.11–0.18 across Google organic, AI Overviews and Gemini: arxiv.org/abs/2604.27790
  7. Kirsten et al., Findings of ACL 2026; audit of Google AI Overviews against organic search and five generative systems: aclanthology.org/2026.findings-acl.526
  8. Zhang, He & Yao, “From Citation Selection to Citation Absorption”, arXiv:2604.25707 (preprint, descriptive); 602 controlled prompts and 21,143 search-layer citations: arxiv.org/abs/2604.25707
  9. Sielinski, “Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement”, Mar 2026, arXiv:2603.08924 (preprint): arxiv.org/abs/2603.08924
  10. Google Search Central, “AI features and your website”, updated 10 Dec 2025 (query fan-out, eligibility; first-party): developers.google.com/search/docs/appearance/ai-features
  11. Vendor study of the top 1,000 pages one major assistant cited during September 2025; roughly a third in categories a business can compete in. Vendor-published, not linked.
  12. Vendor study of 7,683 pages and 47,097 citations, March to June 2026, across four brands; a brand owns roughly 2% of what gets cited about it. Vendor-published, not linked.
FAQ

Frequently asked questions.

Can I report one blended AI visibility number if I also publish the class split?

Publishing both is fine as long as the blended number carries its class mix and the panel version that produced it. The failure mode is not the average itself but an average whose composition changed between periods: adding five easy category-definition prompts to a forty-prompt panel can lift the blended rate by several points without anything about your visibility changing. If the mix is frozen and printed, the blend is a legitimate summary. If the mix is free to move, the blend is an editorial choice wearing a metric's clothes.

Which class should a small team start with?

Start with objection and comparison, because those two are where an error costs a deal and where you are least likely to find out any other way. A wrong answer to "is X safe" or "how does X compare to Y" removes a buyer before they ever contact you, and nothing in your analytics records it. Category definition and how-to are the classes your own pages can move, so they are the right second priority. Shortlist is last for a small team, not because it does not matter but because the work is off-site and slow.

How many prompts do I need per class?

Ten prompts run five times each is the practical floor, and it buys a 95% band of roughly ±14 points at a 50% rate, enough to see a large class-level move and not enough to see a small one. Twelve prompts run ten times gets you to about ±9. Below ten prompts a class rate is a description of those particular prompts rather than of the class. The panel-level number from the same runs is much tighter, which is why the panel stays the headline and the class split stays a diagnostic rather than a KPI.

What do I do with a prompt that belongs to two classes?

Split it into two prompts, one per class, and phrase each the way a buyer would actually ask it on its own. A question like "is X or Y safer for a small team" is comparison, objection and fit at once, and the sources an engine retrieves differ for each of those intents, so a combined prompt produces an answer you cannot attribute to any class and a rate that corrupts two of them. Splitting costs you one extra row in the panel and one extra run per period, which is the cheapest fix available anywhere in this method.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Six classes, six rates.
One number tells you nothing.

Start free trial