Home/Learn GEO/Panel size
Prompt research · Sample size

How big should a prompt panel be, and how many runs each?

Presence in an answer is a coin flip with an unknown bias, so the whole question is ordinary proportion arithmetic. Written out, it says two uncomfortable things about how this field reports numbers.

On this page
Share this
Share on X Share on LinkedIn
The short answer

About 40 prompts run 5 times each, per engine per period, is the smallest panel that produces a number worth publishing: 200 observations, a 95% Wilson band of roughly ±7 points at a 50% presence rate. One run of one prompt carries a band of ±45 points and five runs about ±33, so the reporting unit is the panel, never the individual prompt. The field's critical survey records repeated runs changing 9 to 28% of decisions on the surfaces where sampling can be controlled, which is the variation a single run hides.2

Key takeaways
  • Count observations, not prompts. One observation is one prompt run once on one engine, so n is prompts multiplied by runs, counted separately for every engine.
  • Runs are the cheap purchase up to about five per prompt, then prompts. Past roughly 600 observations per engine per period, extra precision is erased by the engines' own drift.
  • Every band here is a floor, because runs of one prompt cluster and nobody has published the correlation that would correct for it. Publish the panel rate with its n, engine and dates attached.
The formula

What is the arithmetic behind a presence rate?

A presence rate is a sample proportion: every run either names you or it does not. Each run is a Bernoulli draw with an unobservable probability p belonging to one prompt on one engine in one week, and what you see is the fraction of runs that named you. Every band on this page is a 95% Wilson score interval, which Brown, Cai and DasGupta recommend in place of the textbook normal approximation.11 At a 50% rate its half-width is 1.96 divided by twice the square root of (n plus 3.84), where n is the observations behind the number.

Two conventions make it usable before you have any data. Evaluate at p = 0.5, because p(1−p) peaks there and planning at the worst case means the band you promised is never narrower than the band you got. And count observations, not prompts: n is prompts multiplied by runs, separately for every engine.

Run it backwards and it sizes the panel directly. To reach a half-width of w you need n = 0.96/w² − 3.84 observations: about 92 for ±10 points, 192 for ±7, and 380 for ±5. Every n on this page is that arithmetic rather than a research finding, so none is footnoted to a paper; the citations below are for where the arithmetic fails and how far the engines move underneath it.

The bands

How wide is the band at each sample size?

Every row is the same formula at a different n, for one engine in one period.

Observations behind the numbern95% Wilson band at a 50% rateWhat you may honestly say
1 run of 1 prompt1±45 pointsNothing. This is an anecdote
5 runs of 1 prompt5±33 pointsNothing about that prompt alone
30 runs of 1 prompt30±17 pointsLarge per-prompt moves only
40 prompts × 1 run40±15 pointsA rough first baseline
40 prompts × 5 runs200±7 pointsPanel rate, per engine, per period
80 prompts × 5 runs400±5 pointsPanel rate with a tighter band
60 prompts × 10 runs600±4 pointsPanel rate plus class-level splits
60 prompts × 20 runs1,200±2.8 pointsPrecision the engines will not hold still for

Read the first four rows as a warning about dashboards. A per-prompt score with a week-over-week arrow is the default output of most visibility reporting, and at five runs that arrow is drawn through a band of ±33 points: it can point either way while nothing underneath it changed, and thirty runs of that one prompt still leaves ±17.

Read the last row as a budget ceiling. Doubling from 600 to 1,200 observations buys 1.2 points of half-width, less than these systems move on their own between one week and the next. The field's critical survey puts day-to-day source overlap at roughly Jaccard 0.34 to 0.42, with similar levels for repeat runs inside 24 hours, and states the conclusion plainly: visibility is a distribution, and “a point estimate is not a stable indicator.”2

Reporting unit

Why is the panel the reporting unit and never the prompt?

The panel is the reporting unit because it is the smallest object in the design with enough observations behind it to carry a usable interval. A single prompt at a realistic run count carries a band several times wider than the effect anyone is trying to detect, so a per-prompt number is not a small measurement, it is not a measurement. The sentence to publish reads: named in 34% of answers across the panel, ±7, on 40 prompts run 5 times on one engine in the week of 24 August, panel version 3.

Individual prompts still earn their place as exhibits rather than metrics. The answer text on one prompt is the only thing that says why the panel rate is what it is, and reading a handful by hand each period is how misrepresentation gets caught: a Tow Center audit of 1,600 queries across eight engines found the tools answered more than 60% of them incorrectly, which a presence flag records as a win.7

Engines are never pooled either. A SIGIR 2026 study of 11,500 queries measured URL-level Jaccard similarity of 0.11 to 0.18 between three answer surfaces belonging to one company,3 and an earlier audit of 672 phase-specific responses found 355 unique domains of which only 26% were cited by both systems tested.4 A blended cross-engine rate averages populations that barely intersect, and cannot answer the only question the metric exists for: which engine to work on next. Nor is the panel a proxy for search rankings: 53% of consulted domains sat outside the organic top 10 in a 4,706-query AI Overviews audit from September 2025,5 and 29.8% of cited domains were absent from the whole first page in a 55,393-query study from spring 2026.6

Allocation

Should you buy more prompts or more runs?

Buy runs first, up to about five per prompt, then buy prompts, then stop at roughly 600 observations per engine per period. At a fixed total the interval is identical either way, because n is simply prompts multiplied by runs, so the question is what each purchase buys besides the band. Runs buy precision on questions you already decided to ask; prompts buy coverage of questions you might have chosen wrongly, the larger risk early on and the smaller one later.

Runs come first because the first few are unusually cheap. Moving a 40-prompt panel from 1 run to 5 takes the band from ±15 to ±7 for 160 extra collections. Going from 40 prompts to 80 at 5 runs each takes it only from ±7 to ±5, and costs 200. Beyond five runs the curve flattens, and the case for more prompts becomes the case for measuring intents you are currently blind to, which belongs to the class split.

The ceiling is set by the engines, not the arithmetic. One vendor time series covering 230,000 prompts and more than 100 million citations between 14 July and 12 October 2025 recorded one large forum domain's citation rate on one assistant falling from roughly 60% of responses to roughly 10% between early August and mid-September.9 It is vendor-published rather than peer-reviewed, but the survey rates “commercial engines differ from one another and vary over time” at high confidence.2

Cost, concretely: 40 prompts at 5 runs across 5 engines is 1,000 collections a week; 60 prompts at 10 runs is 3,000. Reading them costs unevenly too: the one academic citation taxonomy so far measured mean citations per prompt of 6.88, 12.06 and 16.35 across three assistants.8

The literature

Does published research recommend a run count?

It does, and the one published proposal is higher than five. The 2026 critical survey records a proposal of seven to eight repetitions per prompt as a starting point, and qualifies it in the same breath: the number “is not a universal standard”, deriving from a small universe of Swiss queries and at most ten repetitions.2

What it recommends instead is sequential precision analysis: repeat the measurement until the interval around the quantity you care about is narrow enough for the decision in front of you. That is the table above run backwards. Five runs on 40 prompts answers “did the panel move more than 7 points”; eight runs on the same 40 prompts is 320 observations and a ±5.4 band, and where the collections are affordable it is the better-sourced of the two.

Caveats

Where does this arithmetic stop being right?

It stops being right in three places, and two of them will apply to you. The first is extreme presence rates. A Wilson band stays inside 0 and 1 where the textbook normal approximation does not, which is the reason this page uses it. What it stops being is symmetric: at a 2% rate on 200 observations it runs from 0.8% to 5.0%, so quote both ends rather than a half-width. Where runs cluster or counts are tiny, a bootstrap is the safer instrument, as the 2026 statistical treatment of AI visibility recommends after finding citation distributions power-law shaped and many apparent differences between domains inside the measurement noise floor.1

The second is that runs of the same prompt are not independent draws. They share a query, a retrieval path and often a cache, so they cluster, and clustered observations carry less information than the same count of independent ones. The standard adjustment is a design effect: 1 + (m − 1) × the intraclass correlation, where m is runs per prompt. At 5 runs and a correlation of 0.5 that factor is 3, and 200 observations behave like 67, a band of roughly ±12 rather than ±7. No published study estimates that correlation for AI-answer presence, so spread runs across days instead of firing them in one batch, and read every band here as a floor.

The third is that the formula assumes p held still while you sampled. Over a week that is roughly defensible; over a quarter it is not, and a panel rate compared across a long gap measures your work and the engine's changes together with no way to separate them. Everything above is also a within-period interval around one number; whether a change between two periods is real is a two-proportion calculation with much larger requirements, and belongs to the measuring AI visibility stage.

The honest limit of this article

Every number here is arithmetic applied to an assumption that is only approximately true. Treating presence as an exchangeable Bernoulli draw with a stable probability is what makes the formula work, and clustering within prompts, drift within the week and differences between prompts all violate it by an unmeasured amount. The bands are therefore best-case widths. Nor is there external validation that a panel rate sized this way tracks anything commercial: the field's own survey rates the claim that citation scores predict clicks, conversions or revenue at very low confidence.2

Where a product fits, and where it does not

None of this needs software. The band is one spreadsheet cell, 1.96 divided by twice the square root of (n plus 3.84). What does not scale is the collection: 40 prompts, 5 runs, 5 engines is a thousand answers a week, and at that volume manual gathering stops being tedious and starts being unreliable, because a bored person skips runs and a skipped run is exactly the failure the arithmetic cannot see. Bavior runs a fixed prompt set across five engines on a schedule and logs the sources each answer cited, so the denominator stays whole. It cannot make a small sample mean more than it does, and whether a 4-point move inside a 7-point band counts as a result stays your judgement. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 30 Aug 2026).

Sources, all checked 30 Aug 2026
  1. Sielinski, “Quantifying Uncertainty in AI Visibility”, Mar 2026, arXiv:2603.08924 (preprint); power-law citation distributions, differences inside the noise floor: arxiv.org/abs/2603.08924
  2. “A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035; daily source-level Jaccard 0.34–0.42, 9–28% repeated-decision changes, the seven-to-eight repetitions proposal, the confidence table: arxiv.org/abs/2607.14035
  3. Grossman et al., SIGIR 2026; 11,500 queries; URL-level Jaccard 0.11–0.18 across Google organic, AI Overviews, Gemini: arxiv.org/abs/2604.27790
  4. Li & Sinnamon, 2024; 1,008 responses audited, 672 analysed by phase, 355 unique domains of which 26% cited by both systems
  5. Kirsten et al., Findings of ACL 2026; 4,706-query audit of Google AI Overviews, issued September 2025; 53% of consulted domains outside the organic top 10: aclanthology.org/2026.findings-acl.526
  6. Xu, Iqbal & Montgomery, 2026 (preprint); 55,393 trending queries, 13 March to 21 April 2026; 29.8% of cited domains absent from the first page: arxiv.org/abs/2605.14021
  7. Jaźwińska & Chandrasekar, “AI Search Has a Citation Problem”, Tow Center, Columbia, 6 Mar 2025; eight engines, 1,600 queries, incorrect answers to more than 60% of them: cjr.org
  8. Zhang, He & Yao, “From Citation Selection to Citation Absorption”, arXiv:2604.25707 (preprint); 602 controlled prompts, 21,143 search-layer citations; mean citations per prompt 6.88 / 12.06 / 16.35: arxiv.org/abs/2604.25707
  9. Vendor study of 230,000 prompts and over 100 million citations, 14 July to 12 October 2025, three assistants; one large forum domain's citation rate on one assistant fell from about 60% of responses to about 10% between early August and mid-September 2025. Vendor-published, so described rather than linked.
  10. Vendor index of 126 million US AI search prompts, January to April 2026; one assistant cited 15 sources per response on average against another's 3, and 36 brands held visibility on every platform measured. Vendor-published, so described rather than linked.
  11. Brown, Cai & DasGupta, “Interval Estimation for a Binomial Proportion”, Statistical Science 16(2), 2001; normal-approximation coverage is erratic near 0 and 1: doi.org/10.1214/ss/1009213286
FAQ

Frequently asked questions.

Is 40 prompts a rule, or just a round number?

It is the round number nearest the arithmetic. Reaching a ±7-point Wilson band at a 50% presence rate takes 192 observations, and 40 prompts run 5 times gives 200, so the 40 comes from the band you want and the runs you can afford, not from anything about prompts. Twenty prompts run ten times gives the same 200 observations and the same band, with narrower coverage of buyer intent. Choose the split by how confident you are that you picked the right questions, and the total by the band you are willing to publish.

Why five runs and not three?

Five is where the cheap precision runs out, not a threshold with evidence behind it. On a 40-prompt panel, one run gives a ±15-point band, three runs about ±9 and five runs about ±7; the sixth buys 0.6 of a point and every run after that buys less than half of one. Three runs is a defensible economy if collections are expensive, and it should be stated in the report rather than hidden, because a ±9 band changes which movements you are allowed to call results. What matters more than the exact number is that it stays the same between periods.

Can I run all the collections on the same morning?

Batching them narrows your apparent interval without improving the estimate, so spread the runs across days. Runs fired within minutes of each other share a retrieval state and often a cache, which makes them more alike than two genuinely independent draws would be, so the observations are clustered, and clustered observations carry less information than their count suggests. The correction has a name, the design effect, and no published study supplies the correlation you would need to compute it for AI answers. Spreading the runs is the cheap way to avoid needing it.

How does this differ from working out whether a change was real?

This page sizes the interval around one number; testing a before-and-after difference is a two-proportion comparison that needs several times as many observations. From a 30% base, detecting a 20-point move takes roughly 91 observations per period, a 10-point move roughly 353, and a 5-point move roughly 1,372, so a weekly panel of 200 observations cannot honestly detect anything smaller than a large change. Those figures are the two-proportion formula worked out in Measuring AI visibility on measuring AI visibility, and it is the single most useful thing to read before promising anyone a monthly trend line.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Two hundred observations, or silence.
There is no third option.

Start free trial