Home/Learn GEO/Statistical power
Measuring AI visibility · The calculation

How much statistical power does a before-and-after test need?

One formula decides whether your before-and-after comparison means anything, and it is short enough to check by hand. Every number here follows from it, including the uncomfortable one about weekly reports.

On this page
Share this
Share on X Share on LinkedIn
The short answer

A before-and-after comparison needs enough observations to detect the smallest change you would act on, and the two-proportion sample size formula says how many: n = (1.96 + 0.84)² × [p₁(1−p₁) + p₂(1−p₂)] ÷ δ², at 80% power and a two-sided 5% significance level. From a 30% base that is about 353 observations per period for a 10-point move and 8,400 for 2 points, so a weekly report of 200 observations per engine resolves roughly 13 to 14 points and nothing smaller. Underpowered testing is the normal condition of published research: a 2013 review of 49 meta-analyses covering 730 studies put its median statistical power at 21%.10

Key takeaways
  • Count observations, not prompts. One observation is one prompt run once on one engine, so 60 prompts run six times is 360.
  • The requirement is quadratic: halving the change you want to detect quadruples the data. Observations ≈ 3.5 ÷ δ² is close enough to check in a meeting.
  • A weekly number cannot honestly carry a trend arrow. Publish monthly, pair the design where the panel is stable, and hold back prompts you never act on so engine drift does not read as your effect.
The formula

What is the sample size formula for a before-and-after test?

Comparing two rates uses the ordinary two-proportion formula, and every figure here falls out of it: n = (zα/2 + zβ)² × [p₁(1−p₁) + p₂(1−p₂)] ÷ (p₂ − p₁)². At the conventional settings, a two-sided 5% significance level so zα/2 = 1.96 and 80% power so zβ = 0.84, the leading constant is 2.8² = 7.84.

Each term is something you decide before collecting anything. p₁ is the rate you start from, measured rather than guessed. p₂ is the rate you want to be able to detect, a business decision: the smallest improvement that would change what you fund next quarter. δ = p₂ − p₁ is that difference, written as a decimal. And n is observations in each period, where one observation is one prompt, run once, on one engine.

Work one row by hand so the rest is checkable. From 30% to 40%: 0.30 × 0.70 = 0.21, plus 0.40 × 0.60 = 0.24, sum 0.45, times 7.84 is 3.528, divided by δ² = 0.01 gives 352.8. So 353 observations before and 353 after.

None of this is specific to AI search: it is the arithmetic of any comparison of a binary outcome, and the version above is the smaller of the two in circulation. Casagrande, Pike and Smith’s continuity-corrected formula, which most calculators run, returns a larger n, so read every figure here as a floor.6 Precision on one rate and detection of a difference between two are different calculations, and a panel sized for the first is routinely too small for the second; sizing for precision is Prompt research’s how many prompts and runs.

The numbers

How many observations does each size of change need?

From a 30% starting rate, per engine: about 91 observations per period for a 20-point move, 353 for 10 points, 1,372 for 5 and 8,400 for 2. The middle column is the variance term.

Change to detectp₁(1−p₁) + p₂(1−p₂)Observations per periodA panel that delivers it
30% → 50% (20 points)0.460~9120 prompts × 5 runs
30% → 40% (10 points)0.450~35360 prompts × 6 runs
30% → 35% (5 points)0.438~1,37260 prompts × 23 runs, not weekly
30% → 32% (2 points)0.428~8,400Not achievable at any sane cadence

Two properties matter more than any individual row. The first: the requirement is quadratic in δ, so halving the effect you want to detect quadruples the data. That is why eighteen points between the first row and the last costs a factor of about ninety, and it decides what cadence a programme can afford.

The second is that the variance term barely moves, 0.428 to 0.460 across the table, so almost all the variation comes from δ alone. That gives a shortcut good to within about 5% in this range: observations ≈ 3.5 ÷ δ², with δ as a decimal.

The consequence

What can a weekly report of 200 observations detect?

A weekly report carrying 200 observations per engine can detect a move of roughly 13 to 14 points. Run the shortcut backwards: 3.5 ÷ 200 = 0.0175, whose square root is 0.132; the exact formula from a 30% base gives about 13.5 points. A rise from 30% to 44% is at the edge of detectable; a rise from 30% to 35% is invisible, however confidently the dashboard renders it.

The implication is uncomfortable: a weekly GEO report cannot honestly detect anything but a large move. Anything smaller is a fluctuation, and reporting it as progress is how a programme loses credibility in month four, when the line drops back for no reason anybody can name.

The same arithmetic says how to read a published case study: with the sample size in hand. Generative engine optimization’s best-known claim, that GEO increases visibility by 40%, is listed by the 2026 critical survey as rejected as a general claim, because the figure is “a relative maximum on one metric under a specific configuration” from the founding paper, which measured a position-weighted metric inside a fixed context rather than live-engine visibility.24 Most fail more quietly: when someone reports visibility rising from 22% to 27% after a rewrite, ask how many observations sit behind each figure; below roughly 1,400 per side that comparison cannot separate effect from noise. That is underpowered rather than dishonest, and it applies to anything we publish too.

Type M and Type S

What does a significant result from an underpowered test overstate?

Its own effect size, by a factor you can work out in advance. When power is low, only the estimates that overshot are large enough to clear the significance threshold, so filtering on significance selects for exaggeration. Gelman and Carlin call that distortion the exaggeration ratio, “the factor by which the magnitude of an effect might be overestimated”, and pair it with the Type S error, the probability that a significant estimate carries the wrong sign.9

The practical shape: a comparison powered to detect 13 points, on the occasions it returns a significant result for a true effect of 4 points, reports something far larger than 4. That is not a fault in the test; it is what a filter on significance does to a noisy estimate, and it is why a GEO programme’s first quarter so often produces a number nobody can reproduce.

This is the normal condition of published research rather than a corner case: the 2013 review that established the point put median statistical power across 49 neuroscience meta-analyses, covering 730 studies, at 21%.10 No equivalent audit of AI-visibility case studies exists, and since most are single-run, assume they sit lower. The defence is not a cleverer test: fix the effect you would act on before you look, and treat a first significant result as a hypothesis to re-run rather than a finding to announce.

Three levers

How do you buy detection power without more runs?

Three levers raise what a comparison can detect without adding a run to the schedule: lengthen the period, pool within a prompt class, pair the design. They act on the comparison rather than the panel, so they stack with Prompt research’s prompts-versus-runs allocation.

Lengthen the period. A monthly comparison aggregates four weeks of runs you were already making, so 200 observations become 800 and the detectable effect falls from about 13.5 points to about 6.6. Note the exchange rate: four times the data buys twice the resolution, because the requirement is quadratic. Patience alone cannot rescue the weekly cadence; it has to be abandoned as a reporting unit and kept only as an operational check.

Pool within a prompt class. Compare a whole class of prompts before and after rather than a single prompt, since a content change should show up at class level anyway. The discipline is pooling within a class and never across: classes move independently, so mixing a class where you gained twelve points with four flat ones dilutes a real effect into an undetectable one, and the pooling has cost you power rather than bought it. Pooling across engines is worse, because the sources they retrieve barely overlap.3

Pair the design. Run the identical panel before and after and compare prompt by prompt with McNemar’s test, which uses only the prompts that changed state and so removes prompt-to-prompt variance. Connor’s formula gives the number of pairs: [1.96√ψ + 0.84√(ψ − δ²)]² ÷ δ², where ψ is the share of prompts that flip.7 To detect a 10-point net shift when 20% flip, that is about 155 prompt-pairs, 310 observations against 706 unpaired. The saving comes from stability: largest on engines whose answers move least, smallest on the ones that re-roll heavily between runs.

Where it breaks

When does this calculation stop being valid?

Three assumptions sit underneath the formula, and AI-visibility measurement violates all three in the direction that flatters the numbers. The first is the normal approximation, erratic at rates near 0 or 1: Brown, Cai and DasGupta show its coverage is poor there and does not improve smoothly as n grows.8 If you are named in 4% of answers, the sample size this formula quotes is optimistic; use a Wilson or bootstrap interval, as the March 2026 statistical treatment of AI-visibility measurement recommends after finding citation counts power-law shaped and rankings unstable throughout the frequently cited set.1

The second is independence. Runs fired as one batch within a few minutes are not fresh draws: they meet the same index state, the same cache and the same model checkpoint, so spread them across the whole period. The effective sample size is smaller than the nominal one whenever runs cluster in time, and no design-effect estimate for AI-visibility panels exists, so nobody can tell you by how much.

The third is that the engine did not change between your two measurements. This is violated routinely and it is the most consequential: when it fails you have measured the platform rather than your work. A commercial three-month study of 230,000 prompts and over 100 million citations, collected between mid-July and mid-October 2025, recorded one assistant’s citation rate for a large discussion platform falling from roughly 60% of responses to roughly 10% in six weeks.5 The 2026 critical survey rates “commercial engines differ from one another and vary over time” at high confidence.2 The defence is a holdout set of prompts you never act on, so platform drift shows up separately from your effect; the sibling article on the report itself covers how to build one.

The honest limit of this article

Every number here is a lower bound rather than an estimate. The formula assumes independent draws from a system holding still, and AI answer engines provide neither: runs within a day are correlated, and the engines change underneath the measurement without announcement. Both violations push the same way, so the true requirement is larger than the table says, and no one can say how much larger, because no design-effect estimate for a prompt panel exists. Treat 353 observations for a 10-point move as a floor, and be suspicious of anyone quoting less, ourselves included.

Where a product fits, and where it does not

The whole calculation is four multiplications and a division, and a spreadsheet does it. No software is needed to size a test, and none can raise the power of one already too small. Bavior runs a fixed prompt panel across five engines on a schedule and records every run, including the ones where nothing happened, which is the raw material this calculation consumes. What it does not do is run the significance test, decide your holdout set, or forecast traffic or revenue from citations: the field’s own survey rates that last link at very low confidence.2 The free visibility check and free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).

Sources, all checked 30 Aug 2026
  1. Sielinski, “Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement”, Mar 2026, arXiv:2603.08924 (preprint); bootstrap intervals recommended for citation-share estimates: arxiv.org/abs/2603.08924
  2. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (survey preprint); confidence table with the rejected general claim and “Commercial engines differ from one another and vary over time” rated High: arxiv.org/abs/2607.14035
  3. Grossman et al., “How Generative AI Disrupts Search”, SIGIR 2026, arXiv:2604.27790; 11,500 queries; the paper reports “Jaccard similarities between 0.11 and 0.18” across Google organic, AI Overviews and Gemini sources, summarised in its abstract as under 0.2: arxiv.org/abs/2604.27790
  4. Aggarwal et al., “GEO: Generative Engine Optimization”, KDD ’24; source of the 40% figure, a relative maximum on a position-weighted metric in a fixed context: arxiv.org/abs/2311.09735
  5. Citation study published 10 Nov 2025: 230,000 prompts and over 100 million citations, collected 14 Jul to 12 Oct 2025; one assistant’s citation rate for a large discussion platform fell from ~60% of responses to ~10% in six weeks. Vendor-published, so described rather than linked, per this curriculum’s rule.
  6. Casagrande, Pike, Smith, “An Improved Approximate Formula for Calculating Sample Sizes for Comparing Two Binomial Distributions”, Biometrics 34(3):483, 1978; the continuity-corrected version: doi.org/10.2307/2530613
  7. Connor, “Sample Size for Testing Differences in Proportions for the Paired-Sample Design”, Biometrics 43(1):207, 1987; the paired-design formula: doi.org/10.2307/2531961
  8. Brown, Cai, DasGupta, “Interval Estimation for a Binomial Proportion”, Statistical Science 16(2), 2001; coverage is erratic near 0 and 1: doi.org/10.1214/ss/1009213286
  9. Gelman, Carlin, “Beyond Power Calculations: Assessing Type S (Sign) and Type M (Magnitude) Errors”, Perspectives on Psychological Science 9(6):641–651, 2014; the exaggeration ratio: doi.org/10.1177/1745691614551642
  10. Button et al., “Power failure: why small sample size undermines the reliability of neuroscience”, Nature Reviews Neuroscience 14(5):365–376, 2013; median power 21% across 49 meta-analyses, 730 studies: doi.org/10.1038/nrn3475
FAQ

Frequently asked questions.

How many observations do I need to prove my GEO work made a difference?

It depends entirely on how big a difference you are trying to prove, and the relationship is quadratic. From a 30% starting rate, detecting a 20-point move needs about 91 observations per period, a 10-point move about 353, a 5-point move about 1,372 and a 2-point move about 8,400, per engine, at 80% power and a two-sided 5% significance level. An observation is one prompt run once on one engine, so 60 prompts run six times is 360. Decide the smallest change that would change a funding decision, then size the test to that, rather than collecting data and asking afterwards what it can support.

Can I just report AI visibility weekly and note that it is noisy?

You can report it weekly as an operational check, and you should not draw a trend arrow on it. With about 200 observations per engine per week, the smallest change distinguishable from noise is roughly 13 to 14 points from a 30% base, so almost every week-over-week movement you will see is a fluctuation. The practical arrangement is to run weekly, read the answers weekly for qualitative signal such as new competitors appearing or new false statements about your product, and publish the numbers monthly, where four weeks of the same runs cut the detectable effect to about 6.6 points.

Does a paired before-and-after design really need less data?

Yes, and the saving is large when your panel is stable. A paired design runs the identical prompts before and after and compares them prompt by prompt with McNemar's test, which counts only the prompts that changed state, so prompt-to-prompt variance drops out of the comparison. To detect a 10-point net shift when 20% of prompts flip between measurements, you need about 155 prompt-pairs, 310 observations in total, against 706 for the unpaired design. The saving comes from stability, so it shrinks on engines whose answers re-roll heavily between runs, and it disappears entirely if you change the prompt list between the two measurements.

My rate is 4%. Does this formula still apply?

Not reliably. Near 0 or 1 the normal approximation behind it is erratic, its coverage is poor, and the sample size it quotes is optimistic. At those rates use a Wilson interval or a bootstrap instead, which is what the March 2026 statistical treatment of AI-visibility measurement recommends after finding citation counts to be power-law shaped and rankings unstable throughout the frequently cited set. The practical consequence for a brand sitting at 4% is that almost nothing is detectable week to week, and the honest report says so rather than showing a percentage change on a base that small.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Size the test before you run it.
Then the number means something.

Start free trial