Home/Learn GEO/Tactics that test null
Content for AI citation · Methods

Why do most content tactics test null?

The effect is smaller than the test was ever able to see. A null is a fact about your measurement before it is a fact about the tactic, and telling the two apart is most of the skill.

On this page
Share this
Share on X Share on LinkedIn
The short answer

A null result means the test could not see an effect, which is a statement about the test before it is a statement about the tactic. Expect it as the normal outcome: C-SEO Bench, the NeurIPS 2025 benchmark of ten conversational-SEO methods across six domains, found only three of 54 method-and-domain cases significantly positive, and in its results table every one of those ten methods carries a standard deviation larger than its own mean, usually several times larger.1 Numbers shaped like that cannot resolve a small effect.

Key takeaways
  • Nulls dominate because the effects are small, the same edit helps some queries and hurts others, and the answer changes between identical runs.
  • To claim a tactic does nothing you need an interval excluding the smallest effect you would act on. The field's survey still has to ask studies for intervals and distributions rather than a mean.3
  • A negative result carries far more information. SAGEO Arena, at KDD 2026, found body-text-only rewriting losing at retrieval, reranking and citation for all ten strategies it tested.2
  • Read your own null against untouched control pages, not against last month: two-month page overlap in AI Overviews is 18%, against 45% for organic Google.4
The prior

Why is a null the expected result here?

Because a content edit is a small intervention at the last stage of a pipeline, measured through an instrument that moves on its own.

Start with the size of the thing being looked for. C-SEO Bench, at the NeurIPS 2025 Datasets and Benchmarks Track, ran a total of ten content methods across six domains and more than 1.9k queries. In its retail column the strongest of the ten moved the target document by 0.36 rank places and the weakest by −0.07, while placing the document first in the model's context, which is what ranking does, was worth 2.77.1 A third of a rank place is not merely hard to detect. It is too small to change what you do on Monday.

The instrument moves too. On the surfaces where temperature can be controlled, repeated runs at temperature zero change 9–28% of decisions across 4,706 queries.4 Ask the same question twice and it is not the same experiment, and any effect smaller than that churn is invisible by construction.

Then there is the population. C-SEO Bench reports shifts that “are both positive and negative, and partially cancel each other out”: its strongest retail method helped 26.2% of cases and hurt 12.8%, so “the overall average effect is close to zero, but with high variance.”1 Three of 54 cases came out significantly positive, after a Holm-Bonferroni correction, so those three are not an artefact of testing many methods at once.1

Dispersion

What does a standard deviation larger than the mean tell you?

It tells you that for any single query the outcome is decided by which query it is, not by which tactic you applied. The best content method in that retail column scores 0.36 ±1.47, and even the reference SEO move, the strongest number in the table, is 2.77 ±2.31.1 The spread is the finding: a document can gain four places or lose three under the same intervention.

Three consequences follow. A single page compared before and after on a single prompt is one draw from that distribution, so it carries almost no information about the average, in either direction. A positive average is compatible with a large minority of losses, which is why the 2026 critical survey's methods chapter asks for “absolute effects, relative effects, intervals, and distributions, not merely a mean.”3 And the observations you need grow with the variance and shrink with the square of the effect, so a small effect inside a wide distribution is expensive to see at all.

The arithmetic for that lives elsewhere on this site: the sample-size formula gives the observations per period for a given size of change, and how many prompts and runs turns it into a panel you can operate. Read either before designing a test, not after reading its result.

What a null is not

Does a null mean the tactic does nothing?

No, and the gap between those two sentences is where most GEO reporting goes wrong. “We did not detect an effect” and “we detected that there is no effect” are different claims. The first needs only a test. The second needs an interval narrow enough to exclude every effect size you would have acted on, which is harder to produce and rarer to see published.

The discipline that fixes this costs nothing and happens before the test: write down the smallest change you would act on, then check whether your panel can resolve it. A panel of 30 observations carries a 95% Wilson band of roughly ±17 points at a 50% presence rate, so a genuine 5-point gain reads as a null every time you run it. That result was determined by the size of the test, and knowing so in advance is a reason to skip it.

The mirror-image error is worth naming too. A null on an effect of 0.36 rank places is compatible with the effect being real and with it being worth nothing to you, so a null and a tiny true effect lead to the same decision anyway.

Direction

How is a negative result different from a null?

A null gives you no direction, so your prior is unchanged. A consistent sign across many strategies and stages gives you a direction and a mechanism.

Body-text-only rewritingRetrieval, hit rate at 20After reranking, hit rate at 10Citation rate
Baseline, no rewrite0.581.000.50
Fluency0.57 (−1%)0.91 (−9%)0.50 (−1%)
Statistics0.57 (−2%)0.90 (−10%)0.48 (−4%)
Technical terms0.50 (−14%)0.80 (−20%)0.47 (−6%)
All eight at once0.50 (−14%)0.83 (−17%)0.49 (−2%)
Average of all ten0.53 (−9%)0.84 (−16%)0.47 (−6%)

SAGEO Arena, accepted at KDD 2026, rebuilt the whole pipeline rather than handing a model a fixed context: 171,003 documents and 2,700 queries, with retrieval, reranking and generation all in play. Optimising body text alone, it reports, “consistently degrades visibility across all stages”, and the table above is why that sentence is allowed: all ten strategies fall in all three columns.2 Thirty cells, one sign. The stages are nested rather than independent, so those are not thirty separate trials, but ten different rewrites moving the same way at every stage they touch is a direction, not a coin flip.

The mechanism is stated rather than guessed at, which is what lifts this above a null. The largest retrieval drops belong to the strategies that swap common words for rarer ones, which the authors attribute to “the lexical mismatch between optimized documents and user queries, which typically use common vocabulary”.2 C-SEO Bench found the same asymmetry from the other side: its left-tailed tests show its statistics-adding method reducing rankings in 19 of 24 evaluated settings.1 Where these benchmarks tested for harm they frequently found it, which is a different epistemic situation from failing to find help.

Two honest deductions before anyone quotes that table. The reranking baseline is 1.00 by construction, because “all target documents are selected from the top-10 candidates at the reranking stage”, so that column can only fall.2 And body text only is one of three settings reported; the numbers change when structure is optimised alongside the prose.

Your own test

How do you read a null on your own site?

Six questions, in this order. Most in-house nulls stop at one of the first four.

01

Was the effect ever detectable?

Name the smallest change worth acting on, then check the band your panel gives at that size. If the band is wider than the change, the null was written before the test ran.

02

Did the edit actually ship?

Fetch the changed pages the way a crawler would and confirm the new text is in the response. A rewrite that never became retrievable produces a null about your deploy, and retrieval is where that hides.

03

Was the baseline standing still?

Page overlap across two months runs at 18% for AI Overviews against 45% for organic Google.4 A before-and-after across that gap mostly measures churn, so hold back untouched pages as controls.

04

Did you change one thing?

A bundle that tests null says nothing about its parts and can hide a win cancelling a loss. The all-at-once condition above lands at −14% on retrieval, worse than most of its ingredients.

05

Did you keep the failures?

The survey is blunt about the commonest silent bias: “Outputs without search, without citations, or with errors are outcomes, not data to be discarded.”3 Dropping them inflates every rate.

06

Did you treat runs as independent?

“Queries, sources, and dates are clustered units. Testing all generations as though they were independent underestimates uncertainty.”3 Five runs of one prompt are not five observations.

A team that rewrites forty documentation pages in one release, then compares the month before with the month after on a panel of thirty observations, will report a null nearly every time and will have learned nothing about the rewrite. Three of the six questions above are already answered wrongly: the panel resolves nothing under about seventeen points, eight changes shipped as one, and the comparison period is long enough for the surface to turn over on its own. The same team rewriting twenty pages while holding twenty back, on a panel it sized first, gets an answer worth having even when the answer is that nothing moved.

The decision

What do you do with a tactic that tests null?

Decide on cost and risk rather than on significance, because significance was never the thing you were buying. A tactic that is cheap, harmless and defensible for another reason survives a null intact: question-shaped headings, plain definitions and dated attributed numbers make a page better to read, so a citation effect is a bonus rather than the case for doing them. A tactic that is cheap but carries a documented downside, such as trading the words your readers use for rarer ones, fails on the evidence in the other tail rather than on the null. An expensive tactic that tests null should stop, because its real cost is the constraint you tested nothing on, and Google still describes its generative features as “rooted in our core Search ranking and quality systems”.6

One more property makes the marginal tactic worth less than it measures. The same benchmark reports that as more documents in a candidate set adopt the same method, overall gains decrease, “depicting a congested and zero-sum nature of the problem”.1 A positive result is a depreciating asset; a page that genuinely answers a question is not. A null read properly is not a wasted quarter. It is a map of where your constraint is not, and it usually points upstream, which the evidence review of GEO tactics and relevance and ranking take from there.

The honest limit of this article

Every number above comes from a rig the researchers built. Neither benchmark queried a live commercial engine end to end, both supply their own retrieval, and the verdicts belong to the models they ran on, so “this tactic tests null” is precise only as “it tested null there, then”. C-SEO Bench also tests ranking improvement rather than citation, which are related but not the same outcome. The survey supplying the methods language is a preprint, and the overlap and repeated-run figures reach this page through it. What this article does not claim is that any of these tactics is worthless on your domain. It claims the published tests could not have seen a small effect.

Where a product fits, and where it does not

Nothing on this page needs software. Naming the smallest effect worth acting on, sizing the panel first, holding pages back as controls and changing one thing at a time are free, and they are the whole difference between a null you can read and a null you cannot. What software supplies is the repeated measurement underneath: Bavior runs a fixed prompt set across five engines on a schedule and records which sources each answer cited, so a comparison has a distribution behind it instead of a screenshot. It does not design your experiment, does not compute significance, does not rewrite your pages, and cannot tell you whether a tactic works on your domain. A null in its numbers is a null in the measurement, not a null in the world. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 30 Aug 2026).

Sources, all checked 30 Aug 2026
  1. Puerto, Gubri, Green, Oh, Yun, “C-SEO Bench: Does Conversational SEO Work?”, NeurIPS 2025 Datasets & Benchmarks Track; ten methods, more than 1.9k queries; Table 3 and §6.2, right-tailed Wilcoxon signed-rank tests with Holm-Bonferroni correction: arxiv.org/abs/2506.11097
  2. Kim, Jeong, Kim, Lee, Lee, “SAGEO Arena”, accepted at KDD 2026, arXiv:2602.12187v2; 171,003 documents, 2,700 queries, Table 2 body-text-only columns: arxiv.org/abs/2602.12187
  3. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (preprint); §11.2 and §11.3: arxiv.org/abs/2607.14035
  4. Kirsten et al., Findings of ACL 2026; 4,706 queries; two-month page overlap of 18% for AI Overviews against 45% for organic Google, and 9–28% of decisions changing on repeated runs at temperature zero, as reported in the survey at note 3, §6.2: aclanthology.org/2026.findings-acl.526
  5. Aggarwal et al., “GEO: Generative Engine Optimization”, KDD ’24, arXiv:2311.09735; the source of eight of the ten strategies both later benchmarks re-tested: arxiv.org/abs/2311.09735
  6. Google Search Central, “Google’s Guide to Optimizing for Generative AI Features on Google Search”, updated 10 Jul 2026 (first-party; the “rooted in” sentence): developers.google.com/search/docs/fundamentals/ai-optimization-guide
FAQ

Frequently asked questions.

Does a null result prove a content tactic does not work?

No. A null says the test could not detect an effect, which depends on the effect size, the variance and the number of observations. To claim a tactic does nothing you need an interval that excludes every change you would have acted on, and that is rarer than it sounds. Absent one, the honest report is that the test lacked the resolution to answer, not that the answer was no.

How many observations does a content test need before a null means anything?

Enough that the band around your result is narrower than the smallest change you would act on. At 30 observations the 95% Wilson band is roughly plus or minus 17 points, so anything under that is invisible. The sample-size arithmetic for a given effect size lives in the statistical power article, and the panel design that produces those observations lives in the prompts and runs article.

Why is a negative result more useful than a null?

Because it has a direction and usually a mechanism. SAGEO Arena found all ten body-text-only rewriting strategies falling at retrieval, reranking and citation, and attributed the retrieval losses to lexical mismatch: swapping common words for rarer ones lowers the overlap a retriever scores on. A null leaves your prior where it was; a consistent sign across many strategies changes what you do next.

Should I stop doing something that tested null?

Decide on cost and risk instead. If it is cheap and defensible for another reason, such as headings and definitions that help a human reader, keep it and stop claiming a citation effect for it. If it carries a documented downside, drop it. If it was expensive, stop, because its real cost is the upstream constraint you tested nothing on.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

A null you can read beats
a number you cannot.

Start free trial