Home/Learn GEO/What to promise
GEO program · Expectations

What can you honestly promise a CEO about AI visibility?

The gap between what this work can measure and what an executive wants committed to is where GEO programs get cancelled. This is the wording that closes it without overclaiming, and the five requests to refuse by name.

On this page
Share this
Share on X Share on LinkedIn
The short answer

Promise a measured share of a defined panel of questions, promise the cadence at which you will re-measure it, and refuse everything downstream of that. Two published findings decide where that line falls: the field’s own 2026 critical survey rates the claim that citation scores predict clicks, conversions or revenue at very low confidence, its lowest grade,1 and Pew Research Center’s panel of 900 US adults recorded a click on a source inside a Google AI summary “in just 1% of all visits” to pages carrying one.2

Key takeaways
  • The promise that survives month four is a baseline, a cadence, and movement in mention share and citation share on a published panel. Never a rank, a date, or a number of deals.
  • Every promised figure carries an interval. At five runs the 95% interval is roughly 33 points wide either side, so a headline rate from five runs commits you to nothing defensible.
  • Refuse five requests by name: a ranking, inclusion in any specific answer, a percentage lift, a timeline to results, and revenue attributed to a citation.
  • Say no by substitution: name the number you cannot give, the reason in one clause, then the nearest number you can defend and the runs behind it.
The wording

What is the promise that survives a quarterly review?

Written out it is four sentences, short enough to say in a meeting: “Within two weeks you will have a baseline, which is how often each engine names us in answers to a fixed panel of the questions our buyers actually ask, reported per engine with a margin of error attached. Every month you will see that same panel re-run, with the panel and the runs published. We will work to move mention share and citation share on it, and you will see both readings whichever way they go. We will not promise a rank, a place in any particular answer, a percentage lift, a date, or a revenue number.”

Each clause is defensible for a different reason. The baseline is defensible because it is an observation you make yourself, not a claim about the world. The fixed panel is defensible because it carries its own denominator: it is a share of the questions you chose, which is a real fraction, and no engine publishes prompt volume that would let anybody honestly say “share of the category”. The cadence is defensible because it is work you control rather than an outcome you do not. And the closing refusal is defensible because the 2026 critical survey, reviewing 45 studies, concludes that “no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior”.1

Read the shape of it. Three sentences promise an instrument and a rhythm; one promises a direction of travel with its uncertainty stated. That is smaller than most proposals in this category, and it is the only version still standing when the first flat quarter arrives.

Refusals

What can you never promise, whatever the pressure?

Five of these come up in almost every kickoff, and the sixth arrives the first time an engine says something wrong about the product. Refuse them by name before somebody promises them on your behalf.

01

A ranking

There is no ranked list inside a generated answer to win a place on. Google calls its own AI surfaces “rooted in our core Search ranking and quality systems”,3 and a 4,706-query audit found 53% of the domains an AI Overview consults sit outside the organic top ten.4

02

A place in a specific answer

Repeated runs at temperature zero change 9% to 28% of an engine’s decisions, and page overlap across two months was 18% for AI Overviews against 45% for Google organic.1 One answer is a draw, not a placement.

03

A percentage lift

A NeurIPS 2025 benchmark of ten conversational SEO methods reports that “out of 54 cases, we uncover only three where the ranking improvements are statistically significant”.5 A KDD 2026 arena found ten body-text strategies averaging at or below baseline.6

04

A date for results

Nothing published establishes a time to effect for a white-hat change, and the surface is intermittent: AI Overviews activated on 13.7% of 55,393 trending queries, against 64.7% of the question-shaped ones.7

05

Revenue, or traffic

The 2026 survey grades “citation scores predict clicks, conversions, or revenue” at very low confidence, causality not established.1 In the ordinary case the click never happens, so analytics has no event to record.2

06

That the answer will be right

You can monitor how a product is described and fix the sources behind it. You cannot promise the description. The Tow Center found more than 60% of 1,600 queries answered incorrectly across eight engines, the worst at 94%.8

One number deserves naming directly, because somebody in the room will have read it. The survey’s own confidence table lists “GEO increases visibility by 40%” as rejected as a general claim: the figure is a relative maximum on one metric under one configuration.1 Repeat it to win the budget and you have handed your CEO a number a rival can take apart in one search.

Uncertainty

Why does every promised number need an interval attached?

Because a rate quoted without one is a promise about noise. The arithmetic used across this curriculum is the 95% Wilson half-width at p equals 0.5: 1.96 divided by twice the square root of (n plus 3.84). Five runs gives about ±33 points, thirty about ±17, two hundred about ±7. So “we are named in three of five answers” is 60% with a plausible range from roughly a quarter to nine tenths, and a commitment to lift that to 80% is a commitment to move something nobody can currently see. The sample size arithmetic is worked through separately.

The interval changes the sentence you say out loud, not the work you do. “We are named in 60% of answers on our panel, plus or minus 33 points on five runs each; if you want that margin under ten points I need about a hundred runs per prompt, and here is what it costs.” That turns an argument about how optimistic you are into a decision about how much detection power to buy, which is a decision an executive is equipped to make.

It protects you in the other direction too. An interval attached before a reading falls separates a fluctuation from a failure narrative; added afterwards it reads as what it is.

Saying no

How do you say no to a metric request without sounding evasive?

A refusal only sounds evasive without a substitute. Name the number you cannot give, the reason in one clause, then the nearest number you can defend.

What gets askedWhy it cannot be promisedWhat to say instead, out loud
Can you get us to number one in ChatGPT?There is no ranked list inside a generated answer“There is no number one to win in there. What I can commit to is a measured share of the answers our buyers’ questions produce, per engine, with every run shown to you.”
What lift will we see, and by when?Three of 54 tested cases were significant; no time to effect is published“I will not give you a percentage or a date. I will give you a baseline in two weeks with its margin, and the same panel re-run every month.”
How much pipeline does this produce?Rated very low confidence; the click usually never happens“I cannot attribute pipeline to this, and anyone offering to is modelling rather than measuring. I can tell you how often we are named where the decision is being made.”
Will we show up when a customer asks on Monday?Run-to-run instability, even at temperature zero“Not for any one answer. Across a hundred runs I can give you a rate with an interval, which is the honest version of the same question.”

Two habits decide whether this reads as rigour or as excuse-making. Volunteer the refusals in the kickoff, where they cost nothing, rather than in the review, where they look like a retreat. And attach an option and a price to every no: “not at this sample size, and here is what a bigger one costs” is a proposal, while “that cannot be known” is a shrug.

Then keep one wording everywhere: the same four sentences in the kickoff deck, the board update, the monthly report and any vendor contract you sign. A promise that changes register between the pitch and the report is the one that gets read as spin.

In writing

Which sentences are safe to put in a written commitment?

Five things, plus an exclusions clause. Put the prompt panel in an appendix so the denominator is inspectable, name the engines and the mode each is queried in, state the runs and the interval behind every figure, and write the measurement conditions into the document rather than a footnote: the survey’s minimum checklist for a study asks for “Product, mode, model, date, locale, account, and search enabled”.1 A commitment measured under one set of conditions and reported under another is not the same commitment. Which metrics survive scrutiny is a separate question.

The exclusions clause is the part people skip and the part that saves the program. Write the refusals down as refusals: “This program does not commit to a ranking, to appearing in any specific answer, to a percentage improvement, to a date by which a change becomes visible, or to revenue attributable to a citation.” That costs nothing at kickoff and becomes the document everyone returns to when the mood turns.

Add a re-baseline clause. Engines change underneath the measurement, and the survey rates “commercial engines differ from one another and vary over time” at high confidence.1 State in advance that when a model version, a mode or a surface changes, the baseline resets and the comparison restarts.

The cost of overpromising

What does an overpromised program look like in month four?

It looks like a cancellation for the wrong reason. A team that inherited one such program in month five found the original proposal had committed to a 40% visibility improvement and a share of sourced pipeline, on a panel of eight prompts at one run each. Re-running it, mention rate had in fact risen, from around 25% to around 38%, but at eight runs the interval on each reading is roughly ±30 points, so the honest reading was that nothing measurable had happened either way. The program was cancelled on a promise it should never have made.

The contrast case is duller and it survives. A team that committed only to a panel, a cadence and an interval reported a flat second month, said so in one line, and kept going, because a flat month was in the plan the executive had approved. Nothing about its tactics was better. The number that arrived was inside the range it had already described.

That is the commercial argument for the smaller promise. Overpromising spends patience early, and the bill arrives just as the measurement becomes sensitive enough to show something.

The honest limit of this article

The case for the smaller promise rests on absence of evidence, and absence of evidence is not evidence of absence. That citation scores have not been shown to predict revenue does not mean citations produce none; it means the study that would establish it has not been run. The 1% click figure describes Google AI summaries seen by a US panel in March 2025, not ChatGPT or Perplexity and not this quarter, and the survey beside it is a preprint. Nothing here says AI visibility is not worth funding. It says these specific promises cannot be defended with what is currently published.

Where a product fits, and where it does not

None of this needs software: a fixed panel, a spreadsheet, a stated interval and somebody willing to say no in a meeting will do the whole job. Bavior runs a fixed prompt panel across five engines on a schedule, reports per engine with the run log attached, and records every cited URL so a share-of-voice figure has a denominator you can inspect. What Bavior cannot promise is what you cannot promise: it does not sell a ranking, cannot put a brand into any specific answer, does not forecast traffic or revenue from citations, and will not give you a date. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (checked 30 Aug 2026).

Sources, all checked 30 Aug 2026
  1. Martinez, “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (preprint; Tables 5 and 6, plus the run-to-run and two-month overlap figures from Kirsten et al.): arxiv.org/abs/2607.14035
  2. Pew Research Center, “Google users are less likely to click on links when an AI summary appears in the results”, 22 Jul 2025; 900 US adults, 68,879 Google searches, March 2025 data: pewresearch.org
  3. Google Search Central, “Google’s Guide to Optimizing for Generative AI Features on Google Search”, updated 10 Jul 2026 (first-party): developers.google.com/search/docs/fundamentals/ai-optimization-guide
  4. Kirsten et al., Findings of ACL 2026; 4,706-query audit of Google AI Overviews, data collected September 2025 in the US and Germany: aclanthology.org/2026.findings-acl.526
  5. Puerto, Gubri, Green, Oh, Yun, “C-SEO Bench: Does Conversational SEO Work?”, NeurIPS 2025 Datasets & Benchmarks Track, arXiv:2506.11097: arxiv.org/abs/2506.11097
  6. Kim et al., “SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization”, KDD 2026 (arXiv:2602.12187v2, 7 Aug 2026): arxiv.org/abs/2602.12187
  7. Xu, Iqbal & Montgomery, 2026; 55,393 trending queries collected 13 March to 21 April 2026; AI Overview activation 13.7% overall and 64.7% on question-form queries (preprint): arxiv.org/abs/2605.14021
  8. Jaźwińska & Chandrasekar, “AI Search Has a Citation Problem”, Tow Center for Digital Journalism, Columbia, 6 Mar 2025; 1,600 queries across eight engines: cjr.org
FAQ

Frequently asked questions.

Can I promise my CEO a percentage improvement in AI visibility?

No, and the evidence against it is the field's own. A NeurIPS 2025 benchmark of ten conversational SEO methods found only three of 54 tested cases where the ranking improvement was statistically significant, and a KDD 2026 arena found body-text strategies averaging at or below baseline. Promise a baseline with its margin and a monthly re-run of the same panel instead, and let the direction of travel be reported rather than forecast.

What do I say when the CEO asks how much pipeline this produces?

Say you cannot attribute pipeline to it, that anyone offering to is modelling rather than measuring, and then give the number you do have. The 2026 critical survey rates the claim that citation scores predict clicks, conversions or revenue at very low confidence. Pew recorded a click on a source inside a Google AI summary in just 1% of visits to pages carrying one, so the event analytics needs mostly does not exist.

Is share of voice a safe thing to commit to?

It is safe only with three things attached: the panel it was measured on, the competitor set it was measured against, and its interval. Share of voice on a fixed panel is a real fraction, because you chose the denominator and can publish it. Share of voice in a category is not, because no engine publishes prompt volume, so nobody can build that denominator honestly.

Does refusing to promise results make the program harder to fund?

It makes the first conversation harder and the fourth one survivable. Programs are rarely cancelled because a number was flat; they are cancelled because a number nobody committed to went unmet. A written exclusions clause naming the five refusals costs nothing at kickoff and becomes the document everyone returns to when the mood turns, which is why it belongs in the proposal rather than a later correction.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

The smaller promise is the one still standing.
Start with a baseline that carries its margin.

Start free trial