Home/Learn GEO/First 90 days
GEO program · Operations

What should the first 90 days of a GEO program look like?

A sequence with gates you are willing to fail, ordered by what the measurement can actually support. Month one buys an instrument, not a result.

On this page
Share this
Share on X Share on LinkedIn
The short answer

Spend the first month building the instrument and almost nothing else, because the instrument moves more than the change you hope to detect. Then run the phases in an order where each produces the input the next one needs, with a gate between them you are willing to fail. Repeated runs at temperature zero change 9 to 28% of decisions, and daily source-level Jaccard across four engines sits between 0.34 and 0.42, so a quarter is barely long enough to characterise that noise, let alone beat it.1

Key takeaways
  • A plan promising a measured lift by day 90 promises what the instrument cannot deliver. A quarter delivers a baseline, a clean technical tier, a dated change log and one honest re-read.
  • Weeks 1 to 4 are instrumentation: prompts, runs, logged run metadata, a frozen panel version, a holdout nobody touches.
  • Four gates, each with a stated failure action: the technical pass, the frozen panel, the untouched holdout, the matched re-read.
  • Content waits until week 5, because published content-side effects are small and several are negative, and effects that small have to sit behind an instrument.
The measurement

Why does month one buy an instrument instead of a result?

The reason is not project management convention. It is what the measurement does when you leave it alone. The 2026 critical survey gathers the replication evidence in one place. Kirsten et al., whose 4,706-query audit is peer reviewed, found that repeating a run at temperature zero changes 9 to 28% of decisions.2 Schulte et al., watching four engines for 45 days on a small Swiss query universe, recorded daily source-level Jaccard between 0.34 and 0.42.1

Read the two together. On any given day a large share of the sources an engine cites for your prompt set differ from yesterday’s, with nothing on your side having changed. That is what an intervention has to beat, and it is why a baseline is a distribution rather than a reading. One run of one prompt says little about the engine and a lot about the minute you ran it.

How many prompts and runs buy a usable band is worked out in how many prompts and runs, and what a panel size can and cannot detect in statistical power. Neither is repeated here; this article takes only the conclusion, that the panel is the reporting unit and the individual prompt never is.

The consequence for the calendar is worth saying out loud to whoever set the deadline. Inside 90 days you can establish a baseline you trust, ship the fixes, and re-read the panel once. You cannot also prove that a particular content change caused a particular movement, because one before-and-after against a drifting instrument does not carry that weight. Week one is a far better time to say so than week twelve.

The gates

Which gates does each phase have to pass?

A gate is a condition you are willing to fail. If nothing happens when it fails, it is a milestone, and milestones protect nothing.

GateWhenPasses whenIf it fails
RetrievabilityWeek 4Pages in scope return their answer text to a plain fetch, are indexed and snippet-eligible, and no crawler token you rely on is disallowedStop. A page the engine cannot fetch cannot be measured for anything else
Frozen panelWeek 4Prompt list, runs per engine and logged run metadata are versioned and will not change before week 12Re-baseline. A panel that grew mid-quarter moved its denominator with its numerator
HoldoutWeek 410 to 15 prompts are marked never to be acted on, and everyone doing the work knows which onesProceed, but report movement as coincident with your work, not caused by it
The re-readWeek 12The week-12 run matches the week-4 run on product, mode, model, locale, account and whether search was enabledReport the difference as unmeasured. Unmeasured is not zero

Gate one is the technical pass. Its contents are not this article’s business: the technical GEO checklist lists the items in the order that costs least to run. What belongs here is why it sits first. Google’s documentation states that to be eligible as a supporting link a page “must be indexed and eligible to be shown in Google Search with a snippet”,7 and that its generative features are “rooted in our core Search ranking and quality systems”.6 A page failing that does not produce a weak measurement. It produces none, and every later comparison including it reads an absence as a result.

Gate four looks pedantic until it bites. The survey’s reproducibility guidance lists what a run must record to be comparable to a second one: “Product, mode, model, date, locale, account, and search enabled”.1 Change one of those between week 4 and week 12 and the difference you report silently includes it. Writing them down in week 1 costs nothing; reconstructing them in week 12 is impossible.

The calendar

What does each block of weeks produce?

Ordered so each phase produces the input the next one needs. The output is the real specification; the activity is only how you get there.

01

Weeks 1–2 · Instrument

Write the prompts from real calls, tickets and lost-deal notes, not a keyword tool, then run them on a schedule. Output: a baseline with an interval per engine, and the URLs those engines cited.

02

Week 2 · Audit

Run the technical checklist alongside the baseline, not after it. It is read-only, so it cannot disturb the measurement underneath. Output: a fix list, most of it one sprint.

03

Weeks 3–4 · Fix and freeze

Ship the technical fixes, then version the panel and name the holdout. Output: a retrievable site, and an instrument that no longer moves for any reason you caused.

04

Weeks 5–8 · On-site

Rewrite the highest-intent pages as passages: question headings, the answer in the first two sentences, numbers with denominators. Output: pages eligible at selection, not only at retrieval.

05

Weeks 7–12 · Off-site

Correct wrong third-party facts, make the entity description consistent, answer the threads the engines already cite. Output: movement in which URLs get cited, arriving later than anyone wants.

06

Week 12 · Re-read

Re-run the frozen panel, set the acted-on prompts beside the holdout, and write down the result whichever way it went. Output: a defensible before-and-after, or an honest null.

Content sits at week 5 for a second reason beyond the missing baseline: published effect sizes for content-side intervention are small, and several are negative. C-SEO Bench, a NeurIPS 2025 benchmark of ten methods, reports that most of them “are not only largely ineffective but also frequently have a negative impact on document ranking”, and that out of 54 cases only three showed a statistically significant ranking improvement.4 SAGEO Arena, at KDD 2026, reinstated retrieval and reranking instead of handing the model a fixed context, and found body-text-only optimisation moving three headline averages of 0.58, 1.00 and 0.50 down to 0.53, 0.84 and 0.47: 9%, 16% and 6% the wrong way.5 Effects that small have to sit behind an instrument.

The promise

What can you honestly promise at day 90?

Four deliverables, none of which is a number that went up. A versioned prompt panel with a baseline and an interval per engine. A technical tier that passes a fetch-level audit. A dated log of every change, so the next quarter can attribute rather than guess. And one re-read of the panel with the holdout beside it. That list survives a sceptical room, and it is what a quarter honestly buys.

What you cannot promise is what you will be asked for: a ranking, inclusion in any particular answer, or revenue. The survey’s confidence table rates the claim that citation scores predict clicks, conversions or revenue at very low.1 The traffic end is weak too: Pew tracked 900 US adults across 68,879 Google searches in March 2025 and found a click on a source inside an AI summary happened “in just 1% of all visits”.8 A visibility programme sold as a traffic programme gets judged on a metric it was never measuring.

One further limit is arithmetic rather than judgement. If your panel resolves changes of roughly ten points and the change you shipped is worth three, the week-12 re-read shows nothing, which is a property of the instrument, not evidence about the work. Decide in week 1 what size of movement would change a decision, check it against statistical power, and if the panel cannot see something that small, buy a bigger panel or stop promising to detect it.

Failure modes

Which sequencing mistakes cost you the quarter?

A team that ran the quarter backwards makes the failure concrete. They opened with a six-week content sprint, because content was the part everyone knew how to do, started measuring in week 7, and by week 12 had a rising line they could not defend. The pages had changed before anyone recorded what the engines saw, so the “before” was a memory. The panel had grown from 22 prompts to 60, so the denominator moved with the numerator. There was no holdout, so a platform-side change that lifted everything looked identical to the sprint working. None of the work was bad. The order voided it.

Sequencing that survives review

  • Baseline before any page in scope changes
  • One panel version per quarter, named in the report
  • A holdout of 10 to 15 prompts nobody touches
  • One variable per window: technical, on-site, off-site
  • Per-engine numbers, each with its own interval

Sequencing that voids the result

  • A content sprint in weeks 1 to 4
  • Adding prompts to the panel as they occur to you
  • Shipping technical fixes and content edits the same week
  • One blended score averaged across engines
  • Reading a single run as a fact

The blended-score entry earns its place from measurement rather than taste. A SIGIR 2026 study of 11,500 representative queries measured URL-level Jaccard similarity of 0.11 to 0.18 between Google organic results, AI Overviews and Gemini.3 Jaccard is intersection over union, not the share of cited URLs, and at that level the three surfaces are close to disjoint. An average across them describes no engine anybody can visit.

The handover

What happens on day 91?

The plan ends and a cadence begins, and the two are different shapes. The 90 days exist to build something that then runs indefinitely on a fraction of the attention: a scheduled panel, a monthly report with unchanging columns, an off-site queue that never empties. The recurring rhythm belongs to the off-site weekly workflow, the shape of the report and its holdout comparison to a defensible report, and the standing bill to what GEO costs to run. This article stops at the handover.

One thing crosses that boundary. The gates do not retire at day 90; they become the standing conditions on every later comparison. A quarter that re-baselines, grows the panel, or loses its holdout has started again at week one of this plan, whatever the calendar says.

The honest limit of this article

The sequencing argument rests on evidence about instability, not on evidence that this sequence beats another one. Nobody has run that experiment, and it is hard to see how: the control condition needs a second copy of the same site. Both instability figures come from narrow samples, a small Swiss query universe over 45 days and a 4,706-query audit, and both reach you through a survey rather than independent replication. What they establish is that the measurement moves substantially on its own, which rules out promising a measured lift inside a quarter. It does not establish that these week boundaries are optimal. Treat them as a defensible default and move them when your own baseline says otherwise.

Where a product fits, and where it does not

Most of this plan is calendar discipline, and nobody sells that. The prompts come from your own calls and tickets, the technical fixes are your engineers’, and the freeze and the holdout are decisions somebody must make and then honour. No tool enforces any of it. What a tool can do is the repetitive middle: Bavior runs a fixed prompt panel across five engines on a schedule, records which sources each answer cited, and keeps the run metadata that makes week 12 comparable to week 4; where a cited source is a live discussion thread it drafts a reply on an account you control, which you approve before anything posts. It does not write your panel, does not decide when to freeze it, cannot make an unretrievable page retrievable, and cannot shorten the 90 days. The free AI visibility check and free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 30 Aug 2026).

Sources, all checked 30 Aug 2026
  1. “A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035; the 9–28% repeated-decision figure (Kirsten et al.), daily source-level Jaccard 0.34–0.42 (Schulte et al., four engines, 45 days, a small Swiss query universe), the reproducibility fields and Table 5 (preprint): arxiv.org/abs/2607.14035
  2. Kirsten et al., “Characterizing Web Search in The Age of Generative AI”, Findings of ACL 2026; 4,706 queries, US and Germany, September 2025: aclanthology.org/2026.findings-acl.526
  3. Grossman et al., SIGIR 2026; 11,500 representative queries; URL-level Jaccard 0.11–0.18: arxiv.org/abs/2604.27790
  4. Puerto, Gubri, Green, Oh, Yun, “C-SEO Bench: Does Conversational SEO Work?”, NeurIPS 2025 Datasets & Benchmarks Track, arXiv:2506.11097: arxiv.org/abs/2506.11097
  5. “SAGEO Arena”, KDD 2026, arXiv:2602.12187v2; Table 2, body-text-only averages against baseline: arxiv.org/abs/2602.12187
  6. Google Search Central, “Google’s Guide to Optimizing for Generative AI Features on Google Search”, last updated 10 Jul 2026 (the “rooted in” sentence; first-party): developers.google.com/search/docs/fundamentals/ai-optimization-guide
  7. Google Search Central, “AI features and your website”, last updated 10 Dec 2025 (indexed and snippet-eligible; first-party): developers.google.com/search/docs/appearance/ai-features
  8. Pew Research Center, “Google users are less likely to click on links when an AI summary appears in the results”, 22 Jul 2025; 900 US adults, 68,879 searches, March 2025: pewresearch.org
FAQ

Frequently asked questions.

Can a GEO program prove a measured lift within 90 days?

Rarely, and a plan that promises one is overselling the instrument. Repeated runs at temperature zero change 9 to 28% of decisions, and daily source-level Jaccard across four engines runs 0.34 to 0.42, so a quarter is roughly the time it takes to characterise that drift and read the panel once. Promise a trustworthy baseline, a clean technical tier and an honest re-read instead.

Why is there no content work in the first month?

Because a page that changes before the baseline is taken destroys every later comparison: the "before" then exists only as a memory. There is a second reason as well. Published content-side effects are small and often negative, with C-SEO Bench finding significant ranking improvement in only three of 54 cases, and effects that small are unreadable without an instrument already running.

Why must the prompt panel be frozen until week 12?

Because a panel that grows during the quarter has a denominator that moved with the numerator, and the week-12 series is then not comparable to the week-4 one. Freeze the prompt list, the runs per engine and the recorded run conditions at the end of week 4, version them, and put new questions into the next quarter's panel rather than this one.

What do you report at day 90 when nothing moved?

Report the null, with the panel version, the observation count and the interval beside it, and say what size of movement the panel could have detected. A flat line from an instrument that resolves ten points is not evidence that a three-point change failed; it is evidence that the panel could not see it. That distinction is the difference between a report and a story.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

Month one buys the instrument.
Start it before you start fixing.

Start free trial