Home/Learn GEO/Weekly workflow
Off-site GEO · Practice

The off-site GEO week: what you run, and what a week cannot decide.

Six passes over one artefact, in a fixed order, about two hours. The week collects; the decisions run on a slower clock, and confusing the two is expensive.

On this page
Share this
Share on X Share on LinkedIn
The short answer

The weekly run is one ordered pass over the citation list your prompt panel produced: export it, diff it against last week, bucket every cited URL into pages you own, pages that describe you and pages that ignore you, then read the answers for statements about you that are wrong. It ends in three queues and no conclusions, because a week cannot carry one: the 2026 critical survey reports temperature-zero reruns changing 9 to 28% of decisions, and daily source-level overlap of 0.34 to 0.42.1

Key takeaways
  • Weekly movement in a rate is not evidence. Temperature-zero reruns of the same prompts already change 9 to 28% of decisions, so a few points moved on their own.
  • Events are what a week can catch: a false statement about you, a page that just entered the list, a thread open today and closed next month.
  • Three queues come out, two of them people writing to people: corrections, inclusion candidates, thread replies. Cap the panel at what they can absorb.
  • Decisions wait for the quarter and most will be nulls. One benchmark found three statistically significant ranking improvements out of 54 cases.2
Cadence

Why is a week a collection interval and not a decision interval?

Two reported measurements decide the shape of this workflow and both point the same way. The 2026 critical survey of generative engine optimization reports Kirsten and colleagues finding that repeated runs at temperature zero change 9 to 28% of decisions, and Schulte and colleagues measuring daily source-level Jaccard overlap of 0.34 to 0.42 across four engines over 45 days on a small Swiss query universe.1 Jaccard is intersection over union, so 0.34 means only about a third of the sources cited across two consecutive days appeared on both.

Read together, those numbers say the weekly citation list is a sample of a moving population rather than a state of the world, and that most of what changed since last Tuesday changed without you. A workflow that reacts weekly to a number moving is reacting to its own instrument. So the week is built as collection: it accumulates rows, catches the events one observation can honestly establish, and leaves reactions to a horizon where noise averages out. The band at a given sample size is arithmetic, worked out in how many prompts and runs and statistical power. Running weekly still earns its place, because frequency buys reaction time and the events that expire; it just does not buy significance.

The run

What do you actually run on a Tuesday?

Six passes over one artefact: the URLs your panel’s answers cited this week. The order matters because each pass narrows what the next one reads.

01

Export the week’s citations

One row per run per engine, dated, with the cited URLs and the panel version attached. Which fields keep rows comparable is the ten fields to log; the survey asks for product, mode, model, date, locale, account and whether search was enabled.1

02

Diff against last week

Mark every URL held, new or gone. Expect a third to a half of the list to turn over on its own, so the diff is a reading aid and never a result. What earns a note is a URL new on several engines at once.

03

Bucket into three

Pages you own, third-party pages that describe you, third-party pages that never mention you. What those shares should be is how much of the answer you own; the week needs the split so the next three passes know where to look.

04

Read the answers for errors

The prose, not the URL list. The Tow Center’s eight-engine audit found more than 60% of 1,600 queries answered incorrectly.5 A wrong price or plan name is an event, and it does not wait for the quarter.

05

Check the new third-party pages

For each new page in bucket two, check every claim about you: price, plan names, features, founding date, category. Stale beats hostile by a wide margin, and a stale fact on a page engines already retrieve is the cheapest fix here.

06

Check the open threads

For each newly cited thread, ask whether its question is one you can answer usefully. Why a thread gets selected is how community threads get selected. Answering ones you cannot help with is how this channel gets lost.

Two hours is the honest budget once a panel exists, and most of it goes to passes four, five and six, which are reading. If passes one to three take an afternoon, the log schema is wrong rather than the week. Keep the artefact durable: week eight is worth having only because it compares with week one.

Reading the week

What does a normal week look like, and what needs action?

Most weeks are churn with nothing in them. What separates them is not the size of a movement but whether you are looking at an event or a rate.

What you observeA normal weekThe week it needs action
A third of cited URLs turned overExpected: daily source overlap runs 0.34 to 0.421Never on its own
Your presence rate moved a few pointsExpected: temperature alone moves 9 to 28% of decisions1Never on a weekly delta
An answer says something false about youNot normal, and common5Same week, at the source page
A prompt returned no AI answerOften the query: AI Overviews activate on 13.7% of trending queries, 64.7% of question-form ones4Only if it persists across runs
A new third-party page appears on several enginesRareRead it now, act at the month boundary
A cited thread is open and on topicThe week’s actual opportunityInside your cap, where you can genuinely answer

The second row surprises teams. A presence rate that went from 31% to 36% looks like a result and is usually the same measurement twice, which is why the reporting half of this program publishes monthly and reads weekly; that split is argued in a defensible report. What matters here is that the weekly meeting should not have a rate on its agenda. Give it the event list instead: what was wrong, what is new, what is open.

Output

What comes out of the run, and who does it?

Three queues, ordered by how reliably the work pays back.

A

Corrections

A third-party page says something untrue or stale about you, so a person writes to a person with the correct fact and where it is published. It returns fastest because it repairs a source engines already retrieve. Keeping those pages agreeing is entity consistency.

B

Inclusion candidates

A cited page answers your category’s question without mentioning you. One test decides it: would a fair reviewer name you? If yes, that is outreach with a reason. If no, the fix is product. Why these pages dominate is why third-party pages dominate.

C

Thread replies

A cited thread carries a live question you can answer. Answer it, disclose the affiliation, stay useful to a reader who never buys, and cap the volume, since a run of similar replies reads as a campaign. The line is legitimate versus manipulation.

Two of the three queues are people writing to people and do not scale, and that constraint, not curiosity, should set the size of the panel. A team whose panel had outgrown its reading passes found the symptom in its output rather than its data: the export ran every week and no correction was ever filed, because filing one needed a person nobody had freed up.

Scope

What do you deliberately not do weekly?

Weekly, because one observation supports it

  • Export the citation list and diff it
  • Bucket every cited URL into the three buckets
  • Read the answers for false statements about you
  • Check new third-party pages for stale facts
  • Reply, within a cap, to threads you can answer
  • Log every run, including those where you did not appear

Not weekly, because the measurement cannot carry it

  • Reacting to a presence rate that moved
  • Adding or retiring prompts, which swaps the instrument
  • Calling an outreach or a thread reply a success
  • Averaging engines into one blended score
  • Attributing site traffic to an AI answer
  • Rewriting a page because this week’s citations changed

Every exclusion has a mechanism behind it. Reacting to a rate is ruled out by the reproducibility figures above. Changing the panel swaps the instrument mid-measurement, and when a prompt has earned retirement is worked out in when to retire a prompt. Calling a win is ruled out because effects here are small and usually absent: the NeurIPS 2025 conversational-SEO benchmark reports that out of 54 cases it found only three where the ranking improvements were statistically significant.2 Blending engines is ruled out because they do not retrieve the same pages, with an 11,500-query SIGIR 2026 study measuring URL-level Jaccard of 0.11 to 0.18 across Google organic, AI Overviews and Gemini;3 that argument at the metric level is metrics that hold up. Attributing traffic is ruled out because the click barely exists: Pew recorded one on a source inside an AI summary in just 1% of all visits, across 68,879 searches in March 2025.7 And a list that changed this week may reflect an ordinary ranking change, since Google states its generative features on Search are “rooted in our core Search ranking and quality systems”.8

The slow loop

When does the slow loop actually decide something?

At a boundary set in advance, with the collected weeks in front of you and a holdout slice you never acted on. A quarter suits these queues: a correction filed in week two and a reply posted in week five are both still working through indexes in week eight. The decision it supports is narrow, namely where to point the next quarter’s queue. Whether any of it produced revenue is not on offer, and the survey rates that claim at very low confidence.1

What decides it is a pattern that held across weeks, never a delta between two of them. A third-party page cited in most runs on most engines all quarter is worth a relationship whether or not a rate moved. A cluster of answers getting the same fact wrong points at one source page, not at ten. Neither reading needs a significance test, and both need the weeks of rows the panel produced.

Expect nulls, and write them down as nulls. Verification research found only 51.5% of generated sentences fully supported by their citations, and 74.5% of citations supporting the sentence they were attached to,6 so even an engine’s attribution is noisy evidence about what it read. A quarter with no measurable movement, four corrections and two useful thread replies is ordinary, and saying so is how a panel survives somebody who wants it to say something better.

The honest limit of this workflow

No published study evaluates this workflow against an outcome anybody cares about. Both reproducibility figures are read through a survey rather than the original experiments, and the 0.34 to 0.42 daily overlap comes from four engines over 45 days on a small Swiss query universe, so the direction transfers to another market but the exact value does not. The same survey rates the claim that citation scores predict clicks, conversions or revenue at very low confidence.1

Where a product fits, and where it does not

Passes one to three are bookkeeping and pass six is the one that does not scale by hand. Bavior covers those: it runs a fixed prompt set across five engines on a schedule and records which sources each answer cited, so the export and the diff are done; where a cited URL is a live thread it drafts a reply in your voice on an account you control, which you read, edit or reject before anything posts. It does not file corrections, does not write to editors or publishers, does not judge whether a page would fairly include you, and cannot make a weekly number mean more than the reproducibility figures allow. The free AI visibility check produces the URL list without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 30 Aug 2026).

Sources, all checked 30 Aug 2026
  1. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023 to 2026)”, 15 Jul 2026, arXiv:2607.14035; temperature-zero reruns changing 9 to 28% of decisions (Kirsten et al.); daily source-level Jaccard 0.34 to 0.42, four engines, 45 days, a small Swiss query universe (Schulte et al.); Table 5, citation scores predicting revenue, very low confidence; Table 6, the fields to log: arxiv.org/abs/2607.14035
  2. Puerto et al., “C-SEO Bench: Does Conversational SEO Work?”, NeurIPS 2025 Datasets & Benchmarks Track, arXiv:2506.11097; “Out of 54 cases, we uncover only three where the ranking improvements are statistically significant”: arxiv.org/abs/2506.11097
  3. Grossman et al., SIGIR 2026; 11,500 queries; URL-level Jaccard 0.11 to 0.18 across Google organic, AI Overviews and Gemini: arxiv.org/abs/2604.27790
  4. Xu, Iqbal & Montgomery, “Measuring Google AI Overviews”, 2026 (preprint); 55,393 trending queries, 13 March to 21 April 2026; activation 13.7% overall, 64.7% on question-form queries: arxiv.org/abs/2605.14021
  5. Jaźwińska & Chandrasekar, “AI Search Has a Citation Problem”, Tow Center, Columbia, 6 Mar 2025; 1,600 queries, eight engines, more than 60% answered incorrectly: cjr.org
  6. Liu, Zhang, Liang, “Evaluating Verifiability in Generative Search Engines”, 2023, arXiv:2304.09848; 51.5% of sentences fully supported, 74.5% of citations supporting their sentence: arxiv.org/abs/2304.09848
  7. Pew Research Center, 22 Jul 2025; 900 US adults, 68,879 searches, March 2025; a click on a source inside an AI summary occurred “in just 1% of all visits”: pewresearch.org
  8. Google Search Central, “Google’s Guide to Optimizing for Generative AI Features on Google Search”, last updated 10 Jul 2026; the “rooted in our core Search ranking and quality systems” sentence (first-party): developers.google.com/search/docs/fundamentals/ai-optimization-guide
FAQ

Frequently asked questions.

How long should the weekly run actually take?

About two hours once a panel exists, and most of that is reading rather than exporting. If the export and the bucketing are eating an afternoon, the log schema is the problem and not the week. Cap the panel at a size whose queues you can genuinely work, because a citation list nobody reads produces no corrections, and an unread list is indistinguishable from no measurement at all.

Our presence rate dropped four points this week. What should we do?

Log it and do nothing else yet. Repeated runs of the same prompts at temperature zero already change 9 to 28% of decisions, and daily source-level overlap between consecutive days runs around 0.34 to 0.42, so a few points of weekly movement is roughly what the instrument does at rest. Go and read the answers instead, looking for an event: a false statement, a new page, an open thread.

Can we collect monthly instead of weekly?

Report monthly, and you should, but collecting monthly costs you the events. A false statement about your pricing, a thread that is open right now, and a page that has just entered the citation list are all things one observation establishes and all things that expire. Weekly collection, monthly publication and a quarterly decision is the cadence each part of the evidence can actually support.

The workflow runs every week and nothing happens. What do we fix first?

The queues, not the panel. Most stalled programs collect perfectly and act never: the export runs, the buckets fill, and no correction gets filed because filing one means a person writing to a person. Check whether the last four weeks produced a single outbound action. If they did not, cut the panel until the reading passes fit the week you actually have.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

A week is for collecting.
See what your citation list already says.

Start free trial