Four things travel with every number, or the number cannot be compared with anything: the engine it came from, the observation count behind it, the date range it covers, and the version of the prompt panel that produced it. A figure missing one of them is not a measurement: a reader cannot tell whether this month’s value and last month’s came from the same instrument.
The panel version is the one teams skip, and the one that quietly invalidates a year of history. Adding six prompts changes the denominator of every rate, so a rise from 29% to 34% across two versions may be an artefact of which questions were asked. Use a two-part version, printed beside every figure. A minor bump is a wording fix that does not change what a prompt asks; comparisons across it need only a footnote. A major bump is a prompt added, removed or re-scoped; there you re-baseline rather than compare. Keep the prompt list in version control so anyone can diff two versions.
One scheduling rule follows: never change the panel in the same period as the content or off-site work you want to measure. Do that and the two causes are confounded beyond recovery. Change the panel in a quiet period, take a fresh baseline, then act.
The interval belongs on the figure too. The 2026 critical survey’s minimum checklist for a GEO study asks for product, mode, model, date and locale, whether search was enabled, repeated runs in multiple time windows, retrieval and generation kept apart, and null outcomes retained in the denominator.1 A commercial report meeting that research checklist is simply one an outsider could check.