The reason is not project management convention. It is what the measurement does when you leave it alone. The 2026 critical survey gathers the replication evidence in one place. Kirsten et al., whose 4,706-query audit is peer reviewed, found that repeating a run at temperature zero changes 9 to 28% of decisions.2 Schulte et al., watching four engines for 45 days on a small Swiss query universe, recorded daily source-level Jaccard between 0.34 and 0.42.1
Read the two together. On any given day a large share of the sources an engine cites for your prompt set differ from yesterday’s, with nothing on your side having changed. That is what an intervention has to beat, and it is why a baseline is a distribution rather than a reading. One run of one prompt says little about the engine and a lot about the minute you ran it.
How many prompts and runs buy a usable band is worked out in how many prompts and runs, and what a panel size can and cannot detect in statistical power. Neither is repeated here; this article takes only the conclusion, that the panel is the reporting unit and the individual prompt never is.
The consequence for the calendar is worth saying out loud to whoever set the deadline. Inside 90 days you can establish a baseline you trust, ship the fixes, and re-read the panel once. You cannot also prove that a particular content change caused a particular movement, because one before-and-after against a drifting instrument does not carry that weight. Week one is a far better time to say so than week twelve.