A prompt has stopped discriminating when it returns the same outcome across two consecutive full periods on every engine, and when the run count behind that streak is high enough for the streak to mean something. The second half is the part teams skip, and skipping it retires prompts that were never measured. With zero appearances in n runs, the 95% upper bound on the true rate is roughly 3/n. At 5 runs that ceiling is 45%, at 10 runs 30%, at 30 runs 10% and at 60 runs 5%. A prompt that came back empty ten times is entirely compatible with a real presence rate near one in three.
So set the bar at the bound rather than at the streak: retire on a persistent zero only once the accumulated runs put the ceiling somewhere you genuinely do not care about, which in practice means about 60 observations per engine. The same arithmetic runs upside down, since 60 consecutive appearances put the floor near 95%. Keep two or three saturated prompts in the panel as a tripwire, because if a prompt you always win suddenly drops, something large has happened.
One quieter form of the same failure: a prompt whose presence rate is stable and mid-range while the answer text never changes. That is measuring a fixed retrieval rather than a live competition, and it usually reports a fact about your collection setup rather than the market.