Treat it as routine work rather than an incident. The closest published measurement is the Tow Center audit behind the start list: it tested news attribution rather than brand descriptions, yet found the best of eight engines wrong 37% of the time.6 There is no support queue to file a correction in, and a presence flag has already logged the mention as a win.
The engine is repeating a source, so start from the cited-source list your panel records and find the page carrying the error. Three cases follow, in ascending difficulty: your own page, where a dated correction in the body is the whole job; a third party's page, an outreach task on somebody else's schedule; and a community thread, where the only defensible move is a reply that says who you are.
Then re-measure on the panel, not by asking the engine again the same afternoon. Repeat runs of one query overlap at Jaccard 0.34–0.42,1 so a single re-check after an edit is indistinguishable from the engine's own variance. No published study measures whether a correction propagates or how long it takes, so hold that timeline as unknown rather than slow.
The honest limit of this article
Every grade here comes from one survey's judgement plus three benchmarks, none run end to end on a live commercial engine, and the survey is a preprint. So “keep” means no evidence against it at a low cost, not proven, and “start” means the reasoning is sound rather than the outcome demonstrated. Nothing here has been tested as a package: no published study takes a real site, applies a list like this one, and measures what happened. The ordering is our judgement about where the evidence is strongest; a team with different constraints could sequence it differently.
Where a product fits, and where it does not
The keep and stop columns need no software: they are a backlog conversation and a deletion, both free. The start column has one item that becomes impractical by hand, repeated sampling: a fixed panel run five times across five engines every week is a thousand collections, and manual collection at that volume degrades in the way that ruins the arithmetic, by quietly skipping runs. Bavior runs a fixed prompt set across five engines on a schedule and records which sources each answer cited, which is what makes the acceptance tests above observable, and where a cited source is a live discussion thread it drafts a reply on an account you control for you to approve before anything posts. It does not do the keep column: it will not improve your rankings, fix your indexation, or write your pages. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 29 Aug 2026).