Averaging destroys the signal because the classes have very different base rates, so the blend is as much a property of the panel's composition as of your visibility. Take a 40-prompt panel: 12 category-definition prompts where you are named 70% of the time, 10 shortlist at 15%, 8 comparison at 30%, 4 fit at 25%, 3 objection at 10% and 3 how-to at 45%. The blend is 37.4%. Add five more category-definition prompts you already win, change nothing else, and it reports 41.0%: a 3.6-point rise produced by an editorial decision.
The second reason is that a win is not the same event in each class. On a shortlist prompt, presence anywhere in the list is the whole outcome; on a comparison prompt it is worth nothing if the sentence next to your name is wrong; on a category-definition prompt the win is the correct category rather than a neighbouring one. A single presence rate asks the right question of two classes and the wrong question of four.
The third reason is that the classes are answered from different source populations, so they respond to different work: category-definition and how-to answers lean on pages a company can write, shortlist and objection answers on pages it cannot. One number says something moved without saying which team could have moved it. The prompt research stage builds the panel; this article is the tag on each row.