Home/Learn GEO/The anti-pattern
Prompt research · The anti-pattern

Why is your feature list the wrong prompt panel source?

Because a panel written in your own vocabulary measures whether engines echo your marketing, and they usually will. The number it produces goes up as you publish more, and it can rise in the same quarter your real visibility falls.

On this page
Share this
Share on X Share on LinkedIn
The short answer

A prompt written from your own positioning is a self-referential prompt: its wording appears mostly on pages you own, so there is almost no competing corpus, retrieval returns your page, and the answer names you. A panel built that way reports a presence rate near the top of its range that rises whenever you publish, which measures your content inventory rather than your visibility to a stranger. The one controlled benchmark of the field found only three of 54 cases with a significant ranking gain, while context position one was worth 2.77 rank places in retail against 0.36 for the best content rewrite.3

Key takeaways
  • Self-referential means the wording came from your marketing, not a buyer: a coined term, a feature name only your site uses, or a problem framing that exists because you wrote it.
  • The rate climbs on its own, because official pages are already the largest cited category, publishing adds vocabulary only you own, and panels grow by addition rather than replacement.
  • Two mechanical checks settle most cases: drop your brand name and re-run the prompt, then search the exact phrasing and see whose pages come back.
  • Report attested and self-referential entries as two rates, never one blended headline, and cap branded prompts at about four, scored on accuracy rather than presence.
The mechanism

What is wrong with a panel written from your positioning?

A panel written from your own positioning asks questions only your own pages answer, and then records the fact that your own pages answered them as evidence of visibility. That is a tautology with a percentage attached. A self-referential prompt is one whose wording is drawn from your product marketing rather than from anything a buyer said: it contains a term you coined, a feature name only your site uses, or a framing of the problem that exists because you wrote it. The competing corpus for such a question is close to empty, so retrieval has almost nowhere else to go.

The reason this works so reliably is documented rather than theoretical. Google says its generative features on Search are “rooted in our core Search ranking and quality systems”,8 and that to be eligible a page “must be indexed and eligible to be shown in Google Search with a snippet”.1 Ordinary ranking therefore governs which pages become candidates, and you rank first for your own vocabulary almost by definition. The 2026 critical survey rates query–document relevance and context position as major determinants of what gets used, at high confidence.2 A prompt built from your own words maximises exactly that relevance for exactly one document, which happens to be yours.

Why it drifts up

Why does the number go up on its own?

Three mechanisms push a self-referential panel upward over time, and none of them is your visibility improving.

01

Official pages are already the largest cited category

In a study of 602 controlled prompts producing 21,143 search-layer citations, the sources it classes as official were the single largest category, at 34.22% on ChatGPT and 46.35% on Google. Engines reach for the official page readily.

02

Publishing more covers more of your own phrasing

Every new page adds vocabulary that only you use, which adds prompts you are guaranteed to win if anyone writes them. The rate tracks your publishing calendar rather than any change in how buyers find you.

03

Panels grow by addition, not replacement

Self-referential prompts are the easiest to write, so they are what gets added when a panel expands. The denominator changes composition silently and the headline rises without a single answer changing.

What makes the first mechanism decisive rather than merely helpful is how far position dominates everything else in controlled tests. The 2026 critical survey summarises the one benchmark built for the question: across two tasks, six domains, approximately 1,900 queries and 16,360 documents, only three of 54 method–domain combinations are significantly positive.2 The benchmark itself, C-SEO Bench at NeurIPS 2025, tested ten methods, and in its retail domain moving a source to context position one was worth 2.77 rank places on average against 0.36 for the best content rewrite tested.3 Read that as a statement about self-referential prompts and it is unambiguous: you already occupy position one for your own vocabulary, so the prompt is measuring a lever you cannot lose rather than one you can move.

The arithmetic

What does the corrupted number look like?

A corrupted panel number looks like a headline presence rate climbing from 41% to 45% and then holding at 42% while real visibility falls by a quarter. It is worked below with invented but realistic inputs, as illustrative arithmetic rather than data.

Panel versionReal promptsSelf-referential promptsHeadline presence rate
v1, 40 prompts28 at 20%12 at 90%41%
v2, four more added28 at 20% (unchanged)16 at 90%45%, nothing improved
v3, a quarter later28 at 15% (fallen)16 at 90%42%, still above where you started
The panel you wanted40 at 20%none20%, and it moves when reality does

The v3 row is the one worth sitting with: the real prompts lost a quarter of their presence, and the headline still reads above the starting figure. Nothing in that table requires bad faith: each version was assembled by someone adding reasonable-looking prompts, and every individual number is correct. The corruption is entirely in the composition of the denominator, which is why the source ranking insists on capping any single source at roughly 40% of the panel and on versioning the panel whenever its composition changes.

The defence is arithmetic, not vigilance. Report the attested and self-referential halves as separate rates rather than as one number, and the effect disappears: a slice that never moves and a slice that does are informative side by side, and misleading when averaged. A 2026 statistical treatment of AI-visibility measurement makes the general form of this point, warning that citation distributions follow a power law and that many apparent differences fall within the noise floor of the measurement process.4 A blended headline hides both problems at once.

The test

How do you tell a self-referential prompt from a real one?

Ask three questions about the wording and run two mechanical checks, and the classification takes about fifteen seconds per prompt. Does it contain a word you coined? A feature name, a methodology name, a category label you invented for a positioning deck. Would a buyer who has never heard of you type it? Read it aloud as though you were the buyer; if it only makes sense to someone who has read your homepage, it is self-referential. Can you point to a human who said something close to it? This is the same attestation test that ranks every other source, and it is the one that settles disagreements.

The two mechanical checks catch what judgement misses. Remove your brand name and re-run it: if the prompt retrieves you when your name is absent, the question is genuinely about the category; if it collapses into a generic answer naming five other companies, you were measuring your brand rather than your topic. Search the exact phrasing and see whose pages come back: if every result is a page you own, no competing corpus exists and the panel entry can only tell you whether your own site is indexed. Both checks take a minute and neither requires a tool.

What the checks do not settle is the common borderline case: a phrase you popularised that buyers have genuinely adopted. Treat adoption as empirical rather than flattering. If the phrase reaches your call transcripts in buyers’ mouths it is attested; if it lives only in your marketing and in pages linking to you, it has not been adopted.

Remediation

How do you rebuild a panel that is already contaminated?

Freeze the panel and give it a version number before you edit anything, because once entries change without a version you lose the ability to compare quarters at all. Then classify every entry with the two checks above and store the verdict in the panel itself, so the classification is auditable rather than remembered. Restate history under the new split rather than restating the blend: publish the attested rate back through every earlier run, and let the self-referential rate sit beside it as a second series you never average in.

One caution before anyone announces that the corrected number has moved. Repeated sampling across three generative search platforms found citation distributions follow a power law, with many apparent differences between domains falling inside the measurement noise floor.4 A clean panel still needs several runs before a change counts as a change.

The exception

Should the panel contain any branded prompts?

Yes, a small, fixed, separately reported bucket of them, because branded prompts answer a question no other class can: when an engine describes you, is the description accurate? The evidence says not to assume it is. The Tow Center’s audit of sixteen hundred queries across eight engines found incorrect answers to more than 60% of them, and its earlier test of two hundred quotes found 153 responses partially or entirely incorrect with uncertainty signalled 7 times.5 Those studies measured news attribution rather than product descriptions, so treat them as a reason to check rather than as an estimate of your own error rate.

The rules for that bucket are strict because its presence rate is uninformative by construction. Keep it to about four prompts, score them on accuracy rather than presence, using the three-value outcome from objection prompts, and never let them into the headline number. A vendor analysis of the top 1,000 pages one assistant cited in September 2025 put homepages and landing pages at 23.8%, and found only about a third of the most-cited pages in categories a business could realistically compete in.6 Your own pages get cited readily; that is a fact about engines, not a score for you.

The general principle behind all of this is worth stating flatly, because it applies well beyond prompt panels. A metric you control the inputs to is not a measurement. The moment you can raise a number by writing more of your own copy, the number has stopped describing the outside world, and the 2026 critical survey’s lowest-confidence row is a reminder of how far that gap can run: it rates the claim that citation scores predict clicks, conversions or revenue at very low confidence.2

The honest limit of this article

Nobody has measured how contaminated real prompt panels are, because panels are private and no dataset of them exists. The mechanism argued here is an inference from two measured things, that official pages are the largest single category of cited source and that relevance and context position dominate what gets used, plus the definition of a presence rate. That study defines the category itself and warns its taxonomy contains noisy values, so read 34% to 46% as a soft-edged range. The worked arithmetic uses invented but realistic inputs; its 20% and 90% figures are illustrative, and your own numbers will differ.

Where a product fits, and where it does not

The whole defence here is free and takes an afternoon: label each prompt as attested or self-referential, drop your brand name and re-run the doubtful ones, report the two slices separately, and cap the branded bucket at about four entries. That is a spreadsheet column and a rule, and it decides whether every other number in this stage means anything. Bavior cannot do it for you. It does not write your prompts, cannot know which phrases your buyers actually use, and will run whatever panel you give it, including a bad one. What it does is the execution: a fixed panel across five engines on a schedule, with every cited URL recorded, the fastest way to spot a self-referential entry after the fact; where a cited source is a live thread it drafts a reply on an account you control, and you approve it before anything posts. The free AI visibility check and the free GEO audit run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 30 Aug 2026).

Sources, all re-checked 30 Aug 2026
  1. Google Search Central, “AI features and your website”, updated 10 Dec 2025 (a page must be indexed and snippet-eligible; first-party): developers.google.com/search/docs/appearance/ai-features
  2. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (relevance and context position high confidence; citation scores predicting clicks, conversions or revenue very low; restates C-SEO Bench as approximately 1 900 queries and 16 360 documents): arxiv.org/abs/2607.14035
  3. Puerto, Gubri, Green, Oh & Yun, “C-SEO Bench: Does Conversational SEO Work?”, NeurIPS 2025, arXiv:2506.11097; ten methods, “more than 1.9k queries and 16k documents”; three of 54 cases significant; table 3, GPT-4o-mini, retail: best SEO 2.77 against 0.36: arxiv.org/abs/2506.11097
  4. Sielinski, “Quantifying Uncertainty in AI Visibility”, Mar 2026, arXiv:2603.08924 (preprint; power-law citation distributions and the noise floor across three platforms): arxiv.org/abs/2603.08924
  5. Tow Center, Columbia: Jaźwińska & Chandrasekar, “AI Search Has a Citation Problem”, 6 Mar 2025, sixteen hundred queries, eight engines, over 60% incorrect: cjr.org; and “How ChatGPT Search (Mis)represents Publisher Content”, Nov 2024, two hundred quotes, 153 incorrect, uncertainty acknowledged seven times: cjr.org
  6. Most-cited-pages study, October 2025: the top 1,000 pages one assistant cited that September; homepages and landing pages 23.8%, about a third in categories a business could realistically compete in. Vendor-published; described rather than linked, per this curriculum’s sourcing rule.
  7. Zhang, He & Yao, “From Citation Selection to Citation Absorption”, Apr 2026, arXiv:2604.25707; 602 prompts, 21,143 search-layer citations; official sources 34.22% ChatGPT, 46.35% Google, 44.07% Perplexity (preprint): arxiv.org/abs/2604.25707
  8. Google Search Central, “Google’s guide to optimizing for generative AI features on Google Search”, updated 10 Jul 2026 (the “rooted in our core Search ranking and quality systems” sentence; first-party): developers.google.com/search/docs/fundamentals/ai-optimization-guide
FAQ

Frequently asked questions.

What exactly makes a prompt self-referential?

Its wording comes from your marketing rather than from a buyer, so the only pages that match it are pages you own. Three markers catch nearly all of them: a term you coined, a feature name only your site uses, or a framing of the problem that exists because you wrote it. The quickest confirmation is mechanical: search the exact phrasing and look at whose pages come back. If every result is yours, the prompt has no competing corpus, and measuring your presence in the answer only confirms your own site is indexed.

Is it wrong to track my own brand name in a prompt panel?

No, provided the branded prompts are a small separate bucket scored on accuracy rather than presence, and never folded into the headline. Branded prompts answer a real question, which is whether an engine describes you correctly, and the Tow Center's audit of sixteen hundred queries across eight engines, which found incorrect answers to more than 60% of them, is reason enough to check. What they cannot tell you is whether a stranger finds you, because you will nearly always be present in an answer to a question containing your own name.

Our AI visibility score keeps rising. How do I know it is real?

Split the panel and look at the two rates separately, which takes about an hour and settles the question. Score the entries whose wording came from buyers as one number and the entries whose wording came from your own positioning as another. If the first is flat or falling while the blended headline rises, the movement is composition rather than visibility: a slice you always win growing as a share of the denominator. Also check whether prompts were added without a version bump, because a panel that grows by addition drifts upward on its own.

Can I use my own product vocabulary anywhere in the panel?

Use it where buyers have genuinely adopted it, and test adoption rather than assuming it. A phrase you popularised that now appears in call transcripts in your buyers' own mouths is attested vocabulary and belongs in the panel like any other buyer wording. A phrase that appears only in your marketing and on pages linking to you has not been adopted, however proud you are of it. The evidence lives in your transcripts, so the check is the same one that governs every source in this module: point to a human who said it.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

A number you can raise by writing
is not a measurement.

Start free trial