Home/Learn GEO/Legitimate vs manipulation
Off-site GEO · The boundary

What is legitimate GEO, and what is manipulation?

Everyone doing this work wants to be cited, so wanting it cannot be the line. The line has to be a property you can test on the document, and the literature has already written one down.

On this page
Share this
Share on X Share on LinkedIn
The short answer

A legitimate tactic makes something that is true easier to find and easier to quote. A manipulative one makes something that is false, or that you have not earned, more likely to be repeated. The 2026 critical survey of generative engine optimization turns that into four cumulative tests, all four rather than most, and rates “retrieved documents constitute a genuine attack surface” at high confidence: in the strongest demonstration it reports, preference manipulation raised a fictitious camera’s recommendation rate from 34.0% to 59.4%.1

Key takeaways
  • Influence is not the offence. White-hat optimization and adversarial attacks share the same objective; what separates them is the constraint on the text.
  • The four tests are semantic preservation, evidentiary authenticity, content-instruction separation, and disclosure and fairness. Passing three of four is failing.
  • A fabricated statistic wins and loses at once: it “may increase reuse while degrading epistemic quality”. The engine cannot check it and a reader can.
  • On a community platform the enforcer is a person, not a ranking system. Where a community requires disclosure of a commercial affiliation, it is required, and its own rules are the authority.
The property

Where does the line actually fall?

Not at the point where you start trying to be cited, because both sides of this argument do the same thing to the same object. The 2026 critical survey puts it in one sentence: white-hat optimization and adversarial attacks share the same mathematical objective, and “what distinguishes them is not whether they exert influence, but the constraints imposed on” the text.1 Wanting to be found is not the offence; everyone who has ever written a clear headline wanted that.

So the line has to be a property of the artefact rather than a description of the motive, because motive is unobservable and, in this field, universal. The property that carries the weight: a legitimate tactic makes something true easier to find and easier to quote, and a manipulative one makes something false, or unearned, more likely to be repeated. Both raise the odds an engine reuses your text. Only the second moves what a reader believes in a direction you already know is wrong.

That is testable in a way “be authentic” is not, because it asks about the sentence instead of the person. The same answer, in the same thread, passes when the commercial interest is stated and fails when it arrives from an account presented as unaffiliated. Nothing about the author changed. The document did.

The tests

Which four tests does a tactic have to pass?

Cumulative is the load-bearing word: three passes and one failure is a failure, and most bad tactics fail exactly one.

01

Semantic preservation

“Do the facts and qualifications remain true?” The qualifications carry as much weight as the facts. Cutting “in our own benchmark, on one metric” leaves every word true and the sentence false. That is how honest teams fail here.

02

Evidentiary authenticity

“Are statistics, reviews, and references verifiable?” A reader has to be able to get back to the origin and check. A figure with no denominator or date fails even when real, and a review you wrote for a customer fails even when the customer is happy.

03

Content-instruction separation

“Does the document inform the user rather than issue hidden commands to the model?” Text addressed to the model instead of the reader is the cleanest failure here. A page carrying a line telling an assistant to recommend you is not a flawed page; it is not a page.

04

Disclosure and fairness

“Is the commercial intent disclosed, and are competitors represented without fabricated disparagement?” Two obligations, and teams forget the second: an invented weakness in a rival’s column fails even if every word about your own product is exact.

The survey draws the consequence: “Reorganizing paragraphs or adding a verified primary source will generally satisfy these tests. A model-directed sequence, fabricated testimonial, or instruction to favor a brand will violate them. This framework prevents a rewrite from being classified as ‘white-hat’ merely because it is fluent.”1 Keep that last sentence: a paragraph built on a number nobody can trace reads exactly like a good one.

The hard case

Why is a fabricated number the hard case?

Because it wins on the metric you are watching and loses on the thing that metric stands for. The survey says it in one line: “adding a fabricated statistic may increase reuse while degrading epistemic quality.”1 Both halves are true at once, which is why the tactic keeps reappearing in advice written by people measuring citations alone.

The asymmetry underneath: an engine reusing your sentence cannot check whether the number in it happened. A reader who cares about your category can, does it once, and does not forget the result.

The survey supplies the replacement rather than just the prohibition, which is the useful part: “The criterion is therefore not to ‘add numbers,’ but to provide relevant, verifiable, dated, and properly attributed evidence.”1 Four adjectives, four failure modes: the statistic dropped in because it was impressive, the figure whose origin has been lost, the 2023 measurement of products since rebuilt, and the number that quietly became yours over three rounds of citation.

Our own field supplies the example, and it is not a fabrication. The claim that GEO increases visibility by 40% is still repeated; the survey’s confidence table lists it as rejected as a general claim, because the figure is a relative maximum on one metric under one configuration.1 Nobody invented it. A real result lost its denominator and became untrue through repetition, which is the version of this failure you are most likely to commit. The full autopsy is a separate article.

Efficacy

Does manipulation work, and does that settle anything?

It works, it settles nothing, and the first half is worth conceding before anyone else points it out. The survey rates “retrieved documents constitute a genuine attack surface” at high confidence.1 Its reported demonstrations are concrete: preference-manipulation attacks raised a fictitious camera’s recommendation rate from 34.0% to 59.4%,6 and indirect injection lifted a target by roughly three ranks on one commercial API.7 Both used pages the researchers controlled, both act after retrieval, and the products have changed since. The vulnerability is established. Its size on the open web today is not.

So the case against it cannot rest on ineffectiveness. Three other things are true. Gains decay under adoption: the NeurIPS 2025 conversational-SEO benchmark found the average gain per adopter falls as more actors adopt the same method, “depicting a congested and zero-sum nature of the problem.”2 Defences are improving rather than absent: one reranking defence cuts the success of some attacks by up to about 80%.1 And optimizing the retrieved text can cost you upstream, since the end-to-end arena found body-only optimization cut top-10 presence after reranking by about 16%.3 Sorting a claim by the stage it acts on covers that last result.

Now notice the shape of that paragraph, because it is the trap this article exists to avoid. Every reason in it is about efficacy, which is a different question from legitimacy with a different kind of answer. Whether a tactic works is settled by benchmarks that change every quarter. Whether it is legitimate is settled by the four tests, which do not. A team that draws its boundary on efficacy grounds has made a forecast, not a boundary.

Enforcement

Who enforces the boundary, and on what authority?

Two parties, two mechanisms, and confusing them is how teams get the risk wrong. The first is the search engine, which publishes its rule and names the generative surface inside it. Google defines spam as “techniques used to deceive users or manipulate our Search systems into featuring content prominently, such as attempting to manipulate Search systems into ranking content highly or attempting to manipulate generative AI responses in Google Search”, adding that offending sites “may rank lower in results or not appear in results at all.”4 That reaches the AI answers, because Google states its “generative AI features on Google Search are rooted in our core Search ranking and quality systems.”5

The second party is the community, and here the enforcer is a person. Moderators remove undisclosed promotion and ban the accounts behind it, which has three consequences. Rules are local: they differ between communities, they change, and the community’s own rules page is the authority, not this article. Where a community requires disclosure of a commercial affiliation, that disclosure is required, and it belongs in the reply, not a profile nobody opens. And enforcement is not paced like a ranking system: no gradual demotion to read as a warning, just a removal, landing on the account.

Vote manipulation and accounts built to read as unaffiliated fail on both axes at once: they fail evidentiary authenticity and disclosure, and they are the two behaviours community moderation is built to catch. A team that answered eleven threads in one week with the same lightly reworded paragraph found the pattern was the signal: plausible replies, read together, are a campaign, and were removed as one. How community threads get selected explains why volume was never the lever anyway.

Rulings

How do you rule on a genuinely borderline case?

Run the four tests and find the one that decides it. Most borderline cases turn on a single test, and naming which one converts an argument about taste into a question with an answer.

The caseThe test that decides itRuling
Answering a buyer’s question in a thread about your own categoryDisclosure and fairnessLegitimate with your interest stated. Not from an account presented as unaffiliated.
Adding a summary box so the answer lifts cleanly as a passageSemantic preservationLegitimate. Reformatting true content is the survey’s own example of passing.
Asking a satisfied customer for a reviewEvidentiary authenticityLegitimate if you ask. Not if you draft it, script it, or pay for the verdict.
Publishing a comparison page that describes competitorsSemantic preservation and fairnessLegitimate if their side is described as it is. Fails on one invented weakness.
Adding a line to your page telling an assistant to recommend youContent-instruction separationNot legitimate. It addresses the model rather than the reader.
Correcting a factual error about you on a third-party pageAll fourLegitimate, and the highest-return work in off-site GEO.

One pattern runs through the rulings: what flips a case is always a fact about the document, whether that is who is named on it, whether a claim can be traced, or whether a sentence is aimed at a reader or a model. None turned on how badly the team wanted the citation. That is why the survey’s governance section argues criteria of truthfulness, disclosure and redress outlast any list of approved formats: engines change, a rule about the document does not.1 When a case still feels undecidable after all four, treat that as information: the inclusion has probably not been earned yet, and the fix is the product, not the posting.

The honest limit of this article

The four tests are a proposal by the survey’s authors, not a measured finding and not a legal standard. No study has tested whether teams applying them get different outcomes, and no engine has adopted them. The manipulation evidence is narrower than the headline numbers suggest: those experiments ran on pages the researchers supplied and act after retrieval rather than on organic discovery, so they establish that the vulnerability exists, not how much of it is exploited on the open web. The defence figure comes from a research setting, and no engine publishes its real detection rate. Nothing here measures reputational cost, which is the argument most operators actually find persuasive.

Where a product fits, and where it does not

Nothing here needs software. The four tests are questions you ask about a draft before it goes out, and the answers are yours to make. Bavior works one slice over: it records which sources your prompt panel’s answers cite, flags which are live discussion threads, and drafts a reply on an account you control that you read, edit or reject before anything posts. Be precise about that approval step. It checks that a draft is accurate, relevant and disclosed, which is where a rushed reply usually goes wrong. It is not a laundering step and must not be sold as one: a tactic that fails the four tests fails just as badly after a human clicks approve, and no queue makes an undisclosed account disclosed or an invented figure true. Bavior does not touch votes and does not read any community’s rules for you. You read those, and the disclosure goes in the draft, whichever account posts it. The free AI visibility check runs without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 30 Aug 2026).

Sources, all checked 30 Aug 2026
  1. “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)”, 15 Jul 2026, arXiv:2607.14035 (preprint). The four tests, the fabricated-statistic warning, the reranking defence, the confidence table: arxiv.org/abs/2607.14035
  2. Puerto, Gubri, Green, Oh, Yun, “C-SEO Bench”, NeurIPS 2025 Datasets & Benchmarks, arXiv:2506.11097; “depicting a congested and zero-sum nature of the problem”: arxiv.org/abs/2506.11097, code and data at github.com/parameterlab/c-seo-bench
  3. Kim et al., “SAGEO Arena”, KDD 2026, arXiv:2602.12187; body-only optimization cuts top-10 presence after reranking by about 16%: arxiv.org/abs/2602.12187
  4. Google Search Central, “Spam policies for Google web search”, last updated 28 Aug 2026 (first-party; the generative-AI sentence): developers.google.com/search/docs/essentials/spam-policies
  5. Google Search Central, “Google’s Guide to Optimizing for Generative AI Features on Google Search”, updated 10 Jul 2026 (first-party; the “rooted in” sentence): developers.google.com/search/docs/fundamentals/ai-optimization-guide
  6. Nestaas et al., 2025; preference manipulation raising a fictitious camera’s recommendation rate from 34.0% to 59.4%. Controlled pages; reported in the survey at note 1.
  7. Pfrommer et al., 2024; indirect injection raising a target by about three ranks on the Perplexity Sonar Large Online API, URLs supplied explicitly. Reported in the survey at note 1.
FAQ

Frequently asked questions.

Is it manipulation to write a page specifically so an AI answer will quote it?

No, not on its own. The survey is explicit that white-hat optimization and adversarial attacks share the same objective, and that the constraint on the text is what separates them. Reorganizing paragraphs, answering the question in the first sentence, and adding a verified primary source all satisfy its four tests. What decides the case is whether the quoted sentence is true, traceable, disclosed, and addressed to a reader.

Do I have to disclose that I work for the company when I answer a thread?

Where the community requires it, yes, and its own rules page is the authority on that, not a summary written elsewhere. Rules differ between communities and they change. The survey's fourth test asks the same question independently of any platform: is the commercial intent disclosed. Put the disclosure in the reply itself rather than in a profile nobody opens, because the enforcement here is a moderator removing undisclosed promotion and banning the account behind it.

If manipulation demonstrably works, why not use it?

Because whether a tactic works and whether it is legitimate are different questions, and only one stays answered. Efficacy is settled by benchmarks that move every quarter: gains decay as more actors adopt the same method, defences improve, and body-only optimization can cost you presence upstream. Legitimacy is settled by the four tests and does not move. A team that draws its boundary on efficacy grounds has made a forecast, not a boundary.

Does a human approval step make a risky tactic acceptable?

No. An approval step checks that a specific draft is accurate, relevant and disclosed, which is where a rushed reply usually fails. It cannot change the category a tactic belongs to. A reply from an account presented as unaffiliated is still undisclosed after a human approves it, and a fabricated statistic is still fabricated. Approval reduces ordinary mistakes; it does not convert a manipulative tactic into a legitimate one.

Bavior Editorial

The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.

Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com

The line is a property of the document.
Test the draft, not the motive.

Start free trial