C-SEO Bench is a benchmark for whether rewriting a document to please a generative engine actually works, and its answer is mostly no. It was presented at the NeurIPS 2025 Datasets and Benchmarks Track, covers two tasks, question answering and product recommendation, over six domains, and its code and data are public.1 It is also the result the GEO versus SEO stage is built around.
The abstract states the result without hedging: “most current C-SEO methods are not only largely ineffective but also frequently have a negative impact on document ranking, which is opposite to what is expected. Instead, traditional SEO strategies, those aiming to improve the ranking of the source in the LLM context, are significantly more effective.” The paper's own summary is that making the target document first in the context window “leads to far greater citation ranking gains in the LLM response than any C-SEO method”.
Two details make that harder to dismiss than a single experiment usually is. The first is breadth. The paper reports 54 method and domain cases, enough that scattered wins would show up, and its finding is that “out of 54 cases, we uncover only three where the ranking improvements are statistically significant”, with no method effective at all for the question answering task. The second is the direction of the failures. Several transformations did not merely fail to help, they reduced the document's rank: adding statistics lowered rankings in 19 of the 24 settings where it was measured.1 A tactic that is neutral is cheap; a tactic that is negative is a tax you pay for reading the wrong blog post.