All three get asked as if the answer were a number, and for all three the honest answer is a description of a study design. The shape repeats: a large raw correlation gets shared, the one test built to separate cause from correlation finds far less, and the engine’s own documentation says the thing was never a lever. Content that gets cited covers the tactics on the other side of that line.
Length, schema and llms.txt: what does the evidence actually say?
The three tactics founders ask about most, and the three with the weakest evidence behind them. This page names every study run on each, who ran it, and what its sample could not see.
On this page
Only one of the three has a first-party answer, and it rejects all three: Google’s guide to its generative AI features states there is “no ideal page length”, that “there’s no special schema.org markup you need to add”, and that an LLMS.txt file will “neither harm nor help” because Search ignores it.1 Everything else is observational data that was never designed to answer the question people ask of it, so the popular version of each claim is unsupported rather than disproved. The one before-and-after test with a control group, 1,885 pages adding JSON-LD against 4,000 matched controls, measured AI Overview citations moving −4.6% and the other two surfaces indistinguishable from zero.8
Key takeaways- The only first-party documentation covering all three rejects all three, in the same short list of things to ignore, last updated 10 July 2026.
- The two big length datasets point opposite ways because they measure different outcomes, and neither one contains a single uncited page to compare against.
- Structured data has exactly one matched before-and-after test: a small negative on AI Overviews, nothing elsewhere, against the roughly three-to-one raw correlation that made the tactic popular.
- No engine documents reading llms.txt. Four AI vendors publish one for their own developer docs while none documents consuming one, and server logs show the requests arriving from SEO audit tools.
What is the evidential state of each of the three?
Each row names the strongest test that exists, and the thing its design prevented it from seeing.
| Lever | The strongest test anyone has run | What that test could not see |
|---|---|---|
| Page length | Two observational datasets: 174,048 cited pages, and 18,151 fetched pages scored for influence | No uncited comparison group, so neither can estimate the odds of citation at a given word count |
| Schema markup | One matched difference-in-differences test: 1,885 pages adding JSON-LD against 4,000 controls | Whether schema markup improves AI citation on engines outside the three measured, or over a longer window |
| llms.txt | One cross-section of 37,894 already-cited domains, plus two server-log studies of who requests the file | Whether llms.txt helps a site with no citations at all, since no such site was in any sample |
Does page length change whether you get cited?
Nobody has run the study that would answer that, and Google’s position is that the question is malformed. Its guide to generative AI features, last updated 10 July 2026, files content length under things you can ignore: “There’s no requirement to break your content into tiny pieces for AI to better understand it… There’s no ideal page length, and in the end, make pages for your audience, not just for generative AI search.”1
The number that circulates instead is descriptive. A December 2025 industry analysis of 560,346 AI Overviews and 1,677,876 cited URLs, narrowed to the 174,048 pages whose text could be extracted, reported a mean cited page of 1,282 words against 1,188 for organically ranking pages, with 53.4% under 1,000 words, 30.6% between 1,000 and 2,000, and 16% above 2,000.9 Vendor-published, so described rather than linked.
Read the sample before the median. Every page in it had already been cited, and no group of uncited pages sits alongside it. A dataset with no comparison group can say what cited pages look like and cannot say how much likelier a 2,400-word page is to be cited than a 700-word one. Turning its median into a content brief converts a census into a prescription.
The one relational figure it reports is a Spearman correlation of 0.04 between word count and citation position, with average length flat across positions one to three.
Why does the academic dataset point the other way?
Because it measures a different outcome on a different sample, and the two are not in conflict once lined up. Zhang, He and Yao’s 2026 framework separates citation selection, the engine choosing you, from citation absorption, your wording reaching the answer. Across 602 controlled prompts and 18,151 successfully fetched pages, its top influence quartile averages 1,943.30 words against 169.82 in the bottom quartile, a ratio of 11.44x, and its word-count bins rise steadily to pages beyond 3,000 words with a plateau around 301 to 1,000.3
Look at the bottom quartile before the ratio. An average of 169.82 words is a stub: a tag page, a thin support note, a product blurb. The 11.44x contrast is between a real article and a fragment, and says nothing about whether a 1,200-word article should become a 3,000-word one.
The authors say so themselves. Word count “alone is an incomplete proxy”, and the strongest single correlate they report is not length but an LLM relevance score at r = 0.4322, ahead of answer-citation embedding similarity at 0.3561. Their summary sentence is plain: “Page word count and structural markers matter, but semantic fit is stronger than a simple length signal.”3 Their evidence-container reading is the useful one: length pays when it buys more headings, more distinct claims and more extractable evidence, and does nothing when it buys boilerplate.
Does schema markup move AI citations?
This is the one case where somebody ran the experiment properly, and the correlation and the experiment disagree by the whole size of the claim.
Start with the correlation that made the tactic popular: in the same vendor’s data, AI-cited pages were almost three times likelier to carry JSON-LD than uncited pages. Then the same team went looking for the direction of the arrow.
Published May 2026: 1,885 pages that introduced JSON-LD between August 2025 and March 2026, matched against 4,000 controls and analysed as a difference-in-differences so platform-wide trends cancel. Google AI Overviews moved −4.6%, described as small but statistically significant against the matched controls; Google AI Mode moved +2.4% and ChatGPT +2.2%, both statistically indistinguishable from zero.8 The pitch that schema markup improves AI citation rates is not supported by that test, which is the only one of its kind that exists.
The negative should not be over-read either. Both groups were already declining, the treated pages fell slightly faster, and a handful of extreme outliers dragged that average down; strip them and the groups look roughly the same. Adding markup did nothing measurable, rather than doing harm.
First-party documentation agrees, and both pages are worth quoting, because they carry different wording and one date stamp cannot cover both. The AI features page, last updated 10 December 2025, says “there’s also no special schema.org structured data that you need to add”, while asking that existing markup “matches the visible text on the page”.2 The optimization guide, seven months later, files it under “Overfocusing on structured data” and keeps the real reason to ship it: it “helps with being eligible for rich results on Google Search”.1 Schema still has a job, just not the job it gets sold for.
Does anything actually read your llms.txt?
Nothing documented does, and the file’s author never said it would. The proposal, dated 3 September 2024 and revised to a second version on 10 August 2026, scopes the file to an agent fetching a site at the moment it needs something: llms.txt information “is instead used on demand, when an agent needs information about a topic while assisting a user”, and the expectation was usefulness “mainly… for inference rather than training”.5 Getting cited by a search-grounded answer is not a use case the document contains.
The claim that an llms.txt file helps a page get cited is unsupported by any first-party documentation, contradicted by every server log published so far, and rejected by the only large cross-section anyone has run. Google names the file explicitly: creating and maintaining one “will neither harm nor help your site’s visibility or rankings in Google Search, as Google Search ignores them”.1
The confusing part is that AI companies publish the file. Anthropic, OpenAI, Perplexity and Cloudflare all serve an llms.txt for their own developer documentation, and all four returned content when checked on 30 August 2026.6 Publishing a map for coding agents is a different act from consuming one while grounding a search answer, and none of the four documents doing the second.
Two kinds of measurement exist. A March 2026 cross-section took 37,894 domains cited at least twice in AI answers, drawn from 882 brand snapshots and more than 337,000 citations, and found 13.3% carried the file; adopters averaged 6.8 citations against 6.7, medians identical at 3.0, Mann-Whitney U p = 0.85.7 Its own note that the comparison turns significant at full scale with an effect size of r = −0.065 is the tell: that is how one study gets quoted in both directions.
Server logs ask the cruder, cleaner question of who requests the file. One operator’s month of logs recorded 52 requests for /llms.txt, every one from an SEO audit tool and none from an AI crawler, and roughly 5,000 out of 400 million requests across the wider hosting fleet. A separate 90-day study logged 84 of 62,100 AI bot visits going there, against about 265 for the site’s average content page.10
How do you tell a real test from a correlation?
Four questions settle almost every GEO claim you will be shown, all answerable from the study’s own description. Was anything changed, or were pages that already differed simply compared? Was there a control group, so a platform-wide trend cannot be mistaken for your effect? Who was excluded, and is that the group you belong to? And which outcome was measured, given that position in a list, influence on the wording and probability of citation are three variables that move independently.
The 2026 critical survey makes the same point as a protocol requirement: “Content modifications, internal linking, structured data, domain reputation, and crawler accessibility must be separated.”4 Its confidence table rates the claim that citation scores predict clicks, conversions or revenue as very low, worth holding on to when a dashboard offers a single moving number. What carries over from SEO applies the same test to the levers that survive it.
Four of the studies here are vendor-published and cannot be re-run from outside those companies: no data released, a sampling frame written as a paragraph rather than a methods section, and a vendor selling something that benefits from the conclusion. They are quoted because they are the only tests that exist, not because they are strong. The academic source is a preprint working from 602 prompts. And a null result in a sample of this shape is not proof that no effect exists anywhere: it means nobody has found one, on these engines, in these windows, at a size this design could detect.
All three questions here are decided before any tool is involved: write to the question rather than to a word count, keep structured data for the rich results it earns, and spend the llms.txt hour on a retrieval fix instead. Bavior cannot tell you whether markup or a text file would have helped, and no measurement product can, because sampling answers afterwards cannot separate your change from everything else that moved that week. What it does is narrower: a fixed prompt set across five engines on a schedule, recording which sources each answer cited, which gives you the before-and-after series a real test needs and still no control group. The free GEO audit and the free AI visibility check run without a paid plan; paid plans are from $99/mo billed monthly, or $79.17/mo billed annually (as of 30 Aug 2026).
- Google Search Central, “Google’s Guide to Optimizing for Generative AI Features on Google Search”, last updated 10 Jul 2026 (first-party; page length, LLMS.txt, structured data): developers.google.com/search/docs/fundamentals/ai-optimization-guide
- Google Search Central, “AI features and your website”, last updated 10 Dec 2025 (first-party; a separate page with separate wording: “no special schema.org structured data that you need to add”): developers.google.com/search/docs/appearance/ai-features
- Zhang Kai, He Xinyue, Yao Jingang, “From Citation Selection to Citation Absorption”, arXiv:2604.25707v2, 29 Apr 2026 (preprint); 602 prompts, 21,143 citations, 18,151 fetched pages; §§8.1 and 8.2: arxiv.org/abs/2604.25707
- “Optimizing Visibility in Generative Engines: A Critical Survey”, 15 Jul 2026, arXiv:2607.14035 (preprint); the confound-separation protocol; Table 5 rates citation-to-revenue prediction very low: arxiv.org/abs/2607.14035
- Jeremy Howard, “The /llms.txt file, v2”, dated 3 Sep 2024, modified 10 Aug 2026 (the proposal itself, first-party): github.com/AnswerDotAI/llms-txt
- Four vendors publishing an llms.txt for their own developer documentation, all serving content when fetched 30 Aug 2026: docs.anthropic.com/llms.txt, developers.openai.com/llms.txt, docs.perplexity.ai/llms.txt, developers.cloudflare.com/llms.txt
- llms.txt cross-section, 19 Mar 2026: 37,894 domains cited at least twice, from 882 brand snapshots; 13.3% adoption; 6.8 against 6.7 mean citations, medians 3.0, Mann-Whitney U p = 0.85, r = −0.065 at full scale. Vendor-published; described, not linked.
- Schema difference-in-differences test, 11 May 2026: 1,885 pages adding JSON-LD between Aug 2025 and Mar 2026 against 4,000 matched controls; AI Overviews −4.6%, AI Mode +2.4%, ChatGPT +2.2%. Vendor-published; described, not linked.
- Content-length analysis, 3 Dec 2025: 560,346 AI Overviews, 1,677,876 cited URLs, 174,048 pages with extractable text; mean 1,282 words against 1,188 organic; Spearman 0.04 against citation position. Vendor-published; described, not linked.
- Two server-log studies of who requests /llms.txt: one operator’s month of logs, 52 requests all from SEO audit tools plus roughly 5,000 of 400 million fleet-wide, 5 Mar 2026; and a 90-day study logging 84 of 62,100 AI bot visits, 5 Feb 2026. Described, not linked.
Frequently asked questions.
How long should a page be to get cited by AI search?
There is no evidence-backed number, and Google's own guide states there is no ideal page length. The most-quoted figure, a mean of 1,282 words across 174,048 cited pages, describes pages that were already cited rather than pages that got picked because of their length, and 53.4% of them are under 1,000 words. Write until the question is answered, then stop.
Should I add schema markup to get cited by AI?
Add it for rich results, which is the benefit it actually pays. The one matched before-and-after test, 1,885 pages that added JSON-LD against 4,000 controls, found AI Overview citations moving 4.6% down and the other surfaces indistinguishable from zero, and Google states there is no special schema.org markup you need to add for its AI features. Keep it, claim nothing extra for it.
Do I need an llms.txt file on my site?
No documented search engine reads one. Google states the file will neither harm nor help because Search ignores it, the proposal's own author scopes it to agents fetching a site on demand rather than to search citation, and two server-log studies found the requests arriving almost entirely from SEO audit tools. Publish one if an agent integration wants it, and count nothing toward visibility.
Why do so many GEO guides still recommend all three?
Because each has a large raw correlation behind it and correlations are cheap to produce. Cited pages really are about three times likelier to carry JSON-LD, and cited pages really do average more words than stubs. The step that is missing is a control group, and in the one case where somebody added one the effect disappeared. Ask which study changed something and what it compared against.
The team that researches and maintains Bavior’s writing on Reddit marketing and AI search visibility. Every figure here is attributed to a named source with the date it was checked, and none of our links are affiliate links.
Found a number that looks wrong? Tell us and we will re-check it: support@bavior.com