What actually gets you cited in AI answers
Most GEO advice was written before anyone measured it. So we went back to the papers, the crawler logs and the controlled tests, and sorted the advice by how much evidence sits underneath each piece of it. Some of it survives. Some of it does not, including one thing we sell tooling for.
Quotations beat everything else you can do to a paragraph
The one properly controlled experiment on AI-answer wording is still the KDD 2024 paper from Aggarwal and colleagues at Princeton, GEO: Generative Engine Optimization. They built a benchmark of user queries, rewrote one source per query nine different ways, and measured how much of the generated answer each source ended up supplying. Their headline metric is position-adjusted word count: how many words of the answer came from you, weighted by how early you were cited.
These are the numbers off Table 1, relative to an unmodified page:
| Rewrite | Change in answer share | What it means in practice |
|---|---|---|
| Add quotations | +41% | Quote a named, credible source in the text |
| Add statistics | +31% | Replace vague claims with figures |
| Improve fluency | +28% | Same facts, cleaner sentences |
| Cite sources | +27% | Attribute claims to where they came from |
| Stuff keywords | −8% | Classic SEO, actively harmful here |
Source: Aggarwal et al., GEO: Generative Engine Optimization, KDD 2024, Table 1, position-adjusted word count against the no-optimisation baseline.
They then repeated the test on Perplexity, a real engine rather than their own harness. Quotation addition again came first at 22%, and keyword stuffing again landed about 10% below the baseline. Two engines, same ordering. That is about as close to a replicated result as this field has.
If you are already the answer, this can cost you
Table 2 of the same paper splits those results by where the source ranked in the underlying search engine, and it changes what you should do with them. The gains are not spread evenly. They come off the top and land at the bottom.
| Rewrite | Rank-1 source | Rank-5 source |
|---|---|---|
| Cite sources | −30.3% | +115.1% |
| Add quotations | −22.9% | +99.7% |
| Add statistics | −20.6% | +97.9% |
Source: Aggarwal et al., Table 2, relative change in visibility by the source's search ranking.
A page sitting fifth roughly doubles its share of the answer by adding citations. The page sitting first loses a third of its share doing the same thing. GEO is a challenger's tool. If you are the incumbent answer in your category, the sensible move is to watch what your rivals are publishing rather than rewrite the page that is already winning.
All of that assumes you got fetched
Every number above measures share of an answer for a page that is already in the model's context window. Real engines search first, rerank, and only then write. A 2026 paper, SAGEO Arena, rebuilt that full pipeline over 171,003 documents and 2,700 queries, and found that optimising body text alone reduced average top-20 retrieval presence by around 9%, top-10 presence after reranking by 16%, and final citation by 6%.
Rewriting a page for citability can make it worse at the earlier job of being found. C-SEO Bench reaches the same conclusion from a different angle: most conversational-SEO methods are largely ineffective, frequently damage document ranking, and are beaten by plain traditional SEO aimed at getting the source into the model's context in the first place. It also found gains shrink as more people adopt the same tricks, which is the usual end state of any optimisation everyone runs.
So the ordering matters. Get retrieved, then get quoted. Doing the second at the expense of the first is a net loss.
Three mechanical facts
AI crawlers do not run your JavaScript. Vercel and MERJ instrumented crawler traffic across their network and found that none of the major AI crawlers currently render JavaScript. GPTBot, ClaudeBot and PerplexityBot fetch your script files and treat them as text. Only Google's stack renders, because Gemini borrows Googlebot's infrastructure. If your content arrives client-side, it does not exist for ChatGPT, Claude or Perplexity, whatever your page looks like in a browser. Vercel's numbers also show ChatGPT wasting 34.8% of its fetches on 404s, so a tidy sitemap and honest redirects are worth more than they sound.
Answer in the first third. Kevin Indig analysed 1.2 million ChatGPT responses and validated the pattern against 18,012 verified citations: 44.2% of citations were extracted from the first 30% of the page, 31.1% from the middle, and 24.7% from the last third. Long preambles bury the sentence you want quoted. Same instinct as writing for sub-questions: say the thing, then explain it.
Show the date, and mind your position in the list. A 2026 factorial study, What Gets Cited, ran 252,000 trials across six models over 18 content factors. Four gatekeepers dominated every model: topical match, price, recency and position in the retrieved set. Visible recency earns its place on anything time-sensitive. Formatting, in the same study, had negligible impact.
Mentions predict visibility three times better than links
Ahrefs correlated visibility in AI Overviews against a set of signals across 75,000 brands. Branded web mentions came out at r = 0.664. Backlinks, the metric a whole industry is built on, managed r = 0.218. Correlation is not causation and this is observational data, but the gap is wide enough to change where a budget goes.
This lines up with what the citation data shows about where engines actually source their answers: on Perplexity, brand-owned domains supply well under a third of citations. Being written about somewhere else is doing more work than anything you can do to your own paragraph.
Four things people are still selling that the data kills
Nobody reads your llms.txt. Ahrefs monitored 137,000 domains. Around 28% had published an llms.txt file, and 97% of those files received zero requests in May 2026. Among the handful that did get traffic, the retrieval bots that decide live citations accounted for roughly 1% of requests, well behind SEO audit tools. John Mueller has said no AI system uses it. We wrote about this file before and the picture has only got worse for it.
Keyword stuffing goes backwards. It lands 8% below baseline on the GEO benchmark and about 10% below on Perplexity. Already a bad idea, and generative engines punish it harder than search ever did.
Humanising an AI draft buys you no citation odds. Ahrefs ran AI-content detection over cited pages from a million AI Overviews. 87.8% were a human-AI mix and 3.6% were pure AI, so at least 91.4% of cited content had some machine writing in it. The correlation between how much AI writing a page contained and where it sat in the citation order was 0.017. That is nothing. Write well because your readers can tell. The engines are grading substance.
Formatting-only overhauls do nothing. Bolting FAQ blocks and comparison tables onto a page without adding anything new is the most-sold GEO deliverable, and the one with the least behind it. The 252,000-trial study put formatting in its negligible bucket, and C-SEO Bench found most of these methods flat or backfiring. Structure helps a reader find the answer. It does not manufacture one.
Schema markup does not buy you AI citations
We sell schema tooling, so read this with that in mind. For a long time the argument for schema as an AI-visibility lever rested on a correlation: across six million URLs, pages cited by AI were roughly three times more likely to carry JSON-LD. Then Ahrefs ran the controlled version. They tracked 1,885 pages that added JSON-LD, matched each against control pages on other domains with similar citation levels, and measured what changed.
AI Mode moved 2.4%, ChatGPT 2.2%, both indistinguishable from noise. AI Overviews moved −4.6%, a decline the authors themselves refuse to pin on schema. What the original correlation had captured was sites that invest in schema investing in everything else too. Keep the study's own caveat in view: every page in it already had 100+ AI Overview citations, so it tests pages that were in the consideration set already, and says nothing about whether schema helps a page get crawled and parsed on the way in.
Our position: schema earns Google rich results and keeps your prices, ratings, opening hours and product facts machine-readable for anything parsing your site. Both of those are measurable, which is why we build tools to audit it. Neither of them is a spell for getting quoted by ChatGPT, and anyone selling it as one has run ahead of the evidence. That discipline covers FAQ markup too, which Google retired as a rich result and which plenty of vendors still bill for.
None of this holds still
Semrush ran 230,000 prompts over 13 weeks and watched Wikipedia's share of ChatGPT responses fall from roughly 55% to under 20% inside a single month. Reddit slid over the same window. Whatever an engine's source diet looks like when you read this, it is a snapshot.
Ranking is decoupling too. In July 2025 Ahrefs found 76% of AI Overview citations came from pages in the top 10 for the same query. A later study across 863,000 keywords and four million cited URLs put it at 38%. Meanwhile the cost of losing those citations went up: AI Overviews now correlate with a 58% lower clickthrough rate for position-one results, against 34.5% a year earlier. Ranking first is buying you less, and it is buying you less of something worth less. That is the whole reason a number-one page can be invisible to AI.
In order of how much evidence is behind it
1. Serve your content in the HTML. A mechanism, observed at network scale, rather than a correlation anyone has to interpret. If a crawler needs JavaScript to see your text, three of the four major engines never see it.
2. Put the answer in the first 30% of the page. 44.2% of ChatGPT citations come from there, and the effect matches what attention research predicts.
3. Get mentioned somewhere that is not your site. The strongest observational signal anyone has measured, three times stronger than backlinks.
4. Quote named sources and give figures. Experimentally supported, replicated on a live engine, and worth most if you are not already the top result. Watch that you are not damaging retrieval to get it.
5. Show a visible date on anything time-sensitive. One of four factors that dominated 252,000 trials.
6. Skip the stuffing, the llms.txt file and the formatting-only rebuild. Zero or negative in every test that has looked.
7. Keep schema for rich results and machine legibility. Not for citations. Budget it accordingly.
One thread runs through all seven. Engines reward pages that are easy to fetch and carry something quotable. Every method that failed a test was an attempt to signal quality without having it.
Common questions
Sources: Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan & Deshpande, "GEO: Generative Engine Optimization", KDD 2024 (arXiv:2311.09735); "SAGEO Arena" (arXiv:2602.12187); "C-SEO Bench: Does Conversational SEO Work?" (arXiv:2506.11097); "What Gets Cited: Competitive GEO in AI Answer Engines" (arXiv:2605.25517); Vercel & MERJ, "The rise of the AI crawler"; Kevin Indig's ChatGPT citation-position study via Search Engine Land; Ahrefs studies on brand-mention correlation (75,000 brands), llms.txt adoption (137,000 domains), AI-generated content in AI Overviews, schema markup and AI citations (1,885 pages), AI Overview citations from top-10 pages (863,000 keywords) and AI Overview clickthrough rates (300,000 keywords); Semrush most-cited-domains study (230,000+ prompts, Jul–Oct 2025). Figures were read from the primary reports; where a study reports a range, the range is given.
Thumbnail: the Long Room, Trinity College Dublin, by Diliff, CC BY-SA 4.0, via Wikimedia Commons. Treated versions of the CC BY-SA original are shared under the same licence.
