Wikipedia
Wikipedia is the free encyclopedia run by the Wikimedia Foundation, and in AI search it is the reference layer engines fall back on. Measurements of its share of ChatGPT citations range from 5%[1] to 13.15%[2] depending on the sample, and every large citation study places it at or near the top alongside Reddit. For anyone doing GEO, that position is a constraint rather than an opportunity: the encyclopedia article is usually already there, and the useful question is what a brand can say that it does not.
How much engines cite it
Profound's analysis of about 730,000 ChatGPT conversations carrying citations, sampled in the United States between October and December 2025, found Wikipedia supplying 5% of all citations and appearing in nearly one in six of them[1]. A 5W synthesis built on a Similarweb sample of roughly 600,000 citation events from January to February 2026 put it first at 13.15% of US ChatGPT citations, ahead of Reddit at 11.97%[2], while a figure reported through Contently gives 7.8%[3]. The three disagree because they sample different query mixes and periods, so the shape is trustworthy and the number is not. A March 2026 study of 30 million cited sources across five systems found ChatGPT favouring Wikipedia, Reddit and editorial sites, with Reddit first overall[4].
Wikipedia as a training-data proxy
Wikipedia pageviews are used in language-model research as a measure of how well known something is. A 2026 analysis of the OLMo models and their full Dolma pretraining corpus, covering 7.4 trillion tokens and 2,000 entities, reports that pretraining exposure correlates strongly with Wikipedia popularity, which it treats as validating exposure as a proxy for real-world salience[5]. The same study finds that a model's own popularity judgments track exposure more closely than they track Wikipedia, and that the effect is strongest in larger models and in the long tail[6]. How familiar a brand is to a model is therefore a property of the corpus rather than of the world.
Being cited is not being visited
Wikimedia reported human pageviews in May and June 2025 running about 8% below the same months of 2024 after it revised its bot detection[7], and attributed the fall to search engines answering queries directly with generative AI and to younger readers moving to social video[8]. Crawl load moved the other way: the Foundation said in April 2025 that at least 65% of its most expensive traffic came from bots, which were about 35% of pageviews[9]. Wikipedia is the clearest public case of the zero-click trade, because it publishes both halves of it.
Paid access for AI reuse
Wikimedia Enterprise sells bulk machine-readable access to the same content the free APIs serve. The Foundation's stated reason is that the content is free but the infrastructure is not, and it now rate-limits large-scale reusers of the free endpoints[10]. Enterprise income is capped at 30% of total Foundation revenue[11]. The arrangement is worth watching as a template for content licensing generally: a publisher that cannot stop AI crawlers can still charge for a cleaner pipe.
What it means for a site
Profound's advice on the strength of its own data is that the position is not competable, and that a brand should aim to be the citation after Wikipedia, carrying deeper analysis, current data or expert judgement the encyclopedia cannot hold[12]. That follows from what Wikipedia is: a tertiary source with no original reporting and a lag behind events. The practical read is to keep the encyclopedia entry accurate, treat it as unwinnable ground, and put the effort into fresher material that a fan-out query has a reason to reach for.
Data
Show the numbers (3)
| Point | % |
|---|---|
| Human pageviews, May-Jun 2025 vs 2024 | -8 |
| Share of expensive traffic from bots | 65 |
| Share of pageviews from bots | 35 |
References (12)
- How ChatGPT sources the web - Profound Archive
- Wikipedia and Reddit Now Drive Over 25% of ChatGPT Citations in the US - 5W via PR Newswire Archive
- Top 10 Sources LLMs Cite Most in 2026 - Contently Archive
- AI search engines cite Reddit, YouTube, and LinkedIn most: Study - Search Engine Land Archive
- Pretraining Exposure Explains Popularity Judgments in Large Language Models
- Pretraining Exposure Explains Popularity Judgments in Large Language Models
- New user trends on Wikipedia - Wikimedia Foundation Archive
- New user trends on Wikipedia - Wikimedia Foundation Archive
- How crawlers impact the operations of the Wikimedia projects - Wikimedia Diff Archive
- The cost of 'free': How Wikimedia Enterprise protects Wikipedia - Wikimedia Foundation Archive
- The cost of 'free': How Wikimedia Enterprise protects Wikipedia - Wikimedia Foundation Archive
- How ChatGPT sources the web - Profound Archive
Last updated 2026-09-04. Written and maintained by Baseline Labs.