Skip to content

Wikipedia

Entity in AI search and SEO
Volatile, last checked 2026-09-04 This page carries measured figures that move quickly.Reviewed 2026-09-04

Wikipedia is the free encyclopedia run by the Wikimedia Foundation, and in AI search it is the reference layer engines fall back on. Measurements of its share of ChatGPT citations range from 5%[1] to 13.15%[2] depending on the sample, and every large citation study places it at or near the top alongside Reddit. For anyone doing GEO, that position is a constraint rather than an opportunity: the encyclopedia article is usually already there, and the useful question is what a brand can say that it does not.

How much engines cite it

Profound's analysis of about 730,000 ChatGPT conversations carrying citations, sampled in the United States between October and December 2025, found Wikipedia supplying 5% of all citations and appearing in nearly one in six of them[1]. A 5W synthesis built on a Similarweb sample of roughly 600,000 citation events from January to February 2026 put it first at 13.15% of US ChatGPT citations, ahead of Reddit at 11.97%[2], while a figure reported through Contently gives 7.8%[3]. The three disagree because they sample different query mixes and periods, so the shape is trustworthy and the number is not. A March 2026 study of 30 million cited sources across five systems found ChatGPT favouring Wikipedia, Reddit and editorial sites, with Reddit first overall[4].

Wikipedia as a training-data proxy

Wikipedia pageviews are used in language-model research as a measure of how well known something is. A 2026 analysis of the OLMo models and their full Dolma pretraining corpus, covering 7.4 trillion tokens and 2,000 entities, reports that pretraining exposure correlates strongly with Wikipedia popularity, which it treats as validating exposure as a proxy for real-world salience[5]. The same study finds that a model's own popularity judgments track exposure more closely than they track Wikipedia, and that the effect is strongest in larger models and in the long tail[6]. How familiar a brand is to a model is therefore a property of the corpus rather than of the world.

Being cited is not being visited

Wikimedia reported human pageviews in May and June 2025 running about 8% below the same months of 2024 after it revised its bot detection[7], and attributed the fall to search engines answering queries directly with generative AI and to younger readers moving to social video[8]. Crawl load moved the other way: the Foundation said in April 2025 that at least 65% of its most expensive traffic came from bots, which were about 35% of pageviews[9]. Wikipedia is the clearest public case of the zero-click trade, because it publishes both halves of it.

Paid access for AI reuse

Wikimedia Enterprise sells bulk machine-readable access to the same content the free APIs serve. The Foundation's stated reason is that the content is free but the infrastructure is not, and it now rate-limits large-scale reusers of the free endpoints[10]. Enterprise income is capped at 30% of total Foundation revenue[11]. The arrangement is worth watching as a template for content licensing generally: a publisher that cannot stop AI crawlers can still charge for a cleaner pipe.

What it means for a site

Profound's advice on the strength of its own data is that the position is not competable, and that a brand should aim to be the citation after Wikipedia, carrying deeper analysis, current data or expert judgement the encyclopedia cannot hold[12]. That follows from what Wikipedia is: a tertiary source with no original reporting and a lag behind events. The practical read is to keep the encyclopedia entry accurate, treat it as unwinnable ground, and put the effort into fresher material that a fan-out query has a reason to reach for.

Data

Wikimedia traffic signalsHuman pageviews, May-Jun 2025 vs 2024-8%Share of expensive traffic from bots65%Share of pageviews from bots35%
Wikimedia traffic signals. Figures in %.Sources: wikimediafoundation.org, diff.wikimedia.org
Show the numbers (3)
Point%
Human pageviews, May-Jun 2025 vs 2024-8
Share of expensive traffic from bots65
Share of pageviews from bots35
References (12)
  1. How ChatGPT sources the web - Profound Archive
    Vendor documentation Published 2026-02-03 Retrieved 2026-09-04 Single-source Volatile, last checked 2026-09-04
  2. Wikipedia and Reddit Now Drive Over 25% of ChatGPT Citations in the US - 5W via PR Newswire Archive
    Vendor documentation Published 2026-05-11 Retrieved 2026-09-04 Single-source Volatile, last checked 2026-09-04
  3. Top 10 Sources LLMs Cite Most in 2026 - Contently Archive
    Vendor documentation Published 2026-04-29 Retrieved 2026-09-04 Contested Volatile, last checked 2026-09-04
  4. AI search engines cite Reddit, YouTube, and LinkedIn most: Study - Search Engine Land Archive
    Journalism Published 2026-03-31 Retrieved 2026-09-04 Single-source Volatile, last checked 2026-09-04
  5. Pretraining Exposure Explains Popularity Judgments in Large Language Models
    Academic Published 2026-05-12 Retrieved 2026-09-04 Single-source
  6. Pretraining Exposure Explains Popularity Judgments in Large Language Models
    Academic Published 2026-05-12 Retrieved 2026-09-04 Single-source
  7. New user trends on Wikipedia - Wikimedia Foundation Archive
    Documentation Published 2025-10-17 Retrieved 2026-09-04 Volatile, last checked 2026-09-04
  8. New user trends on Wikipedia - Wikimedia Foundation Archive
    Documentation Published 2025-10-17 Retrieved 2026-09-04
  9. How crawlers impact the operations of the Wikimedia projects - Wikimedia Diff Archive
    Documentation Published 2025-04-01 Retrieved 2026-09-04 Volatile, last checked 2026-09-04
  10. The cost of 'free': How Wikimedia Enterprise protects Wikipedia - Wikimedia Foundation Archive
    Documentation Published 2026-07-16 Retrieved 2026-09-04
  11. The cost of 'free': How Wikimedia Enterprise protects Wikipedia - Wikimedia Foundation Archive
    Documentation Published 2026-07-16 Retrieved 2026-09-04
  12. How ChatGPT sources the web - Profound Archive
    Vendor documentation Published 2026-02-03 Retrieved 2026-09-04

Last updated 2026-09-04. Written and maintained by Baseline Labs.

George
Online
0%