The Reference Layer Beneath the Answers
A volunteer-maintained encyclopedia supplies a large share of what machines say about organisations while its human readership falls and its relationships with AI firms become commercial.
Underneath a substantial share of machine-generated statements about organisations sits one volunteer-maintained corpus. Wikipedia is measured as the most-cited domain inside at least one major assistant, is named by Google as a source of its Knowledge Graph, and ranks near the top of a widely used training corpus by volume. Over the same period its human readership has fallen, most of the load on its servers has become non-human, and the largest AI firms have moved from taking its content freely to paying for structured access to it. This report sets out those measurements and what the arrangement implies for organisations described in it.
What sits on top of it
Profound, a vendor selling AI visibility measurement, analysed 680 million citations between August 2024 and June 2025 and found Wikipedia supplying 7.8% of all ChatGPT citations, the largest single share of any domain on that platform. In its later aggregate across answer engines, covering more than four billion citations to late October 2025, Wikipedia sits third at 1.35% of total citations, behind Reddit and YouTube.
Neither share is stable or uniform across platforms. Semrush, tracking 230,000 prompts and more than 100 million citations between 14 July and 12 October 2025, found that in ChatGPT, Wikipedia dropped from appearing in roughly 55% of AI prompt responses to less than 20% across that quarter, while Google AI Mode cited it in about 2% of responses and Perplexity in 0.8%. Semrush’s unit is the share of responses in which the domain appears, not the share of citations, so its figures cannot be set against Profound’s; the two describe different quantities.
The dependence extends beyond live citation. Google states that Wikipedia is a commonly-cited source for the Knowledge Graph, though not the only one, alongside hundreds of sources across the web and licensed data. No percentage has ever been published for Wikipedia’s share of knowledge panel content, and none should be inferred. In training data the position is similar: The Register, reporting the Washington Post and Allen Institute for AI analysis of Google’s C4 corpus, which ranked 15.7 million domains by token count, records that most of its text comes from Google patents, Wikipedia, and Scribd, placing Wikipedia second by volume. C4’s English portion holds roughly 365 million documents and 156 billion tokens, and the paper describing it confirms that the single-most represented website in the corpus is patents.google.com.
What is happening underneath it
The corpus supporting all of that is being read less by people. The Wikimedia Foundation reported in October 2025 that human pageviews had fallen by roughly 8% compared with the same months in 2024. The figure carries a methodological history that must travel with it: the Foundation observed anomalous traffic from Brazil around May 2025, updated its bot detection, and reclassified its traffic data for March to August 2025 after finding that much of the excess was coming from bots built to evade detection. It cautions that its detection systems apply different rules at different points in time.
The Foundation’s own explanation is unusually direct for a publisher: it believes the declines reflect the impact of generative AI and social media on how people seek information, especially with search engines providing answers directly to searchers, often based on Wikipedia content.
Who reads it, and who pays for it
The composition of the remaining load is the more striking measurement. In April 2025 the Foundation’s infrastructure team reported that at least 65% of the resource-consuming traffic to the website comes from bots, against about 35% of total pageviews, with multimedia bandwidth up 50% since January 2024. The expensive traffic and the human traffic have largely separated.
The commercial relationship has moved in step. Wikimedia Enterprise announced in January 2026 that Amazon, Meta, Microsoft, Mistral AI and Perplexity had joined its roster of partners for the first time, alongside Google, Ecosia, Nomic, Pleias, ProRata and Reef Media. The firms whose products displace visits to the encyclopedia are now paying for access to it.
What this means for an organisation
Three things follow, and the first is a warning against the obvious inference. The measured dependence on this corpus is large but neither fixed nor evenly distributed: a share that moves from 55% of responses to under 20% in a quarter on one platform, and stands at 0.8% on another, does not support treating an encyclopedia entry as a lever. An organisation with no entry is not thereby absent from machine answers, and one with an entry has not thereby secured its description.
The second concerns what an entry actually governs. A reference article is not primarily a source of visibility; it is a source of resolution. It fixes what the organisation is called, which other names refer to it, what category it belongs to, and which dates and figures are treated as settled. Those are the details a system reuses without attribution, and they propagate into knowledge panels and training corpora on timescales that no correction can reverse quickly. An outdated description in this layer is more consequential than an outdated description almost anywhere else, because it is the version other sources check against.
The third is about who maintains it. This layer is written by volunteers under editorial policies an organisation has no standing in, and the same is true of the licensed and crawled sources feeding the Knowledge Graph. An organisation cannot commission its entry, correct it directly without disclosure, or require that it be kept current. What it can do is ensure that the independent published record the volunteers work from — filings, trade coverage, documentation, registry entries — is itself accurate and current, since a reference layer can only be as right as the sources available to the people maintaining it.