TR-2026-003 ·

The Narrow Set of Sources Behind AI Answers

Citations in machine answers are distributed far more unequally than the field of available sources, so a small set of intermediaries decides how a market is described.

The sources an assistant draws on when it answers a question are not a representative sample of the material available to it. Published audits of citation behaviour agree that the distribution is heavily skewed: a small group of domains supplies a disproportionate share of what is cited, while the majority of cited domains appear once and are never seen again. For an organisation this is a structural fact about the market it operates in rather than a fact about its own website. Its representation is settled largely by intermediaries it does not control, and the identity of those intermediaries changes faster than most organisations revise anything.

The shape of the distribution

The most direct measurement of inequality comes from an academic audit rather than a vendor. Mowafak Allaham and Nicholas Diakopoulos of Northwestern University, in a paper published on arXiv in May 2026, ran 712 queries across politics, health and environment through ChatGPT, Copilot, Gemini and Perplexity, and computed a Gini index over the cited domains. They report that the citations in aggregate across all four engines are unequally distributed over sources, with a Gini coefficient of 0.68. Per engine the figures were ChatGPT 0.648, Gemini 0.595, Perplexity 0.563 and Copilot 0.492, against mean citations per response of 14.68, 8.44, 7.92 and 7.26 respectively.

The tail is where the inequality becomes legible. The same audit found that 59.1% of cited domains are cited only once and 16.5% are cited twice, while the top 25 domains take 23.8% of all citations. Three quarters of the domains an assistant touches, in other words, appear once or twice and no more. Being cited is a common event; being cited repeatedly is not.

In one subject area the concentration is sharper still. Kai-Cheng Yang of Northeastern University, in an arXiv paper of July 2025 examining 24,069 conversations and 366,087 citations across twelve models from OpenAI, Perplexity and Google between March and May 2025, found that news accounted for 9.0% of citations and that within that slice the top 20 news sources account for 67.3% of all citations in OpenAI’s models. The equivalent top-20 share was 31.9% for Google and 28.5% for Perplexity. The corresponding Gini coefficients were 0.83 for OpenAI, 0.77 for Perplexity and 0.69 for Google.

Who the intermediaries are

Vendor measurements name the beneficiaries, and their units differ, which matters when the figures are compared. Profound, which sells AI visibility monitoring and therefore has a commercial interest in the subject, published in June 2025 an analysis of 680 million citations gathered between August 2024 and June 2025. Measured as a share of total citations, it reports that Wikipedia supplies 7.8% of ChatGPT’s citations, Reddit 1.8%, and Forbes and G2 1.1% each, with Reddit at 2.2% and YouTube at 1.9% in Google AI Overviews and Reddit at 6.6% in Perplexity. The same analysis records commercial .com domains taking over 80% of citations, with .org second at 11.29%.

Profound’s later and larger dataset, covering more than four billion citations and 300 million responses to late October 2025, gives an aggregate top ten across answer engines: Reddit 3.11%, YouTube 2.13%, Wikipedia 1.35%, Forbes 0.80%, then NerdWallet, TechRadar, Tripadvisor, LinkedIn, Gartner and Quora, none above 0.5%. No single domain approaches dominance in aggregate; concentration is a property of the group, not of any member of it.

Ahrefs, also a vendor in this market, measured the same phenomenon in a different unit. Working from 55.8 million AI Overviews drawn from 590 million keywords in its own index, it reported in May 2025 that the top 50 domains hold 28.90% of all AI Overview mentions, with Reddit at 2,971,746 mentions, English Wikipedia at 2,920,525, Quora at 2,276,494 and YouTube at 2,063,355. These are mention counts, not citation shares, and are not comparable with Profound’s percentages.

The instability that undoes the league table

A ranked list of favoured domains invites the conclusion that an organisation should seek presence on the leaders. The available measurement of change over time argues against treating any such list as durable. Semrush, tracking 230,000 prompts and more than 100 million citations across ChatGPT Search, Google AI Mode and Perplexity between 14 July and 12 October 2025, found substantial movement inside a single quarter. It reports that in ChatGPT, Wikipedia dropped from appearing in roughly 55% of AI prompt responses to less than 20% over that window, while Reddit fell from close to 60% of responses in early August to roughly 10% by mid-September.

Semrush’s unit is share of responses in which a domain appears, not share of total citations, and it cannot be set against Profound’s figures directly. What survives the difference in units is the direction of travel: the composition of the narrow set moved sharply within three months, without any corresponding change in the underlying sources. The same study found Google AI Mode citing LinkedIn in nearly 15% of responses and Wikipedia in about 2%, and Perplexity citing Wikipedia in 0.8%, which is a reminder that the set is not shared across platforms either.

What this means for an organisation

Two findings hold across every study here, and they point in opposite practical directions. Citation is concentrated, so the intermediaries that carry an organisation’s description matter more than the count of places it appears. And the concentration is unstable, so the specific intermediaries that mattered last quarter are a poor guide to next quarter’s.

The reconciliation is to treat the composition of the set as outside an organisation’s control and its own presence in the wider field as inside it. An organisation cannot decide whether ChatGPT favours Reddit or Wikipedia in a given month. It can decide whether an accurate, current description of itself exists in the categories of source that appear in these audits at all — reference works, review and category directories, trade coverage, technical documentation, video — so that a shift in weighting between them is a change in which correct description is used rather than a change from a correct one to none.

The corollary is that a monitoring exercise conducted once produces a snapshot with an unreported expiry. Semrush’s quarter shows movement of tens of percentage points. Any assessment of where an organisation stands with these systems has to be dated, and treated as describing the period measured rather than a settled position.

All reports