How do ChatGPT, Claude, and Gemini choose their sources?
Published on · updated on
Each in its own way, and that is why no single rule applies "to AIs" in general. ChatGPT mostly cites pages placed at the top of its search results and the sources whose name it wrote in its search: it cites about one page out of ten from what it retrieves. Claude reads pages in full and cites about half of them, retaining those that overlap with others and were found by multiple searches. Gemini distributes its response across many sources. And the three do not read the same web: across our five combined corpora, only 1.8% of the cited domains were cited by all three.
For what type of question? It all depends on the question. When it names a brand or a source, all three AIs go look for it. When it names nothing, each AI applies its own way of choosing, described below.
ChatGPT: the top of the list and the sources it names
With ChatGPT, rank in search results dominates. A page placed in the top three positions is cited 6.9 to 8.5 times more often than a page beyond the ninth, across all our corpora: HVAC in French, artificial intelligence in English, about ten sectors in French then in English. In the top three positions, 23 to 40% of pages are cited; beyond the sixteenth, 1%.
ChatGPT also writes the name of the source it wants to consult in its search, in 72 to 91% of its calls depending on the corpus, and it cites this named publisher more: 1.8 to 2.6 times more depending on the corpus. In total, it cites about one page out of ten from what its search retrieves, in our recent corpora. It also more frequently discards thin pages, with fewer than 300 words, and slow pages: in our corpora, the slowest third of pages is cited about half as much as the fastest third.
Claude: it reads everything and retains what overlaps
Claude receives entire pages in its context and cites a large share of them, between 46 and 66% of what it retrieves depending on the corpus. Rank matters little in most of our corpora. What makes a page stand out is being retrieved by two different searches for the same response, then being cited in 88 to 100% of cases, and overlapping with the content of other found pages, being 1.2 to 1.5 times more cited.
At the passage level, Claude retains developed paragraphs that use the words from its search. It is also the only one of the three whose API indicates exactly the cited text, up to 150 characters, which makes its selection rules the most measurable. A page's speed, on the other hand, plays no role with it.
Gemini: a response distributed across many sources
Gemini, queried via its API with Google search, does not expose a list of discarded pages: all the sources it returns are used. It adds topic words and translations to its searches, and distributes its response across many sources without any single one exceeding 6% of the response.
With Gemini, the retained paragraphs also contain the search words, but the API does not provide the exact position of the passage, and the magnitude of the effect remains uncertain. Google, for its part, documents the technique that breaks a question down into sub-topics and launches multiple searches, used by AI Mode and AI Overviews.
Three AIs, three webs
Across our five combined corpora, only 1.8% of the cited domains were cited by all three AIs at the same time. Each AI uses its own search engine, writes its own searches, and applies its own criteria. The same question therefore yields three lists of sources with almost no common ground.
Language widens the gap even further. When we asked the same questions in French then in English, the domains cited in French were found in English for only 1% with Claude, 14% with Gemini, and 23% with ChatGPT. Being cited by one AI tells us nothing about what the next one will do, nor what the same AI will do in another language.
What other studies measure
Ahrefs, on ChatGPT 5.2 in April 2026, also observes that ChatGPT retrieves many addresses and only cites a portion of them, about half according to its method, with sharp differences depending on the source type: 88.46% for search results, 1.93% for Reddit. The discrepancy with our "one page out of ten" is due to what each counts as "retrieved": Ahrefs observes the interface, while we count all the results returned by the API search.
Both measurements agree on the essentials: ChatGPT strictly filters what it finds, and how a page is found matters as much as its content. Furthermore, researchers from the University of Toronto measured in September 2025 that AI engines favor independent media over brand sites, much more so than Google.
What this changes for a business
There is no single optimization "for AIs", but at least three. For ChatGPT, you must be at the top of the search results it consults and be a source whose name it knows. For Claude, you need pages that overlap with what the domain says and address multiple facets of the question. For Gemini, you must be present in Google Search.
A well-structured page serves all three at once: the question in the title and the address, an answer right from the start, paragraphs developed with search vocabulary, content consistent with serious sources, and a fast-loading page. This is what every page in this section aims to apply.
What we do not know
Our measurements focus on the model APIs: GPT-5, Claude Sonnet 4.6 and Gemini 3.5 Flash with web search. Public interfaces and other model versions may behave differently. Our reports describe associations, not the internal workings of the models.
Sources
- Ahrefs, Louise Linehan — "Why ChatGPT cites pages", .
- Anthropic — documentation for Claude's web search tool, .
- Google, Elizabeth Reid — AI Mode update and "query fan-out", .
- Chen, Wang, Chen, Koudas (University of Toronto) — "Generative Engine Optimization: How to Dominate AI Search", .
- FaireDuBruit measurements, from September 4 to 16, 2026 — see "How we measured".