Skip to content
AI Visibility

Do AIs cite PDF documents?

Published on

Rarely, compared to web pages. For ChatGPT, PDFs represent 16% to 21% of the results returned by its searches, but they are cited 2 to 5 times less often than other pages, across all our corpora. Claude returns very few PDFs, around 3% of its results at most. We have found no published study that measures the citation of PDFs by AIs: these figures are our own. An important document published only as a PDF therefore has less chance of being cited than an equivalent web page.

For what type of question? It all depends on the question. If the user is looking for a specific document, a standard, a technical listing, or a report, the AI can look for it as a PDF. On an information or market question, it tends to cite web pages instead.

What we measured for ChatGPT

We counted, in the results returned by ChatGPT's searches, the addresses of PDF documents, and then the share of these documents actually cited. PDFs are numerous in what ChatGPT finds: 21% of the results across a dozen sectors in French, 20% across these sectors in English, 19% on software troubleshooting, and 17% on our questions about visibility in AIs.

Yet they are rarely cited: 5.7% of the returned PDFs compared to 12.1% of other pages for the sectors in French, 5.2% compared to 10.1% in English, 4.9% compared to 11.8% for troubleshooting, and 2.1% compared to 10.8% on our questions from September 15, 2026. A PDF is therefore cited 2 to 5 times less often than a web page found by the same AI. Our first corpora, on HVAC and artificial intelligence, already showed a 3 to 4 times gap.

Claude returns very few PDFs

For Claude, PDFs are almost absent from the results: between 0.5% and 3% of what its search returns, depending on the corpus. This is too little to measure a reliable citation rate. The difference with ChatGPT is therefore primarily due to what each search engine returns, even before the choice of sources to cite.

For Gemini, the API does not distinguish between found pages and cited pages and only provides the domain of the sources, which does not allow us to measure the share of PDFs. Here again, a single rule for "the AIs" is not possible: ChatGPT finds many PDFs and discards them, while Claude finds few.

Why a PDF is cited less

We observe several obstacles, without being able to say which one carries the most weight. A PDF often has neither an exploitable title tag nor a paragraph structure readable by a machine: during our own data collections, some PDFs were read as a single block of text of several tens of thousands of words, and others could not be read at all. Yet the passage cited by an AI is a developed, identifiable paragraph that uses the words of its search.

A PDF is also often long and general, such as an annual report, complete guide, or catalog, whereas the AI looks for a passage that answers a precise question. These explanations remain hypotheses: we measure the gap in citation, not its cause.

What Google and published studies say

Google indexes PDF documents: the format is included in the list of file types it can index, updated on February 3, 2026. And its documentation on AI features indicates that an indexed page eligible for a snippet can serve as a source, with no additional requirements. Therefore, nothing excludes a PDF from Google's AI answers.

We have found no published study, nor any documentation from OpenAI, Anthropic, or Google, that specifically addresses the citation of PDFs by AIs. This is one of the subjects on which our measurements are, to our knowledge, the only ones available.

What this changes for a business

A white paper, a technical listing, or a study published only as a PDF is found, but rarely cited. For important content to be picked up by an AI, it is better to also publish it as a web page, broken down into pages or sections that each answer a specific question, with developed paragraphs, and keep the PDF for download.

The web page carries the question in its title and address, answers from the very first lines, and links to the full document. The PDF remains useful for the reader who wants to read everything; the web page is what the AI can cite.

What we do not know

A document is counted as a PDF when its address ends with ".pdf"; a PDF served under another address is not counted. For ChatGPT, cited documents are identified in the links of the response. Our measurements focus on the models' API, and do not distinguish between types of PDFs.

Sources

  1. Google Search Central — file types indexable by Google, .
  2. Google Search Central — "AI features and your website", .
  3. FaireDuBruit measurements, from September 4 to 16, 2026 — see "How we measured".

Related questions

All questions about visibility in AI · Our offer