ChangelogPricing

How LLMs Choose Sources to Cite: What Each Engine Documents

LLMs choose sources in five steps: decide to search, rewrite the query, retrieve, rerank, and cite. See what OpenAI, Google, and Perplexity document about each.

LLMs choose sources to cite in five steps. The model decides whether to search at all, rewrites your question into one or more search queries, retrieves candidate pages from a search index, narrows those pages down to the passages it will read, and attaches a citation to each sentence a passage supports. Your page has to survive every step, and much of the choosing happens before the model writes a word.

That much is on the record. OpenAI, Google, Perplexity, and Anthropic each describe parts of this pipeline in their own help pages and developer documentation. None of them publishes how it weighs one eligible page against another.

This guide walks through the five steps, quotes what each company says about its own system, and marks the point where the documentation stops. It is the mechanism behind an AI citation. The work that follows from it is in the guide to improving AI search visibility.

Key Takeaways

  • An LLM cites from the pages it retrieved for that one answer, so source selection is a search problem first. A page that is not retrieved cannot be cited.
  • One question becomes several searches. OpenAI says ChatGPT search "typically rewrites your query into one or more targeted queries", and Google calls the same step query fan-out.
  • The final ranking happens on passages, not pages. Perplexity says it scores results "at both the document and sub-document levels".
  • Being read is not being cited. OpenAI's API returns the full list of pages a model consulted and says it is often longer than the list of citations.
  • No engine publishes how it weighs one eligible page against another. Anything more specific comes from outside studies, and those measure outcomes, not rules.

How LLMs Choose Sources: The Five Steps

An LLM with web search runs the same sequence every time it cites something. Each step removes candidates, and a page that drops out at one step never reaches the next.

  1. Search, or answer from memory. The model decides whether the question needs the web. If it does not search, there is no page to cite.
  2. Query rewriting. The model turns the question into one or more search queries. Google calls this query fan-out.
  3. Retrieval. Those queries run against a search index and return candidate pages.
  4. Reranking. The candidates are scored again and cut down to the passages the model will read.
  5. Grounding and citation. The model writes from those passages and links each supported sentence to its source.

Steps 1, 2, and 5 are the best documented, because the engines' APIs return the queries and the citations as data. Steps 3 and 4 are where most of the choosing happens, and outside Perplexity they are thinly documented.

One caution on the sources. Help pages describe the consumer products. Developer documentation describes the API tools, which may not behave exactly like the app your buyer has open. We say which is which as we go.

If you would rather have someone trace this for your own pages, that is what our AI search audit call is for.

A citation needs a retrieved page, so the first cut is whether the model searches at all. Every company documents that decision as the model's own.

  • ChatGPT. OpenAI's help page on ChatGPT search says: "ChatGPT may search the web automatically when your question would benefit from current information."
  • Gemini. Google's Gemini API documentation lists it as a step: "The model analyzes the prompt and determines if a Google Search can improve the answer."
  • Google Search. AI Overviews are "only shown when our systems determine that it is additive to classic Search", according to Google's AI features documentation.
  • Claude. Anthropic's web search tool documentation says: "Claude determines when to search based on the prompt."

Anthropic's page is the most specific about the trigger. Claude searches when a request depends on information that is "current, changing, or outside its training data".

Its examples include "Current prices, rates, scores, or statistics" and "Information about specific organizations, people, or products that might have changed". It answers without searching for "Established facts, math, science fundamentals, or coding concepts".

That split matters for your brand. Take two questions a buyer might ask about a CRM.

"What is a sales pipeline?" is stable knowledge, and a model can answer it with no source in sight. "Which CRM is best for a 20-person sales team this year?" is about products that change, which is the kind of request Anthropic says Claude searches for. Only the second one has citations to win.

Step 2: One Question Becomes Several Searches

The model does not search for what your buyer typed. It writes its own queries, often more than one, and your page competes for those.

OpenAI describes this plainly. When ChatGPT search works with a partner search provider, it "typically rewrites your query into one or more targeted queries that it sends those providers".

OpenAI gives its own example. A researcher asks about the latest on drugs that target CCR8 for cancer, and ChatGPT might first search for "CCR8 immunotherapy drug development 2025".

Then it looks at what came back. OpenAI says that after reviewing the initial results, ChatGPT search "may send additional, more specific queries to other search providers", such as "CHS-114 conference 2025".

Two things stand out. The rewritten query is a string of keywords, not a sentence. And the second query names something the user never typed, because the model found it in the first results and searched again.

The rewrite can also be personal. OpenAI says that with memory enabled, "ChatGPT may use relevant saved memories when rewriting a search query", and that it may share a general location with its search providers.

The other engines document the same step:

  • Google Search. Google calls it query fan-out. Its wording and its worked example are in our Google AI Overview SEO guide, and the AI Mode version is in the Google AI Mode SEO guide.
  • Gemini API. "If needed, the model automatically generates one or multiple search queries and executes them."
  • Claude API. "The API runs the searches and provides Claude with the results. This process can repeat multiple times throughout a single request." Anthropic adds that simple factual questions typically take 1 to 3 searches, and comparative research can take 10 or more.
  • OpenAI API. With a reasoning model, the model "can perform web searches as part of its chain of thought, analyze results, and decide whether to keep searching", according to OpenAI's web search guide.

You can read the rewritten queries

The consumer apps do not always show the queries. The APIs return them.

OpenAI's API records each search action, usually with the queries that were searched. Gemini's response carries a google_search_call step listing the queries the model executed. Anthropic's example response shows the query inside a server_tool_use block.

So run your buyer's question through an API with search switched on and read the queries. That list is what your pages need to be retrievable for.

Treat it as a sample. OpenAI says its Chat Completions search models are "the fine-tuned models and tool used by Search in ChatGPT", but none of the pages cited here promises that an API and its app issue the same queries.

Step 3: Retrieval Decides Who Is a Candidate

The rewritten queries run against a search index, and only pages in that index can come back. Which index depends on the engine, and this is where the documentation differs most.

Google retrieves from its own Search index. Its guide to optimizing for generative AI features describes grounding as "relying on our core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index". To be eligible, a page "must be indexed and eligible to be shown in Google Search with a snippet".

OpenAI uses its own crawler and partners. Its help page says ChatGPT search "sometimes partners with other search providers", and the provider privacy policies it links are Microsoft's and Shopify's.

Its crawler documentation says OAI-SearchBot "is used to surface websites in search results in ChatGPT's search features". Sites that opt out of it "will not be shown in ChatGPT search answers, though can still appear as navigational links". Neither page says which provider handles which query.

Perplexity built its own index. In an engineering article on its search infrastructure, it says: "Our search index tracks over 200 billion unique URLs".

It retrieves two ways at once: "We do not force a choice between lexical and semantic retrieval. Rather, we query the search index via both modalities and merge the results into a hybrid candidate set."

Lexical retrieval matches the words in the query. Semantic retrieval matches the meaning. A page can enter Perplexity's candidate set through either one.

Perplexity is also open that freshness is a trade-off. With a fixed indexing budget, it writes, "refresh operations for existing pages must compete with indexing operations for new unvisited pages". It says machine learning sets the priority, and it does not say what earns a page a faster refresh.

Anthropic documents the tool and not the index. Its page lists what a search result contains, including a page_age field for "When the site was last updated". It does not name the index those results come from.

Access is rarely what keeps you out

The gate in front of retrieval is crawler access, and the AI crawlers guide covers it bot by bot. Our own data says it is not the usual problem.

Of the 2,207 Y Combinator B2B sites we could reach on September 11, 2026, 164 blocked at least one training crawler and none of the three answer-engine crawlers: OAI-SearchBot, Claude-SearchBot, and PerplexityBot. No site did the reverse. The full numbers are in the AI crawler study.

So for a typical startup, robots.txt is not the reason its pages are missing from answers. It allows the fetch, and the page loses later in the pipeline.

Step 4: Reranking Turns Pages Into Passages

Retrieval returns more than a model can read, so a second stage scores the candidates again and keeps the best. What survives is usually a passage, not a page. Perplexity documents this stage in the most detail.

Its pipeline uses "multiple stages of progressively advanced ranking". The early stages "rely on lexical and embedding-based scorers optimized for speed".

Then the expensive models come in: "As the candidate set is gradually winnowed down, we then use more powerful cross-encoder reranker models to perform the final sculpting of the result set."

A cross-encoder reads the query and one candidate text together and returns a single relevance score for the pair. The Sentence Transformers documentation describes the standard pattern: a fast retriever pulls the top candidates, then a cross-encoder re-ranks them "by computing the score for every (query, hit) combination".

In that last stage, what gets scored is how well one piece of text answers one query. And the piece of text is small.

Perplexity says: "we retrieve and score results at both the document and sub-document levels". Its pipeline "is designed to surface the most atomic units possible to the model".

Its stated reason is the model's limits: "AI model brittleness and limited context windows require search results to be presented and ranked at the most granular level possible."

Here is what that means for one of your pages. Say your comparison page runs to 2,000 words, and one paragraph in the middle answers whether the product integrates with HubSpot.

If the model's rewritten query is about HubSpot integration, that paragraph is your candidate on Perplexity's description. It is scored against the query on its own, without the 1,900 words around it.

The other engines document less, and what they document points the same way:

  • Claude API. Newer versions of the tool change how results reach the model: "Claude instead writes and runs code that filters the results first, so only relevant content reaches the context window."
  • OpenAI API. After a search, reasoning models can take two more actions: open_page, and find_in_page, "which represents searching within a page".
  • Google Search. "While responses are being generated, our advanced models identify more supporting web pages".

What no document cited here lists is the set of features these rankers score. Perplexity says its pipelines "must incorporate lexical and semantic signals", and that its ranking models learn from its own answers and "live human feedback". That is as specific as the public record gets.

Writing passages that hold up on their own is a discipline with its own guide: AI content optimization.

Step 5: Grounding Attaches a Citation to a Sentence

Grounding means the model writes from the retrieved passages instead of from memory. A citation is the record of which passage supported which sentence. The APIs return that record as data.

  • Gemini API. "Each url_citation annotation links a text segment (defined by start_index and end_index) to a source URL."
  • OpenAI API. The url_citation annotation "will contain the URL, title and location of the cited source".
  • Claude API. "Citations are always enabled for web search", and each one carries a cited_text field holding "Up to 150 characters of the cited content".

So a citation is narrower than "this page was a source". It ties one stretch of the answer to one page, and in Anthropic's format to one snippet of at most 150 characters.

Being read is not being cited

OpenAI's API makes the gap between step 4 and step 5 visible. Next to the inline citations it offers a sources field.

OpenAI describes it this way: "Unlike inline citations, which show only the most relevant references, sources returns the complete list of URLs the model consulted when forming its response. The number of sources is often greater than the number of citations."

Your page can be retrieved, opened, and read, and still not be linked. OpenAI does not document what separates the two lists beyond "most relevant".

A citation is not a guarantee

The engines say so themselves. OpenAI's help page: "Search results and citations can be incomplete, outdated, or incorrect."

Independent measurement agrees. In Evaluating Verifiability in Generative Search Engines, Nelson Liu, Tianyi Zhang, and Percy Liang audited Bing Chat, NeevaAI, perplexity.ai, and YouChat in 2023.

They found that "a mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentence". The engines have changed since, so read that as proof the failure exists, not as today's rate.

For you, it means an engine can cite your page for something the page does not say. Our glossary entry covers what to do when a citation is wrong.

What Each Engine Documents, Side by Side

This table is the public record as we read it for this guide. "Not documented" means the pages cited above do not describe it, not that the step is missing.

StepOpenAIGooglePerplexityAnthropic
Search or notSearches "when your question would benefit from current information"AI Overviews only when "additive to classic Search". Gemini decides per prompt"uses advanced AI to search the internet in real-time""Claude determines when to search"
Query rewriting"one or more targeted queries"Query fan-outNot documentedSearches "can repeat multiple times" in one request
IndexOwn crawler plus partner search providersGoogle Search indexOwn index, "over 200 billion unique URLs"Not documented
RerankingNot documented. Reasoning models can open a page and search inside itNot documented beyond "identify more supporting web pages"Multi-stage, cross-encoder rerankers, scored at sub-document levelOptional filtering by code Claude writes
CitationInline citations, plus a longer sources listLinks to pages "that support the information in the response""numbered citations linking to the original sources"Always on, with up to 150 characters of cited text

The Perplexity rows for search and citation come from its help page on how Perplexity works.

Where the Documentation Stops

Everything above describes the shape of the pipeline. None of it tells you why one page beat another, and it helps to be exact about what is missing.

  • The weights. No document cited here lists the signals a reranker scores, or how much each one counts.
  • The queries for your question. The engines document that rewriting happens. Which queries your buyer's question turns into is something you observe, one run at a time.
  • The line between read and cited. OpenAI confirms the two lists differ. It does not say what decides it.
  • Authority, freshness, and agreement across sources. Outside studies associate these with being cited, and our AI search ranking entry covers them. They are observed patterns. No engine document cited here states them as rules.

Two more things make the outcome hard to predict from the documentation alone.

The answer moves between runs. In RankZero's own tracking of 96,720 AI answers between June and September 2026, we looked at the 368 prompts that ran 10 or more times. The most-named brand for a prompt appeared in 50.9% of that prompt's runs, so it was missing from about half the answers to the same question.

The search can be personal. OpenAI documents that saved memories and general location can shape the rewritten query. Two buyers typing the same question can trigger different searches.

Outside research fills part of the gap by experiment. In GEO: Generative Engine Optimization, accepted at KDD 2024, Pranjal Aggarwal and co-authors tested content changes on a benchmark of queries.

They report that their methods "can boost visibility by up to 40% in generative engine responses", with results that vary by domain. An experiment like that tells you what moved the outcome on their benchmark. It does not tell you the engine's rules.

What the Pipeline Changes About Your Work

Each step has one check, and the order matters because each step gates the next.

  1. Find the questions that trigger a search. Ask your buyer questions with search on and note which answers carry citations. Only those have a citation to win.
  2. Collect the rewritten queries. Read them from the API responses and check whether you have a page for each one.
  3. Confirm you are in the index each engine retrieves from. That is Google's index for AI Overviews and AI Mode. For ChatGPT, allow OAI-SearchBot and check your pages in Bing, the search engine of Microsoft, one of the providers OpenAI links. For Perplexity, allow PerplexityBot.
  4. Make the passage the unit. Each section should answer one of those queries on its own, without the paragraph above it.
  5. Track cited URLs, not rankings. The AI search tracking guide covers the setup.

Most teams start at step 4, because writing is the part they control. If your pages are not candidates for the rewritten queries, better passages change nothing.

On an audit call we run your buyer questions across ChatGPT, Perplexity, Gemini, and Google AI Overviews and show you which sources get cited instead of you. Get your AI search audit to see where your pages stand.

FAQ

How do LLMs choose which sources to cite? In five steps. The model decides whether to search, rewrites the question into one or more queries, retrieves candidate pages from a search index, reranks them down to passages, and links each sentence it writes to the passage that supports it. A page has to pass all five to be cited.

Do LLMs cite sources from their training data? Not as links. A citation points to a page retrieved for that answer. Anthropic documents that Claude "answers directly without searching when the request draws on stable knowledge", and OpenAI ties citations to answers that use web search. A model can still name your brand from memory, which is a mention and not a citation.

What is query fan-out? It is Google's name for step 2: the model issues several related searches for one question and combines the results. OpenAI documents the same behavior for ChatGPT search as rewriting your query into "one or more targeted queries". Your page can be retrieved for one of those queries without ranking for the question your buyer typed.

Is every page an LLM reads cited in the answer? No. OpenAI's API returns a sources list of every URL the model consulted and says it is often longer than the list of inline citations, which show "only the most relevant references".

Does ChatGPT search use Bing or Google? OpenAI says ChatGPT search "sometimes partners with other search providers" and links the privacy policies of Microsoft and Shopify for them. It also runs its own crawler, OAI-SearchBot. It does not publish which provider answers which query. The practical steps are in our ChatGPT SEO guide.

Can you see which searches an AI ran for a question? Often, through the API. OpenAI, Google, and Anthropic each return search queries in the API response when a model searches, though OpenAI notes its response does not always include them. The consumer apps show less, so the API is the place to look.

Where to Start

Pick five questions your buyers ask before they shortlist a vendor. Run each one with search on and write down three things: whether the model searched, which queries it ran, and which URLs it cited.

That is an afternoon of work, and it shows you how LLMs choose sources in your category instead of in general. You will see which step your pages fail at, and that step is where the work starts.

Or hand it over. Finding the step where your pages drop out, and fixing it, is the work our done-for-you AI SEO service takes on.