On a 15,000-prompt analysis, ChatGPT retrieved 548,534 pages and cited about 15% of them. The other 85% were found and discarded. The reason most brands miss is not crawlability; it is the rerank stage between retrieval and answer assembly. Pages survive that stage when they match the synthetic fan-out queries AI models generate, lead with extractable claims, and pass an authority filter that under-rewards the largest brands.
Retrieval is not the same as citation
Most AI visibility advice still treats "get crawled, get cited" as a continuous process. The data shows it is not. ALM Corp's April 2026 analysis traced 15,000 prompts through ChatGPT's web search and found that the system pulled 548,534 pages into its working set, an average of 36.5 pages per query, but cited only 15% of them in final answers (ALM Corp). That is a quiet but decisive bottleneck. Crawl access lets you compete; it does not win the slot. The work that earns the citation happens after retrieval, in a stage most operators have never measured because their tooling stops at "indexed" or "ranking in Bing." Parse tracks AI visibility across ChatGPT, Google AI Overviews, and Perplexity, and the citation gap shows up consistently across categories.
The five stages between a prompt and an answer
AI models route every prompt through a short, opinionated pipeline. The same five stages apply, with implementation differences, to ChatGPT's web search, Google AI Mode, and Perplexity.
Stage 1 is query parsing: the model rewrites the user's natural-language prompt into a structured intent, often with entities extracted. Stage 2 is fan-out, where the model issues multiple synthetic sub-queries instead of the original phrasing. Stage 3 is hybrid retrieval against an external index (Bing for ChatGPT, Google's own index for AI Mode and AI Overviews, a proprietary live crawl for Perplexity) plus parametric recall from training weights. Stage 4 is reranking, where a separate model scores the retrieved set and discards most of it. Stage 5 is constrained generation, where the answer is assembled from the surviving passages with citation tags attached. Most AI visibility wins or losses happen in stages 2 and 4.
Fan-out turns one prompt into ten searches
The biggest mental shift for SEO operators: the AI model does not search for the prompt you typed. It generates a small fleet of synthetic queries and searches for those. Seer Interactive's Gemini 3 analysis put the average at 10.7 fan-out queries per prompt, with 59% of prompts triggering 5–11 searches (Seer Interactive). ALM Corp's ChatGPT analysis found 89.6% of prompts trigger two or more fan-out queries, and the original 15,000 prompts expanded to 43,233 actual searches (ALM Corp). Google's patent US11663201B2 describes the mechanism: a generative model produces query variants, each variant runs independently, and the result lists are merged with reciprocal rank fusion (Google Patent US11663201B2). The strategic implication: your page has to be findable for queries your buyers will never type.
95% of fan-out queries have no search volume
The corollary is harsher than it sounds. Seer Interactive's Gemini 3 dataset showed 95% of fan-out queries had zero recorded monthly search volume in conventional keyword tools, and Ahrefs reports similar numbers across its own research (Seer Interactive; Ahrefs). That means traditional keyword research, the workflow most SEO teams have run for fifteen years, is structurally blind to the queries that actually drive AI citation. Ahrefs Keyword Explorer, Semrush, Search Console all rely on observed human searches as the input signal. Fan-out queries are generated by a model on the fly, in-context, and look more like fragments of the user's situation than searchable phrases. "best CRM for a five-person legal practice that already uses Clio" is a typical shape. Optimizing for the head term "best CRM" wins fewer fan-out matches than covering the situation-rich variants competitors have not yet thought to write for.
The rerank stage decides what gets cited
After retrieval, the model holds dozens of candidate pages. The rerank stage scores them for relevance, authority, freshness, and extractability, then keeps a handful. ALM Corp's data lets us read what the reranker actually weighted in ChatGPT during the study window. Pages ranking at Google position 1 for at least one fan-out query had a 43.2% citation rate, 3.5× higher than pages outside the top 20 (ALM Corp). Title-word overlap with the query was a strong predictor: pages with 50%+ title-query overlap were cited 20.1% of the time versus 9.3% for pages below 10% overlap. Freshness mattered: pages updated in the last three months were roughly twice as likely to be cited. Domain authority mattered, but inversely from the common intuition: 74% of citations went to sites with DA below 80, and DA 20–40 outscored DA 80–100 by total share. The reranker prefers specific, fresh, well-titled, mid-authority pages over generic top-of-funnel content from the strongest domains.
If you want to see how AI engines describe your own brand, run a free brand check — it takes a minute.
Citation rates by query type
ChatGPT cites at different rates depending on what the user is asking for. ALM Corp's segmentation showed product-discovery queries clear the rerank stage at 18.3% of retrieved pages, the highest cohort in the study, while validation queries (the user is checking a fact they already half-believe) clear at 11.3% (ALM Corp). The asymmetry is operationally useful. If you sell SaaS, the page that earns the citation in "best X for Y use case" is structurally different from the page that earns it in "is X HIPAA compliant." Product-discovery content earns citations through comparison tables, named alternatives, and explicit use-case framing. Validation content earns citations through primary sources, named dates, and direct yes/no leads. The same article rarely wins both.
Why ChatGPT cites pages Google would not rank first
There is a real gap between the Bing rank that gets a page into ChatGPT's retrieval set and the page that survives reranking. Seer Interactive's December 2024 analysis found 87% of SearchGPT citations matched Bing's top results, with most in the top 10 (Seer Interactive). But ALM Corp's parallel run against Google found that only 6.82% of ChatGPT citations appeared in Google's top 10, and 83.39% did not appear in Google's organic results at all. Two effects compound. First, Bing's index is the real retrieval surface for ChatGPT search, so Bing ranking, not Google ranking, governs entry. We have a dedicated breakdown in Bing rankings drive ChatGPT visibility. Second, the reranker has its own opinion. It promotes pages that match the structural pattern of an extractable answer, which often means a smaller, more specific page outranks a flagship guide that owns the Google SERP.
The 32.9% of citations you cannot see without fan-out tracking
There is a category of citations that traditional ranking tools cannot explain at all. ALM Corp found that 32.9% of cited pages appeared exclusively in fan-out query results, never in the search for the original prompt (ALM Corp). The page was retrieved because the AI rewrote the prompt into a sub-query that ranking tools would not normally show. For brand monitoring, this is the gap that explains "we rank for the query but never get cited" and "we get cited but our keyword tool says we do not rank for that." Both are accurate descriptions of the same underlying mechanic: retrieval runs against synthetic queries, not the user's input. Closing this gap requires running prompts directly in the AI surfaces and logging which sources show up, not inferring from SERP data. For more on the prompt-set discipline, see how to build an AI visibility prompt set.
Where citations sit inside a page
The reranker is also picky about where on a page the answer lives. 44.2% of citations were drawn from the first 30% of page content, 31.1% from the middle section, and only 24.7% from the final third (ALM Corp). That maps directly to chunk-based retrieval: the model chunks each page into 200–800 token passages, scores each chunk independently, and pulls the highest-scoring chunk into the answer. A flagship guide that buries the lead under five paragraphs of "what is X" is competing at a structural disadvantage with a smaller post that opens with the answer. For the writing pattern that earns this slot, see the answer capsule technique. The single-sentence operational rule: every H2 should have a self-contained claim in the first sentence so that the chunk built around that heading is independently citable.
What this means for measurement
If your AI visibility dashboard only measures whether your brand appears in answers, it is missing the half of the funnel where the work happens. The reachable metrics for the retrieval-to-citation gap are: how many of your priority prompts retrieve any page on your domain, what fan-out queries trigger your retrievals, what passage on the page gets pulled, and which competing pages survived rerank when yours did not. For most brands, the cheapest first move is to define a 25–50 prompt set, run it weekly in ChatGPT, AI Overviews, and Perplexity, and log every citation with its parent domain and page URL. The next move is to compare those citations against your category map and the gap shows up. The cluster guide which domains AI models cite most covers the source distribution at the category level; the per-page audit pattern lives in content structure for AI citation. For the underlying data on which source domains AI cites most, see Parse's report.
What changes if you actually do this
Three durable shifts. First, keyword tools become a coverage check, not a planning input. They tell you where the head-term traffic is, but the fan-out variants are the real targets. Second, content velocity matters more than asset size: a tighter network of 20 specific pages outperforms two generic flagships, because the surface area for fan-out matches is larger. Third, the page audit changes from "does this rank" to "does this survive rerank." A page that ranks but does not get cited usually fails one of three checks: the lead claim is buried below background, the title and H2s do not match how an AI rewrites the question, or the page leans on brand authority instead of being independently extractable. Fixing any one of those typically moves citation share within two to four weeks on Perplexity and one quarter on ChatGPT and AI Overviews. For the underlying selection mechanic at ChatGPT specifically, see how ChatGPT decides which brands to recommend.
FAQ
What is the difference between AI retrieval and AI citation?
Retrieval is the stage where an AI model pulls candidate pages from an index (Bing for ChatGPT, Google's index for AI Mode, a live crawl for Perplexity). Citation is what survives the reranker after retrieval. On ChatGPT, only about 15% of retrieved pages get cited. The work of AI visibility is increasingly about closing that gap, not about getting indexed.
What is query fan-out and why does it matter?
Query fan-out is the AI model rewriting a single user prompt into multiple synthetic sub-queries (typically 2–11) and running each one independently against its retrieval index. 95% of fan-out queries have zero monthly search volume in conventional keyword tools, which means traditional keyword research is structurally blind to most of what AI actually searches for. Optimizing for the user's question is no longer enough.
If my page ranks #1 on Google, will it get cited by ChatGPT?
Not directly. ChatGPT's web search runs on Bing's index, and only 6.82% of ChatGPT citations appear in Google's top 10 organic results. A Google #1 ranking helps because entity signals correlate, but the proximate driver for ChatGPT is Bing rank plus how well the page survives ChatGPT's reranker on title-query match, freshness, and extractable structure.
Why does domain authority not predict AI citations the way it predicts SEO rankings?
The data shows the opposite of the SEO intuition: 74% of ChatGPT citations go to sites with domain authority below 80, and the DA 20–40 cohort beats the DA 80–100 cohort by total citation share. The reranker rewards specific, extractable answers more than aggregate domain trust. A 50-employee SaaS with sharp comparison pages can out-cite an enterprise with a generic content hub on the same topic.
How fast can content changes show up in AI citations?
It depends on the surface. Perplexity runs a live crawl on every query, so changes can appear in citations within days. ChatGPT's web search reflects new content as Bing indexes it, typically two to six weeks. Google AI Overviews trails the Google organic index by roughly the same lag. Parametric (training-data) recall is on a six-to-twelve-month cycle and is mostly out of reach for direct optimization in the short term.
:::