AI answers are shaped by the sources a model can retrieve, trust, and cite for the specific prompt. In Parse's 2026 citation slice, the winning source graph was broad, not dominated by one domain: the top 50 domains represented only 15.7% of citation references. The operating lesson is to manage source coverage by platform, category, and influenceability.
Parse measured citation-source references from January 2 through June 5, 2026, across ChatGPT, Google AI Overviews, Google AI Mode, and ChatGPT Search. The slice covers 21.2 million citation references, 16,744 prompts, 4.4 million cited source URLs, and 851,051 citation domains. Parse tracks AI visibility across ChatGPT, Google AI Overviews, and Perplexity, analyzing 3.85 million AI responses, 47.25 million citation observations, and 577K+ brands in the public index.
- Source graphs shape AI answers more than any single "AI ranking" does. In Parse's slice, the top 10 domains represented 10.4% of citation references and the top 50 represented 15.7%.
- Google AI Overviews was more concentrated than ChatGPT: its top 50 domains produced 23.0% of citation references, versus 12.8% for ChatGPT.
- Platform source mix diverged sharply. Google AI Overviews drew 14.4% of references from social sources, while ChatGPT drew 3.1%.
- Only 24 of the top 50 domains overlapped between ChatGPT and Google AI Overviews, so one source plan will miss platform-specific gaps.
- The practical workflow is source mapping, influenceability scoring, platform split, and quarterly refresh, not a generic content sprint.
What sources shape AI answers?
The sources that shape AI answers are the pages, domains, profiles, communities, videos, reviews, and reference records a model uses when it turns a prompt into an answer. Some are cited directly. Others shape the model's entity understanding before retrieval. For operators, the direct citation layer is the most measurable layer because it shows which pages the AI surfaced as support.
Google describes AI Overviews and AI Mode as link-backed Search experiences that may issue related searches across subtopics and data sources before generating a response (Google Search Central). OpenAI says ChatGPT search can rewrite a user prompt into search queries, use third-party search providers, and show source links when search is used (OpenAI Help Center). Those mechanics explain why a brand's own site is only one input. The answer is assembled from the broader source graph around the category.
What did Parse measure?
We used Parse production citation metrics from January 2 through June 5, 2026. The data source was PromptCitationSourceDailyMetrics, joined to AI platform and citation-domain metadata, under read-only access. The sample contained 20.5 million metric rows, 21.2 million citation references, 16,744 distinct prompts, 4.4 million unique cited source URLs, and 851,051 citation domains.
The platform split was large enough to support direct comparison. ChatGPT accounted for 13.48 million citation references across 16,164 prompts. Google AI Overviews accounted for 7.34 million citation references across 16,036 prompts. Google AI Mode and ChatGPT Search were present but smaller in this slice, so the main platform contrast below focuses on ChatGPT and Google AI Overviews. The caveat is important: this is a citation-reference study, not a click, conversion, or ranking-causality study.
Which source types matter most?
The broadest category in the current classification is still unknown, at 39.5% of citation references, so no honest source study should pretend every page can be neatly classified. Among classified sources, the largest groups were general web sources at 28.4%, social at 7.1%, SaaS and software sources at 6.8%, news at 3.5%, editorial at 3.4%, ecommerce at 2.9%, directory at 1.3%, and encyclopedia at 1.2%.
The ranked domains make the spread easier to see. Reddit, YouTube, Wikipedia, Forbes, Medium, Facebook, LinkedIn, NerdWallet, TechRadar, Amazon, Gartner, Alibaba, Apple App Store pages, Healthline, Quora, GitHub, Tom's Guide, Zapier, and Consumer Reports all appeared near the top. That is why "publish better blog posts" is too narrow. AI answers pull from community, video, reference, commerce, review, software, publisher, and owned-site surfaces, depending on the query.
If you want to see which sources shape AI answers about your brand, run a free brand check — it takes a minute.
How does ChatGPT differ from Google AI Overviews?
ChatGPT and Google AI Overviews do not use the same source graph. In Parse's slice, ChatGPT's top domains were Reddit, Wikipedia, Forbes, NerdWallet, TechRadar, LinkedIn, Alibaba, Medium, Apple App Store pages, and Healthline. Google AI Overviews led with YouTube, Reddit, Facebook, Amazon, Medium, LinkedIn, Quora, Forbes, Instagram, and NerdWallet.
The source-type split is the operational difference. Google AI Overviews drew 14.4% of citation references from social sources, led by YouTube and Reddit. ChatGPT drew 3.1% from social sources and was more weighted toward unknown, general, SaaS, editorial, news, and encyclopedia categories. This lines up with outside research: Conductor found AI engines have persistent source preferences by intent, while BrightEdge reported meaningfully different source ecosystems across AI engines (Conductor, BrightEdge). Treat platform divergence as a planning input, not a reporting footnote.
Why top-domain lists are not enough
Top-domain lists are useful for orientation, but they can mislead teams into chasing one universal winner. In Parse's full slice, the top 10 domains accounted for 10.4% of citation references and the top 50 accounted for 15.7%. Google AI Overviews was more concentrated, with its top 50 at 23.0%; ChatGPT was flatter, with its top 50 at 12.8%.
Ahrefs' June 2026 Google AI Overviews ranking similarly shows a concentrated top set inside Google's surface, with YouTube, Reddit, Facebook, Wikipedia, Amazon, Quora, TikTok, Walmart, publishers, and commerce sites near the top (Ahrefs). That helps with Google AI Overviews, but it does not tell you whether ChatGPT, Perplexity, Claude, or Gemini will cite the same source. Your source plan should start with the top domains, then move quickly into category-specific and platform-specific gaps.
Which sources can a brand actually influence?
Do not divide the graph into "owned" and "not owned" only. That is too blunt for execution. Use four buckets: controlled, influenceable, relationship-driven, and mostly off-limits. Controlled sources include your site, docs, product pages, help center, schema, feeds, and profiles you own. Influenceable sources include review platforms, directories, partner pages, marketplace pages, community discussions, comparison pages, and listicles. Relationship-driven sources include journalists, analysts, podcasters, creators, and industry publications. Mostly off-limits sources include Wikipedia, government domains, academic repositories, and independent community threads.
OpenAI's crawler documentation makes the controlled layer concrete: OAI-SearchBot determines whether a site can appear in ChatGPT search answers, and GPTBot is a separate training crawler (OpenAI). Google says indexed pages with snippets are eligible for supporting links in AI features, but eligibility is not a guarantee (Google Search Central). Technical access is the floor. Source authority is the work.
How should teams audit their source graph?
Start with prompts, not pages. Pick 50 to 150 revenue-relevant prompts, run them by platform, and export the cited domains for every answer. Then label each cited source by source type, influenceability, competitor presence, and buyer-stage relevance. The output should be a ranked source backlog, not a generic content calendar.
Use this table as the operating model. It keeps the team from treating every source hit as the same kind of task.
| Source layer | What to inspect | Practical next action |
|---|---|---|
| Controlled | Owned pages, docs, help center, product pages, feeds, schema | Make pages crawlable, factual, current, extractable, and internally linked |
| Influenceable | Review sites, directories, marketplaces, community threads, comparison pages | Improve profiles, earn inclusion, refresh claims, and seed useful proof |
| Relationship-driven | Analysts, journalists, creators, podcasts, trade publications | Build evidence-backed pitches and publish data worth citing |
| Mostly off-limits | Wikipedia, government, academic, standards, independent forums | Fix entity accuracy and cite-worthy upstream evidence, but do not treat as direct placement |
The interpretation step matters more than the export. A citation gap only becomes useful when it names the source, the competitor it favors, the platform where the gap appears, and the action owner.
What are the limits of source data?
Source data is not the whole answer. Citations show what the AI exposed as support, not every document retrieved, every training-data influence, or every ranking feature inside the model. A recent arXiv measurement study of Google AI Overviews found that almost 30% of cited domains did not appear in co-displayed first-page results, which means AI source selection is not identical to classic organic ranking (arXiv). That is exactly why source tracking is useful, but it is also why causality claims should stay modest.
There are also data-quality limits. Parse's source-type taxonomy classified 60.5% of references in this slice and left 39.5% as unknown. Some source classes are broad by design. "General" contains very different domains. The right use of this data is directional prioritization, followed by manual inspection of the actual pages your category cites. Pair this piece with which domains AI models cite most for broader domain context and AI citation gap analysis for the gap-closing workflow.
Frequently asked questions
What sources do AI answers use?
AI answers use a mix of retrieved web pages, cited sources, entity records, product data, reviews, communities, videos, publishers, and owned pages. The visible citation layer is the easiest to measure. In Parse's January-June 2026 slice, Reddit, YouTube, Wikipedia, Forbes, Medium, LinkedIn, Amazon, Gartner, Quora, GitHub, and review or commerce domains all appeared in the upper source set.
Do ChatGPT and Google AI Overviews use the same sources?
No. In Parse's slice, only 24 of the top 50 domains overlapped between ChatGPT and Google AI Overviews. ChatGPT's top set was flatter and included Wikipedia, Forbes, and software-related sources. Google AI Overviews was more concentrated and leaned heavily toward YouTube, Reddit, Facebook, Amazon, and other social or commerce surfaces.
Should brands chase the top AI citation domains?
Use top-domain lists for orientation, not as the whole plan. The top 50 domains represented only 15.7% of all citation references in Parse's slice. The better workflow is to find the domains cited for your category prompts, then score each one by influenceability, competitor presence, buyer-stage relevance, and platform.
How do I track source influence in AI answers?
Run a stable prompt set across ChatGPT, Google AI Overviews, Perplexity, and any platform your buyers use. Export cited domains, classify them by source type, and compare the sources attached to your brand against the sources attached to competitors. Track this monthly or quarterly so you can separate a durable source gap from normal model variation.