The archive

Which Websites Do AI Search Engines Cite in 2026

ChatGPT cites Reddit most; Perplexity favors YouTube. Five 2026 studies reveal what AI engines cite and where they disagree for your content strategy.

Written by the IT Master engine, edited by Rodd Azad, Editor-in-Chief14 min read3,005 words
Which Websites Do AI Search Engines Cite in 2026
Which Websites Do AI Search Engines Cite in 2026

There is no single list of websites AI search engines cite, because each engine favours a different slice of the web: Vercite's data shows ChatGPT cites Reddit more than any other domain, Perplexity and Google AI Mode lean on YouTube, and 8.5% of AI Overview citations point back to Google itself.

Geonimo's analysis of 2.1 million cited sources across ChatGPT, AI Mode and Perplexity (Nov 2025–Apr 2026) adds the qualification that matters most: in that pooled dataset, brand-owned corporate pages took 60.6% of citations, Reddit was the most-cited single domain by a factor of three, and 73.5% of all citations went to domains outside the top 100.

The engines disagree on their favourite sources. The long-tail figure is a finding from Geonimo's pooled three-engine dataset, not something every engine has been shown to share.

Which websites do AI search engines cite most?

No single site wins everywhere: ChatGPT cites Reddit most, Perplexity and Google AI Mode cite YouTube most, and AI Overviews lean on Google itself. That split comes from Vercite's capture of 5.31 million citations across five engines, in which 8.5% of AI Overview citations point back to Google, mostly its own Search.

Across 158,847 classified domains, the five engines shared only 9% of their top-100 lists (23 of 253 pooled domains).

Engine Site it cites hardest (per Vercite)
ChatGPT Reddit
Perplexity YouTube
Google AI Mode YouTube
Google AI Overviews Google's own properties, chiefly Search

Zoom out from single domains and a second pattern appears. Geonimo's analysis of 2.1 million sources cited by ChatGPT, AI Mode and Perplexity between November 2025 and April 2026 (567K unique domains, 1,280 projects tracked) found that brand-owned corporate pages made up 60.6% of all citations. Reddit was still the single most-cited domain, by a factor of roughly 3x over the runner-up.

The more useful number for anyone deciding where to publish is the tail. In the same Geonimo dataset, 73.5% of citations went to domains outside the top 100, so the answer engines are not simply recycling a fixed roster of big publishers. Note that 60.6% is a pooled share for a category of pages, while Vercite's 9% is overlap between per-engine top-100 lists; the two measure different things and say nothing about how often any one brand's site is cited.

Overlap between engines is thin by every measure in these studies. SurfacedBy's sample of 127,198 citations to 11,647 sites found seven in ten sources were cited by only one of five engines, and just 2.7% by all five, while Foglift's frozen Q3 2026 panel reports a mean pairwise Jaccard score of 0.094 across five engines. A domain that is a fixture in Perplexity's answers can be invisible to ChatGPT on the identical question.

For a generative search engine optimization plan, read the table as a per-engine bias rather than a ranking to chase: Reddit and YouTube presence matters for ChatGPT and Perplexity respectively, and corporate pages as a category took 60.6% of the citations in Geonimo's three-engine pool. These are descriptive counts, though, and none of the studies isolates why a particular page was picked over another.

Do ChatGPT, Perplexity and Google cite the same sources?

Mostly not: Vercite found only 9% overlap between the five engines' top-100 cited domains, 23 of 253 pooled, across 5.31 million citations. That figure comes from the Vercite citation landscape study, which classified 158,847 domains by source type and publisher across ChatGPT, Perplexity, Gemini, Google AI Overviews and Google AI Mode.

Two further panels reached the same conclusion with different samples and methods, which is what makes the finding hard to dismiss as one dataset's quirk.

Study Sample Overlap measure
Vercite 5.31M citations, 158,847 domains, 5 engines 9% of pooled top-100 domains shared
SurfacedBy 127,198 citations, 11,647 sites, ~16,400 answers 70% of sources cited by one engine only; 2.7% by all five
Foglift Q3 2026 375 responses, 75 prompts, 25 verticals, 5 engines Mean pairwise Jaccard 0.094

SurfacedBy logged its 127,198 citations between 29 March and 27 June 2026 from ChatGPT, Claude, Gemini, Perplexity and Google AI Mode, and seven in ten cited sites appeared in a single engine's output. Foglift's Q3 benchmark sent 75 identical buyer questions to five engines during 1–10 August 2026; three quarters of the combined top-25 domains were exclusive to one engine, and no domain made every engine's top 25.

The one dataset that looks like a contradiction

The Machine Relations Index reports that 71 of the 100 most-cited domains across its 10,661 answer runs were cited by all six engines it tracks. That reads as high agreement, but it is a different metric: whether a heavily cited domain is surfaced measurably by each engine, not whether it ranks in each engine's own top tier. Those two measures are not interchangeable, and neither is a pooled top-100 share or a per-prompt Jaccard score.

A domain cited by all six engines could still sit far down most of their lists, which would be consistent with what Vercite and Foglift describe. The supplied excerpts do not publish the per-engine ranks needed to confirm that, so treat it as a plausible reconciliation rather than a settled one.

The practical takeaway: an "AI visibility" number from a single engine is a weak guide to the other four. Any generative engine optimization work — or any tool claiming to show real citation data — needs per-engine measurement, because the engines are reading largely different webs.

Do domain age, authority or content depth predict AI citations?

Not on current evidence: none of the citation studies in our dossier isolates domain age, a third-party authority score or content depth as a predictor of AI citations. They record which domains were cited and how often, not why.

What the studies do and do not measure

Geonimo's analysis of 2.1 million cited sources found that 73.5% of citations pointed at domains outside its top 100, per the Geonimo study. That is a distribution figure within one pooled dataset, not an authority measurement: the study logged no age or authority scores, so it cannot say whether those long-tail domains scored high or low on any such metric.

Overlap studies are descriptive in the same way. Foglift found that across 375 identical buyer-intent tests, three quarters of the domains in the five engines' combined top-25 lists were exclusive to one engine, while the Machine Relations Index reports 71 of its top 100 domains cited by all six engines it tracks. Both quantify agreement between engines; neither tests what caused a domain to be selected.

Treat any vendor's claim of a proven authority-to-citation correlation as unverified until they show the controlled data behind it.

What is actually scored: source type

Source type is the variable the studies actually score. Geonimo classed 60.6% of all citations in its dataset as brand-owned corporate pages, with Reddit the single most-cited domain by roughly 3x. Vercite, working from 5.31 million citations, reports ChatGPT leaning hardest on Reddit while Perplexity and Google AI Mode cite YouTube most, with each engine's top-100 overlapping only 9%.

Content depth is not scored in any of these datasets. Advice to "write deeper" is reasonable general guidance for E-E-A-T, but it is not a measured citation factor, and we label it as such.

What a real test would need

Testing age or authority properly would take a controlled domain sample: matched topics, logged registration dates and authority scores, identical prompts across engines, and citation outcomes recorded over time. None of the dossier studies were designed that way, and the Foglift benchmark itself notes that provider models change between quarters, making even its own quarter-over-quarter shifts observational rather than causal.

IT Master scores drafts on first-party grounding, novelty and E-E-A-T, with the checks run by models from a different vendor than the one that wrote the draft. We make no claim that those scores predict citations; they are quality gates, and a draft that fails them is not charged and never publishes.

How to measure your own citations with a fixed 24-question basket

Freeze a set of brand-neutral questions, run them across engines on a schedule, and count each cited domain once per response, normalizing URLs to registered domains exactly as Foglift describes in its published benchmarks.

That is the design Foglift used for its frozen panel of 75 brand-neutral questions across 25 verticals, producing 375 responses from five engines between 2026-08-01 and 2026-08-10, per the Foglift Q3 2026 benchmark. A 24-question basket is our smaller adaptation of that method, sized so one person can run it by hand; 24 is a suggestion, not a validated sample size, and it is not the only workable approach.

The basket template

Treat this as a worksheet to fill in, not a result to quote. Nothing below has been run by IT Master against a live panel; it is the structure the published benchmarks share, with our own suggested stage split layered on top.

  • 24 brand-neutral questions, eight per decision stage: problem ("why does X keep failing"), comparison ("X vs Y for a small team"), vendor choice ("best X for Z under a budget"). The three-stage split is our suggestion; Foglift's documented design is 75 buyer questions across 25 verticals.
  • One run per engine with live retrieval enabled — a cached or non-browsing answer produces no citations worth logging.
  • Normalise every cited URL to its registered domain, then count a domain once per response even if the engine linked it four times, exactly as Foglift describes.
  • Log the date and engine on every row. Without the timestamp you cannot tell a July run from an August one later.

Foglift's own caveat applies to any basket, including this one: provider models and indexes changed between its Q2 and Q3 runs, so period-over-period differences are observational rather than controlled causal estimates. Freezing the questions removes one variable; it does not freeze the engines, and re-running the same prompt on the same day can return different sources.

What to compute from the log

Three numbers come straight out of the rows. Your share of citations per engine is your domain's count divided by all domain-responses for that engine. Cross-engine domain overlap is the Jaccard score: domains cited by both engines divided by all distinct domains cited by either — the calculation behind Foglift's reported mean pairwise 0.094 across its five engines. Co-occurrence with a named rival is a separate, purely descriptive set: responses citing both your domain and theirs, divided by responses citing either.

Source-type ranking answers "who outranks me": tag each domain as forum, vendor, publisher, marketplace or video, and count which types sit above yours.

This is also the honest test for "which generative engine optimization tool shows real data?" — a tool that exports the raw rows (date, engine, prompt, cited URL) can be audited against a basket like this; a dashboard that only shows a score cannot. If you want a starting point before you build the schedule, IT Master's free AI visibility check needs no card and no account.

A frozen basket separates a change in your citation share from a change in the questions you asked. It removes one source of variation; engine drift and run-to-run variance remain, so treat what you observe as a trend, not a controlled result.

What does this pattern mean if you want to be cited?

Publish brand-owned pages built on data only your business holds, because corporate pages already take 60.6% of AI citations in the largest pooled dataset and most citations land outside the top 100 domains. In Geonimo's 2.1 million-source sample from ChatGPT, Google AI Mode and Perplexity, 73.5% of citations went to domains outside the top 100.

Those are category-level counts, not a guarantee for any individual site, and none of these studies tests why a given page is chosen. Our reading of them: you do not need to out-rank Reddit; you need a page that answers a specific buyer question with a figure nobody else can quote.

Write for more than one retriever

Engines disagree on sources, so one page rarely earns citations everywhere. SurfacedBy logged 127,198 citations and found seven in ten sources were cited by a single engine, with only 2.7% cited by all five; Vercite's 5.31 million-citation dataset put top-100 overlap at 9%.

A sensible response for generative search engine optimization is to cover a question from several angles — a direct definition, a comparison, a worked example — so different retrievers each have a passage that fits. The studies do not measure whether this works; it is an inference from how little the engines overlap.

Our takeaway: first-party specifics are likelier to be cited than generic depth, and multiple angles are likelier to travel across engines than a single monolithic page. The tactical detail — page structure, citation-friendly phrasing, crawler access — lives in our post "Get cited by ChatGPT & Perplexity: 2026 GEO playbook", so this section stays with the strategy.

How IT Master applies this

IT Master researches each article against the customer's own data rather than a generic corpus, so the checkable figures in a draft come from that business rather than a shared pool. Every draft is then fact-checked by models from a different vendor than the one that wrote it, scored for novelty and E-E-A-T, and scanned for AI tells before it publishes to the customer's site.

A draft that fails those checks is not charged and never publishes. Pricing is a fixed £9–£69 per published article by quality tier, listed at itmaster.uk/pricing; you can create an account without a card and pay only for what goes live.

Which AI search engines exist, and how do they retrieve sources?

Seven engines recur across the citation studies: ChatGPT, Perplexity, Gemini, Google AI Overviews, Google AI Mode, Claude and Copilot, with the xSeek tracker adding Grok. If you want a working list of AI search engines for a visibility programme, those eight are the set the studies and trackers in our source list actually report on.

They do not fetch the web the same way

The engine names matter less than the retrieval backend behind each one. According to xSeek's engine descriptions, Perplexity is built around citations and attaches numbered source links drawn from live web retrieval on every answer, while Gemini is powered by Google Search and cites live results from Google's own index.

Copilot pairs Bing search with GPT models and shows inline numbered references, and ChatGPT uses browsing and retrieval that pull from a broad range of web sources.

So when someone asks "which search engines use AI", the more useful question is which search index each AI engine leans on. A Bing-backed assistant and a Google-backed one are reading different maps of the web before a model ever writes a sentence.

How fast citation activity moves per engine

xSeek's last-seen table, which tracks 15 domains across four models, shows how uneven that activity is. At capture, Reddit had been cited by Perplexity and Gemini that day, by ChatGPT one day earlier, and was marked "Not seen" for Copilot; Wikipedia showed a similar split, cited today by Perplexity and one day earlier by ChatGPT and Gemini.

That per-engine variance is what the overlap studies measure at scale. Machine Relations attributes it to each engine running its own retrieval pipeline, and SurfacedBy found only 2.7% of 11,647 cited websites appeared in all five engines it tested.

Bold takeaway: the studies show low cross-engine overlap and describe separate retrieval pipelines behind it, but they are descriptive counts and do not isolate why one page is cited over another. Treating "AI search" as one channel is still a category error; you are optimising for several indexes at once.

Questions people ask

Which websites do AI search engines cite most?

There is no single winner. Per [Vercite](https://vercite.io/research/citation-landscape), ChatGPT cites Reddit more than any other site, Perplexity and Google AI Mode lean on YouTube, and 8.5% of AI Overview citations point back to Google itself. Across 2.1M sources, Geonimo found Reddit was the most-cited single domain by a factor of 3x, yet brand-owned corporate pages made up 60.6% of all citations in that pooled dataset.

What AI search engines exist?

The engines measured in the studies this article draws on are ChatGPT, Perplexity, Gemini, Google AI Mode and Google AI Overviews; Vercite's citation landscape covers those five, while Geonimo's 2.1M-source study covers ChatGPT, AI Mode and Perplexity. Each pulls from a noticeably different slice of the web, so no one engine's citation pattern predicts another's.

Does ChatGPT cite the same websites as Perplexity and Google AI Mode?

No. Vercite's data shows engine-specific bias: ChatGPT skews to Reddit, Perplexity and AI Mode skew to YouTube, and AI Overviews route a measurable share of citations back to Google Search. Treat each engine as a separate target rather than assuming a page cited in one will surface in the others.

Do AI search engines only cite big, high-authority websites?

Not according to the pooled data. In Geonimo's three-engine dataset of 567K unique domains across 1,280 projects (Nov 2025–Apr 2026), 73.5% of citations went to domains outside its top 100. That is a descriptive share, not a test of authority: none of the studies here isolates domain age, authority scores or content depth, so whether those factors predict citations remains unestablished.

How can I see which AI bots are crawling my website?

Crawling is not the same as being cited, and the citation studies here do not measure it. Your server access logs are the most direct first-party record: filter them by the user-agent strings each AI vendor publishes for its crawlers, then compare that list against which of your pages actually appear as citations in ChatGPT, Perplexity or AI Mode.