We Asked AI 1,226 Questions. Here’s How it Decides Who to Cite.

AI Citation Study: We asked AI 1,226 Questions. Here's How it Decides Who to Cite.

Table of Contents

Generative engine optimisation (GEO) is still taking shape, so it’s no surprise that many marketers aren’t sure how AI search actually works. In client meetings, one question comes up again and again: when ChatGPT, Gemini, or Google AI Overviews answer a question, how do they decide which websites to mention?

To find out, we ran a controlled study of 613 B2B search queries across six sectors, each tested in the US and the UK. We queried ChatGPT (with web search), Google Gemini, Google Search (including AI Overviews), and Bing, then traced more than 11,000 citations back to the exact passages that earned them.

Overall Findings

ChatGPT, Gemini, and Google AI Overviews don’t simply repeat the pages that rank highest on Google, nor do they cite the same sources as one another. Each platform draws from a different mix of company-owned content, review platforms, technology publications, and other sources.

The findings suggest that AI citation is related to traditional SEO but distinct enough to require a dedicated GEO approach. For marketers, that means:

  • Ranking well on Google can help (especially with AI Overviews), but it won’t guarantee visibility in ChatGPT or Gemini.
  • Those models favour content that is clear, comprehensive, well‑structured, and easy to extract, while also leaning heavily on recognised third‑party sources such as review sites and industry publications.
  • In some searches, the AI invented brand names, tools, and features, especially when it couldn’t find strong, up‑to‑date sources. That makes source quality and recency a brand‑safety and trust issue, not just an SEO one.

The sections that follow break down what the data says about source selection, localisation content structure, and organic visibility. We also provide insights into content planning, reputation management, and how to measure your AI visibility across markets.

Let’s explore these tests a little deeper:

  • Ranking well on Google can help (especially with AI Overviews), but it won’t guarantee visibility in ChatGPT or Gemini.
  • Those models favour content that is clear, comprehensive, well‑structured, and easy to extract, while also leaning heavily on recognised third‑party sources such as review sites and industry publications.
  • In some searches, the AI invented brand names, tools, and features, especially when it couldn’t find strong, up‑to‑date sources. That makes source quality and recency a brand‑safety and trust issue, not just an SEO one.

The sections that follow break down what the data says about source selection, localisation content structure, and organic visibility. We also provide insights into content planning, reputation management, and how to measure your AI visibility across markets.

Let’s explore these tests a little deeper:

ChatGPT and Gemini Cite Different Sources

Across 1,226 keyword-and-country tests, ChatGPT and Gemini shared an average of just 5% of their cited domains. This figure uses Jaccard overlap, a measure of how many sources two lists have in common relative to the total number of distinct sources across both lists. 

In the typical test, the overlap was zero: 58% of searches produced completely different citation sets.

Intent Mean overlap
Brand reputation 0.090
Branded 0.063
Client brand 0.115
Commercial 0.035
Head 0.060
Question 0.029
Transactional 0.029
0 %

of searches produced completely different citation sets.of searches produced completely different citation sets.

AI Search Is More Location-Specific Than Organic Search

Higher scores indicate greater overlap between the sources shown in the US and UK. AI tools showed less overlap than Google’s regular organic results, meaning they rebuilt their source lists more extensively for each country.

Interestingly, Gemini showed the greatest degree of localisation.

Source US/UK domain overlap
Google organic (top 10) 0.387
ChatGPT citations 0.267
Bing organic (top 10) 0.221
Gemini citations 0.192
Google AI Overview citations 0.183

Gemini Leans More on Company-Owned Content Than ChatGPT

Most citations in this test came from company-owned content, but Gemini relied on it more heavily. Vendor blogs accounted for 74.6% of Gemini’s citations versus 57.5% of ChatGPT’s.

ChatGPT spread the remaining citations across corporate sites, documentation, news media, and directories more than Gemini did, even though both systems drew equally on independent review sites.

Commercial bias, i.e., the page recommending what it sells, appeared in 75% of ChatGPT’s citations and 82% of Gemini’s.

Page type ChatGPT Gemini
Agency blog 0.003 0.001
Corporate site 0.127 0.033
Directory 0.014 0.013
Documentation 0.038 0.009
Forum / UGC 0.000 0.000
Independent review 0.189 0.189
News media 0.042 0.007
Other 0.012 0.002
Vendor blog 0.575 0.746
0 %

commercial bias in ChatGPT citationscommercial bias in ChatGPT citations

0 %

commercial bias in Gemini citations

What AI-Cited Snippets Have in Common

Most AI-cited passages were concise, self-contained explanations rather than isolated facts, definitions, or table entries. Of the 11,037 citation matches analysed, 89% could be understood without relying heavily on the surrounding page. The typical cited passage was 56 words long, although longer excerpts raised the average to 73 words.

Other observed features:

Format

How cited passages were structured — paragraphs, lists, or tables.

Clarity

Share of passages tagged as clearly explaining the topic.

Coverage

Share of passages pulled from pages that cover the topic comprehensively, not just in passing.

Authority

Share of passages tagged for coming from an authoritative source.

Original data

Share of passages tagged for including unique data not found elsewhere.

Structure

Share of passages tagged for page structure, such as headings and lists, as a citation factor.

The one major difference between the systems was the prevalence of FAQ-style passages. ChatGPT’s cited passages were more likely than Gemini’s to appear under question-based headings (29% versus 16%) and to form the first block of text beneath a heading (53% versus 46%).

Metric ChatGPT Gemini
Matches 3,191 7,846
Self-contained 0.833 0.916
Median words 56.0 56.0
Paragraph share 0.710 0.790
In first quarter 0.440 0.400
Question heading 0.293 0.164
First block after heading 0.527 0.459

AI Cites a Broader Mix of Pages for Low-Volume Keywords

AI didn’t treat all keywords the same way. For high- and mid-volume queries, company-owned content dominated the citations, taking around three‑quarters of the spots (74.0% for high‑volume, 73.8% for mid‑volume keywords). That dominance dropped in the low‑volume group, where vendor content fell to 65.6% of citations.

Lower‑volume keywords brought more variety into the mix. Independent review sites accounted for 15.9% of citations in the low‑volume group, compared with just 8.6% for high‑volume keywords. Directories also showed up more often: 5.2% of citations for low‑volume keywords versus 1.7% for high‑volume.

Page type High Mid Low
Agency blog 0.009 0.001 0.005
Corporate site 0.090 0.067 0.101
Directory 0.017 0.011 0.052
Documentation 0.029 0.016 0.010
Independent review 0.086 0.140 0.159
News media 0.017 0.016 0.014
Other 0.014 0.011 0.004
Vendor blog 0.740 0.738 0.656

Brand Reputation Searches Send Users to Review Platforms

When users asked about a brand’s reputation, AI systems turned to review platforms far more often than they did for other searches. These platforms accounted for 36% of cited URLs in reputation-related queries, compared with 6% across all other intents.

The effect was stronger in ChatGPT, where review sites accounted for 41% of citations, compared with 27% in Gemini. For marketers, this means that reputation management is also an AI visibility issue: the sources that shape an AI-generated answer may sit outside the brand’s own content ecosystem.

When users asked about a brand’s reputation, AI systems turned to review platforms far more often than they did for other searches. These platforms accounted for 36% of cited URLs in reputation-related queries, compared with 6% across all other intents.
The effect was stronger in ChatGPT, where review sites accounted for 41% of citations, compared with 27% in Gemini. For marketers, this means that reputation management is also an AI visibility issue: the sources that shape an AI-generated answer may sit outside the brand’s own content ecosystem.

0 %

of cited URLs in brand-reputation queries came from review platforms, vs 6% for all other intents.

Domain Citations
g2.com 343
gartner.com 207
reddit.com 124
capterra.com 107
trustpilot.com 86
techradar.com 85
forbes.com 79
trustradius.com 73
softwareadvice.com 53
uk.trustpilot.com 52
glassdoor.com 50
nerdwallet.com 37

How Often Each AI System Cited a Source

The table below shows the percentage of searches in each intent category that produced at least one citation from ChatGPT, Gemini, or Google AI Overviews. It measures whether a citation appeared and not how many citations appeared or whether the answer was accurate.

ChatGPT cited a source in every test, partly because web search was enabled for all ChatGPT runs. Gemini’s citation rate varied more by intent: it cited 89% of brand-reputation searches but only 34% of question-based searches. Google AI Overviews appeared most often for branded and client-brand searches, while transactional searches had the lowest AI Overview rate at 67%.

Intent ChatGPT cites Gemini cites AIO present
Brand reputation 1.00 0.89 0.92
Branded 1.00 0.88 0.97
Client brand 0.98 0.79 0.98
Commercial 1.00 0.74 0.91
Head 1.00 0.66 0.83
Question 1.00 0.34 0.93
Transactional 1.00 0.80 0.67

Ranking Higher on Google Improves Your Chances of AI Citation

Ranking well still helps, but AI citation is not simply a copy of Google’s ranking system.

In this test, domains ranking in Google’s top three results were more likely to be cited by ChatGPT, Gemini, and Google AI Overviews than domains ranking lower down the results page. However, this didn’t always guarantee a citation: ChatGPT cited 34.8% of these domains, Gemini cited 21.0%, and Google AI Overviews cited 26.6%.

The relationship weakened further down the rankings. For domains ranking in positions 11–20, citation rates fell to 14.1% for ChatGPT, 8.4% for Gemini, and 12.8% for Google AI Overviews.

Bucket Cited ChatGPT Cited Gemini Cited AIO
1-3 0.348 0.210 0.266
4-10 0.198 0.161 0.181
11-20 0.141 0.084 0.128

ChatGPT’s Citations Are Closer to Google’s Than Bing's

Because ChatGPT has been associated with Bing, one theory is that its citations should overlap more with Bing’s top results than with Google’s. We tested this by comparing ChatGPT’s cited domains with the top 10 results from both search engines.

And guess what? The data does not support this idea.

ChatGPT’s cited domains had an average overlap of 0.104 with Google’s top 10 results, compared with 0.035 with Bing’s. That’s roughly three times as much overlap with Google.

Google AI Overviews showed the strongest relationship with Google’s own rankings (makes sense, I guess), with an average overlap of 0.145. Its overlap with ChatGPT’s sources was 0.043 and with Gemini’s was 0.049.

This means that Google rankings appear to be more relevant to visibility in Google AI Overviews than in chatbot citations. But ranking well is not enough: most AI citations still came from pages outside Google’s top 10.

Comparison Mean Jaccard
ChatGPT vs Google top10 0.104
ChatGPT vs Bing top10 0.035
Gemini vs Google top10 0.116
Gemini vs Bing top10 0.016
AI Overview vs Google top10 0.145

AI-Cited Snippets Tend to Appear Near the Top of the Section

In this test, 80% of the cited passages could be pinned to a specific element on the page, such as a paragraph under a heading or a list item.

The typical cited passage appeared relatively early in the content:

Median Page Depth

How far down the page cited passages typically appeared.

0 %
of the way down the page
Median Distance: Heading to Passage

How close cited passages sat to their nearest heading.

0
words
First Quarter of the Page

Share of cited passages found within the top 25% of the page.

Question Headings

Share of cited passages that appeared under a heading phrased as a question.

First Block After a Heading

Share of cited passages that were the first paragraph or list directly below their heading.

Question Headings

Share of cited passages that appeared under a heading phrased as a question.

In other words, AI systems often cited passages that sat near the top of a section and were close to a clear heading.

Related Searches Reveal Many More Potential Sources

We grouped related searches into “fan-out” groups. Each group contained a median of six related queries for the same topic and market.

  • Using only the main (“seed”) query, the median number of cited domains was 13.
  • When all related queries in the group were combined, the median rose to 60 domains, a 4.7× increase.
  • The overlap between the sources for different queries within the same group was low (mean Jaccard index of 0.10), indicating that each variant surfaced a largely distinct set of domains.

If you expand one target keyword into around five related searches and combine the results, you will see several times more candidate sources than if you look at a single query. 

This matters for content briefs and competitive research because relying on a single search can miss many of the domains that AI systems consider relevant.

0

domains cited by seed query alone

0

domains cited across a full fan-out group

0 x

increase from seed to full group

What Are the Top-Cited Domains?

This test shows how many times ranked domains appeared as a cited source across all tests. Despite the visibility of social and forum platforms, user-generated content (Reddit, LinkedIn, YouTube, Quora, etc.) accounted for only 4.0% of all 24,897 citations. 

B2B AI answers in this test were anchored in edited tech media, review platforms, and company-owned content rather than open forums.

Rank ChatGPT domain ChatGPT citations Gemini domain Gemini citations
1 techradar.com 425 reddit.com 203
2 g2.com 351 g2.com 192
3 forbes.com 270 gartner.com 95
4 reddit.com 243 pcmag.com 83
5 gartner.com 179 sentinelone.com 72
6 salesforce.com 162 forbes.com 70
7 techtarget.com 161 dhl.com 64
8 legalclarity.org 155 airwallex.com 60
9 clutch.co 152 capterra.com 56
10 shopify.com 136 zapier.com 56
11 capterra.com 125 technologyadvice.com 53
12 fitsmallbusiness.com 110 clutch.co 52
13 microsoft.com 109 stripe.com 49
14 nerdwallet.com 109 techradar.com 48
15 semrush.com 105 squareup.com 47

The Domains AI Treats as Core Sources

Known as “anchor domains,” these were cited by three or more related queries within the same fan-out group, indicating they are recurring sources for that topic.

It’s interesting that only 50% of these domains also appeared in Google’s top 10 for their respective groups. In other words, about half of the domains that AI consistently relies on would be invisible to a content brief based only on Google rankings.

0 %

of anchor domains also appear in Google’s top 10 — the other half is invisible to Google-only briefs.

Domain Groups
techradar.com 19
forbes.com 17
techtarget.com 14
reddit.com 11
legalclarity.org 9
microsoft.com 9
shopify.com 9
semrush.com 9
wise.com 7
techrepublic.com 7
fitsmallbusiness.com 7
cyberdefenders.org 7
ahrefs.com 7
quickbooks.intuit.com 6
salesforce.com 6

AI Answers Are Rewritten for Each Country

We compared the full text of ChatGPT’s and Gemini’s answers for the same keyword in the US and UK. We measured how similar the wording was, using a text-similarity metric where 0 means completely different and 1 means identical.

No keyword produced a near-identical answer in both countries. Out of 613 keywords, none had a similarity score above 0.8 for either provider.

Answers diverged most for transactional searches, where currencies, providers and regulations differ.

The metric measures wording, not meaning. That means that answers could still convey similar information while being phrased differently.

For example, here’s what ChatGPT answered when asked about “SEO agency pricing”:

  • US: “Monthly SEO retainer (small business) | $500–$2,000/month”
  • UK: “Basic local SEO | £300–£800 | Small local businesses”

AI systems aren’t simply copying the same answer and swapping a few details. They generate distinct, locally adapted responses for each market.

Provider Mean Median Mostly rewritten (<0.3)
ChatGPT 0.10 0.09 99%
Gemini 0.05 0.05 100%
Intent ChatGPT Gemini
Brand reputation 0.08 0.06
Branded 0.10 0.05
Client brand 0.09 0.05
Commercial 0.10 0.05
Head 0.11 0.06
Question 0.13 0.07
Transactional 0.08 0.04

Global Giants Dominate, but Many Sources Are Local

The test ranked domains by citation count separately for the US and UK.

Key patterns we noticed:

  • The most-cited domains were largely the same in both countries: G2, TechRadar, Reddit, Forbes, and Gartner appeared at or near the top in both markets.
  • The low overall overlap between US and UK sources (from the earlier localisation finding) came from the many smaller, country-specific domains that appeared in one market but not the other.
  • Aggregator and review platforms (for example, G2, Capterra, and Trustpilot) maintained a strong cross-border presence.
  • Content citations, however, were effectively won market by market, with local variants and country-specific domains rising in the UK.

UK-specific top mentions included:

  • uk.trustpilot.com
  • wise.com
  • airwallex.com
  • The .co.uk review ecosystem (capterra.co.uk, softwareadvice.co.uk, ncsc.gov.uk)

For international brands, this means that a strong global domain is not enough. To show up consistently in AI answers across markets, you need market-specific content and a local presence on relevant review and comparison platforms.

Rank US domain US citations UK domain UK citations
1 g2.com 284 g2.com 259
2 reddit.com 234 techradar.com 241
3 techradar.com 232 reddit.com 212
4 forbes.com 183 forbes.com 157
5 gartner.com 122 gartner.com 152
6 clutch.co 106 clutch.co 98
7 capterra.com 103 dhl.com 93
8 legalclarity.org 96 salesforce.com 90
9 salesforce.com 93 shopify.com 85
10 fitsmallbusiness.com 93 wise.com 78

What Cited Pages Have in Common: Structure and Schema

We analysed 8,561 cited pages to understand their structure and technical markup. 

Here’s what we found in the structural anatomy of these 8,561 scraped cited pages:

  • Length and structure: The typical cited page had a median length of 2151 words with 27 headings. AI systems in this test tended to cite long, well-organised content rather than short, lightly structured pages.
  • FAQ-style content: 65% of cited pages included a section with three or more question-style headings, and 19% of all headings across these pages were phrased as questions.
  • Lists and tables: The median cited page contained around 30 list items, and 41% included at least one table. Among independent review pages, 59% featured comparison tables.

Page structure is a given. Pages that were cited five or more times look structurally identical to pages cited once. What separates anchors is coverage and brand, not extra formatting.

Schema markup on the 798 most-cited pages:

Schema type Share of top-cited pages
Any JSON-LD 83%
Organization 72%
Article / BlogPosting 51%
BreadcrumbList 46%
FAQPage (Question/Answer) 30%
Review / Product / AggregateRating 12%
HowTo 1%

In this sample, structured data was common on the most-cited pages: 83% used some form of JSON-LD schema, about half were marked up as articles, and nearly a third carried the FAQPage schema.

To see whether these patterns were specific to AI-cited pages or just a feature of any high-performing page, we compared them with two control groups:

  • Pages cited only by Google AI Overviews.
  • Pages that ranked in Google’s top 10 for the same keywords but were never cited by any AI surface.

The next table shows how the groups differ:

Group Has JSON-LD Has Article Has FAQPage Has Organization Pages
Cited by AI Overview only 0.605 0.444 0.147 0.494 387
Cited by ChatGPT/Gemini 0.826 0.508 0.301 0.716 798
Ranks top-10, never cited 0.582 0.257 0.101 0.405 378

Pages cited by ChatGPT or Gemini were about three times more likely to carry an FAQPage schema and about twice as likely to carry Article schema than pages that ranked well but were never cited. 

This suggests that well-structured, semantically marked-up content is more common among AI-cited pages, although the study does not prove that schema alone causes a page to be cited.

We also broke down schema usage by provider to see whether ChatGPT and Gemini favoured different types of markup.

Group Pages Any JSON-LD Article FAQPage Review/Product Breadcrumb
ChatGPT-cited 635 82% 44% 29% 13% 45%
Gemini-cited 333 87% 69% 32% 10% 47%

There’s no evidence in this data that you can target one chatbot over the other by choosing a specific schema type. Therefore, the goal should be to implement the right schema for your content (articles, FAQs, organisation info, breadcrumbs) rather than trying to optimise separately for ChatGPT and Gemini.

Link Authority: AI Does Not Need Your Backlinks

For many marketers, the assumption is that pages with lots of backlinks will dominate AI citations, just as they often perform well in traditional search. This test suggests otherwise.

We compared backlink data for all 14,116 cited URLs against a control group of 3,915 URLs that ranked in Google’s top 10 for the same keywords but were never cited by any AI surface.

The median AI-cited page had zero page-level backlinks, and 53% had no referring domains at all. That’s a weaker link profile than the Google top-10 pages that were never cited.

Even the most-cited “anchor” pages were lightly linked: those cited more than five times had a median of just three referring domains, compared with zero for once-cited pages.

What does this tell us? Visibility in AI answers appears to depend more on content structure, topical coverage, and domain recognition than on traditional page-level authority signals. 

You do not need a strong backlink profile to be cited, although links may still matter for overall domain strength and traditional search performance.

What Do Cited Pages Rank for Organically?

We examined the organic keyword footprints for 482 of the 600 most-cited pages (US data).

  • Each page ranked for 42 keywords, but only 4 of those were in the top 10.
  • 26% of the top-cited pages ranked for fewer than 10 keywords in total. Even though these pages were heavily cited by AI, they had almost no visible presence in Google.
  • 118 of the 600 pages returned no ranking data at all.
  • AI-cited pages are, on the whole, organically modest: being a Google winner and being an AI source are largely separate games, as we saw earlier in the section on Google rankings and AI citation.

What does this mean for your strategy? You don’t need to dominate Google Search to become a frequent AI source. Many of the most-cited pages in this study had small or even negligible organic footprints. This reinforces the earlier finding that AI citation and traditional rankings are related but distinct: strong SEO helps, but it is not a prerequisite for AI visibility.

Additional Observations

Beyond the main patterns, here are a few technical notes from the data that are worth highlighting:

0 %

Possible Hallucinated Citations

We flagged 152 of 11,037 matches (1.4%) as likely hallucinations, or cases where the cited page did not support the claim and the model’s confidence was low.

While rare, these show that AI answers can occasionally misrepresent sources, so it matters not just whether you’re cited, but how accurately.

0 %

Google AI Overview Trigger Rate

Google triggered an AI Overview in 85% of queries overall, with higher rates for brand and informational intent (92–98%) and lower for transactional queries (67%).

For marketers, this means AI-generated answers are now the default first impression for most non-transactional searches.

0 %

Scrape Coverage

We retrieved 86% of cited pages. Browser-impersonating requests recovered an additional 603 pages that standard requests could not, suggesting some sources are partially blocked to simple crawlers.

This can make AI visibility harder to measure with traditional SEO tooling.

Conclusion: What This Study Means for Marketers

This study shows that AI search is not just a new ranking surface but a distinct information environment. Visibility in ChatGPT, Gemini, and Google AI Overviews depends less on backlinks and raw Google rank and more on:

  • How clearly and completely your pages answer specific questions.
  • How often you appear in the third‑party sources (review sites, publications) that AI models treat as trusted references.
  • How well you’re represented in each market’s local ecosystem, not just in one global “top 10.”

For CMOs, the practical takeaway is to treat AI visibility as its own channel: plan content, reputation, and measurement strategies specifically for GEO, alongside traditional SEO.

Do you know what AI is saying about your brand?
Run a free AI visibility audit.
alan ai2

The findings are based on a single point-in-time snapshot (July 2026). AI citation behavior will shift as models and retrieval systems are updated. Page classification, citation matching, and “why cited” tags were generated by an LLM (gemini-2.5-flash-lite) and validated via spot-checks, not full human-labelled ground truth. Overlap metrics are reported at domain level; URL-level overlap is stricter and therefore lower. In this run, Bing’s API exposed no generative/Copilot elements, so Bing analysis covers organic results only. Review-platform bot protection limited full-page retrieval for some brand queries, so brand routing findings rely on citation inventories rather than detailed page analysis. Taken together, the patterns here should be read as directional signals, not permanent rules, about how AI search currently selects and uses sources.