The Types of Citations in AI Answers: What AI Search Actually Cites

AI answers can cite brand websites, media, Reddit, reviews, official sources and more. Learn the main types of AI citations and how to analyze them.

Published September 5, 2026 · 11 min read

AI citations can come from almost anywhere.

Ask ChatGPT, Gemini or Perplexity about a product and you might see sources from the brand's own website. Ask for the best products in a category and the answer could draw from comparison sites, Reddit discussions, YouTube videos or editorial articles instead.

The source depends on the question.

That makes AI citations more useful than a simple count of how many times your website was cited. They can tell you where AI engines are getting information about your category, which pages influence their answers and where your brand appears – or doesn't.

To understand that, it helps to look at the different types of citations and the role each source plays.

What is an AI citation?

An AI citation is a source that an AI engine connects to an answer.

Depending on the platform, it might appear as an inline link, source card, reference list or another link attached to the response.

A citation isn't the same thing as a brand mention.

An AI answer could recommend your brand without citing your website. The recommendation might instead be supported by a review, comparison article or Reddit discussion.

The opposite can happen too. Your website could be used as a source for factual information even if your brand isn't prominently mentioned in the generated answer.

A mention tells you whether your brand made it into the words the user actually reads. A citation tells you which pages the engine leaned on to build that answer.

Neither substitutes for the other. Reading them together is what starts to explain how a given answer was put together, which matters for understanding AI visibility as a whole.

What types of sources do AI engines cite?

There isn't one fixed set of sources used by AI search.

In our own analysis of more than 5 million citations, we found sources ranging from Reddit and YouTube to government websites, SaaS platforms, publishers, comparison sites and brand-owned domains.

The exact mix also differed substantially between engines.

But instead of thinking about citations as one large pool of URLs, it's more useful to group them by the role they play in an answer.

Brand-owned websites

Sometimes the most appropriate source is the company itself.

That might include:

  • product pages
  • documentation
  • help centers
  • pricing pages
  • research
  • blog posts
  • comparison pages

If someone asks:

What integrations does HubSpot support?

HubSpot's own documentation is a natural source.

But if someone asks:

Is HubSpot good for a 10-person sales team?

The engine may want information beyond what HubSpot says about itself.

That's an important distinction.

Brand-owned content can be very useful for facts about your company and products. It doesn't mean it will be the preferred source for every question involving your brand.

Reviews and comparison sites

Comparison content becomes more relevant when the question involves choosing between products.

For example:

Best CRM for a small business

HubSpot vs Salesforce

Best accounting software for freelancers

These questions require some form of evaluation.

An AI engine could use official product pages to understand features and pricing, but third-party comparisons can provide additional context about differences, strengths and weaknesses.

For brands, these citations are particularly interesting because they sit outside your own website.

If the same comparison page repeatedly appears when a competitor is recommended, you want to know what's on that page.

Does it mention you?

How are you positioned?

Is the information accurate?

Which competitors are included?

The citation itself only tells you where the AI looked. The page can tell you why that source might matter.

Social, video and user-generated content

Reddit, YouTube, forums and other community platforms also appear as AI sources.

They can be particularly useful when a question calls for something official documentation doesn't provide very well: experience.

Think about the difference between:

What battery does this camera use?

and:

Is this camera good for travelling?

The first has a factual answer that a manufacturer can provide.

The second benefits from people who have actually used the product.

That creates a natural role for reviews, videos, forums and community discussions.

This doesn't mean every brand needs to start manufacturing Reddit threads.

It means you should pay attention when user-generated sources consistently appear for important questions in your category.

In our AI Citation Landscape research, for example, we found meaningful differences in how often individual engines turned to sources such as Reddit and YouTube.

The useful question isn't simply whether Reddit is important to AI search.

It's whether Reddit is important for your prompts, in your category, on the engines you care about.

News and editorial content

News publications, magazines, trade publications and specialist blogs can provide another layer of third-party information.

These sources can be useful when an AI answer needs industry context, reporting, expert commentary or information that isn't available directly from the companies involved.

For brands, editorial citations can be interesting even when they don't link to your website.

Imagine an industry publication publishes a comparison of five companies in your category. That article then starts appearing as a source when people ask AI engines which companies they should consider.

Your visibility on that page could matter even though you don't own it.

This is one reason citation analysis extends beyond traditional backlink analysis.

The important relationship isn't necessarily:

Website → link → your website

It can also be:

Website → information about your brand → AI answer

Government and official sources

Some questions need a more authoritative source.

Government agencies, regulators, standards organizations, universities and other official institutions can provide definitions, statistics, rules and primary information.

For example, an AI answer about mortgage rules might use a government source for the regulations while using bank websites for available products and consumer discussions for experiences with those banks.

Each citation is doing a different job.

For most brands, the opportunity isn't to replace an official source. It's to understand which parts of the answer depend on official information and where commercial sources start influencing the response.

Academic and reference sources

Research papers, universities, Wikipedia and other reference sources can provide background knowledge or evidence.

These become particularly relevant for scientific, technical or factual questions.

Again, context matters.

A research paper cited for a medical claim and a Reddit thread cited for customer experience are both technically citations. But treating them as equivalent doesn't tell you much about how the answer was constructed.

The role of the source matters as much as the source itself.

The page type can matter more than the domain

Classifying citations by website is useful, but it only gets you so far.

Consider two citations from the same domain:

example.com/product

and

example.com/research/industry-study

The domain is identical.

The content isn't.

AI engines can cite many different types of pages:

  • Product pages
  • Guides
  • Comparisons
  • Documentation
  • Original research
  • Reviews
  • Category pages
  • News articles
  • Forum discussions
  • Glossaries
  • Videos

Looking at page types can reveal patterns that domain-level reporting misses.

You might discover that a competitor appears across hundreds of cited URLs.

That's interesting.

But discovering that three comparison pages account for most of their visibility on your highest-priority prompts gives you something you can investigate.

Vercite domain citations dashboard showing a stacked chart of citations over time and a table of cited domains with response counts, share, unique URLs and domain rating.
Domain-level reporting is the starting point – notice the Unique URLs column, since a domain's citation count can come from a handful of pages or hundreds of them.

This is where citation analysis starts becoming actionable rather than just another visibility metric.

The prompt changes which sources make sense

Citation patterns also change with the question.

Take four prompts about the same category:

What is CRM software?

What does HubSpot do?

What are the best CRM platforms for a small sales team?

HubSpot vs Salesforce for a startup?

The first asks for a definition.

The second asks about a company.

The third asks for recommendations.

The fourth asks for a direct comparison.

There's no reason to expect the same source mix for all four.

Documentation might work well for the first two. Comparison sites and reviews become more relevant for the latter two. Community discussions might appear when the answer starts discussing real-world experience.

This is why looking at citation data across all your prompts at once can be misleading.

Instead, group prompts by topic or intent and look at the sources within those groups.

You may find that a particular publisher barely appears overall but is highly influential for comparison prompts.

That's much more useful than knowing its total citation count.

Different AI engines use different sources

There's another complication: the source patterns you see in one engine don't necessarily transfer to another.

ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode all have different approaches to search, retrieval and citation.

We've seen this clearly in our own data.

In our analysis of 5.31 million citations, only 23 domains appeared among the top 100 most-cited domains across all five engines.

So there isn't one universal list of websites you need to appear on to "win AI search."

A source might be highly influential in ChatGPT and relatively unimportant in Google AI Mode.

This is one reason citation analysis should happen at the engine level rather than only looking at an aggregate number.

You can explore the full breakdown in our AI Citation Landscape research.

Retrieved doesn't always mean cited

There's another layer behind the citations you see in the final answer.

AI engines can search the web before answering a question.

And they don't necessarily search using the exact words you typed.

A simplified process might look something like:

Prompt → query fan-out → search → retrieved pages → answer generation → citations

Not every engine exposes each of these steps, and the exact mechanics differ between them. Treat this as the shape of the process, not a claim that every engine surfaces its retrieval data.

One prompt can trigger several background searches.

Those searches can discover many pages, but not every page the engine retrieves needs to appear as a citation in the final answer.

That creates an important distinction.

Retrieved but not cited means the engine found or read the page during its research process, but didn't ultimately surface it as a source in the response.

Cited means the page made it through to the sources attached to the answer.

Vercite response breakdown for a single prompt, listing seven sources cited in the answer and nine sources that were retrieved but not cited.
The same prompt, split into what was cited in the answer and what was retrieved but never made it in.

Why does this matter?

Because the two situations point to different problems.

If your page is never retrieved, the engine may not consider it relevant to the searches it's performing.

If your page is retrieved but another source gets cited, you've already passed one hurdle: discovery.

The next question becomes why another page was a better source for the final answer.

Maybe it answered the question more directly.

Maybe it contained stronger evidence.

Maybe it was easier to extract information from.

Maybe it better matched the particular claim the model wanted to support.

We shouldn't assume there's one universal reason. But separating retrieval from citation gives you a much better framework for investigating it.

A citation doesn't mean the source recommends your brand

This is one of the biggest limitations of simply counting citations.

Imagine an AI answer recommends your brand and cites:

example.com/best-crm-software

It's easy to assume that article helped support the recommendation.

Open the page, though, and you might find that it mentions three competitors but never mentions you.

Or perhaps it mentions you negatively.

Or the AI only used the page for a statistic elsewhere in the answer.

The URL can't tell you. We've covered this exact gap – a domain being cited without the brand ever being named – in a dedicated article.

This is why citation analysis shouldn't stop at the domain or page.

You also need to understand what's on the cited page.

Read the content behind the citation

When we analyze citations in Vercite, we don't only want to know which URLs appear.

We also look at the content on those pages.

That makes it possible to investigate questions like:

  • Does this page mention our brand?
  • Which competitors appear?
  • How are the brands described?
  • What information might make this page useful to an AI engine?
  • Are competitors appearing on influential pages where we're absent?
  • Is there outdated or incorrect information about us?
  • Could this page represent a realistic outreach opportunity?

This adds an important layer to citation tracking.

Imagine two pages are cited frequently for your most important prompts.

One already includes your brand prominently.

The other discusses four competitors but doesn't mention you.

The raw citation numbers could look almost identical.

The action you take should be completely different.

What should you actually do with citation data?

You don't need to classify every URL on the internet.

Start with the prompts that matter to your business.

Look at which sources appear when those prompts are run across different AI engines.

Then work down from there:

Which domains keep appearing?

Look for sources that repeatedly influence the topic.

Which individual pages get cited?

A domain may have thousands of pages, but only a handful might matter.

What types of pages are they?

Comparisons, documentation, research, reviews and discussions suggest different opportunities.

What do those pages say?

Check whether your brand appears, which competitors are included and how they're positioned.

When do those sources appear?

A source that's important for informational prompts might be irrelevant for recommendations – and vice versa.

That's when citations become more than another dashboard number.

They start showing you which parts of the web are helping AI engines form answers about your market.

There isn't one type of AI citation that matters most

Brand websites matter.

So do comparison sites, Reddit discussions, YouTube videos, publishers, government sources and research.

Which one matters most depends on what someone is asking, which AI engine they're using and what information the answer needs.

That's why the goal shouldn't simply be to "get more AI citations."

A better goal is to understand the citation landscape around the questions your customers ask.

Find the sources.

Look at the pages.

Read what they say.

And then decide where you actually have an opportunity to improve how your brand appears in AI answers.

Frequently asked questions

What is an AI citation?

An AI citation is a source that an AI engine connects to an answer, shown as an inline link, source card or reference list depending on the platform. It isn't the same as a brand mention: an answer can recommend a brand without citing its website, and a website can be cited without the brand being prominently mentioned in the text.

What's the difference between a brand mention and a citation?

A brand mention is the AI naming your brand in the generated answer. A citation is a source the engine surfaced or used while constructing that answer. They're separate signals – a brand can be mentioned with none of its own pages cited, and a page can be cited without the brand being mentioned at all.

What does 'retrieved but not cited' mean?

It means an AI engine found and read a page while researching a prompt, but didn't ultimately surface it as a source in the final answer. That's different from never being retrieved: retrieval without citation means you've passed the discovery stage, but another source was judged better for the final answer.

William Hollingworth
William HollingworthFounder, Vercite

Builds Vercite and writes most of its research and case studies.

LinkedIn
Keep reading

See how AI engines describe your brand

Vercite tracks your mentions, citations and sentiment across ChatGPT, Gemini, Perplexity and Google AI – on a schedule, not a spot check.

Get started free