GEO · AI Search Visibility

How AI Search Engines Like Perplexity Actually Crawl and Cite Sources

Here's exactly how AI search engines crawl and cite sources: Perplexity, ChatGPT Search and Google AI Overviews each use a different crawler, a different index and a different citation logic, and each rewards different things.

Quick answer

AI search engines crawl and cite sources very differently from each other, and the difference changes what actually gets you cited:

Perplexity runs its own crawler, PerplexityBot, to build a citation index, plus a separate Perplexity-User agent for live browsing at query time. It favors fresh, clearly structured content — pages updated within 30 days get a citation boost, compressing to 48–72 hours for fast-moving topics.

ChatGPT Search uses a different crawler entirely from the one that trains the model. GPTBot handles training data; OAI-SearchBot handles search visibility. Blocking GPTBot does not remove you from ChatGPT Search citations — only blocking OAI-SearchBot does that. ChatGPT Search also leans heavily on Bing's index for real-time retrieval.

Google AI Overviews doesn't crawl separately at all. It draws from the same index as ordinary Google Search, breaks your query into sub-queries (Google calls this "query fan-out"), and narrows 200–500 candidate pages down to the 5–15 it actually cites.

How AI search engines crawl and cite sources differently

Most site owners picture a single AI bot reading the web and deciding who gets quoted. That picture is wrong in a way that costs real visibility. Each major AI search product runs its own crawler, with its own name, its own rules, and its own relationship to the search index behind it. A block rule aimed at "AI crawlers" in general can silently exclude you from one system while leaving you fully exposed to another — or the reverse, blocking the one that actually drives citations while leaving in a crawler that only feeds model training.

How Perplexity actually crawls and cites

Perplexity runs two separate agents. PerplexityBot builds the index Perplexity draws citations from over time. Perplexity-User fetches pages live, triggered by an actual user query, rather than on a standing schedule.

Every Perplexity answer displays its cited sources openly, so appearing in that index brings real, attributable traffic. Perplexity weighs relevance to the query, content freshness, domain credibility, and clarity. Freshness matters more here than on most other platforms: content published or updated in the last 30 days gets a measurable citation boost, and for fast-moving topics that window can shrink to 48 to 72 hours. Pages built around a direct answer near the top, specific verifiable data points, comparison tables, and FAQ content in clear question-and-answer form get cited most consistently.

How ChatGPT Search actually crawls and cites

OpenAI now documents four distinct crawler roles, and conflating them is an easy, costly mistake. GPTBot exists for training data. OAI-SearchBot exists specifically for ChatGPT Search visibility. ChatGPT-User handles fetches triggered directly by a user's request inside a chat. OAI-AdsBot, added in 2026, validates ad landing pages.

  • Blocking GPTBot alone will not remove you from ChatGPT Search citations. That crawler only feeds training. The one that actually matters for search visibility is OAI-SearchBot, and leaving it unblocked is the real requirement for appearing in ChatGPT Search answers.
  • Bing's index sits underneath ChatGPT Search. ChatGPT Search uses Bing as its primary real-time retrieval layer. A page that isn't indexed by Bing is starting from a real disadvantage here, regardless of its Google standing.
  • Domain authority carries real, measured weight. Sites with over 32,000 referring domains are roughly 3.5 times more likely to be cited than sites with fewer than 200. Page speed matters too: pages loading in under 0.4 seconds (First Contentful Paint) average 6.7 citations, against 2.1 for pages loading in over 1.13 seconds.

How Google AI Overviews actually selects sources

Google AI Overviews doesn't run a separate crawl at all. It draws directly from the same index built for ordinary Google Search. When a query comes in, Google's systems use a technique it calls query fan-out: the original question splits into several sub-queries, each retrieved and evaluated separately, before the results merge into one answer.

The selection funnel is genuinely steep. Google typically narrows an initial pool of 200 to 500 candidate documents down to just 5 to 15 sources that actually get cited. Strong E-E-A-T signals (experience, expertise, authoritativeness, trustworthiness) show up in roughly 96% of AI Overview citations, so content without them increasingly fails to surface here regardless of other optimization. Since AI Overviews are generated primarily from the mobile version of a site, mobile-first indexing and page speed both carry real weight, alongside content built for what Google calls "answer extraction" — specific passages clean enough to lift directly into a summary.

What actually predicts getting cited, across all three

Despite the different mechanisms, the same few things show up as genuine predictors everywhere: a clear, direct answer positioned early in the content; specific, verifiable facts rather than vague claims; clean technical performance, especially page speed; consistent entity information (name, location, services, credentials) that removes ambiguity about who you are; and structured data that states those facts in machine-readable form rather than forcing an AI system to infer them from prose.

Frequently Asked Questions

Do Perplexity, ChatGPT Search and Google AI Overviews all use the same crawler?

No. Each runs its own system. Perplexity uses PerplexityBot for indexing and Perplexity-User for live fetches. ChatGPT Search uses OAI-SearchBot, separate from GPTBot's training crawler. Google AI Overviews doesn't crawl separately at all — it reuses the standard Google Search index.

If I block GPTBot, am I excluded from ChatGPT Search?

No. GPTBot only feeds OpenAI's model training. ChatGPT Search citations depend on OAI-SearchBot specifically, which is a separate crawler with a separate purpose. Blocking one does not block the other.

How much does content freshness matter for AI citation?

It matters most on Perplexity, where content updated within the last 30 days gets a measurable citation boost, compressing to 48–72 hours for fast-moving topics. Freshness also factors into ChatGPT Search and Google AI Overviews, though less sharply than on Perplexity.

Does page speed actually affect AI citation rates?

Yes, measurably. Pages with a First Contentful Paint under 0.4 seconds average 6.7 citations in ChatGPT Search research, compared with 2.1 for pages loading over 1.13 seconds. Google AI Overviews also weighs page speed as part of its technical evaluation.

Why does Bing's index matter if my business targets Google users?

Because ChatGPT Search uses Bing as its primary real-time retrieval layer, not Google. A page absent from Bing's index starts at a real disadvantage in ChatGPT Search specifically, regardless of how well it ranks in Google.

What's the single strongest predictor across all three AI search engines?

No single factor guarantees citation, but E-E-A-T signals, a direct early answer, specific verifiable facts, and structured data consistently appear across Perplexity, ChatGPT Search and Google AI Overviews research. Roughly 96% of Google AI Overview citations come from sources with strong E-E-A-T signals specifically.

Sources & References

Not sure which of these crawlers can actually see your business?

Run the free AI Search Scorecard to check your real crawl and citation signals in five minutes — or book a visibility audit for a full breakdown.

Get the free scorecard → See the AI Visibility Audit