AI search engines: how each one finds its sources

·8 min read

AI search engines answer a question by running web searches, reading the pages that come back and writing a reply that links to its sources. They don't share one index. Google's AI Overviews, AI Mode and the Gemini app draw on Google Search, Microsoft Copilot draws on Bing, and ChatGPT, Perplexity and Claude each send crawlers of their own with their own robots.txt tokens. Get indexed in Google and Bing, let OAI-SearchBot, PerplexityBot and Claude-SearchBot in, leave Google-Extended unblocked, and your pages can appear in all of them.

AI search engines compared: sources, crawlers and citations

The biggest difference between AI search engines is whose index they read, because that decides which bot has to reach your pages.

Engine Sources Bot that reaches your pages robots.txt control Citations
ChatGPT search Third-party search providers, partner content and OpenAI's own crawl OAI-SearchBot; ChatGPT-User for live fetches OAI-SearchBot Inline links and a Sources panel
Perplexity Perplexity's own index PerplexityBot; Perplexity-User for live fetches PerplexityBot Numbered citations
Google AI Overviews and AI Mode Google's search index Googlebot Googlebot, snippet controls, Search Console's generative AI setting Supporting links with the answer
Gemini app Google Search, used for grounding Googlebot Google-Extended, plus Googlebot Sources button when the answer has links
Microsoft Copilot Bing search results Bingbot bingbot; NOARCHIVE and NOCACHE meta tags Hyperlinked citations below the answer
Claude Not published; Anthropic also runs its own search crawler Claude-SearchBot; Claude-User for live fetches Claude-SearchBot, Claude-User Direct citations

The robots.txt column is the one to act on. GPTBot and ClaudeBot are missing from it on purpose. They collect training data, and neither OpenAI nor Anthropic ties them to search results.

How each AI answer engine finds its sources

Each one searches an index when the question arrives, and ChatGPT, Perplexity and Claude can also fetch a page live while they write.

OpenAI's launch post for ChatGPT search says it draws on third-party search providers and on content its partners supply, without naming the providers. Its own crawler, OAI-SearchBot, collects pages for ChatGPT's search features. OpenAI's bot documentation says sites that block it won't appear in ChatGPT search answers, apart from plain navigational links, and that robots.txt changes take about 24 hours to reach its systems.

Two other OpenAI agents get mixed up with it. ChatGPT-User visits a page when a person asks ChatGPT about it, and OpenAI warns that robots.txt rules may not apply to it because a user started the request. GPTBot collects training data and has nothing to do with search.

Answers carry inline source links, and a Sources button below the response opens a panel listing everything cited. For what makes ChatGPT pick one page over another, read how to get cited by ChatGPT.

Perplexity

Perplexity answers from its own index. When it opened that index to developers as the Search API in September 2025, it said the index covers hundreds of billions of web pages and runs on the same infrastructure as its answer engine. Perplexity's crawler docs describe PerplexityBot as the bot that surfaces and links sites in those results, and say it isn't used to train foundation models.

Perplexity-User fetches pages live while answering. Because a user asked for the fetch, Perplexity says it generally ignores robots.txt. Every answer comes with numbered citations linked to the source pages.

Google Search's AI features

AI Overviews and AI Mode are built from Google's own index, the one Googlebot fills. Google's documentation on AI features says a page needs only to be indexed and eligible to show with a snippet to be used as a supporting link, with no extra technical requirements. Both features can run several related searches across subtopics, which Google calls query fan-out, and Google says this lets them show a wider set of links than classic search does.

Your controls are the ones you already have. Googlebot rules in robots.txt decide crawling, and nosnippet, data-nosnippet, max-snippet and noindex limit what Google can show. Google-Extended has no effect here. Since August 31, 2026, the Search generative AI setting in Search Console lets any site exclude itself from AI Overviews, AI Mode and Discover's generative features without touching its ranking.

Sundar Pichai said in July 2026 that Google had brought AI Overviews and AI Mode together into one Search experience, which is why they share a row in the table. How Google AI Mode picks sources goes further into source selection.

Gemini app

The Gemini app grounds its answers in Google Search, so it reads the same index Googlebot builds. The control is different, though. Google's crawler documentation says the Google-Extended token governs both Gemini training and grounding in Gemini Apps, and that it has no effect on inclusion or ranking in Google Search.

So if you block Google-Extended to keep your content out of training, you also drop out of the Gemini app's grounded answers, while AI Overviews keep using your pages. Check for that line before you decide Gemini ignores you.

When an answer has sources, Gemini shows a Sources button that opens the links in a side panel. Google notes that not every response includes them.

Microsoft Copilot

Copilot grounds its answers in web search results. Microsoft's documentation for its business Copilot names the Bing search service as the source, while the consumer transparency note says only that Copilot summarizes top web search results, without naming the engine.

Bing reads two meta tags that matter here. NOARCHIVE keeps a page out of Copilot answers, and NOCACHE limits Copilot to the page's URL, title and snippet. Bing's guidelines, rewritten in February 2026, advise against NOCACHE on content you want Copilot to use, Search Engine Journal reported. Copilot lists hyperlinked citations below its answer.

Anthropic doesn't name the search provider behind Claude's web search, though TechCrunch reported in March 2025 that Brave Search had been added to Anthropic's subprocessor list as the feature launched.

What Anthropic does document is its three bots. Claude-SearchBot crawls to improve search result quality, Claude-User fetches a page when someone asks Claude a question, and ClaudeBot collects training data. Anthropic's crawler page says its bots honor robots.txt and that blocking Claude-User stops Claude retrieving your content for user questions. OpenAI and Perplexity say their live fetchers may not follow robots.txt. Anthropic says Claude's web search answers come with direct citations. Our ClaudeBot guide covers all three bots in more depth.

Which robots.txt rules control AI search engines?

Allow the search bots, decide separately about the training bots, and leave Googlebot, Bingbot and Google-Extended unblocked unless you mean to leave those answers. If your robots.txt has no rules for these bots and no blanket Disallow, they're already allowed and you need none of this. If it does, a version that admits the answer engines and opts out of training looks like this:

# Answer engines
User-agent: OAI-SearchBot
User-agent: PerplexityBot
User-agent: Claude-SearchBot
User-agent: Claude-User
Allow: /

# Model training only. Delete this group to allow training.
User-agent: GPTBot
User-agent: ClaudeBot
Disallow: /

One trap. A crawler that finds a group naming it ignores the User-agent: * group completely. If your * group disallows /cart/ or /admin/, copy those lines into the answer engine group too, or those four bots will crawl them. Our robots.txt checker shows which of these AI crawlers your file allows, Google-Extended included, tests any path against the crawler you name, and gives the exact line to change.

Which is the best AI search engine for a site to show up in?

Google's, for most sites, because one crawl feeds three of the six engines. Sundar Pichai said in Google's Q2 2026 earnings remarks that AI Mode passed 1 billion monthly active users after it expanded globally in October 2025, and that the Gemini app has 950 million. AI Overviews sit on ordinary results pages on top of that. A page Google has indexed and can show with a snippet is a candidate for all three.

ChatGPT comes second. OpenAI said in February 2026 that ChatGPT had 900 million weekly active users, and it gates its search answers on its own crawler. A site can rank first on Google and still be missing from ChatGPT search because an old "block all AI bots" rule caught OAI-SearchBot. Check that line before anything else.

Copilot needs little beyond the Bing basics. Confirm in Bing Webmaster Tools that Bing has indexed your important pages, and keep NOARCHIVE off any page you want cited.

Don't plan content around Perplexity or Claude for a typical site. Their search bots cost one robots.txt line each to allow, so allow them and spend your effort on Google and ChatGPT.

Do you need a different plan for each AI powered search engine?

No. Six AI powered search engines reduce to three pipelines. Get indexed in Google, get indexed in Bing, and let the AI companies' search bots in.

Measurement is where they really differ. Search Console's Generative AI performance report, open to all sites since August 31, 2026, shows impressions for AI Overviews and AI Mode combined, by page, country, device and date, and no clicks. Bing's AI Performance report, in public preview since February 2026, counts how often Copilot and Bing's AI summaries cite each of your pages and samples the queries behind those citations. ChatGPT tags the links it cites with utm_source=chatgpt.com, so its visits are easy to filter in analytics. For Perplexity, Claude and the Gemini app, referral traffic and answers you sample yourself are most of what you get.

Check whether AI answers name and cite your brand

AI Answer Visibility asks ChatGPT, Perplexity and Gemini one question your customers ask, with web search on. For each engine it reports whether the answer names your brand and where it ranks among the brands named, whether it cites your site, and which brands and sites it recommends instead. If you don't enter a brand name, it reads one from your home page. When an engine cites no source at all, the report says so, since that usually means the answer came from what the model already knew.

A run costs 120 credits. It doesn't cover Google AI Overviews or AI Mode. AI answers change from one request to the next, so treat each run as a sample and ask the same question again after you change something.

Keep reading