AI crawlers: every AI bot user agent and what it does
·9 min read
AI crawlers are the bots AI companies send to read websites, and each one does one of three jobs. Search bots such as OAI-SearchBot, Claude-SearchBot and PerplexityBot build the indexes that AI answers cite. User fetchers such as ChatGPT-User and Perplexity-User load a page because a person asked about it. Training bots such as GPTBot, ClaudeBot and CCBot collect pages for future models, and they are the only group most sites have a reason to block.
The table below covers 27 tokens from 12 operators, checked against each operator's own documentation in October 2026. Where an operator publishes nothing, the table says so.
The AI crawler list at a glance
Rows are grouped by job. Search bots come first, then user fetchers, training bots and the two tokens that never crawl. "Yes" in the robots.txt column means the operator says the bot follows robots.txt or documents a Disallow rule as the way to block it.
| Token | Operator | Job | Honors robots.txt, per operator | Published IP list |
|---|---|---|---|---|
OAI-SearchBot |
OpenAI | Search index for ChatGPT | Yes | Yes |
Claude-SearchBot |
Anthropic | Search index for Claude | Yes | Yes |
PerplexityBot |
Perplexity | Search index for Perplexity | Yes | Yes |
Meta-WebIndexer |
Meta | Search index for Meta AI | Yes | No |
Amzn-SearchBot |
Amazon | Search in Amazon products | Yes | Yes |
MistralAI-Index |
Mistral | Search index for Mistral | Yes | Yes |
DuckAssistBot |
DuckDuckGo | Live fetches for DuckDuckGo's AI answers | Yes | Yes |
Googlebot |
Google Search, which AI Overviews and AI Mode draw on | Yes | Yes | |
Bingbot |
Microsoft | Bing search, which Copilot draws on | Yes | Yes |
Applebot |
Apple | Siri, Spotlight and Safari search, and may train Apple models | Yes | Yes |
Google-CloudVertexBot |
Crawls a site owner requests for Vertex AI agents | Yes | Yes | |
ChatGPT-User |
OpenAI | User fetch in ChatGPT and GPT Actions | May not apply | Yes |
Claude-User |
Anthropic | User fetch in Claude | Yes | Yes |
Perplexity-User |
Perplexity | User fetch in Perplexity | Generally ignores it | Yes |
Meta-ExternalFetcher |
Meta | User fetch, including agent tasks | May bypass it | No |
Amzn-User |
Amazon | User fetch, such as live Alexa answers | May not follow every rule | Yes |
MistralAI-User |
Mistral | User fetch in Vibe | Yes | Yes |
Google-Agent |
Agents that browse and act for a user | Generally ignores it | Yes | |
GPTBot |
OpenAI | Training | Yes | Yes |
ClaudeBot |
Anthropic | Training | Yes | Yes |
Meta-ExternalAgent |
Meta | Training and product indexing | Yes | No |
Amazonbot |
Amazon | Product improvement, may train Amazon AI models | Yes | Yes |
MistralAI-Training |
Mistral | Training | Yes | No |
CCBot |
Common Crawl | Open web archive that anyone can download | Yes | Yes |
Bytespider |
ByteDance | Toutiao Search, per Chinese docs only | No statement found | No |
Google-Extended |
Gemini training and grounding control, read from Google's own crawls | Token, not a crawler | Not applicable | |
Applebot-Extended |
Apple | Training control for Apple models, read from Applebot's crawls | Token, not a crawler | Not applicable |
Which AI bots should you block?
Block training bots if you object to model training, and leave search bots and user fetchers alone. Block a search bot and that assistant stops finding your pages when it builds an answer. Block a user fetcher that honors robots.txt and it can't open your page even when someone pastes the link. Block a training bot and nothing changes today. Models already trained keep what they learned, and only future training loses your pages.
For a business site, blocking every AI bot usually backfires. An assistant that can't read your pages describes you from whatever else it found. Publishers whose content is the product, such as paywalled news, courses or datasets, have a real case for blocking training.
Google and Apple handle training with a separate token, Google-Extended and Applebot-Extended, that their regular crawlers read. Neither token sends a request of its own, so you will never see one in a log.
AI bot user agents by operator
Each operator's own page is the only reliable source for its tokens.
OpenAI's crawler documentation treats each token as an independent setting, so you can allow OAI-SearchBot and block GPTBot. If you allow both, OpenAI may reuse one crawl for both purposes. Search changes take about 24 hours. ChatGPT-User acts for a person, and OpenAI says robots.txt rules may not apply to it.
Anthropic's help center says its bots honor robots.txt and lists their IP ranges in bots.json. Older guides still name anthropic-ai and Claude-Web, which the current page doesn't mention. Our ClaudeBot guide covers its logs, crawl rate and Crawl-delay.
Perplexity's bot docs say PerplexityBot isn't used for training and recommend allowing it. Perplexity-User generally ignores robots.txt because a person requested the fetch.
Google's common crawlers page describes Google-Extended as a token with no user agent string of its own. It controls Gemini training and grounding, and Google says it has no effect on inclusion or ranking in Google Search, so it won't keep you out of AI Overviews. The Search generative AI setting in Search Console does that, as our guide to turning off Google AI Overviews explains. Google-Agent sits in the user-triggered fetchers list, and Google says those fetchers generally ignore robots.txt.
Microsoft documents no AI-specific crawler. Bing's webmaster guidelines apply one set of rules to Bing search, Copilot and grounding results, so Bingbot is the bot that matters. Control Copilot with robots meta tags. NOARCHIVE keeps a page out of Copilot answers, and NOCACHE limits Copilot to the URL, title and snippet. Bing's 2023 announcement adds that NOARCHIVE content isn't used to train Microsoft's foundation models.
Apple's Applebot page says Applebot data may train Apple's foundation models. Disallow Applebot-Extended to opt out of training and stay in Apple's search features. One quirk catches people out. When your file names Googlebot and not Applebot, Applebot follows the Googlebot rules.
Meta's crawler page says Meta-ExternalFetcher may bypass robots.txt because users request its fetches. Changes take up to 24 hours, and the page lists no IP ranges.
Amazon's Amazonbot page says Amazonbot data may train Amazon AI models, and that Amzn-SearchBot and Amzn-User are not used for training. None of the three supports Crawl-delay.
ByteDance publishes no English Bytespider documentation and no IP list. The only page we found is Chinese help on the Toutiao Search Webmaster Platform, which calls it the Toutiao Search crawler and says nothing about robots.txt. Reporters tie it to training ByteDance's AI models, and our Bytespider guide covers the reports that it ignores robots.txt.
Common Crawl's CCBot page gives the token, a JSON IP list and reverse DNS under crawl.commoncrawl.org. It is a non-profit whose archive anyone can download, AI labs included, so block CCBot if you block training at all.
DuckDuckGo's DuckAssistBot page says the bot's data isn't used to train AI models and that a Disallow takes effect after 72 hours.
Mistral's crawler page gives IP lists for MistralAI-User and MistralAI-Index. MistralAI-Training has none.
robots.txt rules to copy for AI bots
Give the bots you want to block their own group with Disallow: /, and let every other bot fall through to your User-agent: * rules. A group can list several user agents above one set of rules. Our guide to blocking AI crawlers covers what each choice costs you in AI answers.
Allow AI search, block AI training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
User-agent: Amazonbot
User-agent: MistralAI-Training
User-agent: CCBot
User-agent: Bytespider
Disallow: /
User-agent: *
Disallow: /cart/
Disallow: /search
Sitemap: https://example.com/sitemap.xmlSearch bots and user fetchers have no group here, so they follow the * rules. Two lines cost more than training. Google-Extended also controls grounding in Gemini Apps, and Meta-ExternalAgent also indexes for Meta's products. Drop either line if those matter more to you.
Block every AI bot
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Google-CloudVertexBot
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
User-agent: Meta-ExternalFetcher
User-agent: Meta-WebIndexer
User-agent: Amazonbot
User-agent: Amzn-SearchBot
User-agent: Amzn-User
User-agent: MistralAI-User
User-agent: MistralAI-Index
User-agent: MistralAI-Training
User-agent: DuckAssistBot
User-agent: CCBot
User-agent: Bytespider
Disallow: /This leaves Googlebot, Bingbot and Applebot alone, because blocking them takes you out of Google, Bing and Apple search. The user fetchers whose operators say they may skip robots.txt need a CDN or firewall rule as well.
Why adding one rule can unblock a bot
Under RFC 9309, the robots.txt standard, a crawler obeys the group that names it and falls back to User-agent: * only when no group does. It is the easiest robots.txt mistake to make.
User-agent: *
Disallow: /admin/
Disallow: /checkout/
User-agent: GPTBot
Disallow: /drafts/The owner meant to keep GPTBot out of one more folder. Instead, GPTBot now ignores the * group and may fetch /admin/ and /checkout/. Copy every * rule into any bot-specific group you add, including a group you create only for Crawl-delay.
The same standard says matching is case-insensitive, so gptbot works, and several groups that name the same token are merged into one. A robots.txt that returns a 5xx error counts as a full disallow for compliant crawlers, which can hide you from every AI search bot at once.
Do AI bots obey robots.txt?
The major operators say their crawlers do, but robots.txt is a request. RFC 9309 says its rules are "not a form of access authorization", and nothing technical stops a bot from ignoring them.
User fetchers are the first gap. OpenAI, Perplexity, Meta, Amazon and Google all say theirs may skip robots.txt, on the reasoning that a person made the request. Then there are disputes. In August 2025 Cloudflare reported that after sites blocked Perplexity, it crawled them with a generic Chrome user agent from IPs outside its published ranges. Perplexity disputed the report.
Spoofing is the bigger day-to-day problem. Any scraper can put GPTBot in its user agent, so a log line that says GPTBot proves nothing until you check the IP. Bytespider has no IP list at all, so you can't separate it from an impostor.
How to verify a real AI crawler
Check the request's IP against the list the operator publishes, or run a reverse DNS lookup where the operator documents one. OpenAI, Anthropic, Perplexity, DuckDuckGo, Mistral, Common Crawl, Apple, Bing and Google publish JSON files in the same shape, a prefixes array of ipv4Prefix and ipv6Prefix entries. Amazon lists its ranges on web pages.
Take the client IP from the log line. Behind Cloudflare or a load balancer, log
CF-Connecting-IPorX-Forwarded-For, or you will be checking your own proxy.Test it against the operator's list. Swap in the right URL for other bots.
IP=203.0.113.7 # the address from your log curl -s https://openai.com/gptbot.json | python3 -c ' import ipaddress, json, sys ip = ipaddress.ip_address(sys.argv[1]) nets = [p.get("ipv4Prefix") or p.get("ipv6Prefix") for p in json.load(sys.stdin)["prefixes"]] print(any(ip in ipaddress.ip_network(n) for n in nets))' "$IP"For Google, Bing, Apple and Common Crawl, run a reverse lookup, then a forward lookup on the hostname it returns. Both must point at the same IP.
host 66.249.66.1 # 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com. host crawl-66-249-66-1.googlebot.com # crawl-66-249-66-1.googlebot.com has address 66.249.66.1
The hostname should end in googlebot.com, google.com or googleusercontent.com for Google, search.msn.com for Bing, applebot.apple.com for Apple and crawl.commoncrawl.org for Common Crawl. To check a whole log at once, our AI Bot Log Analyzer matches each request's IP against the operator's published list and reports the visits that borrowed a crawler's name.
Blocking AI bots at the CDN
Your CDN can block an AI bot before it ever reads robots.txt, so check there first when a bot you allowed never shows up. Since July 1, 2026, every Cloudflare customer can set AI traffic by category, with separate Search, Agent and Training controls. Since September 15, Block under Training also stops Googlebot, Bingbot and Applebot, search included. The Disallow AI Training setting writes a no-training rule into your robots.txt, keeps those three in for search and blocks the other training bots. Other CDNs and firewalls have their own bot rules, and a block there wins whatever your robots.txt says.
Check your robots.txt for AI crawlers
Our robots.txt checker fetches your live file and shows whether your rules let 14 AI tokens, including GPTBot, ClaudeBot, PerplexityBot and Google-Extended, fetch the URL you enter. For each one it names the rule that decided it and whether the bot fell back to your * group, which is how the precedence mistake above shows up. It also tests any path for any crawler you name, reads Content-Signal lines, and flags duplicate groups, Allow and Disallow ties, a missing * group and an HTML page served as robots.txt.
It doesn't pretend to be the bots, so it can't prove that one obeys your file. It does report when your server answered its own request with a 401, 403 or 429, a sign that a firewall rule may be turning clients away. A run costs 10 credits.