How to block AI crawlers with robots.txt, and when you shouldn't

·7 min read

Here is how to block AI crawlers with robots.txt. Put each bot's token on its own User-agent line, add one Disallow: / under them, and leave your User-agent: * group as it is. Block only the training bots, such as GPTBot, ClaudeBot, Google-Extended and CCBot, and your pages stay out of future training sets while ChatGPT, Claude and Perplexity can still cite them. Add the search bots, such as OAI-SearchBot and PerplexityBot, and you drop out of those answers.

Which bots go in the group decides what the block costs you. Pick the goal first, then copy the matching block below.

Which AI crawlers to block: training, search or user fetchers

Block training bots if you object to models learning from your pages. Leave search bots and user fetchers alone unless you want out of AI answers altogether.

Job What the bot does Main tokens What blocking it costs
Training Collects pages for future models GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot Nothing you can see today
Search Builds the index an assistant cites from OAI-SearchBot, Claude-SearchBot, PerplexityBot, Meta-WebIndexer Your pages stop appearing in that assistant's answers
User fetch Opens a page because a person asked ChatGPT-User, Claude-User, Perplexity-User, Meta-ExternalFetcher The assistant can't read your page even when someone pastes the link

Google-Extended and Applebot-Extended are control tokens. They never send a request. Google and Apple read the rules for them while Googlebot and Applebot crawl as usual. For every token, its operator and its IP list, see our AI crawler list. This post sticks to the blocks.

Use this block if you want your pages kept out of training sets and still want assistants to find and cite them.

User-agent: GPTBot              # OpenAI
User-agent: ClaudeBot           # Anthropic
User-agent: Google-Extended     # Google Gemini
User-agent: Applebot-Extended   # Apple
User-agent: Meta-ExternalAgent  # Meta
User-agent: Amazonbot           # Amazon
User-agent: MistralAI-Training  # Mistral
User-agent: CCBot               # Common Crawl
User-agent: Bytespider          # ByteDance
Disallow: /

OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and Applebot have no group here, so they keep following your * rules and keep feeding search answers.

Three tokens reach past training. Google says Google-Extended also governs grounding, where Gemini Apps pull pages from Google's index to back an answer at prompt time. Blocking it can cost you Gemini app answers, though Google says it has no effect on Google Search inclusion or ranking. Meta-ExternalAgent also indexes content for Meta's products. Amazon says Amazonbot improves its products and may train its models, while its search bot, Amzn-SearchBot, is a separate token that stays allowed. Drop any of the three if that cost matters more to you than the training opt-out.

To protect one section only, such as a paywalled archive, swap Disallow: / for the path, like Disallow: /premium/.

How to block GPTBot or any single vendor

Name only that vendor's tokens. This keeps GPTBot out and leaves ChatGPT search alone:

User-agent: GPTBot
Disallow: /

This blocks OpenAI's training, search and user bots:

User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
Disallow: /

OpenAI's crawler documentation treats each token as a separate setting, and the second block has a price. OpenAI says sites that opt out of OAI-SearchBot "will not be shown in ChatGPT search answers", though they can still show up as navigational links. Our GPTBot guide covers OpenAI's bots in more depth.

The same pattern works for Anthropic, with ClaudeBot, Claude-SearchBot and Claude-User. Perplexity documents no training crawler. Its bot docs say PerplexityBot doesn't crawl for AI foundation models, so blocking it only takes you out of Perplexity's results.

Google is the odd one. Gemini training runs through Google-Extended, but AI Overviews and AI Mode come from Googlebot's ordinary crawl, and blocking Googlebot removes you from Google Search. Robots.txt is the wrong tool for those features. Search Console's Search generative AI setting is the right one, as our guide to turning off AI Overviews explains.

How to block AI crawlers completely

To block AI bots in robots.txt across the board, name the training bots, search bots and user fetchers of every vendor above. This block leaves Googlebot, Bingbot and Applebot alone, because blocking those takes you out of Google, Bing and Apple search.

User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
User-agent: Meta-WebIndexer
User-agent: Meta-ExternalFetcher
User-agent: Amazonbot
User-agent: Amzn-SearchBot
User-agent: Amzn-User
User-agent: MistralAI-Training
User-agent: MistralAI-Index
User-agent: MistralAI-User
User-agent: CCBot
User-agent: Bytespider
Disallow: /

This one has a real price. ChatGPT, Claude, Perplexity, Meta AI and Mistral lose the crawlers they use to find and cite your pages, and the compliant fetchers refuse to open your URL when a buyer pastes it in. The assistants still answer questions about your market. They answer from your competitors' pages and from whatever other sites say about you.

I'd only ship this for a site whose content is the product, like paywalled news, courses or a paid dataset. For a business whose pages are its marketing, it removes you from conversations you can't see and doesn't stop them happening.

Why a wildcard Disallow blocks AI search bots too

Trying to disallow AI crawlers through the * group blocks every search bot you didn't name, because a crawler with no group of its own obeys *. This file is meant to say "Google only, no AI":

User-agent: Googlebot
Allow: /

User-agent: *
Disallow: /

Under RFC 9309, a crawler follows the groups that match its token and falls back to * only when none does. Googlebot has a group and gets in. OAI-SearchBot, Claude-SearchBot, PerplexityBot and Bingbot have none, so they all get Disallow: /. That one group costs you ChatGPT search, Claude, Perplexity and Bing. Applebot gets in by accident, because Apple says it follows your Googlebot rules when the file doesn't name Applebot.

Write it the other way round. Name the bots you want out and keep * permissive.

The rule cuts the other way as well. Once a bot has its own group, it ignores everything under *. A GPTBot group with only Disallow: /drafts/ frees GPTBot from your other rules, so copy your * lines into every bot group that doesn't disallow the whole site.

Bots that may ignore robots.txt

User-triggered fetchers are the gap in every block above. OpenAI says robots.txt rules may not apply to ChatGPT-User. Perplexity says Perplexity-User generally ignores them, and Google says the same of its user-triggered fetchers. Meta says Meta-ExternalFetcher may bypass them, and Amazon says Amzn-User may not follow every directive. Anthropic says its bots, Claude-User included, honor robots.txt.

Keep those tokens in a block-everything file anyway, but don't count on them obeying it. Stopping them, and scrapers that borrow a crawler's name, takes firewall rules and IP checks, which our guide to stopping AI scraping covers.

Your CDN can override robots.txt

A CDN setting can block a bot your robots.txt allows, and it can add Disallow rules you never wrote. Cloudflare is the usual case.

Since September 15, 2026, Cloudflare's Block and Block on pages with ads options also stop mixed-use crawlers such as Googlebot, Bingbot and Applebot, so choosing Block to stop training takes you out of search too. Cloudflare's own advice is to pick Disallow AI Training instead. Its managed robots.txt, which Cloudflare is replacing with Bot Preference Sync, puts Disallow: / groups for GPTBot, ClaudeBot, Google-Extended, CCBot and others at the top of the file it serves.

So the robots.txt crawlers read can differ from the one in your repository. Test the live URL, and check your CDN's bot and firewall settings before you blame a crawler for skipping you.

How long a block takes and how to confirm it worked

Allow about a day. OpenAI says search changes take about 24 hours. Perplexity, Meta and Amazon give the same 24 hours, and Google says it generally caches robots.txt for up to 24 hours. RFC 9309 tells crawlers not to use a cached copy for longer than that unless the file is unreachable. Anthropic gives no figure.

A training block also only looks forward. Anthropic describes it as keeping your future material out of training sets, and neither OpenAI nor Anthropic says a block removes pages they already collected.

To confirm it worked:

  1. Fetch the live file and read what crawlers get, CDN additions included. Check the status too. A 5xx makes compliant crawlers treat the whole site as blocked, and a 4xx makes them treat it as open.

    curl -s https://example.com/robots.txt
    curl -s -o /dev/null -w "%{http_code}\n" https://example.com/robots.txt
  2. Test which rule each token matches, since one stray group can undo the block.

  3. A day or two later, look at the current access log. A compliant bot you blocked should fetch /robots.txt and little else.

    grep -iE "gptbot|claudebot|ccbot" /var/log/nginx/access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head

If page URLs still top that list, check the IPs before you decide the bot ignores you, because anyone can put GPTBot in a user agent. Our AI Bot Log Analyzer checks each request's IP against the list its operator publishes.

Check your robots.txt for AI crawlers

Our robots.txt checker fetches your live robots.txt and shows whether 14 AI tokens, including GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended, may fetch the URL you enter. For each one it names the rule that decided it and says whether the bot fell back to your * group, which is how the wildcard mistake above shows up. Blocked search and user bots come back as warnings and blocked training bots as notes, because one costs you answers and the other is usually a choice.

It also tests any path for any crawler you name, reads Content-Signal lines, and flags syntax errors, duplicate groups, sitewide Disallow rules and a missing Sitemap line. It doesn't pretend to be the bots, so it can't prove one obeys your file, but it reports when your server answers its own request with a 401, 403 or 429. A run costs 10 credits.

Keep reading