How to block AI scrapers when robots.txt is not enough

·7 min read

To block AI scrapers that ignore robots.txt, enforce the block at your CDN or web server, because robots.txt is only a request. In order of effort, check that bots using a known crawler's name come from that operator's published IP ranges, turn on your CDN's AI bot rules, refuse training bots by user agent and rate-limit anything that crawls too fast. None of it stops a headless browser on residential IPs, and blocking too broadly takes you out of ChatGPT search, Perplexity and Claude answers along with the scrapers.

Here is how much each layer stops and what slips past it.

Layer What it stops What gets through Effort
robots.txt Crawlers that honor it, like GPTBot, ClaudeBot and CCBot Anything that ignores it Low
IP verification Scrapers wearing a real crawler's user agent Bots whose operator publishes no IP list Medium
CDN AI bot rules Known AI crawlers the CDN can identify Unknown bots and disguised ones Low
Server user agent rules Bots that name themselves honestly Any bot that lies about its name Medium
Rate limiting Fast crawls from one IP Slow crawls spread across many IPs Medium

Is robots.txt enough to stop AI scraping?

No, but write it first, because GPTBot, ClaudeBot and CCBot follow it and it costs you one text file. RFC 9309, the standard that defines robots.txt, says its rules are not a form of access authorization, so treat it as a sign on the door. Our guide to blocking AI crawlers in robots.txt has the exact blocks for each bot.

What does blocking AI bots cost you?

It depends on which group you block.

Bot group Examples What blocking it costs
Training crawlers GPTBot, ClaudeBot, CCBot Your pages in future model training
Search crawlers OAI-SearchBot, Claude-SearchBot, PerplexityBot Citations in ChatGPT search, Claude and Perplexity answers
User fetchers ChatGPT-User, Claude-User, Perplexity-User Pages read live when someone asks about you or pastes your link

The vendors are blunt about the search group. OpenAI says sites that opt out of OAI-SearchBot are not shown in ChatGPT search answers, apart from navigational links (OpenAI crawler docs). Anthropic says blocking Claude-SearchBot may reduce your visibility and accuracy in search results for Claude users (Anthropic). Perplexity asks sites to allow PerplexityBot so they appear in its results (Perplexity docs).

User fetchers are the awkward group. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores robots.txt, because a person asked for the page. Stopping them takes the enforcement layers below, and it means an assistant can't read your page for a user who is asking about you.

For most business sites, block training crawlers and keep search crawlers and user fetchers. Publishers who sell their archive may reasonably block all three.

How do you verify a bot that claims to be GPTBot?

Check its IP against the list its operator publishes, because anyone can type GPTBot into a user agent header. OpenAI publishes separate JSON files for GPTBot, OAI-SearchBot and ChatGPT-User, Anthropic publishes one at claude.com/crawling/bots.json, and Perplexity links its lists from its bot docs.

Verification matters most for your exceptions. Blocking a fake GPTBot costs nothing, since you were refusing GPTBot anyway. Allowing a fake OAI-SearchBot does cost you, because every rule that exempts it by name also exempts any scraper that copies its user agent. So verify the bots you let in.

On nginx, build the allowed ranges into a file and rebuild it daily from cron.

curl -s https://openai.com/searchbot.json \
  | jq -r '.prefixes[] | (.ipv4Prefix // .ipv6Prefix) + " 1;"' \
  > /etc/nginx/oai-searchbot-ips.conf

Then refuse anything that claims the name from outside those ranges.

geo $oai_searchbot_ip {
    default 0;
    include /etc/nginx/oai-searchbot-ips.conf;
}

map $http_user_agent $claims_oai_searchbot {
    default 0;
    ~*OAI-SearchBot 1;
}

map "$claims_oai_searchbot$oai_searchbot_ip" $fake_oai_searchbot {
    default 0;
    10 1;
}

server {
    if ($fake_oai_searchbot) { return 403; }
}

Behind a CDN or load balancer, set nginx's real_ip_header to CF-Connecting-IP or X-Forwarded-For first, or every request looks like it comes from the proxy. Our AI crawler list covers reverse DNS checks for crawlers that verify that way instead.

How to block AI bots on Cloudflare

Use the AI bot policies under Security Settings, which split AI traffic into Search, Agent and Training and give each one Block, Block on pages with ads, or Allow (Cloudflare docs). Each block covers verified bots with that behavior plus unverified bots that act like them. Cloudflare is retiring the old single Block AI bots setting in favor of these three controls.

Read the Training options carefully. Since September 15, Block and Block on pages with ads also stop Googlebot, Bingbot and Applebot, the crawlers that serve both search and AI, and that includes their search crawling (Cloudflare blog). The Training-only option Disallow AI Training keeps those crawlers for search, publishes a no-training preference in your robots.txt and blocks training-only crawlers from Amazon, Anthropic, Meta and OpenAI. New domains get presets based on whether the site runs ads, so open the page and read what yours is set to.

  1. Go to Security, then Settings, and open Configure AI bot policies.
  2. To refuse training, set Training to Disallow AI Training. Block also cuts off Google and Bing search.
  3. Leave Search on Allow if you want to appear in AI answers.
  4. Set Agent to Allow if assistants should read pages for users who ask about you.
  5. Open AI Crawl Control to block individual crawlers, and check its Robots.txt tab for crawlers that break your rules.
  6. In your WAF custom rules, move the AI Crawl Control rule to the top. Cloudflare adds it last, and a Skip rule above it lets blocked crawlers through (Cloudflare docs).

Without a paid plan, AI Crawl Control recognizes crawlers by user agent string only, so it catches only bots that name themselves. Paid plans use Cloudflare's bot detection instead (AI Crawl Control docs). The separate AI Labyrinth setting adds invisible nofollow links that trap crawlers ignoring no-crawl rules in a maze of endless links, and Cloudflare says they don't affect SEO.

How to block AI scrapers by user agent on your server

Return a 403 to the user agents you don't want, which works on any host and doesn't depend on the bot reading robots.txt. On nginx, a map keeps the list in one place.

map $http_user_agent $ai_training_bot {
    default 0;
    ~*(GPTBot|ClaudeBot|CCBot|Bytespider|meta-externalagent) 1;
}

server {
    if ($ai_training_bot) { return 403; }
}

On Apache, the same rule goes in .htaccess.

RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} (GPTBot|ClaudeBot|CCBot|Bytespider|meta-externalagent) [NC]
RewriteRule ^ - [F]

Match exact tokens. A loose pattern like Claude or bot also catches Claude-SearchBot, Googlebot and your uptime monitor. This layer only catches bots that name themselves, and a scraper sending a Chrome user agent walks straight past it.

Does rate limiting stop AI scrapers?

It stops scrapers that pull pages fast from a single IP and does nothing against a crawl spread thinly across many addresses. Set the limit well above what a reader does and return 429, so a well-behaved crawler slows down instead of failing.

map $oai_searchbot_ip $limit_key {
    1       "";
    default $binary_remote_addr;
}

limit_req_zone $limit_key zone=perip:10m rate=2r/s;

server {
    location / {
        limit_req zone=perip burst=20 nodelay;
        limit_req_status 429;
    }
}

nginx skips requests with an empty key, so the map exempts the verified OpenAI ranges from the previous section. Add the other crawlers you keep the same way. On Cloudflare, every plan includes at least one rate limiting rule. The base plan counts requests per IP over 10 seconds and can exclude verified bots in the rule's expression (Cloudflare docs).

Can you stop a headless browser on residential IPs?

Not reliably. A headless Chrome sends a real browser's user agent and runs your JavaScript, and residential proxy networks route it through home connections shared with real customers, so blocking the IP blocks people. Cloudflare reported in August 2025 that Perplexity, once blocked, crawled with a generic Chrome user agent from IPs outside its published ranges (Cloudflare blog). Perplexity disputed the report. Either way, that traffic looks like a visitor to every rule above.

Cloudflare's Enterprise Bot Management scores each request and catches some of it. The dependable answer is to put the data worth stealing, such as full price lists or the archive you sell, behind an account. Anything on a public page is public.

What's left is paper. Cloudflare's managed robots.txt adds Content Signals for search, ai-input and ai-train, and Cloudflare's policy text calls those restrictions an express reservation of rights under Article 4 of EU Directive 2019/790 (Cloudflare docs). Whether that holds is for a court to decide. Cloudflare's option to charge crawlers per request is in private beta.

Check which AI bots reach your site and which are fake

Paste up to 750,000 characters of your access log into the AI Bot Log Analyzer. It reads Nginx and Apache combined logs, JSON lines from Cloudflare Logpush, Vercel or Caddy, and W3C logs from IIS or CloudFront, and lists each AI crawler and assistant it finds with its purpose, the pages it read and your traffic by day.

It checks every claimed crawler's IP against the list its operator publishes and splits visits into verified, spoofed and unverifiable, with the spoofed IPs listed. Its count of 401, 403 and 429 responses per crawler shows whether your blocks work, and it flags any search crawler or user fetcher you refused. It connects to nothing on your side, and it reports crawlers with no published IP list as unverifiable instead of guessing. A run costs 8 credits.

Keep reading