ClaudeBot: what Anthropic's crawler does on your site

·7 min read

ClaudeBot is the web crawler Anthropic uses to collect public pages that may go into training future Claude models. It is one of three Anthropic bots. Claude-SearchBot indexes pages for Claude's search answers, and Claude-User fetches a page when someone asks Claude a question that needs it. All three follow robots.txt, each answers to its own user agent token, and a rule written for ClaudeBot does nothing to the other two.

The split matters because the costs differ. Blocking ClaudeBot keeps your future pages out of training data. Blocking Claude-SearchBot or Claude-User takes you out of answers people read today. Most site owners who search for "block claudebot" want the first and should leave the other two alone.

What each Anthropic crawler does

Anthropic's help center article on crawling describes three bots and what you give up by disabling each one.

Bot What Anthropic says it does What blocking it costs you robots.txt token
ClaudeBot Collects web content that could be used to train its models Nothing visible today. Your future pages are left out of training data ClaudeBot
Claude-SearchBot Indexes content to make Claude's search results more relevant and accurate Less visibility in Claude's search answers Claude-SearchBot
Claude-User Fetches pages when a Claude user asks a question that needs them Claude can't read your page when someone asks about it Claude-User

Older robots.txt guides still list anthropic-ai and Claude-Web. Neither name appears in Anthropic's current documentation, so write your rules for the three tokens above. Anthropic also asks you to repeat the rules on every subdomain you want covered, because each host reads its own robots.txt.

The ClaudeBot user agent string

ClaudeBot sends this user agent:

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)

Anthropic's help page names the tokens but does not publish full strings. The string above is the one that appears in server logs and in public lists such as the crawler-user-agents project. Claude-SearchBot and Claude-User use the same pattern with their own token, though the contact part at the end varies between sightings, and the same list records a shorter Claude-User string sent from Claude Code.

So match on the token, case-insensitive, and you catch every version. A log search for "claude bot" with a space finds nothing.

Does ClaudeBot respect robots.txt?

Yes, according to Anthropic. Its published crawling principles say the bots honor robots.txt, support the non-standard Crawl-delay rule and do not try to get past CAPTCHAs.

The doubts date from July 2024. iFixit CEO Kyle Wiens posted that ClaudeBot had hit iFixit's servers a million times in 24 hours, and 404 Media reported that his logs showed thousands of requests per minute over several hours. Freelancer.com CEO Matt Barrie told The Information that the bot made 3.5 million visits in about four hours. Engadget's summary reports that the crawling at iFixit stopped once the site added a robots.txt rule for Anthropic's bot, and Anthropic told The Information its crawler "respected that signal."

The lesson from iFixit is narrow. Its terms of service already banned AI training, and that did nothing. A robots.txt line did.

If ClaudeBot seems to ignore your robots.txt today, check the IPs first, because scrapers copy crawler user agents. Then check that the rule matches. A group for anthropic-ai, or a file on www that does not cover the blog subdomain, leaves ClaudeBot unaffected.

Should you block ClaudeBot?

Block ClaudeBot only if your content is what you sell. Leave Claude-SearchBot and Claude-User allowed on almost every site.

Training crawlers send very little traffic back. Cloudflare's crawl data shows Anthropic crawled about 38,000 pages for every visit it referred in July 2025, down from about 286,000 in January. If you publish paywalled reporting, a paid dataset or course material, that is a bad trade, and blocking ClaudeBot is a reasonable call.

For a business that sells a product or a service, the trade looks different. Your pages are marketing. A model that read them during training can describe your company and name you in answers that never trigger a web search. Blocking also only works forward. Anthropic calls it a signal to leave your future material out of training, so nothing comes out of models already trained.

Claude-SearchBot and Claude-User are an easier decision. They control whether Claude can read and cite your page when someone asks about your topic, and Anthropic warns that disabling either one may reduce your visibility in its answers. Crawler access is the first step of generative engine optimization, and these two bots are the Claude part of it.

How to block ClaudeBot in robots.txt

Add a group for the token with Disallow: /. Pick the version below that matches your decision and put it in the robots.txt at the root of each host. If you run WordPress, our WordPress robots.txt guide covers where the file lives and how to edit it.

Block training, keep Claude search and user fetches

User-agent: ClaudeBot
Disallow: /

Claude-SearchBot and Claude-User have no group of their own in this file, so they follow your User-agent: * rules.

Block every Anthropic bot

User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
Disallow: /

Slow ClaudeBot down instead of blocking it

User-agent: *
Disallow: /cart/
Disallow: /search

User-agent: ClaudeBot
Crawl-delay: 10
Disallow: /cart/
Disallow: /search

The usual reading of Crawl-delay is seconds between requests, so a value of 10 holds ClaudeBot to roughly 8,640 requests a day. Copy your * rules into the ClaudeBot group. Under the robots.txt standard, RFC 9309, a crawler follows only the most specific group that names it, so once ClaudeBot has its own group it ignores everything under User-agent: *. Skip the copy and your crawl-delay fix opens paths that were blocked before.

After you publish, run our robots.txt checker to see which rule ClaudeBot, Claude-SearchBot and Claude-User each match for a given path.

How to find ClaudeBot in your server logs

Search the access log for the token. A ClaudeBot request in an Nginx or Apache combined log looks like this:

216.73.216.47 - - [08/Oct/2026:03:14:07 +0000] "GET /pricing HTTP/1.1" 200 18342 "-" "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)"

These commands assume that format. zgrep reads rotated .gz files as well as the live log.

# Requests per day
zgrep -ih "claudebot" /var/log/nginx/access.log* | awk '{print substr($4, 2, 11)}' | sort | uniq -c

# Most requested paths
zgrep -ih "claudebot" /var/log/nginx/access.log* | awk '{print $7}' | sort | uniq -c | sort -rn | head -20

# Status codes ClaudeBot received
zgrep -ih "claudebot" /var/log/nginx/access.log* | awk '{print $9}' | sort | uniq -c | sort -rn

# IPs that claimed to be ClaudeBot
zgrep -ih "claudebot" /var/log/nginx/access.log* | awk '{print $1}' | sort | uniq -c | sort -rn

# All three Anthropic bots
zgrep -ihE "claudebot|claude-searchbot|claude-user" /var/log/nginx/access.log*

How to verify real ClaudeBot traffic

Check the IP, not the user agent. Anthropic publishes its crawler addresses at claude.com/crawling/bots.json and says a crawler whose source IP is on that list is coming from Anthropic. The version dated October 7, 2026 has one /22 block, 216.73.216.0/22, plus a few dozen smaller ranges and single addresses. A request that claims to be ClaudeBot from an IP outside the list is someone else using its name.

Behind Cloudflare, a load balancer or another proxy, your log may record the proxy's address instead of the crawler's. Log the real client IP from CF-Connecting-IP or X-Forwarded-For before you judge anything by IP.

Do not block Anthropic's IPs as a way to opt out. Anthropic warns that IP blocking may not keep you opted out, because it stops the bots from reading your robots.txt.

How much ClaudeBot crawling is normal?

Anthropic publishes no crawl rate, and another site's numbers tell you little about yours. Compare ClaudeBot's daily requests with the number of URLs you actually want crawled, and with Googlebot's requests over the same days. If ClaudeBot requests more URLs a day than your site has pages, week after week, it is re-fetching the same pages or wandering through URLs that should not exist. Thousands of requests a minute, the pattern iFixit reported, is abnormal on any site.

Then look at the paths. Faceted filters, internal search results, calendar pages and session IDs in query strings give every crawler an endless supply of URLs. A Disallow rule for those patterns fixes the volume for every bot at once.

What to do when ClaudeBot crawls too hard

Work through these in order, because the first one often ends the problem.

  1. Confirm it is Anthropic. List the IPs with the command above and check them against bots.json. If they are not, block requests that claim ClaudeBot from unlisted IPs at your firewall or CDN.
  2. Close the crawl traps. Disallow the parameter, search and filter URLs that soak up requests.
  3. Add Crawl-delay to the ClaudeBot group. Anthropic's own example uses 1 second. On a small server, start at 10 and copy your other rules into the group.
  4. Rate-limit by token, not by vendor. If you throttle at the CDN, target ClaudeBot only. Throttling Claude-SearchBot and Claude-User costs you answers.
  5. Contact Anthropic. Email claudebot@anthropic.com from an address at the affected domain. Anthropic asks for that so it can verify the report.

Anthropic does not say how often it rereads robots.txt, so give a change a few days before you escalate.

Check which Anthropic bots reach your site

Our AI Bot Log Analyzer reads a pasted access log of up to 750,000 characters from Apache, Nginx, Caddy, IIS, CloudFront, Cloudflare Logpush or Vercel. It picks out ClaudeBot, Claude-SearchBot and Claude-User along with GPTBot, PerplexityBot and other AI crawlers. Each request's IP is checked against the list its operator publishes, including Anthropic's bots.json. When your log records a proxy address instead of the crawler's, the report says so rather than calling real crawlers impostors.

You get requests per crawler, split into training, AI search and user requests, with counts verified by IP and counts that used a crawler's name from another IP. It lists the requests your server refused with a 401, 403 or 429, the pages that returned 404, 410 or 5xx errors, the pages assistants fetched for users, and requests by day.

Nothing connects to your server. A run costs 8 credits. Filter the log with the zgrep commands above first and the paste covers more days.

Keep reading