Status code 429: when your server turns crawlers away
·4 min read
Status code 429, Too Many Requests, means the server is rate limiting the client because it sent too many requests in a given time. Google's crawlers treat a 429 as a sign your server is overloaded, count it as a server error and slow down crawling of your whole hostname. A few hours of 429s is harmless. Days of them on the same URLs can push pages out of the index, and AI crawlers that hit the same wall can't read or cite you.
What does status code 429 mean?
A 429 says "you, specifically, are asking too fast." RFC 6585 defines it, and says the response should explain the condition and may include a Retry-After header with how long to wait.
HTTP/1.1 429 Too Many Requests
Retry-After: 120
Content-Type: text/plain
Rate limit exceeded. Try again in 2 minutes.Retry-After takes either a number of seconds or an HTTP date. A 503 says the whole server is struggling. A 429 says this one client crossed a limit. For the server-wide version, see our post on status code 503.
How Google treats 429
Google treats a 429 like a 5xx error, so its crawlers back off. Its HTTP status code documentation says 5xx and 429 errors make Google's crawlers temporarily slow down, that pages already indexed stay indexed at first but are eventually dropped, and that crawling speeds up again gradually once the server returns 2xx.
That makes 429 a legitimate emergency brake. Google's guide to reducing the crawl rate lists 500, 503 and 429 as the codes to return when Googlebot is overloading you, and says you can add Retry-After to tell its crawlers when to retry. The same guide sets the limit. Use it for hours or a day or two. If Googlebot sees these codes on the same URL for multiple days, that URL may be dropped from the index.
Two consequences people miss:
- The slowdown applies to the whole hostname, including pages that return 200 just fine.
- Use 429, not 403. A 403 says "forbidden," and Google says not to use it for rate limiting, because it won't read it as a request to slow down. Our post on status code 403 covers what a 403 does instead.
Why your server sends 429 to crawlers by accident
Most crawler 429s aren't a deliberate choice. Somebody set a per-IP limit for scrapers and login abuse, and crawlers tripped it. The usual causes:
- A per-IP rate limit sized for humans. A person clicks a few pages a minute. Googlebot, Bingbot or GPTBot can fetch many pages a second from a handful of IPs, so a limit like 60 requests a minute per IP catches them first.
- CDN or WAF rate limiting rules. A rule that counts requests per IP returns 429 to any client above its threshold unless you exclude verified crawlers.
- Hosting plan limits. Some shared and managed hosts cap requests per IP at the edge, outside anything you configured.
AI crawlers make this worse. ChatGPT-User and Perplexity-User fetch a page when a person asks about it. A limit that never bothered Googlebot can still refuse them, and a refused fetch means the answer cites someone else.
How to find 429s in your logs, by bot
Your access log shows the status code and user agent for every request, so you can count 429s per crawler in one line. For the Apache or Nginx combined format:
awk -F'"' '$3 ~ / 429 / {print $6}' access.log | sort | uniq -c | sort -rn | headThis splits each line on quotes, keeps requests whose status is 429 and prints the user agent with a count. If Googlebot, bingbot, GPTBot or ClaudeBot shows up, your limiter is hitting crawlers. A steady trickle of them usually means a limit set too low.
A user agent is only a claim, so verify the IP before you loosen anything. Scrapers borrow crawler names precisely to slip past rules that trust them. Our guide to log file analysis covers reading logs for bots in more depth.
Safer rate limiting for crawlers
Keep rate limits for anonymous traffic, but exempt crawlers you've verified by IP. Here is an Nginx sketch. Requests with an empty key are not counted by limit_req, so a map can skip known crawlers:
map $http_user_agent $limit_key {
default $binary_remote_addr;
~*(googlebot|bingbot|gptbot|oai-searchbot|claudebot) "";
}
limit_req_zone $limit_key zone=perip:10m rate=10r/s;
server {
location / {
limit_req zone=perip burst=40 nodelay;
limit_req_status 429;
}
}Matching on user agent alone lets impostors through, so pair it with an allow list of the IP ranges the operators publish, or use your CDN's verified bot signal. In Cloudflare rules, the cf.client.bot field is true for requests from known good bots, so you can leave them out of a rate limit rule's expression. A few more habits help:
- Return 429 with
Retry-After, never 403, when you do need to slow a crawler. - Limit expensive endpoints such as search, filters and login, not every static page.
- Set an alert on 429 counts per bot, so a too-tight rule shows up in days, not months.
Check which crawlers your server refuses
The AI Bot Log Analyzer reads access log lines you paste in Apache or Nginx combined format, JSON lines from Cloudflare Logpush, Vercel or Caddy, or W3C logs from IIS and CloudFront. It recognizes AI crawlers and assistants such as GPTBot, ClaudeBot, PerplexityBot and ChatGPT-User, plus Googlebot and bingbot. It counts the 401, 403 and 429 responses each crawler got, lists other error responses by path, and flags visits that use a crawler's name from an IP its operator doesn't publish.
It costs 8 credits per run. It doesn't connect to your server or CDN, so it only sees the lines you paste, and it can't verify crawlers whose operators publish no IP list.