403 Forbidden: why crawlers get blocked when browsers don't
·4 min read
Status code 403, Forbidden, means the server understood the request and refuses to serve it. When a page loads fine in your browser but Googlebot or GPTBot gets a 403, the cause is almost never the page. A firewall, bot rule or CDN bot protection decided the request looks automated. Google treats a 403 like any other 4xx, so the page drops out of the index, and AI crawlers that get one have nothing to read or cite.
What does a 403 status code mean?
A 403 is a refusal. RFC 9110 defines it as a server that "understood the request but refuses to fulfill it," and says a client should not repeat the request with the same credentials.
The three codes people mix up:
| Code | What the server is saying | Typical cause |
|---|---|---|
| 401 Unauthorized | Send valid credentials and try again | Login wall, HTTP basic auth on staging |
| 403 Forbidden | I know what you want and won't give it to you | WAF rule, bot protection, IP block, file permissions |
| 404 Not Found | Nothing here | Deleted or mistyped URL |
The RFC also lets a server answer 404 instead of 403 when it wants to hide that a resource exists. For the rest of the 4xx family, see our HTTP status codes guide.
How Google treats a 403
Google drops the page. Its HTTP status code documentation says Google doesn't use content from URLs that return 4xx codes, and that all 4xx errors except 429 are treated the same. A 403 removes a ranking page the same way a 404 error would.
A 403 is also no way to slow Googlebot down. The same page says 4xx codes other than 429 "have no effect on crawl rate" and tells you not to use 401 and 403 for limiting crawling. To slow it, return 429 or 503 briefly, as our post on status code 429 explains.
Search Console lists these URLs as "Blocked due to access forbidden (403)". Google's help text is blunt. Googlebot never sends credentials, so your server is returning the error incorrectly, and the page won't be indexed.
Why do crawlers get a 403 while browsers get 200?
Because the thing answering isn't your app. Something in front of it refuses requests that look like bots. Check these in order:
- CDN bot protection. Cloudflare, Akamai and Imperva challenge or refuse clients that don't run JavaScript or look like a browser. Crawlers fail that test.
- AI bot settings. Cloudflare's AI bot controls block AI training crawlers when enabled. Its documentation says that from September 15, 2026, bots classed as Training or Agent are blocked by default on pages that display ads, while Search bots stay allowed.
- WAF rules matching the user agent. A scraper rule matching "bot" in the user agent also catches Googlebot and GPTBot.
- IP and country blocks. Crawlers fetch from data centers, so blocking cloud IP ranges or whole countries blocks them.
- Server config. An
.htaccessdeny rule, a security plugin or wrong file permissions. These usually refuse browsers too, so you'd notice sooner.
Nobody on the team sees the 403, because nobody on the team is a bot.
How to diagnose a 403 by user agent
Compare a plain request with ones that claim to be crawlers:
curl -s -o /dev/null -w "%{http_code}\n" https://example.com/pricing
curl -s -o /dev/null -w "%{http_code}\n" \
-A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/120.0.0.0 Safari/537.36" \
https://example.com/pricing
curl -s -o /dev/null -w "%{http_code}\n" -A "GPTBot" https://example.com/pricingA 200 for the first and a 403 for the others means a rule keys on the user agent. Three 200s prove less than they seem. Good bot rules check the source IP, so your laptop can pass while the real Googlebot is refused.
For Google, run the URL Inspection live test in Search Console, which shows what Googlebot actually got. For other crawlers, filter your server or CDN logs by user agent and look for 403s, as our log file analysis guide shows.
How to fix 403s for search and AI crawlers
Fix the rule, not the page.
- In your CDN, turn on the verified bots allowance. Cloudflare excludes verified bots from its default bot configurations, but custom rules can still catch them.
- Review AI bot settings and allow the crawlers you want, such as OAI-SearchBot for ChatGPT search and PerplexityBot. Our AI crawler list has every user agent.
- Rewrite WAF rules that match "bot" in the user agent so they exclude verified crawlers instead of trusting the string.
- If you allow Googlebot by hand, verify it. Google's verification guide says to run a reverse DNS lookup on the IP, confirm the name ends in googlebot.com, google.com or googleusercontent.com, then check that a forward lookup returns the same IP.
- Never use 403 to keep a crawler out of a page you want hidden. Use robots.txt for crawling and noindex for indexing.
Check what crawlers get from your page
Our AI Crawler View fetches a page with a plain HTTP request, as crawlers that skip JavaScript do, then loads it in a headless browser. If the plain request gets a 403 or a bot check page while the browser gets the real page, the report flags it and points you to your CDN or firewall bot rules. When both load, it lists content, headings, prices, structured data, links and metadata that appear only after JavaScript runs.
It sends its own user agent, not GPTBot's or Googlebot's, so a rule aimed at one named bot won't show up. It doesn't read robots.txt. A run costs 8 credits, refunded if the browser render fails.