robots.txt tester: how to test your robots.txt file
·7 min read
Google's robots.txt tester is gone. Google announced on November 15, 2023 that it was retiring the tester in favor of a robots.txt report in Search Console, which shows the file Google fetched and its parse errors but cannot test a URL. To test robots.txt against a URL today, use URL Inspection for pages in your own property, Google's open-source parser for a draft file, curl for the live file and its status code, or a robots.txt checker that runs any path against any user agent, AI crawlers included.
Test the file crawlers receive, not the one you wrote. The two drift apart when it is served from the wrong host or a CDN stitches its own block on top. For the syntax itself, read what robots.txt is and how its rules work.
What happened to Google's robots.txt tester?
Google replaced it with a report. The Google Search Central announcement introduced the robots.txt report under Search Console settings, added related details to the Page indexing report and said the old tester was being sunset.
What you lost is the sandbox. Search Console no longer lets you paste an edited file and try URLs against it before you publish. Google's URL Inspection help still says to "use the robots.txt tester to find the rule", which is out of date.
Search Console robots.txt report: how to read it
The robots.txt report shows each robots.txt file Google found for your site, when it last fetched it and any errors in it. In a Domain property it checks http and https for the top 20 hosts, sorted by crawl rate.
- In Search Console, open Settings and then the robots.txt report.
- Read the fetch status for each file. "Fetched" is healthy. "Not found (404)" means Google treats that host as having no rules. "Any other reason" needs fixing today.
- Click a file to see the last fetched version. Errors stop a rule from being used and warnings don't.
- Click Versions to see fetches from the last 30 days. This is how you find the day a bad file shipped.
- After an urgent fix, open the menu next to the file and choose Request a recrawl. Google suggests it for unblocking important URLs or fixing a fetch error.
A URL-prefix property such as https://example.com/ checks one origin only, so it never shows a broken http:// or www copy. Use a Domain property if you can.
How to test one URL with URL Inspection
URL Inspection is the nearest thing to the old tester, and it only works on URLs inside a property you have verified. Paste the full URL into the bar at the top of Search Console and read the "Crawl allowed?" field. Click Test live URL to check the page as it stands now.
It tells you a page is blocked, not which line blocks it, and it speaks only for Google. A page can pass here and still be closed to GPTBot or ClaudeBot. Google also caches robots.txt for up to 24 hours, so a fix from an hour ago may not show yet.
How to test robots.txt with Google's open-source parser
Run Google's own matcher on your draft before it ships. The google/robotstxt repository holds the C++ parser Google describes as a slightly modified version of Googlebot's production code.
git clone https://github.com/google/robotstxt.git
cd robotstxt && mkdir c-build && cd c-build
cmake .. -DROBOTS_BUILD_TESTS=ON && make
./robots ~/drafts/robots.txt Googlebot https://example.com/blog/postIt prints user-agent 'Googlebot' with URI 'https://example.com/blog/post': ALLOWED and exits with 0 for allowed or 1 for disallowed. Script it over your money pages and fail the deploy when one comes back disallowed.
It matches only the token you pass, without Google's fallbacks such as Googlebot-Image obeying the Googlebot group.
Skip Python's urllib.robotparser here. Its source applies the first matching rule in file order, so given Disallow: / followed by Allow: /blog/, it reports /blog/post as blocked. Google allows it, because the longer rule wins.
How to check the live robots.txt file with curl
curl shows what a crawler receives, status code and redirects included. Check each protocol and host variant, since crawlers read a separate robots.txt for each. Here is GitHub in October 2026.
for u in https://github.com http://github.com https://www.github.com; do
curl -s -o /dev/null -w "%{http_code} $u/robots.txt -> %{redirect_url}\n" "$u/robots.txt"
done200 https://github.com/robots.txt ->
301 http://github.com/robots.txt -> https://github.com/robots.txt
301 https://www.github.com/robots.txt -> https://github.com/robots.txtThat is the healthy pattern. One origin serves the file with a 200 and every other variant redirects to it. Google follows at least five redirect hops, then treats the file as missing. A 200 that returns HTML usually means a catch-all route answered and your real rules never reach the crawler.
Then request the file the way an AI crawler would, with OpenAI's published GPTBot user agent.
curl -s -o /dev/null -w "%{http_code}\n" \
-A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" \
https://example.com/robots.txtA 403 or 429 here while your browser gets a 200 means a firewall acts on the name before robots.txt is read. Some bot rules also check the source IP, so your laptop can be refused where the real crawler gets in. Access logs settle it, and the AI Bot Log Analyzer shows which AI crawlers visited, which requests failed and which were impostors.
How to test robots.txt for GPTBot and other AI crawlers
Test the same path once per crawler token. A compliant crawler obeys only the most specific group that names it and ignores the * group, so an exception added for everyone can miss the bots you care about.
GitHub's file shows this. Its first group names five AI crawler tokens. Its * group, which Googlebot uses because there is no Googlebot group, opens with an exception for achievement pages that the AI group lacks.
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ClaudeBot
User-agent: anthropic-ai
User-agent: PerplexityBot
Crawl-delay: 1
Allow: /$
Allow: /pricing
# ...
Disallow: /*?tab=*
User-agent: Bytespider
Disallow: /
User-agent: *
Allow: /*?tab=achievements&achievement=*
# ...
Disallow: /*?tab=*Each crawler gets a different answer.
| Path | Googlebot | GPTBot | Bytespider |
|---|---|---|---|
/torvalds/linux |
Allowed | Allowed | Blocked by Disallow: / |
/octocat?tab=achievements&achievement=pull-shark |
Allowed by the longer Allow |
Blocked by Disallow: /*?tab=* |
Blocked |
/torvalds/linux/tree/master |
Blocked by Disallow: /*/tree/ |
Blocked by the same rule | Blocked |
Only GitHub knows whether that gap is deliberate. Copy each exception into every group that needs it and test every token. Our list of AI crawler user agents has the tokens.
Google-Extended is different. Google's crawler documentation says it has no user agent string of its own, so curl can't imitate it and your logs won't show it. It is a token Google reads to decide whether Gemini may train on or ground answers in your pages, and it doesn't affect Google Search. Test it as a rule match only.
Robots.txt validator checklist: what testing catches
Testing should catch each of these. The rules come from Google's robots.txt specification.
Rules on the wrong host or protocol
Robots.txt covers only the protocol, host and port it is served from. A file at https://www.example.com/robots.txt does nothing for https://example.com or shop.example.com. Every subdomain that serves pages needs its own file.
A 5xx on robots.txt
A server error on robots.txt is worse than having no file. Google reads a 4xx, other than 429, as no restrictions. A 5xx makes Google stop crawling the whole site for the first 12 hours while it retries. For the next 30 days it uses the last good copy. After that it acts as if there is no file when the rest of the site is up, and keeps crawling stopped when it isn't. A robots.txt generated by your app breaks with your app, so fix the underlying 500 error or serve the file statically.
Allow and Disallow rules that conflict
When both match, Google applies the rule with the longest path, and on a tie the less restrictive one. File order doesn't matter.
| URL | Matching rules | Google applies |
|---|---|---|
/page |
Allow: /p, Disallow: / |
Allow, the longer path |
/page.htm |
Allow: /page, Disallow: /*.htm |
Disallow, the longer path |
/page.php5 |
Allow: /page, Disallow: /*.ph |
Allow, a tie |
/ |
Allow: /$, Disallow: / |
Allow |
/page.htm |
Allow: /$, Disallow: / |
Disallow, since /$ matches only the root |
Wildcard surprises
* matches any run of characters, $ anchors the end of the URL, and paths are case-sensitive. These trip people up:
Disallow: /*.php$does not block/filename.php?id=1or/filename.php/.Disallow: /fishdoes not block/Fish.aspor/catfish.Disallow: /fish*is the same rule asDisallow: /fish.- Parameter position matters. GitHub pairs
Disallow: /*?tab=*withDisallow: /*&tab=*because the first only catchestabright after the?.
A CDN block on top of your file
When you turn it on, Cloudflare's managed robots.txt prepends its own block to the robots.txt your origin serves, so the live file differs from the one in your repo.
# BEGIN Cloudflare Managed content
User-Agent: *
Content-signal: search=yes, ai-train=no, use=reference
Allow: /
User-agent: GPTBot
Disallow: /
# ...
# END Cloudflare Managed ContentThe documented block disallows eight crawlers, GPTBot, ClaudeBot and Google-Extended among them, and leaves OAI-SearchBot and PerplexityBot alone. If you meant to welcome training crawlers, the live file now says otherwise. Cloudflare plans to replace this setting with Bot Preference Sync, which writes managed rules whenever an AI bot policy in the dashboard is set to block, so a dashboard change can still rewrite your file. Test the deployed URL, never the copy in Git.
A staging Disallow left on production
Disallow: / under User-agent: * is right for staging and a disaster on production. It ships when both environments build from one static file. Generate robots.txt from an environment variable and have the deploy fail on a sitewide block in production. Better still, put staging behind a login so it never needs the rule.
Test your robots.txt for search and AI crawlers
Our Robots.txt & AI Crawler Checker fetches the live file for the URL you enter and checks it three ways. The audit lists lines Google can't parse, Allow and Disallow lines with identical paths, sitewide blocks, Sitemap lines that aren't absolute URLs, files over 500 KiB, HTML served in place of the file and a 5xx that pauses crawling. The path test takes any path and user agent, Googlebot by default, and returns allowed or disallowed with the group that applied and the rule that won.
The AI section covers 14 crawler tokens, including GPTBot, ClaudeBot, PerplexityBot and Google-Extended. It separates training blocks from search blocks, reads any Content-Signal line and flags a 401, 403 or 429 on your page without posing as any crawler. Each blocked crawler comes with the rule that blocked it. It can't prove a crawler obeys, and a run costs 10 credits.