robots.txt Disallow: syntax, examples and common mistakes
·7 min read
A robots.txt Disallow rule tells the crawlers named in its group not to fetch any URL whose path starts with the rule's value. Disallow: / blocks every URL on the host, an empty Disallow: blocks nothing, and Disallow: /account/ blocks one folder, inside which a longer Allow rule can reopen a single page. Disallow controls crawling only, so a blocked URL can still show up in Google when other pages link to it.
Everything below follows Google's robots.txt specification and RFC 9309, the IETF standard. For the basics, such as groups, the * and $ wildcards and how crawlers settle a conflict, start with what robots.txt is and how it works. This post covers the Disallow line itself and the ways it goes wrong on real sites.
Robots.txt disallow all vs an empty Disallow
Disallow: / and Disallow: differ by one character and mean opposite things. Every URL path starts with /, so a rule of / is a prefix of every URL on the host and blocks the whole site. A Disallow with no value matches nothing. Google's spec says crawlers ignore a rule without a path, and RFC 9309 says a URL that matches no rule is allowed.
# Disallow all: blocks every URL for crawlers without a group of their own
User-agent: *
Disallow: /# Allow all: same effect as having no robots.txt
User-agent: *
Disallow:I'd avoid the empty form. It looks like a line someone forgot to finish, and the obvious "fix" blocks the whole site. If you mean to allow everything, write Allow: / or put a comment above the empty line saying it is deliberate.
Two more readings catch people out. Disallow: / does not mean "block the home page". To block only the root URL, anchor it with Disallow: /$. And Disallow: /* is the same rule as Disallow: /, because Google ignores a trailing wildcard.
A disallow all for a single crawler works the same way under that crawler's own User-agent line. That is how sites turn away one AI training bot, and our guide to blocking AI crawlers covers which ones are worth it.
How to keep a staging disallow all off production
Generate robots.txt per environment, and make the production deploy fail when the live file disallows everything under User-agent: *. The usual accident is a single static robots.txt in the repo, set to Disallow: / while the site was in staging, that ships with the launch or the next release. Nothing breaks visibly. Pages load fine for people, Googlebot stops fetching them, and the first sign is a climbing "Blocked by robots.txt" count in Search Console days later.
This script fetches the live file and exits with an error when a User-agent: * group contains Disallow: / or Disallow: /*. Run it as the last step of every production deploy, and on a schedule in case someone edits the file by hand.
#!/usr/bin/env bash
# Fails when the live robots.txt disallows the whole site for User-agent: *
set -euo pipefail
site="${1:?usage: check-robots.sh https://www.example.com}"
curl -fsSL "${site%/}/robots.txt" | awk '
{ sub(/\r$/, ""); sub(/#.*/, "") }
tolower($0) ~ /^[ \t]*user-agent[ \t]*:/ {
if (rules) { star = 0; rules = 0 }
agent = $0
sub(/^[^:]*:[ \t]*/, "", agent)
sub(/[ \t]+$/, "", agent)
if (agent == "*") star = 1
next
}
tolower($0) ~ /^[ \t]*(allow|disallow)[ \t]*:/ {
rules = 1
if (star && tolower($0) ~ /^[ \t]*disallow[ \t]*:[ \t]*\/\*?[ \t]*$/) blocked = 1
}
END {
if (blocked) { print "robots.txt disallows the whole site for User-agent: *"; exit 1 }
}
'curl -f also fails the step when robots.txt returns an error status. You want that too, because Google pauses crawling of a site whose robots.txt answers with a server error.
The script catches the mistake after it ships. To stop it shipping, build the file from the environment. In Next.js, app/robots.ts can return different rules per environment:
import type { MetadataRoute } from "next";
export default function robots(): MetadataRoute.Robots {
if (process.env.SITE_ENV === "staging") {
return { rules: { userAgent: "*", disallow: "/" } };
}
return {
rules: { userAgent: "*", disallow: ["/account/", "/cart/"] },
sitemap: "https://www.example.com/sitemap.xml",
};
}The test is === "staging" on purpose, so a missing variable leaves production open instead of blocking it. The Next.js docs say robots.ts is cached by default unless it uses a request-time API, which means the variable is read when the build runs. If you build once and promote that build from staging to production, staging's rules go with it. Put staging behind a password as well, so the Disallow is a backstop and not the only lock.
Allow vs disallow: open one page inside a blocked folder
Add an Allow rule with a longer path to the same group, and it wins for the URLs it matches. Say the account area should stay out of search but the signup page should rank:
User-agent: *
Disallow: /account/
Allow: /account/signup/account/signup and /account/signup?plan=pro are now crawlable, while /account/settings and /account/billing/invoices stay blocked. Allow paths are prefixes too, so this also opens /account/signup-complete and anything under /account/signup/. To open exactly one URL, end the rule with $:
User-agent: *
Disallow: /account/
Allow: /account/signup$Now /account/signup?plan=pro is blocked again, because $ marks the end of the URL and the query string is part of the URL. Pick the version that fits how people link to the page. If your ads point at the signup page with tracking parameters, the anchored rule blocks every one of those URLs.
The carve-out only works inside the group that holds the Disallow. Put it in a separate User-agent: Googlebot group and Googlebot follows that group alone, ignoring your * rules, so the whole account area opens to it.
What the trailing slash changes in a Disallow rule
Without a trailing slash, a rule blocks every path that starts with those characters. With one, it blocks only what sits inside the folder. Google's spec shows the same split with /fish and /fish/.
| URL | Disallow: /blog |
Disallow: /blog/ |
|---|---|---|
/blog |
Blocked | Allowed |
/blog/ |
Blocked | Blocked |
/blog/robots-txt-guide |
Blocked | Blocked |
/blog-post-template |
Blocked | Allowed |
/blogroll |
Blocked | Allowed |
/blog.html |
Blocked | Allowed |
People usually mean the folder. The slashless version is the one that quietly blocks a /blog-post-template landing page nobody meant to touch. If you want the folder and its bare URL but nothing else that shares the prefix, spell out each case:
User-agent: *
Disallow: /blog/
Disallow: /blog$
Disallow: /blog?How to disallow URLs with query parameters
Rules match the path and query string together, so Disallow: /*?sort= blocks any URL whose query string starts with sort=. The ? is a literal character and the * covers whatever path comes before it. A parameter in second place follows an & instead, so /shoes?color=red&sort=price gets through unless you also add Disallow: /*&sort=.
The harder question is which parameters deserve a Disallow. Block the ones that multiply URLs without changing the content, such as sort orders, view modes, session IDs and stacked filters. Those can turn one category page into thousands of crawlable variants.
Leave tracking parameters like ?utm_source= alone. They arrive in links from newsletters and social posts, and Google's duplicate URL guide says not to use robots.txt for canonicalization, because a disallowed URL can still be indexed without its content. Google can't read the canonical tag on a page it may not fetch. Let those URLs be crawled and let rel="canonical" point Google at the clean version.
Don't disallow the CSS and JavaScript your pages need
Leave every stylesheet, script and image that a page needs to render open to Googlebot. Google's robots.txt introduction says you can block resource files only when pages load fine without them, and not when their absence makes a page harder for Google to understand. Few sites block their CSS on purpose. It happens through rules aimed at something else:
Disallow: /_next/on a Next.js site blocks/_next/static/, where the framework serves its scripts and stylesheets.Disallow: /*?on WordPress blocks stylesheets and scripts too, because WordPress adds?ver=to the files it enqueues.Disallow: /assets/on a site built with Vite blocks its scripts and stylesheets, because Vite writes them to/assets/by default.
If a broad rule has to stay, carve the file types back out. Longer rules win, so these Allow lines beat Disallow: /*?:
User-agent: *
Disallow: /*?
Allow: /*.css
Allow: /*.jsAllow: /*.css matches /wp-content/themes/site/style.css?ver=6.6, because the pattern only needs .css somewhere in the URL.
Disallow vs noindex: why a blocked page can stay in Google
Disallow stops crawling, and it does not remove a URL from Google's index. Google's robots.txt introduction says it can still index a disallowed URL that other pages link to and show it in results with no description. Adding noindex to that page does nothing while the Disallow stays, since Google's guide to blocking indexing says noindex only works when robots.txt doesn't block the page. The crawler never fetches the page, so it never sees the tag.
If you want a page out of Google, remove the Disallow first and let Googlebot read the noindex. Noindex in robots.txt walks through the deindexing steps, and blocked by robots.txt explains the Search Console statuses each mistake produces.
How long a Disallow change takes to work
Expect up to a day. Google's spec says it generally caches robots.txt for up to 24 hours, and RFC 9309 tells crawlers not to use a cached copy for more than 24 hours unless the file is unreachable. A new Disallow can let Googlebot through for most of a day, and deleting an accidental Disallow: / won't bring it back within the hour.
Google may also lengthen or shorten that cache based on the Cache-Control: max-age header your server sends with robots.txt, so keep it short. Check it with curl:
curl -sI https://www.example.com/robots.txt | grep -i cache-controlWhen a fix is urgent, open the robots.txt report in Search Console and use Request a recrawl, which Google's robots.txt update guide recommends when the cache needs refreshing sooner. That refreshes the rules only. Google still has to come back and crawl each URL the old rule blocked.
Check your robots.txt Disallow rules
Our Robots.txt and AI Crawler Checker fetches the live robots.txt for the host of the URL you enter, so it reads the file your server and CDN actually serve, not the copy in your repo. The audit flags any group that disallows the whole site with Disallow: /, /* or *, an Allow and a Disallow with the same pattern, rules placed before any User-agent line, and paths that start with neither / nor *, such as Disallow: blog, which never match anything.
The path test checks the path of the URL you enter, query string included, for Googlebot unless you name another crawler, and shows which group matched and the winning rule. Enter https://www.example.com/shoes?color=red&sort=price to see whether your parameter rules catch it, or a stylesheet URL to confirm it is open. The AI check reports access for 14 AI crawler tokens, marking a blocked search or user-triggered crawler as a warning and a blocked training crawler as an info note. A run costs 10 credits.