What is robots.txt? A plain guide with examples

·7 min read

Robots.txt is a plain text file at the root of a website, such as https://example.com/robots.txt, that tells crawlers what they may fetch. It holds groups of rules, and each group names a crawler with a User-agent line and then lists paths with Disallow and Allow. Googlebot, Bingbot, GPTBot and other well-behaved bots read it before they crawl. It controls crawling only, so it neither locks a page nor removes one from Google's index.

Martijn Koster defined the format in 1994, and the IETF standardized it as RFC 9309 in September 2022. Google's robots.txt specification adds the details of how its own crawlers read it.

What is robots.txt and where does the file go?

A robots.txt file lives at /robots.txt on one host and covers only that host, protocol and port. Crawlers look nowhere else, so a file at https://example.com/blog/robots.txt does nothing.

URL Covered by https://example.com/robots.txt?
https://example.com/pricing Yes
http://example.com/pricing No, different protocol
https://www.example.com/pricing No, different host
https://shop.example.com/ No, a subdomain needs its own file
https://example.com:8443/ No, different port

Most sites redirect every variant to one host, so you maintain one robots.txt file, plus one per subdomain such as a shop or docs site. Name it robots.txt in lowercase and save it as UTF-8 plain text.

Robots.txt syntax: user-agent, disallow, allow and sitemap

Google supports four fields and ignores every other line. Each line is a field, a colon and a value, and # starts a comment.

Field What it does Example
User-agent Starts a group and names the crawler it applies to. * means any crawler. User-agent: Googlebot
Disallow A path the crawler should not fetch. An empty value blocks nothing. Disallow: /admin/
Allow A path the crawler may fetch inside a disallowed one. Allow: /admin/help/
Sitemap The full URL of an XML sitemap. It belongs to no group. Sitemap: https://example.com/sitemap.xml

A minimal user-agent and disallow file looks like this:

# Rules for every crawler without a group of its own
User-agent: *
Disallow: /admin/
Allow: /admin/help/

Sitemap: https://example.com/sitemap.xml

Paths are case-sensitive, so Disallow: /Admin/ leaves /admin/ open. Field names and crawler names are not. Every path is a prefix, so Disallow: /admin also blocks /admin-guide, while Disallow: /admin/ blocks only the folder. Several User-agent lines in a row share the rules under them.

Google ignores Crawl-delay and Noindex lines. Anthropic and Common Crawl say their bots honor Crawl-delay, but it never slows Googlebot. A Sitemap line takes an absolute URL and can repeat.

Wildcards: * and $

* matches any run of characters, including none, and $ marks the end of the URL. Matching includes the query string.

Rule Matches Doesn't match
Disallow: /*.pdf$ /guides/robots.pdf /guides/robots.pdf?download=1
Disallow: /*? /shoes?color=red /shoes
Allow: /$ / Every other URL

A trailing * does nothing. /private* is the same rule as /private, because every rule is already a prefix.

How crawlers choose a user-agent group

A crawler follows exactly one group, the one that names it most specifically, and ignores the others. If no group names it, it falls back to User-agent: *. If there is no * group either, nothing is blocked for it.

This catches people out:

User-agent: *
Disallow: /admin/
Disallow: /search?

User-agent: Googlebot
Disallow: /drafts/

Googlebot reads only its own group here, so it may crawl /admin/ and /search?q=shoes. Google never combines a named group with the * group, so copy every general rule you still want into each named group. Two groups that name the same crawler do get merged.

How crawlers settle allow and disallow conflicts

The longest matching rule wins, counted in characters of the rule's path. When an Allow and a Disallow match with the same length, Allow wins. Line order doesn't matter to Google or any RFC 9309 parser.

User-agent: *
Disallow: /blog/
Allow: /blog/robots-txt-guide

Both rules match /blog/robots-txt-guide. The Allow path is 22 characters and the Disallow path 6, so the page stays open. /blog/another-post stays blocked.

Wildcards count toward the length. In Google's own example, Allow: /page and Disallow: /*.htm both match /page.htm, and the Disallow wins because /*.htm is six characters against five.

What robots.txt does not do

Robots.txt asks crawlers not to fetch URLs. It doesn't protect them or take them out of search results.

It isn't access control. RFC 9309 says its rules "are not a form of access authorization." Anyone can open a disallowed URL, and a public list of private paths shows people where to look. Put private areas behind a login.

It doesn't remove pages from Google. Google can still index a blocked URL that other pages link to and show it with no description, as its robots.txt introduction explains. To keep a page out, add a noindex meta tag or X-Robots-Tag header and leave the page crawlable. Google must fetch the page to read the tag, so a Disallow hides the tag. A Noindex line in robots.txt does nothing.

Not every bot obeys it. Google says respectable crawlers follow it and others might not. Stop those at your server or CDN.

What happens when robots.txt is missing, too big or down

A missing robots.txt is fine, because Google treats a 404 as permission to crawl everything. A server error is the case that hurts.

Response for /robots.txt What Google does
2xx Applies the rules in the file
3xx Follows at least five redirects, then treats it as a 404
4xx other than 429, including 401, 403 and 404 Crawls as if there were no robots.txt, so nothing is blocked
5xx, timeout or network error Stops crawling the site for 12 hours, then uses the last good copy for up to 30 days
Still failing after 30 days Acts as if there is no robots.txt if the site is otherwise up, or stops crawling if it isn't

A 403 on robots.txt, say from a firewall rule, tells Google you have no rules at all. A robots.txt that returns 500 errors pauses Google's crawling of the whole site, while RFC 9309 tells other crawlers to assume a complete disallow. Serve it as a static file your CDN can answer even when the app is down.

Google reads the first 500 KiB of the file and ignores the rest. If yours nears that, you are probably listing URLs one by one where a wildcard pattern would do. Google also caches robots.txt for up to 24 hours, so a fix can take a day to reach Googlebot.

Robots.txt examples for common sites

Each robots.txt example below is a complete file. Change the paths and the sitemap URL, then test the result before you publish, because one wrong character can block a whole folder. Our guide to testing robots.txt shows how.

Allow everything

User-agent: *
Disallow:

Sitemap: https://www.example.com/sitemap.xml

An empty Disallow blocks nothing, so this file works like having none, plus a sitemap pointer. If your site has no crawl traps, stop here.

A typical business site or blog

User-agent: *
Disallow: /admin/
Disallow: /search?

Sitemap: https://www.example.com/sitemap.xml

/search? catches internal search results such as /search?q=pricing, which can spawn endless thin pages, and leaves /search-tips alone. Keep CSS, JavaScript and image folders open, because Google renders pages and a blocked stylesheet or script can change what it sees. WordPress builds its own default file, which our WordPress robots.txt guide walks through.

A staging site

# Served only on staging.example.com
User-agent: *
Disallow: /

Disallow: / asks every compliant crawler to stay off the staging host, and since rules are per host, it doesn't touch www.example.com. Treat it as a backstop. Anyone with the link can still load staging, and a leaked staging URL can still appear in Google without a description. Put staging behind a password, and check after each deploy that this file didn't ship to production.

An online store

User-agent: *
Disallow: /search?
Disallow: /*?*sort=
Disallow: /*?*color=
Disallow: /*?*size=
Disallow: /*?*price=

Sitemap: https://shop.example.com/sitemap.xml

Sort orders and filters turn one category page into thousands of URL variants, and Google's faceted navigation guide recommends disallowing filter URLs. The /*?* prefix matches the parameter anywhere in the query string, so /shoes?size=9&color=red is blocked while /shoes and every product page stay open. If you built a landing page on a filter URL on purpose, give it a longer Allow rule such as Allow: /shoes?color=red$.

AI companies run separate bots for training and for search, so you can turn one away and keep the other.

# Training crawlers and control tokens
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
Disallow: /

# Everyone else, including OAI-SearchBot, Claude-SearchBot and PerplexityBot
User-agent: *
Disallow: /admin/

Sitemap: https://www.example.com/sitemap.xml

GPTBot and ClaudeBot collect pages that may train OpenAI's and Anthropic's models, and CCBot builds Common Crawl's public web archive. OpenAI's crawler docs say each of its bot settings is independent, so blocking GPTBot leaves you eligible for ChatGPT search through OAI-SearchBot. Google-Extended is a control token with no crawler of its own. It tells Google not to use your pages for Gemini training and grounding, and Google's crawler documentation says it has no effect on inclusion or ranking in Google Search. It won't keep you out of AI Overviews either, which is a separate setting in Search Console.

Our AI crawler list covers every token and what each one does, and how to block AI crawlers weighs the cost of each block.

Check your robots.txt file

Our Robots.txt and AI Crawler Checker fetches /robots.txt from your site, parses it by RFC 9309 and Google's rules, and runs three checks in one pass. The audit finds misspelled fields, rules before any User-agent line, paths that don't start with / or *, and Noindex or Crawl-delay lines Google ignores. It also flags groups that disallow the whole site, an Allow and a Disallow with the same pattern, a crawler named in more than one group, a missing * group, Sitemap URLs that aren't absolute, a file over 500 KiB and an HTML page served in its place. If the fetch fails, it tells you what Google does with that status.

The path test takes one URL and one crawler, Googlebot unless you name another, and shows the group and the exact rule that decide it. The AI check reports which of 14 AI crawler tokens your rules allow, including GPTBot, ClaudeBot, PerplexityBot and Google-Extended, with training bots reported apart from search bots. It reads your rules without pretending to be those crawlers, so it can't prove a bot obeys them. A run costs 10 credits.

Keep reading