X-Robots-Tag: the HTTP header version of meta robots

·4 min read

X-Robots-Tag is an HTTP response header that carries the same indexing rules as the robots meta tag, such as noindex or nosnippet. Google accepts every rule in the X-Robots-Tag header that works in the meta tag, so you can control PDFs, images and whole hostnames without touching HTML. Because it lives in server or CDN config, it is also the noindex people forget they set.

The Apache and Nginx snippets for PDFs are in noindex in robots.txt. This post covers directives, crawler targeting, conflicts, and CDN setup.

X-Robots-Tag directives the header accepts

The header takes the full list from Google's robots meta tag documentation, as a comma-separated list. Header name, crawler name and values are all case-insensitive.

Directive What Google does
noindex Keeps the URL out of search results
nofollow Doesn't follow links on the page
none Same as noindex, nofollow
nosnippet No text snippet or video preview
max-snippet:[n] Caps the snippet at n characters. 0 means none, -1 no limit
max-image-preview:[setting] none, standard or large
max-video-preview:[n] Caps video previews at n seconds. -1 no limit
noimageindex Doesn't index images on the page
unavailable_after:[date] Drops the URL after that date. RFC 822, RFC 850 and ISO 8601 all work
indexifembedded Allows content embedded in an iframe to be indexed. Only works alongside noindex
notranslate No translated title link or snippet

Two rules you will still see in old configs, noarchive and nocache, do nothing in Google Search now. Delete them when you find them. The snippet rules have their own guide in nosnippet and max-snippet.

unavailable_after is the one the header handles better than HTML. A press release PDF or an event page can carry its own expiry:

X-Robots-Tag: unavailable_after: 2026-12-31T23:59:59+00:00

Targeting one crawler and combining X-Robots-Tag headers

Put a crawler name and a colon before the rules to scope them. Rules without a name apply to every crawler.

X-Robots-Tag: googlebot: noindex
X-Robots-Tag: max-image-preview:large

Googlebot gets noindex, and every crawler gets the large image preview. Several X-Robots-Tag headers in one response work the same as one comma-separated header.

When rules conflict, the more restrictive one applies. A response with X-Robots-Tag: all from your origin and X-Robots-Tag: noindex added by the CDN is a noindexed page. A scoped googlebot: index can't undo an unscoped noindex either, because Google adds up the restrictions. The header and the meta tag also combine, so a clean <meta name="robots"> doesn't help when a header says noindex.

How to set the X-Robots-Tag header on a CDN or host

You set the header in a rule that matches a path or hostname.

Cloudflare. In Rules, create a Response Header Transform Rule. Set the filter, for example a hostname or a URI path ending in .pdf, then pick Set static with header name X-Robots-Tag and value noindex. Set static replaces any X-Robots-Tag your origin sent. Add static appends another one, and then the most restrictive value wins. Cloudflare's response header modification docs explain both.

Vercel. Add a headers entry to vercel.json. The has condition can match the request's host, which matters for the staging setup below.

{
  "headers": [
    {
      "source": "/downloads/(.*)",
      "headers": [{ "key": "X-Robots-Tag", "value": "noindex" }]
    }
  ]
}

Netlify. Use a _headers file in the publish directory, or [[headers]] in netlify.toml. Netlify's headers docs note that these headers don't apply to proxied URLs or to responses from functions and edge functions. Those must send the header themselves.

/downloads/*
  X-Robots-Tag: noindex

X-Robots-Tag noindex for a whole staging host

A header on every response of the staging hostname keeps it out of Google without editing a single template. Scope it by host, never by path, or it ships to production with the next deploy.

On Vercel, match the staging host:

{
  "headers": [
    {
      "source": "/(.*)",
      "has": [{ "type": "host", "value": "staging.example.com" }],
      "headers": [{ "key": "X-Robots-Tag", "value": "noindex" }]
    }
  ]
}

Vercel already sends X-Robots-Tag: noindex on preview deployments, per its preview indexing guide. It drops that header when you assign a custom domain to a non-production branch, which is exactly how a staging. subdomain gets indexed. On Cloudflare, filter the transform rule on the staging hostname. Netlify's header files are global across deploy contexts, so give staging its own site, or copy a staging-only _headers file into the publish folder in that context's build command.

Noindex hides staging from search. It doesn't hide it from people. If staging holds anything private, put it behind a login.

How to check the X-Robots-Tag header with curl

Request the headers only and filter for the one you want:

curl -sI https://example.com/page/ | grep -i x-robots-tag

Check the final URL, not a redirect. Then test with a Googlebot user agent, because some CDN rules match on it, and test a few paths, because a rule written for /downloads/* can match more than you meant.

The usual accident is a CDN rule nobody remembers. A staging transform rule whose filter matched all hostnames, or a "block indexing" toggle left on after a launch, puts noindex on live pages while the HTML looks clean. Search Console then reports excluded by noindex tag and the CMS shows nothing wrong. If curl shows the header and your origin config doesn't set it, look at the CDN.

Google only reads the header on URLs it may crawl, so a robots.txt Disallow hides it.

Check a URL's X-Robots-Tag and indexing signals

Our Indexing & Canonical Checker fetches the URL, follows redirects and reads the X-Robots-Tag on the final response. When the header blocks indexing, it quotes the exact header value, so you can tell a CDN-added noindex from a meta tag. It applies crawler-scoped values, such as googlebot: noindex, only to the crawler they name, and reports snippet controls for Google and Bing separately.

The same run reads robots.txt for Googlebot, flags a noindex that a Disallow hides, and checks the canonical and its target. It costs 8 credits per URL. It doesn't render JavaScript and can't confirm whether Google has indexed the page. URL Inspection in Search Console does that.

Keep reading