Noindex in robots.txt: why it doesn't work and what to use
·7 min read
A Noindex: line in robots.txt does nothing in Google. Google stopped reading noindex in robots.txt on September 1, 2019, so a URL you "noindexed" that way is as indexable as any other page on your site. To keep a page out of search, leave it crawlable and put noindex in a meta robots tag or an X-Robots-Tag response header.
The costlier mistake usually comes next. Someone learns the robots.txt line is dead, adds a noindex tag to the page and keeps the Disallow to be safe. That pairing hides the tag from Google, and the URL can sit in results for months.
Why noindex in robots.txt stopped working
Google never documented noindex as a robots.txt rule, and in 2019 it switched off the code that still read it. On July 2, 2019, Gary Illyes posted A note on unsupported rules in robots.txt on the Google Search Central blog. From September 1, 2019, Google would retire its handling of rules it had never published, and the post named noindex, nofollow and crawl-delay.
The post went up a day after Google open-sourced its production robots.txt parser, as part of a push to make robots.txt a formal standard. That standard is RFC 9309, published in September 2022. It defines three rules, user-agent, allow and disallow. Crawlers may read other lines such as Sitemap, but nothing in the standard makes noindex one of them.
So this file, which still turns up on older sites, keeps nothing out of Google:
User-agent: *
Disallow: /internal/
Noindex: /old-landing-pages/The Disallow stops Googlebot from fetching /internal/. The Noindex line is ignored, and /old-landing-pages/ stays eligible for search. Our robots.txt checker flags lines like it as unsupported, along with Crawl-delay, which Google also ignores. For what the file can and can't do, start with what robots.txt is.
Noindex vs disallow: what each one controls
Disallow controls crawling and noindex controls indexing, and Google treats those as two separate decisions. A Disallow rule tells a crawler not to fetch a URL. A noindex rule tells a search engine not to show a URL it has fetched.
Links are what make the difference matter. If other pages link to a disallowed URL, Google can index the address without ever reading the page. Its robots.txt introduction says that result shows no description. Searchers get a bare URL, and you get no say in what it looks like.
Reach for Disallow when the problem is crawl volume, such as filter combinations or calendar pages that go on forever. Reach for noindex when the problem is a page showing up in search. For the rule syntax, see our guide to robots.txt Disallow.
The Disallow plus noindex trap
A noindex tag on a page that robots.txt blocks does nothing, because Google never fetches the page and never sees the tag. Google's guide to blocking indexing states the condition directly. For noindex to work, robots.txt must not block the page.
Picture a staging copy at /beta/ that picks up a few links and gets indexed. Someone adds Disallow: /beta/ and a noindex tag on the same afternoon. Googlebot stops fetching /beta/, never reads the noindex, and the indexed URLs stay put as results with no description. Search Console files them under "Indexed, though blocked by robots.txt". If your problem is the reverse and pages you want indexed are stuck behind a rule, read blocked by robots.txt.
Headers fall into the same hole. Google's robots meta tag documentation says a crawler blocked by robots.txt never finds the page's indexing rules, so an X-Robots-Tag on a blocked PDF is invisible too.
What to use instead of robots.txt noindex
Pick the method by the goal, because each one solves a different problem.
| Goal | Use | Notes |
|---|---|---|
| Keep an HTML page out of search | Meta robots noindex |
The page must stay crawlable |
| Keep a PDF, image or video out of search | X-Robots-Tag: noindex header |
Set it in the server or CDN config |
| Keep content away from everyone | Password or login | Stops crawlers and people |
| Remove a page that is gone for good | 404 or 410 status | Google drops it after it recrawls |
| Hide a URL from Google today | Search Console Removals tool | Lasts about six months, so pair it with a row above |
| Cut crawling of URLs nobody searches for | robots.txt Disallow | Controls crawling, not indexing |
For staging sites, the password beats every other row. It keeps out Google, AI crawlers and anyone who guesses the URL. And if someone forgets to remove it at launch, the whole team finds out within minutes, which never happens with a leftover noindex.
The Removals tool gets misused the most. Google's Removals help page says a request lasts about six months and that the tool alone won't remove a page for good. It hides the URL while noindex, a 404 or a password does the permanent work.
If you want a page indexed but not quoted in snippets or AI answers, none of these apply. That is a job for nosnippet.
Meta robots noindex for HTML pages
A meta robots noindex tag in the page's <head> is the standard way to keep an HTML page out of search engines.
<meta name="robots" content="noindex">The robots name applies to every search engine that reads the tag. Use name="googlebot" to address only Google.
Check what Google received rather than what your CMS settings say. Themes, SEO plugins and tag managers can each write a robots tag, and when two of them disagree, Google's documentation says the more restrictive rule applies. A stray noindex from a theme beats the "index" your plugin sets. URL Inspection in Search Console shows the HTML Googlebot fetched.
X-Robots-Tag noindex for PDFs and other files
The X-Robots-Tag header does the meta tag's job for files that have no HTML head, such as PDFs, images and video. Google accepts any rule in the header that works in the meta tag.
Apache, with mod_headers enabled, in the site config or .htaccess:
<FilesMatch "\.pdf$">
Header set X-Robots-Tag "noindex"
</FilesMatch>Nginx, inside the server block:
location ~* \.pdf$ {
add_header X-Robots-Tag "noindex";
}Nginx has one catch. A location block with its own add_header stops inheriting the add_header lines from the server level, so security headers set there disappear for those PDFs unless you repeat them inside the block. On Nginx 1.29.3 or later, add_header_inherit merge; in the block keeps them (Nginx docs).
Then confirm the header reaches the outside world, since a CDN or cache in front of the server can drop it:
curl -sI https://example.com/files/price-list.pdf | grep -i x-robots-tagHow to deindex a page that robots.txt blocks
Unblock the page first, then noindex it, then wait for Google to see the tag. Doing it in any other order leaves the page stuck.
- Remove the Disallow rule that matches the URL, or add a more specific
Allowfor it. Test the path for Googlebot afterward, since a broader rule elsewhere in the file can still match. - Add noindex with a meta tag or an
X-Robots-Tagheader, and confirm it is in the live response. - Get Google to recrawl. For a few URLs, use URL Inspection and request indexing. For hundreds, keep them in your XML sitemap until they drop out. Google's guide to blocking indexing warns that a revisit can take months.
- If the pages must disappear now, file a Removals request to cover the wait.
- Watch the Page indexing report until the URLs move from indexed to excluded by noindex. Then take them out of the sitemap.
- Put the Disallow back only if crawl load is a real problem.
That last step is where I'd push back. With the Disallow back, Google can't fetch the page, so it can't see the noindex anymore. If links to the URL remain, Google can index it again as a bare address, and you are back under "Indexed, though blocked by robots.txt". A noindexed page costs an occasional crawl, which most sites never notice. Leave the noindex in place and skip the Disallow, unless you run millions of parameter URLs where crawl volume is a measurable cost.
Check whether a page can be indexed
Our Indexing & Canonical Checker reads the signals that decide whether Google may index a URL. It follows redirects to the final URL and reads the HTTP status. It fetches robots.txt for that origin and finds the rule that matches the path for Googlebot. It reads the meta robots tag and the X-Robots-Tag header, loads the canonical target to check that it is indexable, and reads snippet controls such as nosnippet and max-snippet for Google and Bing.
When a page carries noindex but robots.txt blocks Googlebot, the report says Google can never see the noindex and tells you to allow crawling. For every signal that blocks the page, it names the tag, header or rule to change. A run costs 8 credits. It reads the HTML your server sends without running JavaScript, and it can't confirm that Google has the URL in its index, which is what URL Inspection is for.