Excluded by noindex tag: find and fix the cause

·4 min read

"Excluded by noindex tag" means Google crawled the URL, found a noindex rule in the HTML or in an X-Robots-Tag response header, and left the page out of its index. If you meant to hide the page, nothing is wrong. If you didn't, find which system added the noindex, remove it at that source, then ask Google to recrawl with URL Inspection or Validate fix.

The hard part is the middle step. The place you check first is rarely the one that set it.

What excluded by 'noindex' tag means

The status sits in the Page indexing report under "Why pages aren't indexed". Google's Page indexing help documents it as "URL marked 'noindex'" and says Google hit a noindex directive when it tried to index the page. Its advice is short. If you don't want the page indexed, you're done. If you do, remove the tag or the HTTP header.

The status covers headers as well as tags, despite the name. It only appears when Google could fetch the page. A URL blocked by robots.txt lands under blocked by robots.txt instead, and its noindex is never read.

For the meta tag and header syntax, see noindex in robots.txt. This post is about finding who put it there.

Where an unexpected noindex comes from

Most surprise noindex rules come from a setting someone changed for a good reason and forgot.

Source What it looks like Where it lives
CMS visibility setting Every page carries noindex WordPress Settings > Reading, "Discourage search engines"
SEO plugin, per post One page or a handful The plugin's box on the post editor
SEO plugin, per type or taxonomy All tags, categories, authors or a custom post type The plugin's global search appearance settings
Server or CDN header X-Robots-Tag: noindex, often on a whole path or file type Web server config, CDN rules, edge functions
Staging leftover Whole site, right after a launch or migration Environment variables, a staging config pushed live
JavaScript Tag absent in source, present after render Tag manager, a client-side SEO component

WordPress shows how the first row works. When the "Discourage search engines" box is ticked, core adds a robots meta tag with noindex through wp_robots_noindex. The robots.txt side of that same checkbox is covered in robots.txt for WordPress.

The taxonomy row bites stores hardest. Someone noindexes tag archives to clean up thin pages, the shop uses tags for product collections, and a few hundred collection pages drop out of Google.

How to find which source set the noindex

Check the response headers, then the raw HTML, then the rendered HTML, in that order. Each check rules out a layer.

  1. Headers. Run the command below. If X-Robots-Tag shows noindex, the server, CDN or framework set it, and no plugin setting will change it.
  2. Raw HTML. Open view-source and search for "noindex". A hit in the <head> comes from the CMS, theme or plugin. Yoast and Rank Math both print an HTML comment with their name around the tags they write.
  3. Rendered HTML. If the source is clean, inspect the live DOM in DevTools or run URL Inspection's live test and view the tested page. A noindex that appears only there was added by JavaScript. The difference between those two views is explained in view rendered source vs raw HTML.
curl -sI -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" https://example.com/page/ | grep -i x-robots-tag

Send a Googlebot user agent, because some CDN and bot rules give crawlers different headers than browsers.

The JavaScript case needs a warning. Google's JavaScript SEO basics says that when Google finds noindex it may skip rendering, so using JavaScript to remove a noindex from the original HTML may not work. Keep noindex out of the server response for any page you want indexed.

How to fix it and validate the fix

Remove the noindex where it was set, confirm the live response is clean, then tell Google.

  1. Change the setting, config rule or component that added it, or the plugin will write it back.
  2. Purge the CDN and page cache, then rerun the header and source checks from the outside.
  3. In URL Inspection, run Test live URL. Under Indexing allowed, the noindex should be gone. Request indexing for the pages that matter most.
  4. For many URLs, open the issue in the Page indexing report and click Validate fix. Google's help says validation typically takes up to about two weeks and sometimes longer.

When excluded by noindex tag is correct

Leave the status alone when the URLs listed are pages you never wanted in search. Typical examples are cart and checkout pages, internal search results, thank-you pages, login screens and thin archive pages.

Open the report's example URLs. If every one is a page you meant to hide, the report is working and the count can grow without harm. Act only when a page you want ranking turns up.

Do take noindexed URLs out of your XML sitemap. Listing a page you also tell Google to drop sends a mixed signal.

Find the noindex source on your page

Our Indexing & Canonical Checker fetches the URL, follows redirects, and reads the meta robots tag and the X-Robots-Tag header separately. A header noindex is reported with the exact header value, and a meta noindex with the tag that applies to Googlebot, so you know which layer to fix. It also reads robots.txt for Googlebot, flags a noindex that a Disallow hides, and checks the canonical and its target.

A run costs 8 credits. It reads the HTML your server sends without running JavaScript, so a noindex added by script won't show up, and it can't confirm what Google has in its index. Use URL Inspection for both of those.

Keep reading