"Blocked by robots.txt" in Search Console: how to fix it

·7 min read

"Blocked by robots.txt" in Search Console means Googlebot found a URL, but a Disallow rule in your robots.txt told it not to crawl, so Google left the page out of its index. "Indexed, though blocked by robots.txt" means Google indexed the URL anyway, without reading it, because links pointed there. If you want the page in Google, remove or narrow the rule that matches it. If you want it gone, remove the rule and add a noindex tag, because Google cannot obey a noindex it is not allowed to read.

What does blocked by robots.txt mean?

It is a reason in the Page indexing report for URLs that Googlebot knows about but was told not to fetch. Google found the URL through a link or a sitemap, read your robots.txt, matched a Disallow and stopped. Google's Page indexing help lists it as "URL blocked by robots.txt" and adds that the block does not guarantee the page stays out of the index.

The status speaks for Googlebot alone. A Disallow aimed at GPTBot, ClaudeBot or Google-Extended never shows up here, so a clean report tells you nothing about AI crawlers.

What "Indexed, though blocked by robots.txt" means

It is a warning. Google indexed the URL without crawling it, because pages on your site or someone else's link to it. Google cannot read the page, so it knows the URL and the anchor text pointing at it and little else. In results, the snippet often reads "No information is available for this page."

Robots.txt controls crawling, not indexing, and this warning is what you get when a site uses it for the second job.

Should you fix every URL blocked by robots.txt?

No. Google's help says URLs blocked by robots.txt are probably blocked on purpose. Cart and checkout pages, internal search results, sorted and filtered copies of category pages and admin paths all belong here, and a count of 40,000 means nothing if that is all it holds.

Look for the exceptions instead:

  • Switch the filter above the chart from "All known pages" to "All submitted pages". A URL in your sitemap that robots.txt blocks is a contradiction you sent Google yourself, and each one needs a decision.
  • Scan the examples for pages you sell from or rank with, such as product, pricing, blog and docs URLs.
  • Read the chart. A step up on one date usually lines up with a deploy.

For "Indexed, though blocked", leave /search?q= pages and filter URLs alone. Unblocking thousands of them so Google can read a noindex on each costs crawling for very little. Fix it when a stranger should not see the URL in results, or when it ranks for your brand with a blank snippet.

How to find the robots.txt rule that blocks a URL

Work from the URL back to the line. Google's help still points to the robots.txt tester, which Google retired in 2023, so you need a few tools instead.

  1. In Search Console, open Indexing, then Pages, click the "Blocked by robots.txt" row and copy an example URL. The table shows up to 1,000 of them.
  2. Paste it into URL Inspection. "Crawl allowed?" reads "No: blocked by robots.txt". Click Test live URL to see whether that is still true today.
  3. Open Settings, then the robots.txt report. It shows the file Google last fetched for each host and when. Compare it with your source file. A difference means a CDN, a plugin or an old deploy is serving something else.
  4. Test the path against the live file. Googlebot follows the most specific group that names it, User-agent: Googlebot if one exists, otherwise *. Inside that group the longest matching rule wins, and Allow wins a tie.

URL Inspection tells you a URL is blocked, not which line did it, so step 4 is where the answer comes from. The robots.txt tester guide covers each testing method, including Google's open-source parser for checking a draft.

Common causes of a URL blocked by robots.txt

A staging Disallow: / shipped to production

Staging sites carry this file to stay out of Google:

User-agent: *
Disallow: /

Copy the staging build to production with the file inside, and Googlebot is refused on every URL it tries afterwards. The Blocked count climbs from the day of the deploy. Generate robots.txt per environment from a variable, or put staging behind a password so it needs no file at all.

A short prefix that catches more than you meant

Rules match from the start of the path, so a short prefix blocks every path that begins with it.

User-agent: *
Disallow: /p

Someone wanted to hide /p/12345 print views. This rule also blocks /pricing, /partners and /products/. Write Disallow: /p/ instead. Matching is case-sensitive too, so Disallow: /Admin leaves /admin open.

Query string patterns

User-agent: *
Disallow: /*?

This blocks every URL with a query string. That is fine when your real pages have clean paths. On a store whose product pages live at /item?id=4821, or a blog that paginates with ?page=2, it blocks the catalog. Name the parameters you mean:

User-agent: *
Disallow: /*?sort=
Disallow: /*&sort=
Disallow: /*?sessionid=

The robots.txt Disallow guide covers wildcards and the $ anchor in more depth.

A CMS or plugin setting

Check what your platform writes before you blame it. WordPress's "Discourage search engines from indexing this site" checkbox used to add Disallow: /, but since version 5.3 it adds a noindex meta tag instead, so those pages land under a noindex reason in the report. On WordPress the usual robots.txt culprit is a physical file in the web root, which overrides whatever you set in Yoast or Rank Math. Robots.txt for WordPress explains how the two files interact. On any platform, flip a visibility switch and then load /robots.txt to see what changed.

A robots.txt your CDN edits

Your CDN can add lines you never wrote. In August 2026 Cloudflare announced Bot Preference Sync, which customers on its managed robots.txt setting will move to. It turns the AI bot policies in your dashboard into robots.txt rules and prepends them to your existing file, and Cloudflare says it will be on by default for new customers. Those rules target AI bots, but the file crawlers receive is no longer the file in your repo. Read the top of the live one:

curl -s https://example.com/robots.txt | head -20

How to fix blocked by robots.txt when you want the page indexed

  1. Delete the rule or narrow it so the URL no longer matches. When the Disallow protects other paths, add a longer Allow instead of deleting it:

    User-agent: *
    Disallow: /account/
    Allow: /account/signup
  2. Deploy, purge the CDN cache for /robots.txt and fetch it to confirm the live file changed.

  3. In the robots.txt report, open the menu next to the file and choose Request a recrawl. Google caches robots.txt for up to 24 hours, and it suggests a recrawl request for exactly this case.

  4. Run Test live URL in URL Inspection. "Crawl allowed?" should now say Yes. Click Request indexing for the pages that matter most.

  5. Start validation in the Page indexing report, as described below.

How to fix indexed, though blocked by robots.txt

Decide what you want first, because the two fixes go in opposite directions.

If you want the page indexed properly, follow the steps above. Once Googlebot can crawl it, the URL leaves the warning and gets a real snippet.

If you want it out of Google, unblock it and tell Google not to index it. Add the tag to the page:

<meta name="robots" content="noindex">

For PDFs and other files without HTML, send the header instead:

HTTP/1.1 200 OK
Content-Type: application/pdf
X-Robots-Tag: noindex

Then remove the Disallow that covers the URL. Googlebot recrawls, sees the noindex and drops the page, which moves to a noindex reason in the report. Leave the Disallow off afterwards. Put it back and Google stops seeing the noindex, so the next link to that URL can return it to this warning. The guide to noindex vs robots.txt explains which one to use where.

Need it gone today? The Removals tool hides a URL for about six months while the noindex does the permanent work. For anything private, a password beats both.

How to validate the fix and how long it takes

Click Validate fix only after every URL you meant to unblock is fixed. Open the reason row in the Page indexing report and click Validate fix once.

Google checks a sample of URLs first. If the problem is still there, validation stops without changing state. Otherwise it moves from Started to Looking good to Passed, or to Failed if enough URLs still show the issue. Google says validation typically takes up to about two weeks and sometimes much longer. If you filtered the report to one sitemap, validation covers only the URLs in that sitemap.

Skip validation for URLs you kept blocked on purpose. And do not wait two weeks to learn whether a fix worked. The live test in URL Inspection answers in seconds, and validation only updates the report.

Find the blocking rule on your own site

Our Robots.txt & AI Crawler Checker fetches the live robots.txt for the URL you enter and runs three checks on it. The path test takes any path and any user agent, Googlebot by default, and shows the group and the rule that decide whether the path is allowed. The audit flags syntax errors, a sitewide Disallow, duplicate groups, Allow and Disallow rules with identical patterns, a missing * group, a Noindex line that Google ignores, a missing Sitemap line, and a robots.txt that serves an HTML page or cannot be reached. The third check reports which of 14 AI crawler tokens, including GPTBot, ClaudeBot and PerplexityBot, robots.txt allows.

It reads the file your server sends now, not the copy Google cached, and it cannot see your Search Console data, so use it next to URL Inspection rather than instead of it. A run costs 10 credits.

Keep reading