robots.txt for WordPress: the right setup
·8 min read
WordPress serves a virtual robots.txt that blocks /wp-admin/, allows admin-ajax.php and, since WordPress 5.5, points crawlers to your sitemap. For most sites that default is nearly right. Edit it with Yoast, Rank Math or All in One SEO, or upload a real robots.txt file to your web root, which replaces the virtual one. Block crawl traps such as WooCommerce's add-to-cart links, never block /wp-content/uploads/ or theme files, and decide on purpose which AI crawlers you let in.
The default WordPress robots.txt
WordPress builds robots.txt on each request, so there is no file to find in your site's root folder. Open https://yoursite.com/robots.txt on a fresh install and you get this:
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://example.com/wp-sitemap.xmlThe core function do_robots() writes the first three lines. User-agent: * addresses every crawler. Disallow: /wp-admin/ keeps bots out of the dashboard, which needs a login anyway. Allow: /wp-admin/admin-ajax.php reopens the one admin file that themes and plugins call from public pages for AJAX features. WordPress builds both paths from your admin URL, so a site installed in a subfolder shows /wp/wp-admin/ instead.
The Sitemap line came with WordPress 5.5, which added built-in XML sitemaps at /wp-sitemap.xml. Core adds the line only while the site is public. SEO plugins usually replace the core sitemap with their own.
How to edit robots.txt in WordPress
Pick one method and stay with it, because a physical file silently overrides the other three.
Yoast SEO
- Go to Yoast SEO, then Tools, then File editor.
- Click "Create robots.txt file" if you don't have one yet.
- Edit the rules and save.
Yoast writes a real file to your web root, so WordPress needs permission to write there. The File editor menu disappears when file editing is disabled in WordPress. In that case, use FTP.
Rank Math
Switch on Advanced Mode, then go to Rank Math SEO, General Settings, Edit robots.txt. Rank Math edits the virtual file, so its rules only take effect when no physical robots.txt exists. Delete that file first over FTP or your host's file manager.
All in One SEO
Go to All in One SEO, Tools, Robots.txt Editor and turn on Enable Custom Robots.txt. AIOSEO stores your rules in the database and adds them to the virtual output. It writes no file.
Upload a physical robots.txt file
Create a plain text file named robots.txt and upload it to the folder that answers for https://yoursite.com/, usually public_html. Apache and Nginx serve a file that exists on disk before they hand the request to WordPress, so once it is there, do_robots() never runs. That is also why Rank Math or AIOSEO edits sometimes seem to vanish. The plugin saved them, and an old file on the server wins.
A WordPress robots.txt example for most sites
For a blog or a brochure site, keep the default and point the Sitemap line at the sitemap you actually serve.
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://example.com/sitemap_index.xmlCore's sitemap lives at /wp-sitemap.xml. Yoast and Rank Math serve /sitemap_index.xml. Open the URL in a browser before you paste it, and check it again whenever you change SEO plugins.
What you leave out matters more. Skip Disallow: /?s= for internal search. Current WordPress sends a noindex robots meta tag on search result pages, and a robots.txt block would stop Google from ever seeing it. Core also marks ?replytocom= comment links noindex, so the old rule for those is dead weight.
Robots.txt for WooCommerce
WooCommerce needs a few more lines, because sort orders, filters and add-to-cart links turn each product list into endless URL variations.
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /*?*add-to-cart=
Disallow: /*?*orderby=
Disallow: /*?*filter_
Disallow: /*?*min_price=
Disallow: /*?*max_price=
Disallow: /*?*rating_filter=
Sitemap: https://example.com/sitemap_index.xmladd-to-cart=covers the links behind "Add to cart" buttons, such as/shop/?add-to-cart=123. Each crawl of one asks WooCommerce to add a product to a cart.orderby=covers the sort dropdown, which gives every category page one copy per sort order.filter_covers attribute filters such as?filter_color=red. Combined filters multiply into thousands of URLs.min_price=andmax_price=cover the price slider, andrating_filter=the star rating filter.
The /*?* prefix matches the parameter anywhere in the query string. It is the pattern Google uses in its faceted navigation guide, which recommends robots.txt for filter URLs you don't need indexed. If you built landing pages on filter URLs on purpose, drop the filter_ line.
Leave /cart/, /checkout/ and /my-account/ out of robots.txt. WooCommerce already sends noindex on those three pages. A Disallow would hide that tag, and since every product page links to the cart, Google could still index the bare URL.
What not to block in WordPress
Don't block anything a browser needs to draw the page. Google's robots.txt introduction warns against blocking resources when the page is harder to understand without them.
/wp-content/uploads/holds every image and PDF in your media library. Block it and Google can't crawl your images./wp-content/themes/and/wp-content/plugins/serve your CSS and JavaScript. Without them Google renders an unstyled, broken layout./wp-includes/ships jQuery and the block library styles that many themes load on every page./wp-admin/admin-ajax.phpkeeps its Allow line, or content loaded through it can fail to render for Google.
Robots.txt is also the wrong tool for hiding a page. Google can index a blocked URL that other pages link to, without a description. Use noindex instead, which every SEO plugin sets per page.
What "Discourage search engines" actually does
The checkbox under Settings, Reading adds a noindex, nofollow robots meta tag to every page and turns off the core XML sitemap. It no longer touches your robots.txt rules. WordPress dropped the old Disallow: / line in version 5.3 in favor of the meta tag.
So with the box ticked, robots.txt still lets every crawler in, minus the Sitemap line. Google has to fetch each page to see the noindex, and AI crawlers get no robots.txt instruction to stay away. WordPress's own settings screen says it is up to search engines to honor the request. Put staging sites behind a password instead.
On a live site, confirm the box is off after launch. The Indexability Checker checks a URL's robots.txt rules for Google and Bing, its meta robots tag and its X-Robots-Tag header.
How to allow or block AI crawlers in WordPress robots.txt
Add a group for each crawler's user agent token, after deciding which are training bots and which are search bots.
| Token | Operator | What it does | If you block it |
|---|---|---|---|
| GPTBot | OpenAI | Crawls content that may train OpenAI models | Out of training, still eligible for ChatGPT search |
| OAI-SearchBot | OpenAI | Fetches pages for ChatGPT search | Out of ChatGPT search answers, apart from navigational links |
| ClaudeBot | Anthropic | Crawls content for model training | Future content is excluded from training |
| Claude-SearchBot | Anthropic | Crawls to improve Claude's search results | Can reduce your visibility in Claude's answers |
| PerplexityBot | Perplexity | Indexes pages for Perplexity search, not training | Your pages can drop out of Perplexity results |
| Google-Extended | A control token with no crawler of its own, for Gemini training and grounding | No effect on Google Search inclusion or ranking |
Training bots feed future models. Search bots fetch pages so an assistant can quote and link them today. OpenAI's crawler docs treat GPTBot and OAI-SearchBot as independent settings, and Anthropic's crawler article splits its bots the same way. Our ClaudeBot guide covers Anthropic's three.
To block training and stay in AI search:
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://example.com/sitemap_index.xmlOAI-SearchBot, Claude-SearchBot and PerplexityBot have no group of their own here, so they follow * and stay allowed. To block every AI crawler, add them to the first group:
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: Google-Extended
Disallow: /If you sell anything, allow the search bots. Being fetchable is the first step to getting cited by ChatGPT and the other assistants. Blocking training bots is a licensing decision, and it costs little search visibility because OpenAI and Anthropic run separate bots for search.
Watch two catches. A crawler that finds a group naming it ignores the * group entirely, as Google's robots.txt spec spells out. Give OAI-SearchBot its own Allow: / group and it stops reading your WooCommerce rules, so copy any rule you still need into that group. And user-triggered fetchers like ChatGPT-User and Perplexity-User load a page because a person asked, and both OpenAI and Perplexity say robots.txt may not apply to them.
How to test your robots.txt
Test it from the outside, the way a crawler sees it.
- Open
https://yoursite.com/robots.txtin a private window. You should see plain text. An HTML page means crawlers find no rules at all. - Compare it with what you saved. If your edits are missing, look for a physical file in the web root, then purge your page cache and CDN.
- In Google Search Console, open the robots.txt report under Settings. It shows the version Google last fetched, when and any errors. After a change, request a recrawl there.
- Run URL Inspection on an important page to confirm Googlebot can fetch it.
Give changes time to land. Google caches robots.txt for up to 24 hours, and OpenAI says its search crawler takes about a day to adjust. Robots.txt is a request, so to see which AI bots actually visit, paste an access log into the AI Bot Log Analyzer.
Check your WordPress robots.txt for AI crawlers
Our Robots.txt and AI Crawler Checker fetches /robots.txt from your site and runs three checks in one pass. The audit flags misspelled fields, rules that come before any User-agent line, paths that don't start with / or *, Noindex lines Google ignores, groups that disallow the whole site, Sitemap URLs that aren't absolute and an HTML page served in place of the file. The path test checks one URL for one crawler, Googlebot unless you name another, and shows the group and the exact rule that decides it.
The AI check covers 14 crawler tokens, including every one in the table above. It reports training blocks apart from search blocks, reads Content-Signal lines, and notes when your server answers our fetch with 401, 403 or 429, a sign that a firewall may stop bots your robots.txt allows. It does not pretend to be those crawlers or prove they obey. A run costs 10 credits.