How to add your sitemap to robots.txt

·7 min read

To add your sitemap to robots.txt, put one line anywhere in the file with the sitemap's full URL, such as Sitemap: https://www.example.com/sitemap.xml. The robots.txt sitemap line belongs to no user-agent group, so every crawler that reads the file sees it. Add one line per sitemap, or a single line for your sitemap index. Then submit the same URL in Google Search Console and Bing Webmaster Tools, because Google's Sitemaps report only tracks sitemaps you submit there.

Writing the line takes a minute. It breaks months later, when an SEO plugin swap moves the sitemap or someone ships the staging robots.txt to production. Nothing visibly fails, so nobody opens the file again.

Robots.txt sitemap syntax

A Sitemap line is the field name, a colon and the absolute URL of a sitemap or sitemap index. A complete file with an index and a video sitemap looks like this:

User-agent: *
Disallow: /checkout/
Disallow: /search

Sitemap: https://www.example.com/sitemap_index.xml
Sitemap: https://www.example.com/video-sitemap.xml

The sitemaps.org protocol and Google's robots.txt specification set the rules.

  • The URL must be absolute, with protocol and host. Sitemap: /sitemap.xml is not valid.
  • Google treats the field name as case-insensitive, so sitemap: works too. The URL itself is case-sensitive.
  • You can repeat the line. Google sets no limit on how many.
  • A sitemap index needs only its own line, not one for every sitemap inside it.

That last rule decides how I write the file. If your CMS splits sitemaps by post type or by month, list the index. Then robots.txt never changes when a new child sitemap appears.

Where the sitemap directive goes in the file

The sitemap directive can go anywhere, because it sits outside every user-agent group. sitemaps.org says it is independent of the user-agent line, and Google says any crawler may follow it unless the sitemap URL itself is disallowed.

So you can't scope a sitemap to one crawler. A Sitemap line placed under User-agent: Googlebot still applies to every crawler that supports the field. Shopify's default template prints the sitemap with each user-agent group, so one file can repeat the same line several times. The copies are harmless.

RFC 9309, the robots.txt standard, doesn't define Sitemap at all. It names sitemaps as an example of other records a crawler may read, as long as reading them doesn't interfere with the allow and disallow rules.

I put Sitemap lines at the bottom after a blank line, where nobody editing the last group mistakes them for part of it. Also check that no Disallow rule covers the sitemap's path, or Search Console reports "Couldn't fetch", which our guide to checking your XML sitemap covers.

Can the sitemap live on another host?

Yes. Pointing to it from the site's robots.txt is one way to prove you own it. sitemaps.org normally limits a sitemap to URLs on its own host and protocol, below its own folder. The cross-submit rule is the exception. When the robots.txt on www.example.com points to a sitemap on another host, that counts as the owner vouching for it, so the sitemap may list www.example.com URLs. Google's spec agrees the sitemap doesn't have to share a host with robots.txt.

Headless and static sites that write sitemaps to object storage use this:

# https://www.example.com/robots.txt
User-agent: *
Disallow: /api/

Sitemap: https://static.example-assets.com/sitemaps/www-example-com.xml

Keep one host's URLs per sitemap and reference each site's sitemap from its own robots.txt, which is one of two setups in Google's sitemap guide. The other is verifying every site in Search Console. Still, serve the sitemap from the site root if you can. Google recommends the root, since a sitemap you don't submit in Search Console only covers URLs below its own folder.

Sitemap in robots.txt on WordPress and hosted platforms

Most platforms write the Sitemap line for you, so your job is to confirm it names the sitemap you actually serve.

WordPress 5.5 added core sitemaps at /wp-sitemap.xml and a reference to them in its virtual robots.txt, per the release announcement. SEO plugins can replace the core sitemap. Yoast does, and the robots.txt that Yoast creates points to /sitemap_index.xml instead. After any plugin change, open /robots.txt and click the URL. Our WordPress robots.txt guide covers each plugin.

Shopify generates sitemap.xml at the root of each store domain and prints the Sitemap line from its robots.txt.liquid template, which you can edit to add another sitemap. Each international domain gets its own sitemap to submit. On Wix, the robots.txt file includes your sitemap URL by default, and the editor has a Reset to Default option.

A CDN can change the file too. Cloudflare's managed robots.txt prepends its own AI crawler rules, so the live file differs from your repo copy. Whatever builds the file, read the live one:

curl -s https://www.example.com/robots.txt | grep -in sitemap

Which crawlers read the Sitemap line?

Search engines do. Google's spec says Google, Bing and other major search engines support the field, and Bing's July 2025 post on sitemaps in AI-powered search asks site owners to list the sitemap in robots.txt for automatic discovery.

AI crawlers make no such promise. The crawler docs from OpenAI, Anthropic and Perplexity explain how to allow or block their bots in robots.txt and never mention the Sitemap line. That doesn't mean GPTBot ignores it. It means you can't count on it. Your server logs show which bots fetch your sitemap URLs, and our log file analysis guide shows how to pull those requests out.

Should you still submit your sitemap?

Yes. The robots.txt line gets the sitemap discovered. Submitting it gets you a report.

Google picks up a robots.txt sitemap the next time it crawls robots.txt, but the Sitemaps report lists only sitemaps submitted through the report or the API. Without a row you see no status, no last read date and no discovered page count, so a sitemap that starts failing with "Couldn't fetch" fails quietly. You can submit one Google already found.

Bing's post gives the same reason. Bing fetches a submitted sitemap right away and then rechecks it, typically at least once a day, and Webmaster Tools shows the status, last read date and processing errors.

Don't rely on pinging. Google announced in June 2023 that it was retiring the sitemap ping endpoint, and pings now return 404. None of the remaining routes is a guarantee either, because Google calls a submitted sitemap only a hint.

Common robots.txt sitemap mistakes

Most broken Sitemap lines fit one of these patterns, and each one survives a quick look at the file.

Mistake What it looks like Fix
Relative URL Sitemap: /sitemap.xml Write the full URL with https:// and the host
Old protocol or host http://example.com/sitemap.xml on an https://www site Use the URL the sitemap finally loads from
URL redirects /sitemap.xml returns a 301 to /sitemap_index.xml Point the line at the final URL
URL returns 404 Still names /wp-sitemap.xml after a plugin moved it Update the line and resubmit
Staging host Sitemap: https://staging.example.com/sitemap.xml Generate robots.txt per environment
Line glued to a rule Disallow: /tmp/Sitemap: https://... Fix the line break in the template

A relative URL breaks both specs, and since RFC 9309 doesn't define the field, no standard tells a crawler how to resolve a bare path. Some may guess. Don't make them.

The http or non-www URL usually still loads because it redirects, which makes it the next row. Every reader pays an extra request, and the wrong address stays in your file. Load the sitemap in a browser and copy the final URL.

A staging host is worse than it looks. Under the cross-submit rule, your production robots.txt now vouches for a staging sitemap. Either the fetch fails behind a password, or crawlers get a list of staging URLs.

The glued line is real. A large Shopify store I checked this week has a custom Disallow rule with a Sitemap line stuck to its end, twice. A parser reads that as one Disallow rule with a strange path, so the sitemap URL inside it does nothing. A correct Sitemap line higher in the file saves them. The grep command above shows the problem in one line of output.

Check the Sitemap lines in your robots.txt

Our Robots.txt & AI Crawler Checker fetches the live robots.txt for the host of the URL you enter, after your CMS and CDN have finished with it. It lists every Sitemap line, marks whether each is an absolute http or https URL, and notes when a file with user-agent groups declares no sitemap. The same run audits the rules, tests the path of your URL for Googlebot or a crawler you name, and shows which search and AI crawlers can get in. A run costs 10 credits.

It doesn't fetch the sitemaps, so a line pointing at a 404 or a redirect still passes. For that, paste the sitemap URL into the XML Sitemap Validator. It checks the XML, namespace, size, entry count, out-of-scope URLs and lastmod dates against sitemaps.org and Google's limits, for 10 credits. It reads only the file you give it, not robots.txt or the child sitemaps of an index.

Keep reading