How to check and validate your XML sitemap
·7 min read
To check a sitemap, open its URL and confirm it returns HTTP 200 with XML whose root element is urlset or sitemapindex. Then run it through a sitemap checker that tests the XML, the namespace, the 50,000 URL and 50 MB limits, absolute URLs and lastmod dates, and submit it in Search Console's Sitemaps report to see whether Google could fetch and read it. A sitemap full of redirects, noindexed pages and dates that change on every deploy passes both, so audit what it lists too.
What a valid XML sitemap looks like
A valid sitemap is a UTF-8 XML file with a urlset root in the sitemaps.org namespace and one <url> element per page, each holding a <loc>. <lastmod> is optional and worth adding when it is accurate. Google ignores <changefreq> and <priority>, so I leave them out.
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://www.example.com/</loc>
<lastmod>2026-09-28</lastmod>
</url>
<url>
<loc>https://www.example.com/pricing</loc>
<lastmod>2026-10-02T14:30:00+00:00</lastmod>
</url>
</urlset>Get four details right.
- Every
<loc>is an absolute URL with its protocol. Writehttps://www.example.com/pricing, never/pricing. - Values must be entity-escaped. An ampersand in a URL is written
&, and one raw&makes the whole file malformed. - The namespace is exactly
http://www.sitemaps.org/schemas/sitemap/0.9. It starts with http, not https, and that is correct. lastmoduses W3C Datetime. A bare date like2026-09-28works. If you add a time, add a time zone too.
Sitemap index example
A sitemap index lists other sitemaps instead of pages. You need one once a site passes 50,000 URLs.
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://www.example.com/sitemaps/pages.xml</loc>
<lastmod>2026-10-01</lastmod>
</sitemap>
<sitemap>
<loc>https://www.example.com/sitemaps/products-1.xml.gz</loc>
<lastmod>2026-10-08</lastmod>
</sitemap>
</sitemapindex>An index lists sitemaps, never other indexes, and Google rejects a nested one. Google's guide to large sitemaps also requires the child sitemaps to sit on the same site as the index, in its folder or deeper. Put the index at the root and give every child a full URL.
Sitemap limits: URLs, file size and encoding
One sitemap file holds at most 50,000 URLs and 50 MB uncompressed, and it must be UTF-8. Those numbers come from the sitemaps.org protocol, and Google's sitemap docs apply the same ones.
| Rule | Limit |
|---|---|
| URLs per sitemap | 50,000 |
| Size per sitemap | 50 MB uncompressed |
| Sitemaps per index file | 50,000 |
| Index files per site in Search Console | 500 |
Length of one <loc> |
Under 2,048 characters |
| URLs | Absolute, same protocol and host, under the sitemap's folder |
Gzip does not raise the size limit, because the 50 MB counts after decompression. Hit either limit and you split the file and list the parts in an index.
A sitemap at https://www.example.com/blog/sitemap.xml covers only URLs under /blog/, and a sitemap on https://www. cannot list http:// or non-www URLs. Google reports these as "URL not allowed". A root sitemap avoids the folder problem.
What Google actually uses from your sitemap
Google reads <loc>, uses <lastmod> when it trusts the dates, and ignores <priority> and <changefreq>. Setting every page to priority 1.0 changes nothing.
Trust is the catch with lastmod. Google's docs say it uses the value if it is "consistently and verifiably" accurate. A 2023 Search Central post says a new date is warranted when the main text, structured data or links change, not for a footer or sidebar edit. If your CMS cannot tell when the home page or a category page last changed, leave lastmod off those URLs. Google says that is fine.
The same post retired the sitemap ping endpoint, which now returns 404. Submit through Search Console or robots.txt instead.
A sitemap is also no indexing guarantee. Google's help page says a URL found in a sitemap may never be crawled or indexed. If Google knows your URLs and still leaves them out, the cause is usually the pages, as covered in Discovered, currently not indexed.
How to add your sitemap to robots.txt
Add a Sitemap: line with the full URL anywhere in robots.txt. It does not belong to any User-agent group, and you can list as many sitemaps as you need.
User-agent: *
Disallow: /cart/
Sitemap: https://www.example.com/sitemap_index.xml
Sitemap: https://www.example.com/sitemaps/news.xmlSubmit it in Search Console as well. Google finds the robots.txt line on its next robots.txt crawl, but the Sitemaps report lists only sitemaps submitted there or through the API.
Make sure robots.txt does not block the sitemap itself. Google respects robots.txt when fetching sitemaps, so a Disallow: /sitemaps/ rule produces "Couldn't fetch". Our robots.txt checker tests a path against Googlebot's rules and flags a robots.txt with no Sitemap line.
How to fix "Couldn't fetch" and "Sitemap could not be read"
"Couldn't fetch" means Google never got the file, so nothing inside it was read. On that sitemap's details page, the same failure reads "Sitemap could not be read". Fix the fetch before anything else.
| Status in the Sitemaps report | What it means |
|---|---|
| Success | Fetched and read with no errors |
| Has errors, or Sitemap had X errors | Fetched, but parts failed. URLs that parsed are still queued. |
| Couldn't fetch | Google could not retrieve the file |
Work through a couldn't fetch sitemap in this order:
- Open the sitemap URL in a private window. It should return 200 with XML, with no login, no redirect to the home page and no HTML error page. A 404 means the submitted URL is wrong, often www against non-www.
- Run URL Inspection on the sitemap URL with a live test. Under Page availability, Crawl allowed should say Yes and Page fetch should say Successful. If crawling is not allowed, remove the robots.txt rule that blocks the path.
- If your browser loads the file but the live test fails, check your firewall or CDN. A bot challenge or 403 rule can block Googlebot while the site works for you.
- Open the Manual actions report. Google does not read sitemaps on a site with an unresolved manual action.
- Resubmit. Google retries a failed sitemap for a few days and then stops, so a fixed file needs a fresh submission.
If all of that passes, Google's remaining explanations are a temporary server error or low crawl demand, so wait a few days.
When the status is Has errors, the details page names each problem. The common ones, in the wording of Google's Sitemaps report help page:
| Error | Usual cause | Fix |
|---|---|---|
| Parsing error | A raw &, < or quote in a URL |
Escape it, for example & |
| Unsupported format | Missing or mistyped namespace, curly quotes | Exact 0.9 namespace, straight quotes |
| Invalid date | A non-W3C date, or a time with no zone | 2026-10-02 or 2026-10-02T14:30:00+00:00 |
| URL not allowed | Another host or protocol, or above the sitemap's folder | Fix the URLs or move the sitemap to the root |
| Too many URLs | Over 50,000 URLs | Split the file and add an index |
Sitemap problems a validator won't catch
A sitemap validator checks the file, not the pages it lists. Each problem below is valid XML that tells Google something your site contradicts.
| Problem | Usual cause | Fix |
|---|---|---|
| Redirected URLs | Old URLs left after a migration | List the final destination |
| Noindexed URLs | Tag pages or thank-you pages the CMS also noindexes | Drop the URL or the noindex |
| Non-canonical URLs | Parameter variants, pages canonicalized elsewhere | List only the canonical |
| 404 and 410 URLs | Deleted pages the generator kept | Remove them |
| lastmod changes on every build | The generator stamps the deploy time | Use the content's last edit date |
| http URLs on an https site | A hard-coded http base URL | Fix the base URL |
| Sitemap missing from robots.txt | Nobody added the line | Add a Sitemap: line |
Google's docs say a sitemap should list the URLs you want shown in results, which means canonical URLs only. Variants feed reports like Duplicate without user-selected canonical.
The lastmod problem does the quietest damage. When every URL claims it changed this morning, the dates are no longer consistently accurate, and Google can stop using them, including for the pages that did change.
To find the rest, sample. Paste 20 <loc> values from each child sitemap into our bulk HTTP status checker, and treat anything other than 200 as a URL to replace or remove. For noindex and canonical problems, run a few listed URLs through the indexability checker, or filter Search Console's Page indexing report to one sitemap and read why each excluded URL was left out.
Validate your sitemap with our sitemap checker
Our XML Sitemap Validator fetches one sitemap or sitemap index and checks it against the sitemaps.org protocol and Google's limits. It parses the XML strictly and reports the line and column where parsing stopped, and the spot where a raw & sits. It checks the namespace, UTF-8 encoding and the 50,000 entry and 50 MB limits. Then it reads every entry, up to 50,000, for a missing or relative <loc>, URLs over 2,048 characters, duplicates, URLs on another host or protocol or outside the sitemap's folder, lastmod values that are not W3C Datetime or fall in the future, and invalid changefreq and priority values. It also flags a bot check or rate limit served instead of the file, and recognizes an HTML page, feed or plain-text list without validating it.
You get a verdict, entry counts, each finding with its fix, and the first 25 entries as a sample. It reads up to 10 MB of plain XML, inflates gzip files up to the 50 MB limit, and says when only part of a larger file was checked. A run costs 10 credits.
It does not open the child sitemaps in an index, so run it once per child. It does not request the listed pages either, so the redirects, noindex tags and 404s above need the status and indexability checks.