XML sitemap guide: format, limits and examples
·4 min read
An XML sitemap is a UTF-8 file that lists the URLs you want search engines to crawl, inside a urlset element with one <url> and <loc> per page. Only those three tags are required. Add <lastmod> when your dates are accurate, skip <priority> and <changefreq> because Google ignores them, and keep each file under 50,000 URLs and 50 MB uncompressed. Past either limit, split the URLs across several sitemaps and list them in a sitemap index.
This post is about building one. If you already have a sitemap and Search Console reports errors, read how to check and validate your XML sitemap instead.
A minimal XML sitemap example
This is the smallest sitemap worth shipping. It declares the namespace, lists two pages and gives each a last modified date.
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://www.example.com/</loc>
</url>
<url>
<loc>https://www.example.com/guides/returns</loc>
<lastmod>2026-09-14</lastmod>
</url>
</urlset>The home page has no lastmod on purpose, which I explain below. Here is what the sitemaps.org protocol requires.
| Tag | Required? | What Google does with it |
|---|---|---|
<urlset> |
Yes | Root element, must carry the 0.9 namespace |
<url> |
Yes | One per page |
<loc> |
Yes | Absolute URL, entity-escaped |
<lastmod> |
No | Used as a crawl scheduling signal if accurate |
<changefreq> |
No | Ignored |
<priority> |
No | Ignored |
Google's sitemap guide adds two rules. Use fully qualified absolute URLs, and encode the file as UTF-8.
Sitemap lastmod, priority and changefreq
Google uses lastmod only when it is "consistently and verifiably" accurate. A 2023 Search Central post spells out what counts. Update it when the main text, structured data or links change. Leave it alone for a sidebar or footer tweak. If your dates claim yesterday for a page untouched in years, Google stops believing them.
The same post says lastmod can go on all pages or only on the ones you are confident about. A home page or category page that just aggregates other pages often has no honest edit date, and Google says leaving lastmod off is fine. That is why my example omits it.
The post also confirms Google still doesn't use changefreq or priority at all. I delete both. They add bytes and imply control you don't have.
Sitemap index and size limits
One sitemap holds at most 50,000 URLs and 50 MB uncompressed. Gzip shrinks the transfer, not the limit. Above either number, write several sitemaps and point to them from a sitemap index.
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://www.example.com/sitemaps/pages.xml</loc>
</sitemap>
<sitemap>
<loc>https://www.example.com/sitemaps/products-1.xml</loc>
<lastmod>2026-10-08</lastmod>
</sitemap>
</sitemapindex>In an index, <sitemapindex>, <sitemap> and <loc> are required and <lastmod> is optional. I split by content type even on small sites. Search Console's Page indexing report can filter by submitted sitemap, so with products and blog posts in separate files a problem in one section is easy to spot.
Which URLs belong in an XML sitemap
List the URLs you want in search results, which Google's guide says outright. In practice that means pages that:
- return 200, not a redirect or a 404
- are the canonical version, not a parameter or print variant
- have no noindex tag and aren't blocked in robots.txt
Everything else sends a mixed signal. A sitemap that lists a URL whose canonical points elsewhere asks Google to index a page you told it not to prefer. Generators get this wrong most often with tag archives, paginated pages and pages your SEO plugin already noindexes.
A clean sitemap still doesn't force indexing. It helps Google find URLs. Our guide on how to get Google to index your site covers the rest.
Image, video, news and hreflang extensions
Extensions add namespaced tags inside each <url>. Use them only when they describe content Google would otherwise miss.
- Image sitemaps use the namespace
http://www.google.com/schemas/sitemap-image/1.1, with<image:image>and<image:loc>, up to 1,000 images per URL (Google docs). - Video sitemaps use their own namespace for title, description and thumbnail per video.
- News sitemaps list only articles created in the last two days, at most 1,000
news:newstags per sitemap (Google docs). - hreflang annotations are
xhtml:linkelements with thehttp://www.w3.org/1999/xhtmlnamespace, where each URL lists every language version including itself (Google docs).
How to generate an XML sitemap
Generate it from your content source, never by hand. A hand-written file is out of date by the next publish.
CMS. WordPress has served a core sitemap at /wp-sitemap.xml since 5.5, and SEO plugins such as Yoast replace it with their own index. Shopify builds sitemap.xml for each store domain. Check which one is live before you submit anything.
Next.js. Add app/sitemap.ts and return an array typed as MetadataRoute.Sitemap. The Next.js docs call it a special route handler, cached by default.
import type { MetadataRoute } from "next";
import { getPublishedGuides } from "@/lib/guides";
export default async function sitemap(): Promise<MetadataRoute.Sitemap> {
const guides = await getPublishedGuides();
return [
{ url: "https://www.example.com/" },
...guides.map((guide) => ({
url: `https://www.example.com/guides/${guide.slug}`,
lastModified: guide.updatedAt,
})),
];
}Pass the content's real edit date, never new Date(). The docs' own example uses new Date(), which stamps every URL with the build time and wrecks the trust described above. For large sites, generateSitemaps splits the output into /sitemap/[id].xml files.
Static generators. Hugo writes sitemap.xml to the root of public by default. Front matter can set changefreq and priority, which you can leave empty.
Once the file is live, reference it in robots.txt and submit it in Search Console. Our post on adding your sitemap to robots.txt has the syntax.
Validate your XML sitemap
Our XML Sitemap Validator fetches one sitemap or sitemap index and checks it against the sitemaps.org protocol and Google's limits. It parses the XML strictly and points to where parsing failed, then checks the namespace, UTF-8, the 50,000 entry and 50 MB limits, missing or relative <loc> values, duplicates, URLs outside the sitemap's scope, and lastmod dates that aren't W3C Datetime or sit in the future. A run costs 10 credits and returns a verdict, entry counts and each finding with its fix.
It reads only the file you give it. It doesn't fetch the child sitemaps of an index, validate image, video or news tags, or request the listed pages, so it can't tell you which URLs redirect or carry noindex.