An XML sitemap is a file listing the URLs on a site that you want search engines to know about, in a format built for machines rather than people. It is a discovery aid: a direct inventory handed to crawlers, so they are not relying solely on following links to find everything you publish.

What it is not is a ranking factor, or a guarantee of anything. Including a URL in a sitemap does not make it index, and omitting one does not hide it. A sitemap helps search engines find pages and notice changes. Whether those pages then get indexed depends on whether they deserve to be.

What it contains

Each entry has a <loc>, the full absolute URL, and optionally a <lastmod> date indicating when the content last changed meaningfully.

The older <priority> and <changefreq> elements are ignored by Google entirely, which has confirmed as much. They are harmless but pointless, and any tool that presents them as a tuning lever is selling you a dial connected to nothing.

<lastmod> is different and genuinely useful, on one condition: it has to be accurate. Google uses it if it finds it trustworthy, and ignores it entirely if the site stamps every URL with today’s date or bumps dates on trivial edits. A sitemap that claims everything changed this morning is telling Google nothing, and it teaches Google to stop believing the field.

Size limits and sitemap index files

A single sitemap file can hold up to 50,000 URLs and must not exceed 50MB uncompressed. Larger sites split across multiple files and publish a sitemap index, a sitemap of sitemaps, which is what you submit.

Most WordPress SEO plugins do this automatically, splitting by content type. Our own site’s index references separate sitemaps for posts, pages, work, and glossary entries, which is also a convenient way to check whether an entire content type has gone missing.

The rules for what goes in

A sitemap should list only URLs you actually want indexed. In practice that means every entry should:

Return a 200 status. No redirects, no 404s.

Be indexable. No noindex tag, not blocked in robots.txt.

Be the canonical version. If a URL canonicalizes to a different address, the canonical belongs in the sitemap, not the variant.

A sitemap full of redirected, missing or noindexed URLs is worse than a smaller accurate one. It sends crawlers to dead ends and reduces the confidence search engines place in the file as a whole.

Submitting it

Google Search Console, under Sitemaps. This also gives you the report described below.

Bing Webmaster Tools, which is worth doing and often forgotten, and which also feeds other surfaces.

The Sitemap: directive in robots.txt, which any crawler can read without you submitting anything. This is how crawlers you never thought to tell find it, and it belongs in every robots.txt.

A failure that cost us months

Worth telling in full, because it is invisible unless you look for it specifically.

Our sitemaps had been corrected in WordPress. Every check we ran confirmed the right URLs were being generated. Indexing stayed stubbornly flat anyway. What was actually happening: the host’s edge cache was serving an outdated copy of the sitemap at the bare URL, the exact address Googlebot requests. Every test we ran added a query string, which bypassed the cache and returned the correct current file. The version Google received and the version we kept checking were different documents.

Purging the host cache fixed it, and indexing climbed steadily over the following weeks.

The generalizable lesson: verify what is served at the exact URL a crawler requests, with no parameters appended, and re-check after anything that touches caching. This applies equally to robots.txt and to redirects. It is now a standing step in every technical audit we run.

Reading the Search Console sitemap report

The sitemap report shows how many URLs were discovered from the file, and the page indexing report shows how many are actually indexed. The gap between those two numbers is the interesting part.

A large gap usually means one of three things: the pages are thin or duplicative and Google chose not to index them, they are blocked or noindexed despite being listed, or the site is new or low-authority enough that Google is still working through them. The report’s own status reasons distinguish these, and “Discovered, currently not indexed” in particular usually points at crawl priority rather than a technical fault. Our entry on indexability covers how to read those statuses.

Common mistakes

Listing URLs that redirect, 404, or carry noindex. The most frequent problem, and usually a sign the sitemap is generated from a stale source.

Inaccurate lastmod dates, which get the field ignored.

Never submitting it, or submitting it once and never looking at the report again.

Blocking the sitemap in robots.txt, which happens more often than it should via an over-broad rule.

Forgetting it after a migration, so it still lists the old URL structure.

Treating it as a substitute for internal links. A page reachable only via the sitemap is still an orphan. Sitemaps aid discovery; internal links communicate importance and pass signals. A page nothing links to looks unimportant no matter how prominently it is listed. See crawlability.

Other sitemap types

Beyond pages, there are image sitemaps for sites where image search matters, video sitemaps carrying duration and thumbnail data, and news sitemaps for publishers in Google News. Most business sites need none of these; ecommerce and media sites often benefit from the first two.

Where this fits

Keeping sitemaps accurate, cache-free and submitted is routine work in our technical SEO services, and checking what is actually served at the sitemap URL is one of the first things we do on any SEO audit, for the reason described above. If your pages are not being indexed and you cannot see why, book a discovery call.