How Google reads a sitemap.xml
A sitemap.xml is a discovery aid, not a ranking signal. Google reads the file to find URLs it might otherwise miss (deep pages, orphaned pages, brand-new pages) and to prioritize its crawl budget. The sitemap does not tell Google which pages are important, only which pages exist.
When Googlebot fetches a sitemap:
- URL discovery. Every
<loc> entry is added to Google's crawl queue if it is not already known. - Last-modified check. If
<lastmod> is present and trustworthy, Google uses it to decide whether to recrawl a page sooner. Google has publicly stated it ignores <priority> and <changefreq> (source: Google Search Central, Sitemaps overview). - Canonicalization. Google matches sitemap URLs against its known canonical for each page. If your sitemap lists a non-canonical URL, it is a weak signal that the sitemap URL is the intended canonical.
- Coverage reporting. Every URL in the sitemap appears in Google Search Console's Coverage report as "Discovered" or "Excluded" with a reason — this is the single best debugging tool for indexation issues.
Bottom line: submit a sitemap for discovery and coverage debugging. Do not expect it to boost rankings on its own.
Sitemap size limits and how to split
The sitemaps.org protocol sets two hard limits per file:
- 50,000 URLs per sitemap file.
- 50 MB uncompressed per sitemap file (was 10 MB before 2016).
Whichever limit you hit first is the ceiling. A site with 500,000 URLs cannot use a single sitemap — it needs to split into multiple files.
Sitemap index file
When you have more than one sitemap, you create a sitemap index file that lists each child sitemap. The index itself is limited to 50,000 sitemaps and 50 MB — meaning a single index can reference up to 2.5 billion URLs.
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://example.com/sitemap-posts.xml</loc>
<lastmod>2026-07-11</lastmod>
</sitemap>
<sitemap>
<loc>https://example.com/sitemap-products.xml</loc>
<lastmod>2026-07-11</lastmod>
</sitemap>
</sitemapindex>
Practical splitting patterns
- By content type — one sitemap per section (posts, products, categories, users). Easiest to debug when a section stops indexing.
- By date — one sitemap per year or month. Useful for news sites; makes Google recrawl only fresh files.
- By numeric range —
sitemap-1.xml covers URL IDs 1-50,000, sitemap-2.xml covers 50,001-100,000, etc. Simple, mechanical.
If your sitemap exceeds 50 MB before hitting 50,000 URLs, gzip it (sitemap.xml.gz). Google and Bing both accept gzipped sitemaps.
sitemap.xml vs robots.txt vs llms.txt
Three files sit at the root of most sites. They do different jobs.
| File | Tells crawlers | Read by |
|---|
sitemap.xml | Which URLs exist and when they last changed | Google, Bing, Yandex, most SEO crawlers |
robots.txt | Which URLs crawlers must NOT fetch (allow/disallow rules) and where to find the sitemap | Every well-behaved crawler, including LLM crawlers |
llms.txt | A curated summary of your site's content for large language models | Emerging — no major LLM has confirmed production use yet |
Practical setup:
- Publish
sitemap.xml (or a sitemap index) at your site root. - Add a
Sitemap: line to robots.txt pointing at the sitemap URL. This is how most crawlers discover it before you submit anywhere. - Submit the sitemap in Google Search Console and Bing Webmaster Tools for coverage reporting.
- Optionally publish
llms.txt for LLM discoverability — but treat it as experimental, not a replacement for a sitemap.
Do not use robots.txt to hide pages from Google's index. A blocked page can still be indexed by URL alone if it is linked from elsewhere. Use <meta name="robots" content="noindex"> on the page itself.