Rudra Analyzer

How to diagnose XML sitemap problems, step by step

A sitemap should be a clean list of the pages you want indexed. How to check that yours loads, parses and lists only live, indexable URLs — and how to fix it when it doesn't.

Rudra Techno Team 7 min read

An XML sitemap is a list of the URLs you'd like search engines to crawl and index. It doesn't guarantee indexing, but a clean sitemap helps search engines find new and updated pages, and a messy one wastes their attention on redirects, errors and pages you've asked them not to index. If you're new to the file itself, start with robots.txt and sitemap.xml explained.

Sitemap problems fall into two groups: the file itself (can it be found and read?) and its contents (are the URLs in it worth listing?). Check them in that order.

Part 1: can the sitemap be found and read?

Is it where search engines look?

There's no single required location, so make it discoverable. Add an absolute Sitemap: https://example.com/sitemap.xml line to your robots.txt, and submit the sitemap in Google Search Console and Bing Webmaster Tools. Many platforms generate one at /sitemap.xml, /sitemap_index.xml or /wp-sitemap.xml.

Does it return 200 with XML?

Open the sitemap URL in a browser and check the status in developer tools. Common failures: a 404 after a plugin change, a redirect to the homepage, a login page, or a "soft" error — a 200 response that's actually an HTML error page. The response should be XML (application/xml or text/xml), or gzip-compressed XML for .xml.gz files.

Does it parse?

The root element must be <urlset> (a list of pages) or <sitemapindex> (a list of other sitemaps). Stray whitespace or PHP warnings before the <?xml declaration, unescaped & characters in URLs (they must be written &amp;) and truncated output from a timeout are the usual culprits. A browser that shows "XML parsing error" is telling you exactly where to look.

Is it within the limits?

One sitemap file may hold at most 50,000 URLs and 50 MB uncompressed. Bigger sites split their URLs across several files and list them in a sitemap index. Make sure every child sitemap the index names also loads.

Part 2: are the right URLs listed?

Every URL in your sitemap should be the final, canonical, indexable address of a page that returns 200. Anything else sends a mixed message. Look for:

  • URLs that return 404, 410 or 5xx. Deleted pages left in a static sitemap, or products that went out of stock and were removed.
  • URLs that redirect. Often the old form of a URL after a migration, or the http:// or non-www version. List the destination instead.
  • noindex pages. Listing a page and telling search engines not to index it contradicts itself. See what noindex means.
  • Non-canonical URLs. Parameter versions or duplicates whose canonical points elsewhere. More in how to fix canonical problems.
  • Other hosts. A sitemap may only list URLs on its own host unless you've verified cross-site submission; staging or CDN hostnames are a classic leak.
  • Duplicates. The same URL listed twice, often with and without a trailing slash.
  • URLs blocked by robots.txt. Search engines can't crawl what you've disallowed, so listing it is pointless.

lastmod dates

<lastmod> should use the W3C Datetime format, such as 2026-09-24 or 2026-09-24T10:30:00+00:00, and reflect the last meaningful change to the page. Google says it uses lastmod when it's consistently accurate, and ignores priority and changefreq. A sitemap that stamps every URL with today's date on every build teaches search engines to ignore the field.

Check your robots.txt and sitemap

Rudra's free SEO checker fetches your robots.txt and declared sitemaps, and reports whether they load and parse, list the page you tested, and carry valid lastmod dates.

Checks one page · Free · No sign-up needed

How to check it with Rudra

The single-page SEO checker

Rudra's free SEO checker reads your robots.txt and the sitemaps it declares (up to three, following a sitemap index to its first few child sitemaps), falling back to common locations when none is declared. It reports sitemaps that can't be reached, don't parse, don't list the page you tested or lack lastmod dates, plus informational notes for repeated URLs, URLs on other hosts and lastmod values that aren't valid dates. It also flags robots.txt syntax problems and a rule that blocks the whole site.

A site scan

Whether listed URLs actually work can only be answered by visiting them. A Rudra site scan (free account, started from your dashboard) records the URLs in your sitemap and compares them with what the crawl found at each address: sitemap URLs that returned 4xx or 5xx, URLs marked noindex, URLs that redirected and URLs on other hosts. URLs the crawl didn't reach, because of the page limit of your plan or robots rules, are counted as "not crawled" rather than reported as broken — so a small crawl of a large sitemap will show many of those, and that's expected.

Search Console

The Sitemaps report shows whether Google could read each file and how many URLs it discovered. The Page indexing report, filtered to a sitemap, tells you which listed URLs aren't indexed and why.

Fixing the common causes

  1. Generate the sitemap from your CMS or framework, not by hand, so deleted and unpublished pages drop out automatically.
  2. Exclude noindex, redirected and canonicalised-away pages in your SEO plugin or sitemap generator settings.
  3. Use your preferred URL form everywhere — protocol, host and trailing slash — so the sitemap, canonicals and internal links agree.
  4. Only update lastmod when content changes.
  5. Resubmit after big changes such as a migration, and watch the Sitemaps report for read errors.

Frequently asked questions

Do I need a sitemap?

Small, well-linked sites can be crawled fine without one. A sitemap helps most on large sites, new sites with few links pointing to them, and sites that add or update pages often.

Does listing a page in my sitemap get it indexed?

No. A sitemap helps search engines discover URLs; whether they index a page depends on its quality, duplication and signals like noindex or canonical tags.

Should images or PDFs be in my sitemap?

You can list them, and image sitemap extensions exist, but the priority is that every important HTML page is listed with its canonical URL.

Why does the site scan say many sitemap URLs were not crawled?

The crawl stops at your plan's page limit and respects robots.txt, so it may not reach every URL your sitemap lists. Those URLs aren't treated as broken — they simply weren't verified in that crawl.

Keep reading

Find out what's holding your website back

Run a free check on any public page. No sign-up needed for a basic check.