freeXML Sitemap Validator

XML sitemap validator: is your sitemap telling the truth?

Enter a domain or a sitemap URL. We find the sitemap, validate it against the protocol, and spot-check 50 of its URLs for errors, redirects and noindex.

Try

What it checks

The file, then the URLs inside it.

We look for Sitemap: lines in robots.txt first, then try /sitemap.xml and /sitemap_index.xml. For a sitemap index we open up to five child sitemaps and show the rest in the tree.

  1. 1Structure

    A <urlset> or <sitemapindex> root with the sitemaps.org namespace, and extensions like image, video, news and hreflang noted.

  2. 2Limits

    No more than 50,000 URLs and 50MB uncompressed per file.

  3. 3URLs

    Absolute, on the sitemap's own host, using https, and listed once.

  4. 4lastmod

    W3C datetime format, not in the future, and not the same stamp on every URL.

  5. 5Discovery

    Whether robots.txt points to the sitemap, so crawlers that never see Search Console still find it.

  6. 6Sampled URLs

    50 URLs spread through the file, checked for errors, redirects, noindex and canonicals pointing elsewhere.

Fixing it

A sitemap should only list pages you want indexed.

The most common problem isn't broken XML. It's a sitemap that lists the wrong URLs: old pages that now redirect, filtered or paginated duplicates that canonicalise to another page, and pages someone marked noindex. Each one tells search engines your sitemap can't be trusted.

Generate it from the same source as your canonicals. If your CMS or framework builds the sitemap, make it skip anything with noindex or a canonical pointing elsewhere, and list final URLs rather than ones that redirect. Trace a redirecting URL with the redirect checker.

Keep lastmod honest. A lastmod that changes on every build, or sits in the future, teaches Google to ignore it.

Reference it in robots.txt. Test that your robots.txt is valid and doesn't block the pages you list with the robots.txt tester.

AI search

Sitemaps help AI crawlers find your pages too.

Crawlers from OpenAI, Anthropic, Perplexity and others discover pages through links and through sitemaps referenced in robots.txt. They don't get Search Console submissions. A clean sitemap, listed in robots.txt, with accurate lastmod dates is the simplest way to tell every crawler which pages exist and which changed recently.

FAQ

Good questions.

Keep going

Related free tools

All 24 tools

Found problems? We fix them, then get you named in AI answers.

These tools show what's broken. In 20 minutes we'll show you who ChatGPT recommends instead of you, and the three fixes we'd make first.

Book a call

Or get the free AI visibility report by email.

no strings, really