freeXML Sitemap Validator
XML sitemap validator: is your sitemap telling the truth?
Try
What it checks
The file, then the URLs inside it.
We look for Sitemap: lines in robots.txt first, then try /sitemap.xml and /sitemap_index.xml. For a sitemap index we open up to five child sitemaps and show the rest in the tree.
1Structure
A <urlset> or <sitemapindex> root with the sitemaps.org namespace, and extensions like image, video, news and hreflang noted.
2Limits
No more than 50,000 URLs and 50MB uncompressed per file.
3URLs
Absolute, on the sitemap's own host, using https, and listed once.
4lastmod
W3C datetime format, not in the future, and not the same stamp on every URL.
5Discovery
Whether robots.txt points to the sitemap, so crawlers that never see Search Console still find it.
6Sampled URLs
50 URLs spread through the file, checked for errors, redirects, noindex and canonicals pointing elsewhere.
Fixing it
A sitemap should only list pages you want indexed.
The most common problem isn't broken XML. It's a sitemap that lists the wrong URLs: old pages that now redirect, filtered or paginated duplicates that canonicalise to another page, and pages someone marked noindex. Each one tells search engines your sitemap can't be trusted.
Generate it from the same source as your canonicals. If your CMS or framework builds the sitemap, make it skip anything with noindex or a canonical pointing elsewhere, and list final URLs rather than ones that redirect. Trace a redirecting URL with the redirect checker.
Keep lastmod honest. A lastmod that changes on every build, or sits in the future, teaches Google to ignore it.
Reference it in robots.txt. Test that your robots.txt is valid and doesn't block the pages you list with the robots.txt tester.
AI search
Sitemaps help AI crawlers find your pages too.
Crawlers from OpenAI, Anthropic, Perplexity and others discover pages through links and through sitemaps referenced in robots.txt. They don't get Search Console submissions. A clean sitemap, listed in robots.txt, with accurate lastmod dates is the simplest way to tell every crawler which pages exist and which changed recently.
FAQ
Good questions.
Keep going
Related free tools
Found problems? We fix them, then get you named in AI answers.
These tools show what's broken. In 20 minutes we'll show you who ChatGPT recommends instead of you, and the three fixes we'd make first.
Book a callOr get the free AI visibility report by email.
no strings, really