Skip to content

free tool · no login

shahid ali › tools › sitemap checker

Sitemap Checker and XML Sitemap Validator

Type a domain or a sitemap address. The checker finds every sitemap, follows sitemap indexes down to the last file, validates the XML against the sitemap rules, checks every listed URL, and opens a sample of the pages to see what Google would find there. Each finding comes with a plain explanation, the fix, and the guide that goes deeper.

  • 98 checks across discovery, XML, URLs, dates, image, video and news extensions, hreflang and live pages.
  • Reads .xml and .xml.gz sitemaps, indexes of any size, 40 files per round with a button to carry on.
  • Score, findings by severity, a table per sitemap and per URL, and two CSV downloads.

Check a sitemap

  1. 01 Type a domain or a sitemap address

    A domain finds every sitemap for you: the Sitemap: lines in robots.txt, then the usual addresses used by WordPress, Yoast, Rank Math, Shopify, Wix, Squarespace and Blogger. A full address checks that one file and everything it lists.

    The sample is spread evenly across every sitemap, up to 500 pages. Each page is opened once by a reader on this domain that keeps no copy and writes no log.

  2. 02 Run the check

    nothing checked yet

What it checks

Every line below is a finding the report can show, in the same words. Errors stop pages being crawled or indexed, warnings waste crawling or confuse Google, and notes are worth knowing but cost nothing.

Finding the sitemap

Fetching each file

XML validity and limits

URL hygiene

lastmod, changefreq and priority

Image, video and news extensions

  • error Image entry without image:loc
  • warning Image address is not a full URL
  • note Retired image tags
  • warning More than 1,000 images on one URL
  • error Video entry missing a required field
  • warning Video address is not a full URL
  • warning Video duration out of range
  • warning Video description too long
  • warning content_loc is the page itself
  • warning Video date in the wrong format
  • error News entry missing a required field
  • error News language code not valid
  • error News date in the wrong format
  • warning News article older than two days
  • note Retired news tags
  • error More than 1,000 news entries

hreflang alternates

Live page checks

A sitemap is one half of what crawlers read. The robots.txt testerchecks the other half, rule by rule, for Googlebot and the AI crawlers. When one listed page comes back noindexed or canonicalised elsewhere, the indexability checkergoes through that page in full, redirect chain and canonical target included. And once the sitemap is clean, thesitemap index coverage checker tells you how much of it Google actually kept, straight from Search Console.

Where your data actually goes

Browsers are not allowed to read another site's files, so two small readers on this domain do it./api/sitemap-check reads robots.txt and one sitemap file per request, unzips it if needed, and returns the entries. /api/sitemap-health opens the sampled pages and returns only the status, redirect target, robots directives, canonical and title.

Neither writes a log line or stores anything. The report is built in this tab, so reloading destroys it, and the CSV files are made in the browser. Full detail on the privacy page.

The limits, stated plainly

  • 40 sitemap files per round. A bigger index gets a button to read the next 40, and the report grows as you go.
  • 150,000 URLs per run. Past that the tab gets slow, so the run stops reading new files and says so. Check a single child sitemap for the rest.
  • 50 MB per file, unzipped, which is the protocol limit. A larger file is read up to 50 MB and reported as too large.
  • Three levels of index. An index inside an index is followed and flagged, because Google will not follow it.
  • 50 live pages by default, 500 at most, opened a few at a time so the site is never hammered. A page that does not answer in 12 seconds gets one gentler second try.
  • robots.txt is matched for Googlebot on the host it came from, with Google's precedence rules.
  • It reads what the site serves to this reader. A site that treats Googlebot differently can show Google something else.

Straight answers

What is the difference between a sitemap checker and a sitemap validator?

A validator asks whether the file follows the rules: well-formed XML, the right root element and namespace, a <loc> in every entry, under 50,000 URLs and 50 MB. A checker asks whether the URLs inside are worth listing: do they load, are they on the right host, are they blocked or marked noindex. This tool does both in one run, because a sitemap can be perfectly valid and still point Google at pages it will never index.

Do I need to know where my sitemap is?

No. Type the domain and the checker reads the Sitemap: lines in robots.txt, then tries the addresses the common platforms use: /sitemap.xml, /sitemap_index.xml for Yoast and Rank Math, /wp-sitemap.xml for WordPress itself, and a few more. If you already know the address, paste it and only that file and the files it lists are read.

Why does it only open a sample of the pages?

Reading the sitemap files is cheap, so every listed URL gets the file checks. Opening each page is not: it costs the site you are checking a real request. The default is 50 pages spread evenly across all the sitemaps, and you can raise it to 500. A pattern that affects a whole template shows up in an even sample long before you would need to open every page.

Does Google use changefreq and priority?

No. Google has said it ignores both, so the report lists them as notes, not faults. They do no harm. What Google does use is an accurate lastmod, which is why the checker looks hard at it: a date in the future, the same date on every URL, or every date stamped in the hour before the file was fetched all teach Google to stop trusting the field.

My sitemap is valid. Why are pages from it still not indexed?

Because a sitemap only asks. If the pages it lists redirect, return errors, carry noindex or point their canonical somewhere else, Google declines, and the live page checks here show which of those applies. If the sample comes back clean, the reason is on the page itself, and the guide to every Search Console status is the next place to look.

Is anything I check stored?

No. Two small readers on this domain fetch what your browser is not allowed to read directly, and neither writes a log line or keeps a copy. The report is built in your tab, so closing it destroys it. The CSV downloads are made in the browser too.

Send me the findings CSV and I will tell you which one cause is behind most of it, free.Send it on WhatsApp.