@pipeworx/sitemap-stats

Connect: https://gateway.pipeworx.io/sitemap-stats/mcp · Install: one-click buttons

Tools: 1

Reads a domain’s sitemap(s) and reports their STRUCTURE, not their content — URL counts by path prefix, the fingerprint that separates a normal site from a programmatic content farm (a farm concentrates a large majority of its URLs under one or two path prefixes).

Tools

  • sitemap_stats({ domain, max_urls? }) — resolves a domain’s sitemap(s) (robots.txt Sitemap: directives, falling back to /sitemap.xml then /sitemap_index.xml), follows one level of sitemap-index nesting, and returns total_urls, sitemap_count, a top-20 path-prefix histogram (top_prefixes, each with count and share), top_prefix_share (the single number that separates a normal site from a generated one), and lastmod_min/lastmod_max where present. Caps total URLs read (default 150,000, hard ceiling 500,000 via max_urls) and always sets capped: true with a cap_reason when it stops early, so a truncated read never looks like a small site. A domain with no sitemap returns found: false with a reason — not an error.

Auth

Keyless. No vendor, no key.

Data sources

  • The target domain’s own robots.txt and sitemap XML files, read live on every call — a direct pass-through, not a copy.

Traps found while building this

  • Sitemap protocol caps a single file at 50,000 URLs, so real multi-file sites are common — the tool follows a <sitemapindex> one level down to its child <sitemap> files, but does not recurse further (nested indexes-of-indexes are non-standard and rare).
  • A raw .xml.gz served as a static file carries no Content-Encoding: gzip header, so the runtime’s automatic decompression never triggers — decoded by hand via Workers-native DecompressionStream('gzip') when the URL ends in .gz or the response’s content-type mentions gzip.
  • The URL cap applies to total URLs read across every file, not per file, and a run that hits it stops mid-file rather than skipping whole files — so total_urls under a cap is always the true count of what was actually read, never silently smaller than that.
  • robots.txt Sitemap: lines are supposed to be absolute per spec but aren’t always in the wild; relative ones are resolved against the domain’s origin.

Tools

  • sitemap_stats — Fetch a domain’s sitemap(s) and report their STRUCTURE — URL counts by path prefix — not their content. This is the fingerprint that separates a normal site from a programmatic content farm: a farm co

Tools

  • sitemap_stats — Fetch a domain's sitemap(s) and report their STRUCTURE — URL counts by path prefix — not their content. This is the fingerprint that separates a normal site from a programmatic content farm: a farm co

Regenerated from source · build September 3, 2026