@pipeworx/sitemap-stats
Connect: https://gateway.pipeworx.io/sitemap-stats/mcp · Install: one-click buttons
Tools: 1
Reads a domain’s sitemap(s) and reports their STRUCTURE, not their content — URL counts by path prefix, the fingerprint that separates a normal site from a programmatic content farm (a farm concentrates a large majority of its URLs under one or two path prefixes).
Tools
sitemap_stats({ domain, max_urls? })— resolves a domain’s sitemap(s) (robots.txtSitemap:directives, falling back to/sitemap.xmlthen/sitemap_index.xml), follows one level of sitemap-index nesting, and returnstotal_urls,sitemap_count, a top-20 path-prefix histogram (top_prefixes, each withcountandshare),top_prefix_share(the single number that separates a normal site from a generated one), andlastmod_min/lastmod_maxwhere present. Caps total URLs read (default 150,000, hard ceiling 500,000 viamax_urls) and always setscapped: truewith acap_reasonwhen it stops early, so a truncated read never looks like a small site. A domain with no sitemap returnsfound: falsewith areason— not an error.
Auth
Keyless. No vendor, no key.
Data sources
- The target domain’s own
robots.txtand sitemap XML files, read live on every call — a direct pass-through, not a copy.
Traps found while building this
- Sitemap protocol caps a single file at 50,000 URLs, so real multi-file sites
are common — the tool follows a
<sitemapindex>one level down to its child<sitemap>files, but does not recurse further (nested indexes-of-indexes are non-standard and rare). - A raw
.xml.gzserved as a static file carries noContent-Encoding: gzipheader, so the runtime’s automatic decompression never triggers — decoded by hand via Workers-nativeDecompressionStream('gzip')when the URL ends in.gzor the response’scontent-typementions gzip. - The URL cap applies to total URLs read across every file, not per file, and
a run that hits it stops mid-file rather than skipping whole files — so
total_urlsunder a cap is always the true count of what was actually read, never silently smaller than that. robots.txtSitemap:lines are supposed to be absolute per spec but aren’t always in the wild; relative ones are resolved against the domain’s origin.
Tools
- sitemap_stats — Fetch a domain’s sitemap(s) and report their STRUCTURE — URL counts by path prefix — not their content. This is the fingerprint that separates a normal site from a programmatic content farm: a farm co
Tools
-
sitemap_stats— Fetch a domain's sitemap(s) and report their STRUCTURE — URL counts by path prefix — not their content. This is the fingerprint that separates a normal site from a programmatic content farm: a farm co