fetch_sitemap
Fetches sitemap XML or index files to extract URLs with metadata, follows nested sitemap indexes, decompresses gzip, and returns partial results with warnings if a child sitemap fails.
Instructions
Fetch a sitemap.xml (or sitemap-index) and return the contained URLs with their lastmod / changefreq / priority. Follows sitemap-index chaining up to max_depth levels. Gzipped .xml.gz payloads are auto-decompressed. Partial failures (one child sitemap 500s while others work) are returned under 'warnings' without aborting the whole request. SSRF-protected by default.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Sitemap URL (sitemap.xml, sitemap.xml.gz, or a sitemap index) | |
| max_urls | No | Cap on total URLs returned (default 5000) | |
| max_bytes | No | Max bytes to read per sitemap response (default 20MiB) | |
| max_depth | No | How many sitemap-index levels to follow (default 1). 0 keeps the top-level index flat and only returns its childSitemaps list. | |
| timeout_ms | No | ||
| user_agent | No | ||
| max_sitemaps | No | Cap on sitemap documents fetched -- the index plus every child (default 50, max 1000). Children past the cap are listed under childSitemaps, unfetched. | |
| max_redirects | No | ||
| allow_private_hosts | No | Allow loopback / private / link-local targets for this call (default false). Refused unless the server operator launched fetch-mcp with FETCH_MCP_ALLOW_PRIVATE_HOSTS=1 -- SSRF protection stays on by default either way. |