Skip to main content
Glama

fetch_sitemap

Read-onlyIdempotent

Fetches sitemap XML or index files to extract URLs with metadata, follows nested sitemap indexes, decompresses gzip, and returns partial results with warnings if a child sitemap fails.

Instructions

Fetch a sitemap.xml (or sitemap-index) and return the contained URLs with their lastmod / changefreq / priority. Follows sitemap-index chaining up to max_depth levels. Gzipped .xml.gz payloads are auto-decompressed. Partial failures (one child sitemap 500s while others work) are returned under 'warnings' without aborting the whole request. SSRF-protected by default.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesSitemap URL (sitemap.xml, sitemap.xml.gz, or a sitemap index)
max_urlsNoCap on total URLs returned (default 5000)
max_bytesNoMax bytes to read per sitemap response (default 20MiB)
max_depthNoHow many sitemap-index levels to follow (default 1). 0 keeps the top-level index flat and only returns its childSitemaps list.
timeout_msNo
user_agentNo
max_sitemapsNoCap on sitemap documents fetched -- the index plus every child (default 50, max 1000). Children past the cap are listed under childSitemaps, unfetched.
max_redirectsNo
allow_private_hostsNoAllow loopback / private / link-local targets for this call (default false). Refused unless the server operator launched fetch-mcp with FETCH_MCP_ALLOW_PRIVATE_HOSTS=1 -- SSRF protection stays on by default either way.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changedv0.8.0
    • addedInput schema / properties / allow_private_hosts / description
      Added value: +"Allow loopback / private / link-local targets for this call (default false). Refused unless the server operator launched fetch-mcp with FETCH_MCP_ALLOW_PRIVATE_HOSTS=1 -- SSRF protection stays on by default either way."
    • addedInput schema / properties / max_sitemaps
      Added value: +{
      +  "description": "Cap on sitemap documents fetched -- the index plus every child (default 50, max 1000). Children past the cap are listed under childSitemaps, unfetched.",
      +  "maximum": 1000,
      +  "minimum": 1,
      +  "type": "integer"
      +}
  2. First observedv0.6.1

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses substantial behavioral details beyond the annotations: it follows index chaining up to max_depth, auto-decompresses .gz payloads, returns partial failures as 'warnings' without aborting the request, and is SSRF-protected by default. This goes well beyond the readOnly/idempotent hints and gives the agent a clear model of what happens during execution.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: the main purpose is front-loaded, followed by a few brief, high-value behavior notes (chaining, decompression, partial failures, SSRF). No redundant sentences or fluff; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 9 parameters and no output schema, the description covers the core return data (URLs, lastmod, etc.), explains special behaviors like warnings and childSitemaps, and touches on SSRF safety. It does not fully describe the output JSON structure, but it gives enough context to understand what to expect, especially since 'warnings' is mentioned. The missing parameter details are partially covered by schema, so the overall picture is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers 67% of parameters with descriptions. The tool description adds meaningful context for max_depth by explaining chaining behavior, and indirectly clarifies the url format (supports .gz). However, it does not add semantic value for the less-documented parameters (timeout_ms, user_agent, max_redirects), which remain self-explanatory but unaddressed. Overall the description adds some value but does not fully compensate for the missing schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Fetch', the resource ('sitemap.xml or sitemap-index'), and what is returned (URLs with lastmod/changefreq/priority). It also mentions distinguishing features like index chaining, which separates it from generic HTTP fetch tools and the other sitemap-adjacent tools (fetch_robots, fetch_meta) — it is obviously the sitemap-specific tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly compare to sibling tools or state when to prefer this over alternatives. It implies usage via its focus on sitemap parsing, but there is no direct 'use this when you need sitemap URLs' or a note about when not to use it (e.g., for non-sitemap XML). The guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.