Skip to main content
Glama
competlab

competlab-mcp-server

by competlab

check_sitemap

Read-only

Scan any domain's sitemap to discover all declared URLs, categorize them by section, and report depth, freshness, and per-category counts.

Instructions

Live sitemap analysis for any domain — discovers URLs, categorizes them by section, and reports depth, freshness, and per-category counts. Reads both the conventional /sitemap.xml and every sitemap robots.txt declares, deduplicated, so a site publishing several returns all of them. status: 'partial' means the scan did not cover the whole corpus — NOT that the site is broken. It has two distinct causes and they are not interchangeable: the scan hit its own limits, or a sitemap the site declares could not be read. Only the second populates unreadSitemaps, and those entries ARE the site's own defect (a relative URL in robots.txt, a redirect off its domain, a server that refused). Never report a count from a partial scan as the site's total. Categories include programmatic: templated pages generated from a database or pattern, assigned only to groups of 25+ sibling URLs with machine-generated slugs. It describes how pages are generated, not their purpose — example URLs ship in insights.sampleUrlsByCategory.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to scan, e.g. example.com
sitemapUrlNoOptional full URL to a specific sitemap (with http:// or https:// prefix). Short-circuits discovery.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv4.0.1
    • removedInput schema / additionalProperties
      Removed value: -false
  2. Addedv1.2.0

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare a safe read of an open world; the description adds substantial behavior beyond them — deduplication across /sitemap.xml and robots.txt-declared sitemaps, the two distinct causes of status: 'partial', and that only the unreadSitemaps cause reflects a site defect. This is exactly the kind of context an agent cannot infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and every sentence carries interpretive weight, but the partial-status and programmatic-category explanations are dense and somewhat verbose for a description. Efficient overall, with minor room to tighten.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden itself — status semantics, unreadSitemaps, category definitions, and where sample URLs live (insights.sampleUrlsByCategory). An agent has what it needs to call and correctly interpret the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented, including that sitemapUrl short-circuits discovery. The description adds no additional parameter syntax or format detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence gives a specific verb+resource+scope: 'Live sitemap analysis for any domain — discovers URLs, categorizes them by section, and reports depth, freshness, and per-category counts.' It is instantly distinguishable from siblings like check_ai_crawlers or fetch_url.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description is rich on how to interpret results but never states when to reach for this tool versus alternatives, nor any exclusions. Usage is implied by the domain-scan framing rather than stated explicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.