Sitemap MCP Server
Mapa del sitio del servidor MCP
Descubra la arquitectura del sitio web y analice su estructura obteniendo, analizando y visualizando mapas del sitio desde cualquier URL. Descubra páginas ocultas y extraiga jerarquías organizadas sin necesidad de exploración manual.
Incluye plantillas de indicaciones listas para usar para Claude Desktop que le permiten analizar sitios web, verificar el estado del mapa del sitio, extraer URL, encontrar contenido faltante y crear visualizaciones con solo ingresar una URL.
Manifestación
Obtenga respuestas a preguntas sobre cualquier sitio web que aproveche el poder de los mapas del sitio.
Haga clic en el botón "adjuntar" junto al botón de herramientas:
Luego seleccione visualize_sitemap :
Ahora entramos en windsurf.com:
Y obtenemos una visualización del mapa del sitio:
Related MCP server: jcrawl4ai-mcp-server
Instalación
Asegúrese de que el uv esté instalado.
Instalación en Claude Desktop, Cursor o Windsurf
Agregue esta entrada a su claude_desktop_config.json , configuración del cursor, etc.:
{
"mcpServers": {
"sitemap": {
"command": "uvx",
"args": ["sitemap-mcp-server"],
"env": { "TRANSPORT": "stdio" }
}
}
}Reinicia Claude si se está ejecutando. Para el cursor, simplemente presiona "Actualizar" o activa el servidor MCP en la configuración.
Instalación mediante herrería
Para instalar el mapa del sitio para Claude Desktop automáticamente a través de Smithery :
npx -y @smithery/cli install @mugoosse/sitemap --client claudeInspector de MCP
npx @modelcontextprotocol/inspector env TRANSPORT=stdio uvx sitemap-mcp-serverAbra el Inspector MCP en http://127.0.0.1:6274 , seleccione el transporte stdio y conéctese al servidor MCP.
# Start the server
uvx sitemap-mcp-server
# Start the MCP Inspector in a separate terminal
npx @modelcontextprotocol/inspector connect http://127.0.0.1:8050Abra el Inspector MCP en http://127.0.0.1:6274 , seleccione el transporte sse y conéctese al servidor MCP.
Transporte SSE
Si desea utilizar el transporte SSE, siga estos pasos:
Iniciar el servidor:
uvx sitemap-mcp-serverConfigure su cliente MCP, por ejemplo Cursor:
{
"mcpServers": {
"sitemap": {
"transport": "sse",
"url": "http://localhost:8050/sse"
}
}
}Desarrollo local
Para obtener instrucciones sobre cómo crear y ejecutar el proyecto desde la fuente, consulte la guía DEVELOPERS.md .
Uso
Herramientas
Las siguientes herramientas están disponibles a través del servidor MCP:
get_sitemap_tree : obtiene y analiza el árbol del mapa del sitio desde la URL de un sitio web
Argumentos:
url(URL del sitio web),include_pages(opcional, booleano)Devuelve: representación JSON de la estructura de árbol del mapa del sitio
get_sitemap_pages : obtiene todas las páginas del mapa del sitio de un sitio web con opciones de filtrado
Argumentos:
url(URL del sitio web),limit(opcional),include_metadata(opcional),route(opcional),sitemap_url(opcional),cursor(opcional)Devuelve: lista JSON de páginas con metadatos de paginación
get_sitemap_stats - Obtener estadísticas sobre el mapa del sitio de un sitio web
Argumentos:
url(URL del sitio web)Devoluciones: objeto JSON con estadísticas del mapa del sitio, incluidos recuentos de páginas, fechas de modificación y detalles del submapa del sitio
parse_sitemap_content - Analiza un mapa del sitio directamente desde su contenido XML o de texto
Argumentos:
content(contenido XML del mapa del sitio),include_pages(opcional, booleano)Devuelve: representación JSON del mapa del sitio analizado
Indicaciones
El servidor incluye indicaciones listas para usar que aparecen como plantillas en Claude Desktop. Después de instalar el servidor, verá estas plantillas en el menú "Plantillas" (haga clic en el icono + junto a la entrada del mensaje):
Analizar el mapa del sitio : proporciona un análisis integral de la estructura del mapa del sitio de un sitio web.
Comprobar el estado del mapa del sitio : evalúa el SEO y las métricas de salud de un mapa del sitio
Extraer URL del mapa del sitio : extrae y filtra URL específicas de un mapa del sitio
Encontrar contenido faltante en el mapa del sitio : identifica lagunas de contenido en el mapa del sitio de un sitio web
Visualizar la estructura del mapa del sitio : crea un diagrama de Mermaid.js que visualiza la estructura del mapa del sitio
Para utilizar estas indicaciones:
Haga clic en el ícono + junto a la entrada del mensaje en Claude Desktop
Seleccione la plantilla deseada de la lista
Complete la URL del sitio web cuando se le solicite
Claude ejecutará el análisis apropiado del mapa del sitio.
Ejemplos
Obtenga un mapa del sitio completo
{
"name": "get_sitemap_tree",
"arguments": {
"url": "https://example.com",
"include_pages": true
}
}Obtener páginas con filtrado y paginación
Filtrar por ruta
{
"name": "get_sitemap_pages",
"arguments": {
"url": "https://example.com",
"limit": 100,
"include_metadata": true,
"route": "/blog/"
}
}Filtrar por submapa del sitio específico
{
"name": "get_sitemap_pages",
"arguments": {
"url": "https://example.com",
"limit": 100,
"include_metadata": true,
"sitemap_url": "https://example.com/blog-sitemap.xml"
}
}Paginación basada en cursor
El servidor implementa la paginación basada en cursor MCP para gestionar mapas de sitios grandes de manera eficiente:
Solicitud inicial:
{
"name": "get_sitemap_pages",
"arguments": {
"url": "https://example.com",
"limit": 50
}
}Respuesta con paginación:
{
"base_url": "https://example.com",
"pages": [...], // First batch of pages
"limit": 50,
"nextCursor": "eyJwYWdlIjoxfQ=="
}Solicitud posterior con cursor:
{
"name": "get_sitemap_pages",
"arguments": {
"url": "https://example.com",
"limit": 50,
"cursor": "eyJwYWdlIjoxfQ=="
}
}Cuando no haya más resultados, el campo nextCursor estará ausente de la respuesta.
Obtener estadísticas del mapa del sitio
{
"name": "get_sitemap_stats",
"arguments": {
"url": "https://example.com"
}
}La respuesta incluye estadísticas totales y estadísticas detalladas de cada submapa del sitio:
{
"total": {
"url": "https://example.com",
"page_count": 150,
"sitemap_count": 3,
"sitemap_types": ["WebsiteSitemap", "NewsSitemap"],
"priority_stats": {
"min": 0.1,
"max": 1.0,
"avg": 0.65
},
"last_modified_count": 120
},
"subsitemaps": [
{
"url": "https://example.com/sitemap.xml",
"type": "WebsiteSitemap",
"page_count": 100,
"priority_stats": {
"min": 0.3,
"max": 1.0,
"avg": 0.7
},
"last_modified_count": 80
},
{
"url": "https://example.com/blog/sitemap.xml",
"type": "WebsiteSitemap",
"page_count": 50,
"priority_stats": {
"min": 0.1,
"max": 0.9,
"avg": 0.5
},
"last_modified_count": 40
}
]
}Esto permite a los clientes de MCP comprender qué submapas de sitio podrían ser de interés para una investigación más exhaustiva. Posteriormente, puede usar el parámetro sitemap_url en get_sitemap_pages para filtrar las páginas de un submapa de sitio específico.
Analizar el contenido del mapa del sitio directamente
{
"name": "parse_sitemap_content",
"arguments": {
"content": "<?xml version=\"1.0\" encoding=\"UTF-8\"?><urlset xmlns=\"http://www.sitemaps.org/schemas/sitemap/0.9\"><url><loc>https://example.com/</loc></url></urlset>",
"include_pages": true
}
}Expresiones de gratitud
Este servidor MCP aprovecha la biblioteca de análisis de mapas de sitios definitiva
Construido con el SDK de Python del Protocolo de Contexto de Modelo
Licencia
Este proyecto está licenciado bajo la Licencia MIT. Consulte el archivo de LICENCIA para más detalles.
Available Tools
4 toolsget_sitemap_pagesB
Get all pages from a website's sitemap with optional limits and filtering options. Supports cursor-based pagination.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | Pagination cursor for fetching the next page of results | |
| include_metadata | No | Whether to include additional page metadata (priority, lastmod, etc.) | |
| limit | No | Maximum number of pages to return per page (0 for default of 100) | |
| route | No | Optional route path to filter pages by (e.g., '/blog') | |
| sitemap_url | No | Optional URL of a specific sitemap to get pages from | |
| url | Yes | The URL of the website homepage (e.g., https://example.com) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'cursor-based pagination' which is valuable behavioral information, but doesn't address other important aspects like rate limits, authentication requirements, error conditions, or what happens when no sitemap exists. The description adds some context but leaves significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely efficient - just two sentences that convey the core functionality and key behavioral characteristic (pagination). Every word earns its place with zero redundancy or unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 6 parameters, no annotations, and no output schema, the description provides basic functionality but lacks important context. It doesn't explain what the output looks like, how pagination works in practice, or what happens with edge cases. The description is complete enough to understand what the tool does at a high level, but insufficient for confident usage without additional trial-and-error.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents all 6 parameters. The description mentions 'optional limits and filtering options' which aligns with parameters like 'limit' and 'route', but adds no additional semantic meaning beyond what's in the schema. This meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get all pages') and resource ('from a website's sitemap'), making the purpose immediately understandable. However, it doesn't explicitly differentiate from sibling tools like 'get_sitemap_stats' or 'get_sitemap_tree', which likely provide different types of sitemap information.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions optional limits and filtering, but provides no guidance on when to use this tool versus alternatives like 'get_sitemap_stats' or 'parse_sitemap_content'. There's no indication of prerequisites, typical use cases, or scenarios where other tools might be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sitemap_statsC
Get comprehensive statistics about a website's sitemap structure
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL of the website homepage (e.g., https://example.com) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. While 'Get comprehensive statistics' implies a read-only operation, it doesn't specify what 'comprehensive statistics' includes, whether it requires authentication, rate limits, error conditions, or how it interacts with the sitemap (e.g., fetching vs. analyzing). This leaves significant gaps for a tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It is appropriately sized and front-loaded, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description is incomplete for a tool that presumably returns statistical data. It doesn't explain what 'comprehensive statistics' entails (e.g., counts, sizes, formats), how results are structured, or any behavioral traits. This leaves the agent with insufficient context to use the tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, clearly documenting the single 'url' parameter. The description adds no additional parameter semantics beyond what the schema provides, such as format examples or constraints. With high schema coverage, the baseline score of 3 is appropriate as the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get comprehensive statistics') and resource ('about a website's sitemap structure'), making the purpose immediately understandable. However, it doesn't explicitly differentiate this tool from its siblings (get_sitemap_pages, get_sitemap_tree, parse_sitemap_content), which all relate to sitemaps but likely serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus its siblings or alternatives. It doesn't mention prerequisites, constraints, or scenarios where this tool is preferred over others, leaving the agent to infer usage from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sitemap_treeC
Fetch and parse the sitemap tree from a website URL
| Name | Required | Description | Default |
|---|---|---|---|
| include_pages | No | Whether to include page details in the response | |
| url | Yes | The URL of the website homepage (e.g., https://example.com) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'fetch and parse,' implying network interaction and data processing, but lacks details on error handling, rate limits, authentication needs, or what the parsed tree structure looks like. For a tool that interacts with external websites, this omission is significant and leaves key behavioral aspects unclear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's function without unnecessary words. It is front-loaded with the core action ('fetch and parse') and resource ('sitemap tree'), making it easy to grasp quickly. Every part of the sentence contributes to understanding the tool's purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (involving network fetching and parsing), lack of annotations, and absence of an output schema, the description is incomplete. It doesn't address what the parsed tree output entails, potential errors (e.g., invalid URLs or sitemap formats), or performance considerations. For a tool that likely returns structured data, more context is needed to guide effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with clear documentation for both parameters ('url' and 'include_pages'). The description adds no additional parameter semantics beyond what the schema provides, such as explaining how 'include_pages' affects the parsed tree or providing examples of valid URL formats. Given the high schema coverage, a baseline score of 3 is appropriate, as the description doesn't compensate but also doesn't detract.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('fetch and parse') and resource ('sitemap tree from a website URL'), making the tool's purpose understandable. However, it doesn't explicitly differentiate from sibling tools like 'get_sitemap_pages' or 'parse_sitemap_content', which likely handle similar sitemap-related tasks, leaving some ambiguity about when to choose this specific tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It mentions fetching and parsing a sitemap tree, but doesn't specify scenarios where this is preferred over siblings like 'get_sitemap_pages' (which might retrieve individual pages) or 'parse_sitemap_content' (which might handle raw sitemap data). Without such context, users must infer usage from tool names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
parse_sitemap_contentC
Parse a sitemap directly from its XML or text content
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | The content of the sitemap (XML, text, etc.) | |
| include_pages | No | Whether to include page details in the response |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool parses sitemap content but doesn't mention error handling, output format, performance implications, or any side effects. For a tool with 2 parameters and no output schema, this is inadequate, as it leaves key behavioral traits unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence: 'Parse a sitemap directly from its XML or text content.' It is front-loaded with the core action and resource, with no wasted words. This makes it highly concise and well-structured for quick understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (2 parameters, no output schema, no annotations), the description is incomplete. It lacks details on what the parsed output looks like, how errors are handled, or any behavioral context. Without annotations or an output schema, the description should provide more context to be fully helpful, but it falls short.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the input schema already documents both parameters ('content' and 'include_pages') thoroughly. The description adds no additional semantic details beyond what the schema provides, such as examples or constraints. Thus, it meets the baseline of 3, as the schema handles the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Parse a sitemap directly from its XML or text content.' It specifies the verb ('parse') and resource ('sitemap'), and mentions the input format ('XML or text content'). However, it doesn't explicitly differentiate from sibling tools like 'get_sitemap_pages' or 'get_sitemap_tree,' which might have overlapping functionality, so it doesn't reach a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools like 'get_sitemap_pages' or 'get_sitemap_stats,' nor does it specify prerequisites or exclusions. This lack of context leaves the agent without clear usage instructions, scoring a 2 for minimal guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.0.0- First observed
get_sitemap_pages - First observed
get_sitemap_stats - First observed
get_sitemap_tree - First observed
parse_sitemap_content
TDQS
Scored across 4 tools
The tools have mostly distinct purposes with clear boundaries: get_sitemap_pages retrieves individual pages, get_sitemap_stats provides analytics, get_sitemap_tree handles hierarchical structure, and parse_sitemap_content processes raw content. However, get_sitemap_pages and get_sitemap_tree could potentially overlap in some use cases as both involve fetching sitemap data from a URL, though their outputs differ significantly.
All tool names follow a consistent verb_noun pattern with 'get_' or 'parse_' prefixes and snake_case formatting. This uniformity makes the tool set predictable and easy to understand at a glance, with no deviations in naming conventions.
With 4 tools, this server is well-scoped for its sitemap-focused purpose. Each tool serves a distinct function without redundancy, and the count is appropriate for covering core operations like retrieval, analysis, structure parsing, and content processing in this domain.
The tool set provides comprehensive coverage for sitemap operations, including fetching pages, analyzing statistics, parsing trees, and handling raw content. A minor gap exists in update or modification capabilities (e.g., editing or generating sitemaps), but this is reasonable for a read-focused server, and agents can work around this limitation.
Maintenance
Related MCP Connectors
MCP server for web extraction and rendering via AceDataCloud WebExtrator
MCP server for Google search results via SERP API
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Related MCP Servers
- FlicenseBqualityDmaintenanceAn MCP Server for Web scraping and Crawling, built using Crawl4AI224-
- AlicenseNot gradedqualityCmaintenanceJava implementation of MCP Server for Crawl4ai4MIT
- AlicenseNot gradedqualityDmaintenanceMCP server for intelligent web crawling using a ReAct agent. Adaptively explores websites and returns structured analysis.Apache 2.0
- FlicenseNot gradedqualityDmaintenancePython MCP server that scrapes web pages with JS rendering, structured metadata, tables, PDFs, screenshots, and multi-page crawling.-