Sitemap MCP Server
网站地图 MCP 服务器
通过从任意 URL 获取、解析和可视化站点地图,探索网站架构并分析网站结构。无需手动探索,即可发现隐藏页面并提取有序的层次结构。
包括适用于 Claude Desktop 的现成提示模板,让您可以分析网站、检查站点地图健康状况、提取 URL、查找缺失内容以及仅通过 URL 输入创建可视化效果。
演示
利用站点地图的功能获取有关任何网站的问题的答案。
点击工具按钮旁边的“附加”按钮:
然后选择visualize_sitemap :
现在我们输入windsurf.com:
我们得到了站点地图的可视化效果:
Related MCP server: jcrawl4ai-mcp-server
安装
确保已安装uv 。
在 Claude Desktop、Cursor 或 Windsurf 中安装
将此条目添加到您的claude_desktop_config.json 、光标设置等中:
{
"mcpServers": {
"sitemap": {
"command": "uvx",
"args": ["sitemap-mcp-server"],
"env": { "TRANSPORT": "stdio" }
}
}
}如果 Claude 正在运行,请重启它。对于 Cursor,只需按刷新并/或在设置中启用 MCP 服务器即可。
通过 Smithery 安装
要通过Smithery自动安装 Claude Desktop 的站点地图:
npx -y @smithery/cli install @mugoosse/sitemap --client claudeMCP 检查器
npx @modelcontextprotocol/inspector env TRANSPORT=stdio uvx sitemap-mcp-server打开 MCP Inspector,网址为http://127.0.0.1:6274 ,选择stdio transport,然后连接到 MCP 服务器。
# Start the server
uvx sitemap-mcp-server
# Start the MCP Inspector in a separate terminal
npx @modelcontextprotocol/inspector connect http://127.0.0.1:8050打开 MCP 检查器http://127.0.0.1:6274 ,选择sse transport,然后连接到 MCP 服务器。
上交所运输
如果您想使用 SSE 传输,请按照以下步骤操作:
启动服务器:
uvx sitemap-mcp-server配置您的 MCP 客户端,例如光标:
{
"mcpServers": {
"sitemap": {
"transport": "sse",
"url": "http://localhost:8050/sse"
}
}
}本地开发
有关从源代码构建和运行项目的说明,请参阅DEVELOPERS.md指南。
用法
工具
可通过 MCP 服务器使用以下工具:
get_sitemap_tree - 从网站 URL 获取并解析站点地图树
参数:
url(网站 URL)、include_pages(可选,布尔值)返回:站点地图树结构的 JSON 表示
get_sitemap_pages - 使用过滤选项获取网站站点地图中的所有页面
参数:
url(网站 URL)、limit(可选)、include_metadata(可选)、route(可选)、sitemap_url(可选)、cursor(可选)返回:带有分页元数据的页面的 JSON 列表
get_sitemap_stats - 获取网站站点地图的统计信息
参数:
url(网站 URL)返回:带有站点地图统计信息的 JSON 对象,包括页面计数、修改日期和子站点地图详细信息
parse_sitemap_content - 直接从 XML 或文本内容解析站点地图
参数:
content(站点地图 XML 内容)、include_pages(可选,布尔值)返回:已解析站点地图的 JSON 表示
提示
服务器包含现成的提示,这些提示在 Claude Desktop 中以模板形式显示。安装服务器后,您将在“模板”菜单中看到这些模板(点击消息输入旁边的 + 图标):
分析站点地图:提供网站站点地图的全面结构分析
检查站点地图健康状况:评估站点地图的 SEO 和健康指标
从站点地图中提取 URL :从站点地图中提取并过滤特定的 URL
在站点地图中查找缺失内容:识别网站站点地图中的内容空白
可视化站点地图结构:创建 Mermaid.js 图表可视化站点地图结构
要使用这些提示:
单击 Claude Desktop 中消息输入旁边的 + 图标
从列表中选择所需的模板
出现提示时填写网站网址
Claude 将执行适当的站点地图分析
示例
获取完整的站点地图
{
"name": "get_sitemap_tree",
"arguments": {
"url": "https://example.com",
"include_pages": true
}
}获取具有过滤和分页功能的页面
按路线过滤
{
"name": "get_sitemap_pages",
"arguments": {
"url": "https://example.com",
"limit": 100,
"include_metadata": true,
"route": "/blog/"
}
}按特定子站点地图过滤
{
"name": "get_sitemap_pages",
"arguments": {
"url": "https://example.com",
"limit": 100,
"include_metadata": true,
"sitemap_url": "https://example.com/blog-sitemap.xml"
}
}基于游标的分页
服务器实现了基于 MCP 游标的分页,以高效处理大型站点地图:
初始请求:
{
"name": "get_sitemap_pages",
"arguments": {
"url": "https://example.com",
"limit": 50
}
}分页响应:
{
"base_url": "https://example.com",
"pages": [...], // First batch of pages
"limit": 50,
"nextCursor": "eyJwYWdlIjoxfQ=="
}带有游标的后续请求:
{
"name": "get_sitemap_pages",
"arguments": {
"url": "https://example.com",
"limit": 50,
"cursor": "eyJwYWdlIjoxfQ=="
}
}当没有更多结果时, nextCursor字段将不会出现在响应中。
获取站点地图统计信息
{
"name": "get_sitemap_stats",
"arguments": {
"url": "https://example.com"
}
}响应包括每个子站点地图的总体统计信息和详细统计信息:
{
"total": {
"url": "https://example.com",
"page_count": 150,
"sitemap_count": 3,
"sitemap_types": ["WebsiteSitemap", "NewsSitemap"],
"priority_stats": {
"min": 0.1,
"max": 1.0,
"avg": 0.65
},
"last_modified_count": 120
},
"subsitemaps": [
{
"url": "https://example.com/sitemap.xml",
"type": "WebsiteSitemap",
"page_count": 100,
"priority_stats": {
"min": 0.3,
"max": 1.0,
"avg": 0.7
},
"last_modified_count": 80
},
{
"url": "https://example.com/blog/sitemap.xml",
"type": "WebsiteSitemap",
"page_count": 50,
"priority_stats": {
"min": 0.1,
"max": 0.9,
"avg": 0.5
},
"last_modified_count": 40
}
]
}这样,MCP 客户端就能了解哪些子站点地图可能值得进一步研究。然后,您可以使用get_sitemap_pages中的sitemap_url参数来筛选特定子站点地图中的页面。
直接解析站点地图内容
{
"name": "parse_sitemap_content",
"arguments": {
"content": "<?xml version=\"1.0\" encoding=\"UTF-8\"?><urlset xmlns=\"http://www.sitemaps.org/schemas/sitemap/0.9\"><url><loc>https://example.com/</loc></url></urlset>",
"include_pages": true
}
}致谢
此 MCP 服务器利用了ultimate-sitemap-parser库
使用模型上下文协议Python SDK 构建
执照
本项目遵循 MIT 许可证。详情请参阅LICENSE文件。
Available Tools
4 toolsget_sitemap_pagesB
Get all pages from a website's sitemap with optional limits and filtering options. Supports cursor-based pagination.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | Pagination cursor for fetching the next page of results | |
| include_metadata | No | Whether to include additional page metadata (priority, lastmod, etc.) | |
| limit | No | Maximum number of pages to return per page (0 for default of 100) | |
| route | No | Optional route path to filter pages by (e.g., '/blog') | |
| sitemap_url | No | Optional URL of a specific sitemap to get pages from | |
| url | Yes | The URL of the website homepage (e.g., https://example.com) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'cursor-based pagination' which is valuable behavioral information, but doesn't address other important aspects like rate limits, authentication requirements, error conditions, or what happens when no sitemap exists. The description adds some context but leaves significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely efficient - just two sentences that convey the core functionality and key behavioral characteristic (pagination). Every word earns its place with zero redundancy or unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 6 parameters, no annotations, and no output schema, the description provides basic functionality but lacks important context. It doesn't explain what the output looks like, how pagination works in practice, or what happens with edge cases. The description is complete enough to understand what the tool does at a high level, but insufficient for confident usage without additional trial-and-error.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents all 6 parameters. The description mentions 'optional limits and filtering options' which aligns with parameters like 'limit' and 'route', but adds no additional semantic meaning beyond what's in the schema. This meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get all pages') and resource ('from a website's sitemap'), making the purpose immediately understandable. However, it doesn't explicitly differentiate from sibling tools like 'get_sitemap_stats' or 'get_sitemap_tree', which likely provide different types of sitemap information.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions optional limits and filtering, but provides no guidance on when to use this tool versus alternatives like 'get_sitemap_stats' or 'parse_sitemap_content'. There's no indication of prerequisites, typical use cases, or scenarios where other tools might be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sitemap_statsC
Get comprehensive statistics about a website's sitemap structure
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL of the website homepage (e.g., https://example.com) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. While 'Get comprehensive statistics' implies a read-only operation, it doesn't specify what 'comprehensive statistics' includes, whether it requires authentication, rate limits, error conditions, or how it interacts with the sitemap (e.g., fetching vs. analyzing). This leaves significant gaps for a tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It is appropriately sized and front-loaded, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description is incomplete for a tool that presumably returns statistical data. It doesn't explain what 'comprehensive statistics' entails (e.g., counts, sizes, formats), how results are structured, or any behavioral traits. This leaves the agent with insufficient context to use the tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, clearly documenting the single 'url' parameter. The description adds no additional parameter semantics beyond what the schema provides, such as format examples or constraints. With high schema coverage, the baseline score of 3 is appropriate as the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get comprehensive statistics') and resource ('about a website's sitemap structure'), making the purpose immediately understandable. However, it doesn't explicitly differentiate this tool from its siblings (get_sitemap_pages, get_sitemap_tree, parse_sitemap_content), which all relate to sitemaps but likely serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus its siblings or alternatives. It doesn't mention prerequisites, constraints, or scenarios where this tool is preferred over others, leaving the agent to infer usage from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sitemap_treeC
Fetch and parse the sitemap tree from a website URL
| Name | Required | Description | Default |
|---|---|---|---|
| include_pages | No | Whether to include page details in the response | |
| url | Yes | The URL of the website homepage (e.g., https://example.com) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'fetch and parse,' implying network interaction and data processing, but lacks details on error handling, rate limits, authentication needs, or what the parsed tree structure looks like. For a tool that interacts with external websites, this omission is significant and leaves key behavioral aspects unclear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's function without unnecessary words. It is front-loaded with the core action ('fetch and parse') and resource ('sitemap tree'), making it easy to grasp quickly. Every part of the sentence contributes to understanding the tool's purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (involving network fetching and parsing), lack of annotations, and absence of an output schema, the description is incomplete. It doesn't address what the parsed tree output entails, potential errors (e.g., invalid URLs or sitemap formats), or performance considerations. For a tool that likely returns structured data, more context is needed to guide effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with clear documentation for both parameters ('url' and 'include_pages'). The description adds no additional parameter semantics beyond what the schema provides, such as explaining how 'include_pages' affects the parsed tree or providing examples of valid URL formats. Given the high schema coverage, a baseline score of 3 is appropriate, as the description doesn't compensate but also doesn't detract.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('fetch and parse') and resource ('sitemap tree from a website URL'), making the tool's purpose understandable. However, it doesn't explicitly differentiate from sibling tools like 'get_sitemap_pages' or 'parse_sitemap_content', which likely handle similar sitemap-related tasks, leaving some ambiguity about when to choose this specific tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It mentions fetching and parsing a sitemap tree, but doesn't specify scenarios where this is preferred over siblings like 'get_sitemap_pages' (which might retrieve individual pages) or 'parse_sitemap_content' (which might handle raw sitemap data). Without such context, users must infer usage from tool names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
parse_sitemap_contentC
Parse a sitemap directly from its XML or text content
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | The content of the sitemap (XML, text, etc.) | |
| include_pages | No | Whether to include page details in the response |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool parses sitemap content but doesn't mention error handling, output format, performance implications, or any side effects. For a tool with 2 parameters and no output schema, this is inadequate, as it leaves key behavioral traits unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence: 'Parse a sitemap directly from its XML or text content.' It is front-loaded with the core action and resource, with no wasted words. This makes it highly concise and well-structured for quick understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (2 parameters, no output schema, no annotations), the description is incomplete. It lacks details on what the parsed output looks like, how errors are handled, or any behavioral context. Without annotations or an output schema, the description should provide more context to be fully helpful, but it falls short.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the input schema already documents both parameters ('content' and 'include_pages') thoroughly. The description adds no additional semantic details beyond what the schema provides, such as examples or constraints. Thus, it meets the baseline of 3, as the schema handles the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Parse a sitemap directly from its XML or text content.' It specifies the verb ('parse') and resource ('sitemap'), and mentions the input format ('XML or text content'). However, it doesn't explicitly differentiate from sibling tools like 'get_sitemap_pages' or 'get_sitemap_tree,' which might have overlapping functionality, so it doesn't reach a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools like 'get_sitemap_pages' or 'get_sitemap_stats,' nor does it specify prerequisites or exclusions. This lack of context leaves the agent without clear usage instructions, scoring a 2 for minimal guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.0.0- First observed
get_sitemap_pages - First observed
get_sitemap_stats - First observed
get_sitemap_tree - First observed
parse_sitemap_content
TDQS
Scored across 4 tools
The tools have mostly distinct purposes with clear boundaries: get_sitemap_pages retrieves individual pages, get_sitemap_stats provides analytics, get_sitemap_tree handles hierarchical structure, and parse_sitemap_content processes raw content. However, get_sitemap_pages and get_sitemap_tree could potentially overlap in some use cases as both involve fetching sitemap data from a URL, though their outputs differ significantly.
All tool names follow a consistent verb_noun pattern with 'get_' or 'parse_' prefixes and snake_case formatting. This uniformity makes the tool set predictable and easy to understand at a glance, with no deviations in naming conventions.
With 4 tools, this server is well-scoped for its sitemap-focused purpose. Each tool serves a distinct function without redundancy, and the count is appropriate for covering core operations like retrieval, analysis, structure parsing, and content processing in this domain.
The tool set provides comprehensive coverage for sitemap operations, including fetching pages, analyzing statistics, parsing trees, and handling raw content. A minor gap exists in update or modification capabilities (e.g., editing or generating sitemaps), but this is reasonable for a read-focused server, and agents can work around this limitation.
Maintenance
Related MCP Connectors
MCP server for web extraction and rendering via AceDataCloud WebExtrator
MCP server for Google search results via SERP API
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Related MCP Servers
- FlicenseBqualityDmaintenanceAn MCP Server for Web scraping and Crawling, built using Crawl4AI224-
- AlicenseNot gradedqualityCmaintenanceJava implementation of MCP Server for Crawl4ai4MIT
- AlicenseNot gradedqualityDmaintenanceMCP server for intelligent web crawling using a ReAct agent. Adaptively explores websites and returns structured analysis.Apache 2.0
- FlicenseNot gradedqualityDmaintenancePython MCP server that scrapes web pages with JS rendering, structured metadata, tables, PDFs, screenshots, and multi-page crawling.-