Skip to main content
Glama
mugoosse

Sitemap MCP Server

by mugoosse

사이트맵 MCP 서버

모든 URL에서 사이트맵을 가져오고, 파싱하고, 시각화하여 웹사이트 아키텍처를 파악하고 사이트 구조를 분석하세요. 숨겨진 페이지를 발견하고, 수동 탐색 없이 체계적인 계층 구조를 추출할 수 있습니다.

Claude Desktop용으로 즉시 사용 가능한 프롬프트 템플릿이 포함되어 있어, 이를 통해 웹사이트 분석, 사이트맵 상태 확인, URL 추출, 누락된 콘텐츠 찾기, URL 입력만으로 시각화 생성 등의 작업을 할 수 있습니다.

특허파이파이파이썬 버전상태 대장간 배지

데모

사이트맵의 힘을 활용한 모든 웹사이트에 대한 질문에 대한 답변을 받아보세요.

도구 버튼 옆에 있는 "첨부" 버튼을 클릭하세요.

영상

그런 다음 visualize_sitemap 선택합니다.

이제 windsurf.com에 접속해 보겠습니다.

영상

그러면 사이트맵의 시각화를 볼 수 있습니다.

Related MCP server: jcrawl4ai-mcp-server

설치

uv가 설치되어 있는지 확인하세요.

Claude Desktop, Cursor 또는 Windsurf에 설치

claude_desktop_config.json , 커서 설정 등에 이 항목을 추가하세요.

지엑스피1

Claude가 실행 중이면 다시 시작하세요. 커서를 변경하려면 새로 고침을 누르거나 설정에서 MCP 서버를 활성화하세요.

Smithery를 통해 설치

Smithery를 통해 Claude Desktop용 사이트맵을 자동으로 설치하려면:

npx -y @smithery/cli install @mugoosse/sitemap --client claude

MCP 검사관

npx @modelcontextprotocol/inspector env TRANSPORT=stdio uvx sitemap-mcp-server

http://127.0.0.1:6274 에서 MCP Inspector를 열고 stdio 전송을 선택한 다음 MCP 서버에 연결합니다.

# Start the server
uvx sitemap-mcp-server

# Start the MCP Inspector in a separate terminal
npx @modelcontextprotocol/inspector connect http://127.0.0.1:8050

http://127.0.0.1:6274 에서 MCP Inspector를 열고 sse transport를 선택한 후 MCP 서버에 연결합니다.

SSE 운송

SSE 전송을 사용하려면 다음 단계를 따르세요.

  1. 서버를 시작합니다:

uvx sitemap-mcp-server
  1. MCP 클라이언트 구성(예: 커서):

{
  "mcpServers": {
    "sitemap": {
      "transport": "sse",
      "url": "http://localhost:8050/sse"
    }
  }
}

지역 개발

소스에서 프로젝트를 빌드하고 실행하는 방법에 대한 지침은 DEVELOPERS.md 가이드를 참조하세요.

용법

도구

다음 도구는 MCP 서버를 통해 사용할 수 있습니다.

  • get_sitemap_tree - 웹사이트 URL에서 사이트맵 트리를 가져와서 구문 분석합니다.

    • 인수: url (웹사이트 URL), include_pages (선택 사항, 부울)

    • 반환: 사이트맵 트리 구조의 JSON 표현

  • get_sitemap_pages - 필터링 옵션을 사용하여 웹사이트 사이트맵의 모든 페이지를 가져옵니다.

    • 인수: url (웹사이트 URL), limit (선택 사항), include_metadata (선택 사항), route (선택 사항), sitemap_url (선택 사항), cursor (선택 사항)

    • 반환: 페이지 매김 메타데이터가 포함된 페이지의 JSON 목록

  • get_sitemap_stats - 웹사이트 사이트맵에 대한 통계를 가져옵니다.

    • 인수: url (웹사이트 URL)

    • 반환: 페이지 수, 수정 날짜, 하위 사이트맵 세부 정보를 포함한 사이트맵 통계가 포함된 JSON 객체

  • parse_sitemap_content - XML 또는 텍스트 콘텐츠에서 사이트맵을 직접 구문 분석합니다.

    • 인수: content (사이트맵 XML 콘텐츠), include_pages (선택 사항, 부울)

    • 반환: 구문 분석된 사이트맵의 JSON 표현

프롬프트

서버에는 Claude Desktop에서 템플릿으로 표시되는 즉시 사용 가능한 프롬프트가 포함되어 있습니다. 서버를 설치하면 "템플릿" 메뉴에서 다음 템플릿을 확인할 수 있습니다(메시지 입력란 옆의 + 아이콘 클릭).

  • 사이트맵 분석 : 웹사이트 사이트맵의 포괄적인 구조 분석을 제공합니다.

  • 사이트맵 상태 확인 : 사이트맵의 SEO 및 상태 측정 항목을 평가합니다.

  • 사이트맵에서 URL 추출 : 사이트맵에서 특정 URL을 추출하고 필터링합니다.

  • 사이트맵에서 누락된 콘텐츠 찾기 : 웹사이트 사이트맵의 콘텐츠 간격을 식별합니다.

  • 사이트맵 구조 시각화 : 사이트맵 구조의 Mermaid.js 다이어그램 시각화를 생성합니다.

이러한 프롬프트를 사용하려면:

  1. Claude Desktop의 메시지 입력 옆에 있는 + 아이콘을 클릭하세요.

  2. 목록에서 원하는 템플릿을 선택하세요

  3. 메시지가 표시되면 웹사이트 URL을 입력하세요

  4. Claude는 적절한 사이트맵 분석을 실행합니다.

예시

전체 사이트맵 가져오기

{
  "name": "get_sitemap_tree",
  "arguments": {
    "url": "https://example.com",
    "include_pages": true
  }
}

필터링 및 페이지 매김을 사용하여 페이지 가져오기

경로별 필터링
{
  "name": "get_sitemap_pages",
  "arguments": {
    "url": "https://example.com",
    "limit": 100,
    "include_metadata": true,
    "route": "/blog/"
  }
}
특정 하위 사이트맵으로 필터링
{
  "name": "get_sitemap_pages",
  "arguments": {
    "url": "https://example.com",
    "limit": 100,
    "include_metadata": true,
    "sitemap_url": "https://example.com/blog-sitemap.xml"
  }
}
커서 기반 페이지 매김

서버는 대규모 사이트맵을 효율적으로 처리하기 위해 MCP 커서 기반 페이지 매김을 구현합니다.

초기 요청:

{
  "name": "get_sitemap_pages",
  "arguments": {
    "url": "https://example.com",
    "limit": 50
  }
}

페이지 번호가 포함된 응답:

{
  "base_url": "https://example.com",
  "pages": [...],  // First batch of pages
  "limit": 50,
  "nextCursor": "eyJwYWdlIjoxfQ=="
}

커서를 사용한 후속 요청:

{
  "name": "get_sitemap_pages",
  "arguments": {
    "url": "https://example.com",
    "limit": 50,
    "cursor": "eyJwYWdlIjoxfQ=="
  }
}

더 이상 결과가 없으면 nextCursor 필드가 응답에서 사라집니다.

사이트맵 통계 가져오기

{
  "name": "get_sitemap_stats",
  "arguments": {
    "url": "https://example.com"
  }
}

응답에는 전체 통계와 각 하위 사이트맵에 대한 자세한 통계가 모두 포함됩니다.

{
  "total": {
    "url": "https://example.com",
    "page_count": 150,
    "sitemap_count": 3,
    "sitemap_types": ["WebsiteSitemap", "NewsSitemap"],
    "priority_stats": {
      "min": 0.1,
      "max": 1.0,
      "avg": 0.65
    },
    "last_modified_count": 120
  },
  "subsitemaps": [
    {
      "url": "https://example.com/sitemap.xml",
      "type": "WebsiteSitemap",
      "page_count": 100,
      "priority_stats": {
        "min": 0.3,
        "max": 1.0,
        "avg": 0.7
      },
      "last_modified_count": 80
    },
    {
      "url": "https://example.com/blog/sitemap.xml",
      "type": "WebsiteSitemap",
      "page_count": 50,
      "priority_stats": {
        "min": 0.1,
        "max": 0.9,
        "avg": 0.5
      },
      "last_modified_count": 40
    }
  ]
}

이를 통해 MCP 클라이언트는 추가 조사가 필요한 하위 사이트맵을 파악할 수 있습니다. 그런 다음 get_sitemap_pages 의 sitemap_url 매개변수를 사용하여 특정 하위 사이트맵의 페이지를 필터링할 수 있습니다.

사이트맵 콘텐츠 직접 구문 분석

{
  "name": "parse_sitemap_content",
  "arguments": {
    "content": "<?xml version=\"1.0\" encoding=\"UTF-8\"?><urlset xmlns=\"http://www.sitemaps.org/schemas/sitemap/0.9\"><url><loc>https://example.com/</loc></url></urlset>",
    "include_pages": true
  }
}

감사의 말

특허

이 프로젝트는 MIT 라이선스에 따라 라이선스가 부여됩니다. 자세한 내용은 라이선스 파일을 참조하세요.

Available Tools

4 tools
get_sitemap_pagesB

Get all pages from a website's sitemap with optional limits and filtering options. Supports cursor-based pagination.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoPagination cursor for fetching the next page of results
include_metadataNoWhether to include additional page metadata (priority, lastmod, etc.)
limitNoMaximum number of pages to return per page (0 for default of 100)
routeNoOptional route path to filter pages by (e.g., '/blog')
sitemap_urlNoOptional URL of a specific sitemap to get pages from
urlYesThe URL of the website homepage (e.g., https://example.com)

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'cursor-based pagination' which is valuable behavioral information, but doesn't address other important aspects like rate limits, authentication requirements, error conditions, or what happens when no sitemap exists. The description adds some context but leaves significant gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely efficient - just two sentences that convey the core functionality and key behavioral characteristic (pagination). Every word earns its place with zero redundancy or unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 6 parameters, no annotations, and no output schema, the description provides basic functionality but lacks important context. It doesn't explain what the output looks like, how pagination works in practice, or what happens with edge cases. The description is complete enough to understand what the tool does at a high level, but insufficient for confident usage without additional trial-and-error.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents all 6 parameters. The description mentions 'optional limits and filtering options' which aligns with parameters like 'limit' and 'route', but adds no additional semantic meaning beyond what's in the schema. This meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get all pages') and resource ('from a website's sitemap'), making the purpose immediately understandable. However, it doesn't explicitly differentiate from sibling tools like 'get_sitemap_stats' or 'get_sitemap_tree', which likely provide different types of sitemap information.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions optional limits and filtering, but provides no guidance on when to use this tool versus alternatives like 'get_sitemap_stats' or 'parse_sitemap_content'. There's no indication of prerequisites, typical use cases, or scenarios where other tools might be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_sitemap_statsC

Get comprehensive statistics about a website's sitemap structure

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL of the website homepage (e.g., https://example.com)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. While 'Get comprehensive statistics' implies a read-only operation, it doesn't specify what 'comprehensive statistics' includes, whether it requires authentication, rate limits, error conditions, or how it interacts with the sitemap (e.g., fetching vs. analyzing). This leaves significant gaps for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It is appropriately sized and front-loaded, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description is incomplete for a tool that presumably returns statistical data. It doesn't explain what 'comprehensive statistics' entails (e.g., counts, sizes, formats), how results are structured, or any behavioral traits. This leaves the agent with insufficient context to use the tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, clearly documenting the single 'url' parameter. The description adds no additional parameter semantics beyond what the schema provides, such as format examples or constraints. With high schema coverage, the baseline score of 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get comprehensive statistics') and resource ('about a website's sitemap structure'), making the purpose immediately understandable. However, it doesn't explicitly differentiate this tool from its siblings (get_sitemap_pages, get_sitemap_tree, parse_sitemap_content), which all relate to sitemaps but likely serve different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus its siblings or alternatives. It doesn't mention prerequisites, constraints, or scenarios where this tool is preferred over others, leaving the agent to infer usage from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_sitemap_treeC

Fetch and parse the sitemap tree from a website URL

ParametersJSON Schema
NameRequiredDescriptionDefault
include_pagesNoWhether to include page details in the response
urlYesThe URL of the website homepage (e.g., https://example.com)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'fetch and parse,' implying network interaction and data processing, but lacks details on error handling, rate limits, authentication needs, or what the parsed tree structure looks like. For a tool that interacts with external websites, this omission is significant and leaves key behavioral aspects unclear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's function without unnecessary words. It is front-loaded with the core action ('fetch and parse') and resource ('sitemap tree'), making it easy to grasp quickly. Every part of the sentence contributes to understanding the tool's purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (involving network fetching and parsing), lack of annotations, and absence of an output schema, the description is incomplete. It doesn't address what the parsed tree output entails, potential errors (e.g., invalid URLs or sitemap formats), or performance considerations. For a tool that likely returns structured data, more context is needed to guide effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with clear documentation for both parameters ('url' and 'include_pages'). The description adds no additional parameter semantics beyond what the schema provides, such as explaining how 'include_pages' affects the parsed tree or providing examples of valid URL formats. Given the high schema coverage, a baseline score of 3 is appropriate, as the description doesn't compensate but also doesn't detract.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('fetch and parse') and resource ('sitemap tree from a website URL'), making the tool's purpose understandable. However, it doesn't explicitly differentiate from sibling tools like 'get_sitemap_pages' or 'parse_sitemap_content', which likely handle similar sitemap-related tasks, leaving some ambiguity about when to choose this specific tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It mentions fetching and parsing a sitemap tree, but doesn't specify scenarios where this is preferred over siblings like 'get_sitemap_pages' (which might retrieve individual pages) or 'parse_sitemap_content' (which might handle raw sitemap data). Without such context, users must infer usage from tool names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

parse_sitemap_contentC

Parse a sitemap directly from its XML or text content

ParametersJSON Schema
NameRequiredDescriptionDefault
contentYesThe content of the sitemap (XML, text, etc.)
include_pagesNoWhether to include page details in the response

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool parses sitemap content but doesn't mention error handling, output format, performance implications, or any side effects. For a tool with 2 parameters and no output schema, this is inadequate, as it leaves key behavioral traits unspecified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence: 'Parse a sitemap directly from its XML or text content.' It is front-loaded with the core action and resource, with no wasted words. This makes it highly concise and well-structured for quick understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (2 parameters, no output schema, no annotations), the description is incomplete. It lacks details on what the parsed output looks like, how errors are handled, or any behavioral context. Without annotations or an output schema, the description should provide more context to be fully helpful, but it falls short.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the input schema already documents both parameters ('content' and 'include_pages') thoroughly. The description adds no additional semantic details beyond what the schema provides, such as examples or constraints. Thus, it meets the baseline of 3, as the schema handles the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Parse a sitemap directly from its XML or text content.' It specifies the verb ('parse') and resource ('sitemap'), and mentions the input format ('XML or text content'). However, it doesn't explicitly differentiate from sibling tools like 'get_sitemap_pages' or 'get_sitemap_tree,' which might have overlapping functionality, so it doesn't reach a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools like 'get_sitemap_pages' or 'get_sitemap_stats,' nor does it specify prerequisites or exclusions. This lack of context leaves the agent without clear usage instructions, scoring a 2 for minimal guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv1.0.0
    • First observedget_sitemap_pages
    • First observedget_sitemap_stats
    • First observedget_sitemap_tree
    • First observedparse_sitemap_content

TDQS

B3.4/5.0

Scored across 4 tools

Disambiguation4/5

The tools have mostly distinct purposes with clear boundaries: get_sitemap_pages retrieves individual pages, get_sitemap_stats provides analytics, get_sitemap_tree handles hierarchical structure, and parse_sitemap_content processes raw content. However, get_sitemap_pages and get_sitemap_tree could potentially overlap in some use cases as both involve fetching sitemap data from a URL, though their outputs differ significantly.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern with 'get_' or 'parse_' prefixes and snake_case formatting. This uniformity makes the tool set predictable and easy to understand at a glance, with no deviations in naming conventions.

Tool Count5/5

With 4 tools, this server is well-scoped for its sitemap-focused purpose. Each tool serves a distinct function without redundancy, and the count is appropriate for covering core operations like retrieval, analysis, structure parsing, and content processing in this domain.

Completeness4/5

The tool set provides comprehensive coverage for sitemap operations, including fetching pages, analyzing statistics, parsing trees, and handling raw content. A minor gap exists in update or modification capabilities (e.g., editing or generating sitemaps), but this is reasonable for a read-focused server, and agents can work around this limitation.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers