seo-audit-mcp
seo-audit-mcp
Claude(또는 모든 MCP 클라이언트)가 라이브 웹사이트의 기술적 SEO를 감사할 수 있게 해주는 MCP 서버입니다: 사이트맵 적용 범위, 페이지별 문제, 리다이렉트 체인.
평범한 언어로 요청하세요 — "mortgagecalculatortools.com을 감사하고 Google에 절대 알려지지 않은 페이지를 알려줘" — 그러면 모델이 도구를 호출하고, 사이트를 크롤링하며, 구체적인 답변을 제공합니다.
해결하는 문제
사이트의 sitemap.xml은 Google에 어떤 페이지가 존재하는지 알리는 방법입니다. 페이지가 거기서 누락되어도 오류도 경고도 발생하지 않습니다. 단지 노출이 누적되지 않을 뿐입니다. 수동으로 확인하려면 파일 시스템 목록과 XML 파일을 비교(diff)해야 하므로, 실제로는 아무도 하지 않습니다.
사례 연구: 25페이지 누락이 사실은 정확했다
이 도구를 처음 적용한 사이트에는 디스크에 HTML 파일 125개, 사이트맵에 URL 100개가 있었다. 25페이지 누락은 버그로 기록되어 누군가에게 배정되는 종류의 발견입니다.
한 번의 sitemap_coverage 호출로 누락이 드러났고, 샘플에 대한 한 번의 audit_urls 호출로 설명되었습니다: 25개 페이지 모두 <meta name="robots" content="noindex, follow">를 포함하고 있었습니다. 이들은 의도적으로 색인에서 제외된 두 콘텐츠 클러스터였고, 사이트맵이 이를 생략한 것은 완전히 정확했습니다. 이후 파일 시스템과 대조해 확인한 결과: 디스크에 25개의 noindex 페이지가 있었고, 사이트맵에서 빠진 것도 같은 25개였으며, 잘못 포함된 noindex 페이지는 0개였습니다. 완벽한 일관성입니다.
그것이 유용한 결과입니다. 단순한 커버리지 숫자("125 vs 100")만으로는 결함으로 읽혀 누군가의 하루를 소모하게 합니다. 하지만 커버리지에 페이지별 noindex 상태를 더하면 그 질문은 1분 만에 해결됩니다. 이 도구는 실제 누락을 찾는 것만큼이나 잘못된 경보를 제거하는 데도 가치가 있습니다 — 그래서 audit_urls는 단순히 URL 개수만 세지 않고 페이지마다 noindex를 보고합니다.
Related MCP server: web-audit-mcp
도구
도구 | 기능 |
|
|
| URL을 동시에 크롤링하고 페이지별 문제를 보고합니다: 깨진 상태, 리다이렉트 체인, 누락/과도한 길이의 |
| 알고 있는 URL 목록과 사이트맵을 비교(diff)합니다 → 사이트맵에서 누락된 것, 선언되었지만 죽은(dead) 것 |
| 리다이렉트 체인을 추적하고, 다중 홉 체인과 4xx/5xx로 끝나는 체인에 플래그를 표시합니다 — URL 구조 변경 후 사용하세요 |
모든 도구는 페이지별 issues 목록과 집계된 issue_summary를 포함한 구조화된 JSON을 반환하므로, 모델은 원시 HTML을 다시 읽는 대신 숫자를 기준으로 추론할 수 있습니다.
설치
pip install -e .Python 3.10+가 필요합니다. 의존성: mcp>=2.0.0, httpx.
Claude Code에 연결하기
프로젝트의 .mcp.json에 추가하세요(전역 사용 시 ~/.claude.json):
{
"mcpServers": {
"seo-audit": {
"command": "python",
"args": ["-m", "seo_audit_mcp"]
}
}
}Claude Desktop의 경우 동일한 블록을 claude_desktop_config.json에 넣습니다.
그런 다음 그냥 요청하세요:
https://example.com/sitemap.xml의 사이트맵을 가져와서 처음 20개 URL을 감사하고, 이슈를 빈도별로 요약해 줘.
직접 실행하기
python -m seo_audit_mcp # stdio transport설계 노트
강조할 가치가 있는 세 가지 결정 사항이 있습니다. 이는 데모와 클라이언트 프로덕션 사이트에 사용할 수 있는 도구의 차이를 만듭니다:
크롤링은 구조적으로 속도가 제한됩니다. fetch_many는 최대 16개의 동시 요청으로 제한된 asyncio.Semaphore 뒤에서 실행되며, 모든 도구는 입력을 제한합니다. 그 상한선이 없다면 500-URL 사이트맵은 한 번에 500개의 소켓을 열고 대상 호스트에 공격으로 읽힐 수 있습니다. 크롤링은 대상 사이트가 부담하는 비용이므로, 상한선은 도구 표면에서 위로 구성할 수 없습니다.
가져오기 실패가 실행을 중단시키지 않습니다. fetch_one은 httpx.HTTPError를 잡아 예외를 발생시키는 대신 반환된 PageAudit에 기록합니다. 200-URL 크롤링에서 호스트 하나가 죽어도 199개의 좋은 결과를 잃는 대신 한 행만 저하됩니다.
파싱은 의도적으로 관대합니다. 실제 HTML은 잘못된 형식인 경우가 많아서 크롤링 중에 예외를 던지는 엄격한 파서는 오히려 부담입니다. 추출기들은 예외를 던지는 대신 None을 반환하는 허용적인 정규표현식입니다. 단, 함정은 처리합니다: <script>와 <style> 본문은 단어 수 계산과 제목 추출 전에 제거되므로 JS 문자열 리터럴 안의 <h1>은 제목으로 계산되지 않으며, 상대 canonical은 페이지 URL을 기준으로 해석됩니다.
normalize_url은 의도적으로 후행 슬래시를 제거하지 않습니다: /a와 /a/는 실제로 다른 페이지일 수 있으며, 이를 합치면 이 도구가 드러내기 위해 존재하는 중복 콘텐츠 문제가 숨겨질 수 있습니다.
테스트
pip install -e ".[dev]"
pytest테스트 스위트는 네트워크 없이 실행됩니다 — HTTP는 httpx.MockTransport를 통해 실행되므로 CI와 비행기 안에서도 작동합니다. 프로덕션에서 문제가 되는 파싱 엣지 케이스를 다룹니다: 스크립트에 포함된 제목, 네임스페이스 없는 사이트맵, 상대 canonical, 스타일이 적용된 HTML 404를 200 상태로 반환하는 사이트맵 URL, 그리고 "제목 누락"으로 잘못 보고되는 비-HTML 콘텐츠 타입.
라이선스
MIT
Available Tools
4 toolsaudit_urlsA
Crawl a list of URLs and report per-page technical SEO issues: broken status codes, redirect chains, missing or over-length titles and meta descriptions, missing or duplicate H1, missing canonical, noindex, and thin content. Returns a per-URL breakdown plus an issue summary.
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | ||
| concurrency | No | ||
| timeout_seconds | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral disclosure burden. It usefully states that the tool crawls URLs and returns a per-URL breakdown plus an issue summary, but it does not disclose operational behaviors such as crawl duration, rate limiting, redirect-following details, or auth/network requirements. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single information-dense sentence that front-loads the action and resource, then lists issue categories and the return shape. It avoids repetition and wastes no words, though the enumeration makes it slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description does not need to detail return values, and it does give a useful high-level summary. However, with zero annotations, zero schema descriptions, and no usage guidance, the description leaves concurrency/timeout semantics and tool-selection boundaries undocumented, making it only moderately complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it never names or explains urls, concurrency, or timeout_seconds. The parameter names are somewhat self-explanatory, yet the description adds no detail about what concurrency or timeout_seconds control, how URLs should be formatted, or whether limits apply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource pair ('Crawl a list of URLs and report per-page technical SEO issues') and then enumerates the exact issue categories. This clearly distinguishes it from siblings like fetch_sitemap and sitemap_coverage by centering on per-URL technical SEO auditing, even though it overlaps with check_redirects on redirect chains.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies its use case by listing SEO checks, but it never explicitly states when to prefer this over check_redirects, fetch_sitemap, or sitemap_coverage, nor does it give any 'when not to use' guidance. An agent can infer the purpose but must decide on selection criteria without direct help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_redirectsA
Trace the redirect chain for each URL and flag chains longer than one hop, redirect loops, and URLs that resolve to a 4xx/5xx. Use after a site migration or a URL-structure change.
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It clearly discloses what the tool does: traces chains, flags one-hop violations, loops, and error responses. It could add operational details like network cost or rate limits, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, and the main behavior is front-loaded. The second sentence supplies a practical trigger for use. Every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with an output schema present, the description covers what the tool does and when to use it. It does not fully cover input format details, but the missing information is minor given the low complexity and available output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds little about the 'urls' parameter beyond saying 'for each URL.' It does not clarify expected URL format (absolute vs relative), whether schemes are required, or any limits on list length, so the agent gets almost no parameter guidance beyond the bare schema type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Trace the redirect chain for each URL' and names concrete detection outcomes (long chains, loops, 4xx/5xx). This is distinct from siblings like fetch_sitemap or audit_urls, making the tool's purpose immediately clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use after a site migration or a URL-structure change,' which gives a clear context for when this tool is appropriate. It does not mention alternatives or when not to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch_sitemapA
Fetch and parse a sitemap.xml, following sitemap-index nesting, and return every page URL it declares. Call this first when auditing a site you do not have a URL list for.
| Name | Required | Description | Default |
|---|---|---|---|
| sitemap_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. It discloses that the tool follows sitemap-index nesting and returns every declared page URL, which are meaningful behavioral details beyond what the schema shows. It does not mention failure modes or network behavior, but for a straightforward fetch-and-parse read operation this is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The first sentence states what the tool does and the second gives usage guidance. The most important behavioral details are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with an output schema present, the description is complete: it explains what the tool does, how it behaves with index nesting, what it returns, and when to call it. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. It indirectly clarifies that sitemap_url should point to a sitemap.xml and that index nesting is followed, but it does not explicitly describe the URL format, required scheme, or example values. Because the single parameter is highly self-evident from the tool name and description, this is adequate but not enriched.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Fetch and parse a sitemap.xml' and the specific outcome: 'return every page URL it declares.' It also mentions the non-obvious behavior of following sitemap-index nesting, which distinguishes this from simply fetching one XML file. The phrase 'Call this first when auditing a site' also separates it from siblings like audit_urls and sitemap_coverage.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit guidance: 'Call this first when auditing a site you do not have a URL list for.' This clearly tells an agent when to use it. However, it does not name alternative tools or explicitly state when not to use it, so it falls just short of full usage differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sitemap_coverageA
Compare a sitemap against a list of URLs you know exist (e.g. from the filesystem or a crawl) and report which are missing from the sitemap and which the sitemap declares but are unreachable. Missing pages are pages Google is never told about.
| Name | Required | Description | Default |
|---|---|---|---|
| verify | No | ||
| known_urls | Yes | ||
| sitemap_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. It does disclose that the tool checks reachability and reports missing/unreachable URLs, which is meaningful. However, it does not explain the 'verify' behavior, whether network requests are made to each known URL, or any side effects such as rate-limit impact or request costs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences cover the operation, inputs, outputs, and practical significance with no filler. The core comparison is front-loaded, and every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so describing return values is not necessary. The description is adequate for a straightforward comparison tool, but it omits the semantics of the optional 'verify' parameter and does not provide guidance on how this tool relates to siblings such as audit_urls or check_redirects, which would help an agent choose correctly in more ambiguous cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds some meaning by indicating 'sitemap' maps to sitemap_url and 'list of URLs you know exist' maps to known_urls. However, the 'verify' parameter is entirely unexplained despite being a schema property with a default value, and the mapping from description to parameters remains implicit rather than explicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('compare') with clear resources: a sitemap and a list of known URLs. It precisely defines the two reported outcomes—URLs missing from the sitemap and sitemap entries that are unreachable—and the final sentence explains why this matters. It is clearly distinct from siblings like fetch_sitemap or audit_urls.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear scenario for when to use the tool: when you have a sitemap and a separate list of URLs known to exist, such as from a filesystem or crawl. It does not explicitly name alternatives or exclusion conditions, but the context is sufficient for an agent to identify this as the coverage-comparison tool among the siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.0.0- First observed
audit_urls - First observed
check_redirects - First observed
fetch_sitemap - First observed
sitemap_coverage
TDQS
Scored across 4 tools
Each tool targets a distinct phase of an SEO audit: sitemap fetching, page-level auditing, sitemap coverage comparison, and redirect tracing. The main overlap is that audit_urls already reports redirect chains and broken status codes, which overlaps with check_redirects.
Three tools follow a clear verb_noun pattern (fetch_sitemap, audit_urls, check_redirects), but sitemap_coverage is a noun_noun exception. The inconsistent name is still readable and does not create real confusion.
Four tools is a well-scoped size for a focused SEO audit server. Each tool has a clear job, and there is no redundant filler or overwhelming number of endpoints.
The server covers sitemap parsing, on-page/technical issue auditing, sitemap coverage, and redirects, but it lacks a site-crawling or internal-link-discovery tool, which is needed to find URLs not listed in a sitemap. This is a notable gap for a full audit, though the core workflow is usable with an existing URL list.
Maintenance
Related MCP Connectors
- CrawlieOAuthapp.crawlie
Technical SEO + GEO (AI-search) site audits: hosted crawls, prioritized fixes, report diffs.
Audit public webpages and supplied markup for HTML, CSS, SEO, JSON-LD, and link issues.
Audit any site's AI visibility from your assistant: crawler access, rendering, and schema.
- seegeoOAuthcom.see-geo
Audit any website for AI visibility: graded report, findings with fixes, AI crawler access check.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables SEO auditing and site analysis by crawling websites, identifying issues, and generating reports like sitemaps and markdown exports.59 npm4MIT
- FlicenseNot gradedqualityDmaintenanceEnables auditing of websites for performance, SEO, accessibility, security, and mobile readiness, with tools to validate URLs, run page audits, save results, and retrieve reports.1-
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to crawl and audit websites for SEO issues, returning structured JSON reports with errors, warnings, and key statistics.MIT
- FlicenseNot gradedqualityDmaintenanceAudits any website for SEO issues, providing scored health checks, schema validation, and performance analysis through AI assistants.-