seo-audit-mcp
seo-audit-mcp
一个 MCP 服务器,让 Claude(或任何 MCP 客户端)能够审计在线网站的技术 SEO:站点地图覆盖、每个页面上的问题以及重定向链。
用日常语言提问 — "审计 mortgagecalculatortools.com,并告诉我哪些页面 Google 从未被告知" — 模型就会调用工具、爬取网站,并用具体细节作答。
它解决的问题
网站的 sitemap.xml 是你告诉 Google 哪些页面存在的方式。当某个页面缺失时,不会有报错,也不会有警告——该页面只是永远无法积累展示。手动检查意味着要将文件系统列表与 XML 文件进行对比,因此实际上没有人会这么做。
案例研究:一个 25 页的缺口,结果证明是正确的
最初指向的这个网站在磁盘上有 125 个 HTML 文件,其 sitemap 中有 100 个 URL。一个 25 页的缺口——这种发现通常会被当作 bug 记录下来并分配给某人。
一次 sitemap_coverage 调用暴露了这个缺口,对样本进行的 audit_urls 调用解释了原因:这 25 个页面全都带有 <meta name=\"robots\" content=\"noindex, follow\">。它们是两个刻意取消索引的内容集群,sitemap 完全正确地省略了它们。随后对照文件系统验证:磁盘上有 25 个 noindex 页面,同样的 25 个在 sitemap 中缺席,零个 noindex 页面被错误包含。完全一致。
这才是最有用的结果。单独的覆盖数字("125 对 100")看起来像个缺陷,会白白消耗某人一天的时间;而覆盖情况加上逐页的 noindex 状态,则能在一分钟内结束讨论。这个工具在消除误报方面,与发现真实缺口一样有价值——这就是为什么 audit_urls 会逐页报告 noindex,而不是只统计 URL。
Related MCP server: web-audit-mcp
工具
工具 | 功能 |
| 获取 |
| 并发爬取 URL 并报告每个页面上的问题:损坏的状态、重定向链、缺失/过长的 |
| 将 sitemap 与你知道存在的 URL 列表进行对比 → 找出 sitemap 中缺失的内容,以及已声明但 已失效的内容 |
| 跟踪重定向链,标记多跳链和以 4xx/5xx 结尾的链——在 URL 结构调整后使用 |
每个工具都会返回结构化 JSON,其中包含每个页面的 issues 列表和聚合的 issue_summary,这样模型就可以基于计数进行推理,而无需重新阅读原始 HTML。
安装
pip install -e .需要 Python 3.10+。依赖:mcp>=2.0.0、httpx。
连接到 Claude Code
在你的项目中添加到 .mcp.json(或 ~/.claude.json 用于全局使用):
{
"mcpServers": {
"seo-audit": {
"command": "python",
"args": ["-m", "seo_audit_mcp"]
}
}
}对于 Claude Desktop,同一内容块放在 claude_desktop_config.json 中。
然后直接提问:
获取 https://example.com/sitemap.xml 的 sitemap,审计前 20 个 URL,并按频率总结问题。
直接运行
python -m seo_audit_mcp # stdio transport设计说明
有三个值得特别指出的设计决策,因为它们是演示与实际指向客户生产环境的版本之间的区别:
爬虫在结构上就是限速的。 fetch_many 运行在一个上限为 16 个并发请求的 asyncio.Semaphore 之后,并且每个工具都会限制其输入。如果没有这个上限,一个 500 个 URL 的 sitemap 会一次打开 500 个 socket,看起来就像是对目标主机的攻击。爬取是由 目标 站点承担的成本,因此无法从工具层面向上配置这个上限。
没有一次抓取失败会中止整个运行。 fetch_one 捕获 httpx.HTTPError 并将其记录在返回的 PageAudit 上,而不是抛出异常。在 200 个 URL 的爬取中,一个死主机会使一行降级,而不会丢掉 199 个正确结果。
解析是故意宽松的。 现实中的 HTML 经常格式错误,以至于严格的解析器在爬取中途抛错是一个负担。提取器使用宽容的正则表达式,返回 None 而不是抛出——但陷阱已被处理:在计词和标题提取之前会剥离 <script> 和 <style> 主体,因此 JS 字符串字面量中的 <h1> 不会被计为标题,并且相对 canonical 会相对于页面 URL 进行解析。
normalize_url 故意 不 去除末尾斜杠:/a 和 /a/ 可能是完全不同的页面,而将它们折叠起来会掩盖这个工具要揭示的重复内容问题。
测试
pip install -e ".[dev]"
pytest测试套件无需网络——HTTP 通过 httpx.MockTransport 进行模拟,因此它可以在 CI 和飞机上运行。它覆盖了生产中会踩坑的解析边界情况:脚本嵌入的标题、无命名空间的 sitemap、相对 canonical、返回带样式 HTML 404 但状态码为 200 的 sitemap URL,以及将非 HTML 内容类型错误报告为"缺少标题"的页面。
许可证
MIT
Available Tools
4 toolsaudit_urlsA
Crawl a list of URLs and report per-page technical SEO issues: broken status codes, redirect chains, missing or over-length titles and meta descriptions, missing or duplicate H1, missing canonical, noindex, and thin content. Returns a per-URL breakdown plus an issue summary.
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | ||
| concurrency | No | ||
| timeout_seconds | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral disclosure burden. It usefully states that the tool crawls URLs and returns a per-URL breakdown plus an issue summary, but it does not disclose operational behaviors such as crawl duration, rate limiting, redirect-following details, or auth/network requirements. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single information-dense sentence that front-loads the action and resource, then lists issue categories and the return shape. It avoids repetition and wastes no words, though the enumeration makes it slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description does not need to detail return values, and it does give a useful high-level summary. However, with zero annotations, zero schema descriptions, and no usage guidance, the description leaves concurrency/timeout semantics and tool-selection boundaries undocumented, making it only moderately complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it never names or explains urls, concurrency, or timeout_seconds. The parameter names are somewhat self-explanatory, yet the description adds no detail about what concurrency or timeout_seconds control, how URLs should be formatted, or whether limits apply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource pair ('Crawl a list of URLs and report per-page technical SEO issues') and then enumerates the exact issue categories. This clearly distinguishes it from siblings like fetch_sitemap and sitemap_coverage by centering on per-URL technical SEO auditing, even though it overlaps with check_redirects on redirect chains.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies its use case by listing SEO checks, but it never explicitly states when to prefer this over check_redirects, fetch_sitemap, or sitemap_coverage, nor does it give any 'when not to use' guidance. An agent can infer the purpose but must decide on selection criteria without direct help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_redirectsA
Trace the redirect chain for each URL and flag chains longer than one hop, redirect loops, and URLs that resolve to a 4xx/5xx. Use after a site migration or a URL-structure change.
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It clearly discloses what the tool does: traces chains, flags one-hop violations, loops, and error responses. It could add operational details like network cost or rate limits, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, and the main behavior is front-loaded. The second sentence supplies a practical trigger for use. Every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with an output schema present, the description covers what the tool does and when to use it. It does not fully cover input format details, but the missing information is minor given the low complexity and available output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds little about the 'urls' parameter beyond saying 'for each URL.' It does not clarify expected URL format (absolute vs relative), whether schemes are required, or any limits on list length, so the agent gets almost no parameter guidance beyond the bare schema type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Trace the redirect chain for each URL' and names concrete detection outcomes (long chains, loops, 4xx/5xx). This is distinct from siblings like fetch_sitemap or audit_urls, making the tool's purpose immediately clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use after a site migration or a URL-structure change,' which gives a clear context for when this tool is appropriate. It does not mention alternatives or when not to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch_sitemapA
Fetch and parse a sitemap.xml, following sitemap-index nesting, and return every page URL it declares. Call this first when auditing a site you do not have a URL list for.
| Name | Required | Description | Default |
|---|---|---|---|
| sitemap_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. It discloses that the tool follows sitemap-index nesting and returns every declared page URL, which are meaningful behavioral details beyond what the schema shows. It does not mention failure modes or network behavior, but for a straightforward fetch-and-parse read operation this is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The first sentence states what the tool does and the second gives usage guidance. The most important behavioral details are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with an output schema present, the description is complete: it explains what the tool does, how it behaves with index nesting, what it returns, and when to call it. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. It indirectly clarifies that sitemap_url should point to a sitemap.xml and that index nesting is followed, but it does not explicitly describe the URL format, required scheme, or example values. Because the single parameter is highly self-evident from the tool name and description, this is adequate but not enriched.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Fetch and parse a sitemap.xml' and the specific outcome: 'return every page URL it declares.' It also mentions the non-obvious behavior of following sitemap-index nesting, which distinguishes this from simply fetching one XML file. The phrase 'Call this first when auditing a site' also separates it from siblings like audit_urls and sitemap_coverage.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit guidance: 'Call this first when auditing a site you do not have a URL list for.' This clearly tells an agent when to use it. However, it does not name alternative tools or explicitly state when not to use it, so it falls just short of full usage differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sitemap_coverageA
Compare a sitemap against a list of URLs you know exist (e.g. from the filesystem or a crawl) and report which are missing from the sitemap and which the sitemap declares but are unreachable. Missing pages are pages Google is never told about.
| Name | Required | Description | Default |
|---|---|---|---|
| verify | No | ||
| known_urls | Yes | ||
| sitemap_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. It does disclose that the tool checks reachability and reports missing/unreachable URLs, which is meaningful. However, it does not explain the 'verify' behavior, whether network requests are made to each known URL, or any side effects such as rate-limit impact or request costs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences cover the operation, inputs, outputs, and practical significance with no filler. The core comparison is front-loaded, and every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so describing return values is not necessary. The description is adequate for a straightforward comparison tool, but it omits the semantics of the optional 'verify' parameter and does not provide guidance on how this tool relates to siblings such as audit_urls or check_redirects, which would help an agent choose correctly in more ambiguous cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds some meaning by indicating 'sitemap' maps to sitemap_url and 'list of URLs you know exist' maps to known_urls. However, the 'verify' parameter is entirely unexplained despite being a schema property with a default value, and the mapping from description to parameters remains implicit rather than explicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('compare') with clear resources: a sitemap and a list of known URLs. It precisely defines the two reported outcomes—URLs missing from the sitemap and sitemap entries that are unreachable—and the final sentence explains why this matters. It is clearly distinct from siblings like fetch_sitemap or audit_urls.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear scenario for when to use the tool: when you have a sitemap and a separate list of URLs known to exist, such as from a filesystem or crawl. It does not explicitly name alternatives or exclusion conditions, but the context is sufficient for an agent to identify this as the coverage-comparison tool among the siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.0.0- First observed
audit_urls - First observed
check_redirects - First observed
fetch_sitemap - First observed
sitemap_coverage
TDQS
Scored across 4 tools
Each tool targets a distinct phase of an SEO audit: sitemap fetching, page-level auditing, sitemap coverage comparison, and redirect tracing. The main overlap is that audit_urls already reports redirect chains and broken status codes, which overlaps with check_redirects.
Three tools follow a clear verb_noun pattern (fetch_sitemap, audit_urls, check_redirects), but sitemap_coverage is a noun_noun exception. The inconsistent name is still readable and does not create real confusion.
Four tools is a well-scoped size for a focused SEO audit server. Each tool has a clear job, and there is no redundant filler or overwhelming number of endpoints.
The server covers sitemap parsing, on-page/technical issue auditing, sitemap coverage, and redirects, but it lacks a site-crawling or internal-link-discovery tool, which is needed to find URLs not listed in a sitemap. This is a notable gap for a full audit, though the core workflow is usable with an existing URL list.
Maintenance
Related MCP Connectors
- CrawlieOAuthapp.crawlie
Technical SEO + GEO (AI-search) site audits: hosted crawls, prioritized fixes, report diffs.
Audit public webpages and supplied markup for HTML, CSS, SEO, JSON-LD, and link issues.
Audit any site's AI visibility from your assistant: crawler access, rendering, and schema.
- seegeoOAuthcom.see-geo
Audit any website for AI visibility: graded report, findings with fixes, AI crawler access check.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables SEO auditing and site analysis by crawling websites, identifying issues, and generating reports like sitemaps and markdown exports.59 npm4MIT
- FlicenseNot gradedqualityDmaintenanceEnables auditing of websites for performance, SEO, accessibility, security, and mobile readiness, with tools to validate URLs, run page audits, save results, and retrieve reports.1-
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to crawl and audit websites for SEO issues, returning structured JSON reports with errors, warnings, and key statistics.MIT
- FlicenseNot gradedqualityDmaintenanceAudits any website for SEO issues, providing scored health checks, schema validation, and performance analysis through AI assistants.-