Skip to main content
Glama

seo-audit-mcp

MCP-сервер, который даёт Claude (или любому MCP-клиенту) возможность аудировать техническое SEO живого веб-сайта: покрытие sitemap, проблемы на страницах и цепочки редиректов.

Спросите простым языком — «проверь mortgagecalculatortools.com и скажи, о каких страницах Google никогда не узнает» — и модель вызовет инструменты, обойдёт сайт и ответит с подробностями.


Проблема, которую это решает

sitemap.xml сайта — это способ сообщить Google, какие страницы существуют. Когда страница в нём отсутствует, не возникает ни ошибок, ни предупреждений — страница просто никогда не накапливает показы. Проверка вручную означает сравнение списка файлов в файловой системе с XML-файлом, поэтому на практике никто этого не делает.

Кейс: расхождение в 25 страниц, оказавшееся корректным

Первый сайт, на который это было направлено, имел 125 HTML-файлов на диске и 100 URL-адресов в своём sitemap. Расхождение в 25 страниц — это как раз та находка, которую оформляют как баг и назначают ответственному.

Один вызов sitemap_coverage выявил расхождение, а один вызов audit_urls на выборке объяснил его: каждая из 25 страниц содержала <meta name="robots" content="noindex, follow">. Это были два намеренно деиндексируемых контентных кластера, и sitemap был совершенно прав, опустив их. Затем проверили файловую систему: 25 noindex страниц на диске, те же 25 отсутствуют в sitemap, и ни одна noindex страница не была ошибочно включена. Идеальная согласованность.

Вот это и есть полезный результат. Одно только число покрытия («125 против 100») выглядит как дефект и стоит кому-то целого дня; покрытие плюс статус noindex по каждой странице закрывает вопрос за минуту. Этот инструмент настолько же ценен за ложные тревоги, которые он предотвращает, насколько и за реальные пробелы, которые он находит — именно поэтому audit_urls сообщает noindex по каждой странице, а не просто подсчитывает URL-адреса.


Related MCP server: web-audit-mcp

Инструменты

Инструмент

Что делает

fetch_sitemap

Загружает sitemap.xml, следует по вложенности sitemap-index (максимальная глубина 3), возвращает все объявленные URL-адреса без дубликатов

audit_urls

Обходит URL-адреса конкурентно и сообщает о проблемах на страницах: нерабочий статус, цепочки редиректов, отсутствующий/слишком длинный <title> и meta description, отсутствующий или дублирующийся <h1>, отсутствующий canonical, noindex, тонкий контент

sitemap_coverage

Сравнивает sitemap со списком URL-адресов, которые, как вы знаете, существуют → что отсутствует в sitemap, что объявлено, но мертво

check_redirects

Прослеживает цепочки редиректов, помечает многошаговые цепочки и цепочки, заканчивающиеся 4xx/5xx — используйте после изменения структуры URL-адресов

Каждый инструмент возвращает структурированный JSON со списком issues для каждой страницы и агрегированным issue_summary, чтобы модель могла рассуждать о количественных показателях, а не перечитывать сырой HTML.


Установка

pip install -e .

Требуется Python 3.10+. Зависимости: mcp>=2.0.0, httpx.

Подключение к Claude Code

Добавьте в .mcp.json в вашем проекте (или в ~/.claude.json для глобального использования):

{
  "mcpServers": {
    "seo-audit": {
      "command": "python",
      "args": ["-m", "seo_audit_mcp"]
    }
  }
}

Для Claude Desktop тот же блок помещается в claude_desktop_config.json.

Затем просто спросите:

Загрузите sitemap для https://example.com/sitemap.xml, проверьте первые 20 URL-адресов и обобщите проблемы по частоте.

Запуск напрямую

python -m seo_audit_mcp        # stdio transport

Заметки о проектировании

Три решения, о которых стоит сказать, потому что они отличают демо-версию от того, что можно показать на продакшен-сайте клиента:

Сканирование ограничено по скорости самой архитектурой. fetch_many работает за asyncio.Semaphore с пределом в 16 одновременных запросов, и каждый инструмент ограничивает свои входные данные. Sitemap на 500 URL-адресов без этого потолка открыл бы 500 сокетов одновременно и был бы воспринят как атака на целевой хост. Сканирование — это стоимость, которую платит целевой сайт, поэтому потолок нельзя увеличить на уровне инструментов.

Ни один сбой загрузки не прерывает выполнение. fetch_one перехватывает httpx.HTTPError и записывает её в возвращаемый PageAudit, а не выбрасывает исключение. Один недоступный хост при сканировании 200 URL-адресов ухудшает одну строку, а не приводит к потере 199 хороших результатов.

Разбор намеренно снисходительный. Реальный HTML достаточно часто бывает некорректным, так что строгий парсер, выбрасывающий исключение в середине сканирования, становится обузой. Экстракторы — это нестрогие регулярные выражения, которые возвращают None, а не выбрасывают ошибку, но с учётом ловушек: содержимое <script> и <style> удаляется перед подсчётом слов и извлечением заголовков, поэтому <h1> внутри строкового литерала JS не считается заголовком, а относительные canonical разрешаются относительно URL-адреса страницы.

normalize_url намеренно не удаляет завершающие слэши: /a и /a/ могут быть действительно разными страницами, и их схлопывание скрыло бы проблемы с дублирующимся контентом, для выявления которых и существует этот инструмент.


Тесты

pip install -e ".[dev]"
pytest

Набор тестов не использует сеть — HTTP тестируется через httpx.MockTransport, поэтому он работает в CI и в самолёте. Он покрывает крайние случаи разбора, которые доставляют проблемы в продакшене: заголовки, встроенные в скрипты, sitemap без пространства имён, относительные canonical, URL-адреса sitemap, возвращающие стилизованную HTML-ошибку 404 со статусом 200, а также ошибочное отнесение не-HTML типов контента к страницам «без заголовка».

Лицензия

MIT

Available Tools

4 tools
audit_urlsA

Crawl a list of URLs and report per-page technical SEO issues: broken status codes, redirect chains, missing or over-length titles and meta descriptions, missing or duplicate H1, missing canonical, noindex, and thin content. Returns a per-URL breakdown plus an issue summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYes
concurrencyNo
timeout_secondsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral disclosure burden. It usefully states that the tool crawls URLs and returns a per-URL breakdown plus an issue summary, but it does not disclose operational behaviors such as crawl duration, rate limiting, redirect-following details, or auth/network requirements. This is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single information-dense sentence that front-loads the action and resource, then lists issue categories and the return shape. It avoids repetition and wastes no words, though the enumeration makes it slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description does not need to detail return values, and it does give a useful high-level summary. However, with zero annotations, zero schema descriptions, and no usage guidance, the description leaves concurrency/timeout semantics and tool-selection boundaries undocumented, making it only moderately complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it never names or explains urls, concurrency, or timeout_seconds. The parameter names are somewhat self-explanatory, yet the description adds no detail about what concurrency or timeout_seconds control, how URLs should be formatted, or whether limits apply.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb-resource pair ('Crawl a list of URLs and report per-page technical SEO issues') and then enumerates the exact issue categories. This clearly distinguishes it from siblings like fetch_sitemap and sitemap_coverage by centering on per-URL technical SEO auditing, even though it overlaps with check_redirects on redirect chains.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies its use case by listing SEO checks, but it never explicitly states when to prefer this over check_redirects, fetch_sitemap, or sitemap_coverage, nor does it give any 'when not to use' guidance. An agent can infer the purpose but must decide on selection criteria without direct help.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_redirectsA

Trace the redirect chain for each URL and flag chains longer than one hop, redirect loops, and URLs that resolve to a 4xx/5xx. Use after a site migration or a URL-structure change.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It clearly discloses what the tool does: traces chains, flags one-hop violations, loops, and error responses. It could add operational details like network cost or rate limits, but the core behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, and the main behavior is front-loaded. The second sentence supplies a practical trigger for use. Every part earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with an output schema present, the description covers what the tool does and when to use it. It does not fully cover input format details, but the missing information is minor given the low complexity and available output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description adds little about the 'urls' parameter beyond saying 'for each URL.' It does not clarify expected URL format (absolute vs relative), whether schemes are required, or any limits on list length, so the agent gets almost no parameter guidance beyond the bare schema type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Trace the redirect chain for each URL' and names concrete detection outcomes (long chains, loops, 4xx/5xx). This is distinct from siblings like fetch_sitemap or audit_urls, making the tool's purpose immediately clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use after a site migration or a URL-structure change,' which gives a clear context for when this tool is appropriate. It does not mention alternatives or when not to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch_sitemapA

Fetch and parse a sitemap.xml, following sitemap-index nesting, and return every page URL it declares. Call this first when auditing a site you do not have a URL list for.

ParametersJSON Schema
NameRequiredDescriptionDefault
sitemap_urlYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden. It discloses that the tool follows sitemap-index nesting and returns every declared page URL, which are meaningful behavioral details beyond what the schema shows. It does not mention failure modes or network behavior, but for a straightforward fetch-and-parse read operation this is reasonably transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The first sentence states what the tool does and the second gives usage guidance. The most important behavioral details are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with an output schema present, the description is complete: it explains what the tool does, how it behaves with index nesting, what it returns, and when to call it. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, so the description must compensate. It indirectly clarifies that sitemap_url should point to a sitemap.xml and that index nesting is followed, but it does not explicitly describe the URL format, required scheme, or example values. Because the single parameter is highly self-evident from the tool name and description, this is adequate but not enriched.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Fetch and parse a sitemap.xml' and the specific outcome: 'return every page URL it declares.' It also mentions the non-obvious behavior of following sitemap-index nesting, which distinguishes this from simply fetching one XML file. The phrase 'Call this first when auditing a site' also separates it from siblings like audit_urls and sitemap_coverage.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit guidance: 'Call this first when auditing a site you do not have a URL list for.' This clearly tells an agent when to use it. However, it does not name alternative tools or explicitly state when not to use it, so it falls just short of full usage differentiation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sitemap_coverageA

Compare a sitemap against a list of URLs you know exist (e.g. from the filesystem or a crawl) and report which are missing from the sitemap and which the sitemap declares but are unreachable. Missing pages are pages Google is never told about.

ParametersJSON Schema
NameRequiredDescriptionDefault
verifyNo
known_urlsYes
sitemap_urlYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral disclosure burden. It does disclose that the tool checks reachability and reports missing/unreachable URLs, which is meaningful. However, it does not explain the 'verify' behavior, whether network requests are made to each known URL, or any side effects such as rate-limit impact or request costs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences cover the operation, inputs, outputs, and practical significance with no filler. The core comparison is front-loaded, and every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so describing return values is not necessary. The description is adequate for a straightforward comparison tool, but it omits the semantics of the optional 'verify' parameter and does not provide guidance on how this tool relates to siblings such as audit_urls or check_redirects, which would help an agent choose correctly in more ambiguous cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds some meaning by indicating 'sitemap' maps to sitemap_url and 'list of URLs you know exist' maps to known_urls. However, the 'verify' parameter is entirely unexplained despite being a schema property with a default value, and the mapping from description to parameters remains implicit rather than explicit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('compare') with clear resources: a sitemap and a list of known URLs. It precisely defines the two reported outcomes—URLs missing from the sitemap and sitemap entries that are unreachable—and the final sentence explains why this matters. It is clearly distinct from siblings like fetch_sitemap or audit_urls.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear scenario for when to use the tool: when you have a sitemap and a separate list of URLs known to exist, such as from a filesystem or crawl. It does not explicitly name alternatives or exclusion conditions, but the context is sufficient for an agent to identify this as the coverage-comparison tool among the siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv1.0.0
    • First observedaudit_urls
    • First observedcheck_redirects
    • First observedfetch_sitemap
    • First observedsitemap_coverage

TDQS

A3.8/5.0

Scored across 4 tools

Disambiguation4/5

Each tool targets a distinct phase of an SEO audit: sitemap fetching, page-level auditing, sitemap coverage comparison, and redirect tracing. The main overlap is that audit_urls already reports redirect chains and broken status codes, which overlaps with check_redirects.

Naming Consistency4/5

Three tools follow a clear verb_noun pattern (fetch_sitemap, audit_urls, check_redirects), but sitemap_coverage is a noun_noun exception. The inconsistent name is still readable and does not create real confusion.

Tool Count5/5

Four tools is a well-scoped size for a focused SEO audit server. Each tool has a clear job, and there is no redundant filler or overwhelming number of endpoints.

Completeness3/5

The server covers sitemap parsing, on-page/technical issue auditing, sitemap coverage, and redirects, but it lacks a site-crawling or internal-link-discovery tool, which is needed to find URLs not listed in a sitemap. This is a notable gap for a full audit, though the core workflow is usable with an existing URL list.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers