Skip to main content
Glama
CrawlCove

crawlcove-mcp

Official
by CrawlCove

crawlcove-mcp

An MCP server that gives Claude, Cursor, Claude Code and any other MCP client your site's SEO crawl data — ask "which pages are missing meta descriptions?", "what links to the 404s?", or "crawl staging and tell me what's wrong" and get answers from a real crawl, not a guess.

Setup

Requires Node 18+. Nothing to install: MCP clients run the server with npx straight from GitHub (the npm package is coming).

Claude Desktop — claude_desktop_config.json:

{
  "mcpServers": {
    "crawlcove": {
      "command": "npx",
      "args": ["-y", "github:CrawlCove/crawlcove-mcp"]
    }
  }
}

Claude Code:

claude mcp add crawlcove -- npx -y github:CrawlCove/crawlcove-mcp

Cursor — .cursor/mcp.json in your project (or the global one):

{
  "mcpServers": {
    "crawlcove": {
      "command": "npx",
      "args": ["-y", "github:CrawlCove/crawlcove-mcp"]
    }
  }
}

The first start takes a few seconds while npx fetches the package; after that it is cached.

Related MCP server: Screaming Frog MCP Server

Tools

Tool

What it does

crawl_site {url, maxPages?, ignoreRobots?}

Live crawl of a site (same-origin, breadth-first, robots.txt respected), up to 200 pages. Becomes the active dataset.

load_export {path}

Load a Crawl Cove desktop app JSON export (Reports → Export → JSON) or a crawlcove-cli JSON result from disk. No page cap.

get_issues {type?, limit?}

The SEO issues in the active dataset. Types: broken-page, missing-title, long-title, duplicate-title, missing-meta-description, long-meta-description, missing-h1, multiple-h1, noindex, redirect-chain, missing-canonical.

get_page {url}

Everything recorded about one URL: status, title, meta description, H1 count, canonical, robots, redirect hops, broken links out.

list_broken_links {limit?}

Broken internal links as source page → broken target, so you know which page to fix.

list_pages {status?, indexable?, urlContains?, limit?}

Pages with status and title, filtered.

Things to ask once it is connected:

  • "Crawl https://staging.example.com and list every issue."

  • "Which pages are missing meta descriptions? Draft one for each."

  • "What links to the 404s?"

  • "Load ~/Downloads/crawl-export.json and show me the noindex pages."

Where the data comes from

  • Live crawls use the same crawler as crawlcove-cli: same-origin links only, robots.txt honoured (a User-agent: crawlcove-cli group is respected over *), a hard cap of 200 pages so an assistant cannot accidentally hammer a site. Broken-link sources are available because the crawler keeps the link graph.

  • Desktop exports follow crawlcove-export-spec and carry more per page (depth, word count, content type, X-Robots-Tag-aware indexability) with no page cap — but no link graph, so list_broken_links lists the broken pages and points you at the app for inlinks.

Works with CrawlCove

This server is the assistant-facing half of Crawl Cove, a desktop SEO crawler for Windows and Mac. For anything past 200 pages, for history over time, for Search Console data next to the crawl, or to fix findings in bulk, run the crawl in the desktop app and hand its export to load_export.

This repo has its own page on crawlcove.com: Crawl Cove MCP server.

Development

npm ci && npm run build && npm test

dist/ is committed (it is what npx github:… runs); CI fails if it is stale. The server logic is in src/server.ts; the pure dataset layer (src/dataset.ts) has no MCP dependency and is exported for scripting.

  • crawlcove-js — crawlcove-export, a typed JavaScript/TypeScript library to load, query and convert Crawl Cove exports.

  • crawlcove-sheets — Google Sheets add-on that turns a Crawl Cove export into an audit workbook (issues by type, pages by status, title/meta flags).

  • crawlcove-sf-import — convert a Screaming Frog export into the Crawl Cove export format, with a report of what carried over.

  • crawlcove-schema-validator — validate a page's JSON-LD against Google's required and recommended rich-result properties.

  • crawlcove-hreflang-checker — check a page's or a sitemap's hreflang tags: codes, self-reference, x-default and return tags.

  • crawlcove-cli — the command line crawler this server uses for live crawls.

  • crawlcove-action — the same checks as a GitHub Action on every PR.

  • crawlcove-export-spec — the JSON Schema for the desktop export load_export reads.

  • crawl-cove-connector — WordPress plugin that applies Crawl Cove's approved fixes to Yoast, Rank Math, SEOPress or AIOSEO.

  • crawlcove-redirect-chain-checker — follow every hop of a URL’s redirects; flags chains, loops, HTTPS downgrades and meta refreshes.

  • crawlcove-sitemap-validator — validate an XML sitemap or sitemap index against the protocol and search-engine limits.

  • crawlcove-robots-txt-tester — lint a robots.txt and test which URLs each crawler may fetch, with the deciding line.

License

MIT — see LICENSE.

Available Tools

6 tools
crawl_siteCrawl a siteA

Crawl a website live (same-origin, breadth-first, robots.txt respected) and make the result the active dataset for the other tools. Capped at 200 pages; for a full-site audit with history and Search Console data, use the Crawl Cove desktop app and load its export with load_export.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesStart URL, e.g. https://example.com/
maxPagesNoPage cap (default 50, max 200)
ignoreRobotsNoCrawl URLs robots.txt disallows. Only for sites you own.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses that the crawl is live, same-origin, breadth-first, robots.txt-respecting, capped at 200 pages, and that it mutates shared state by making the result the active dataset. It omits things like expected runtime, rate limiting, and failure behavior on unparseable or unreachable URLs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no filler: the core behavior lands first, then the cap and the escalation path. Every clause carries information an agent needs (scope constraints, side effect, limit, alternative).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-annotation, no-output-schema mutation tool, the description covers what the crawl does, its constraints, its limit, and its effect on the rest of the toolset. It leaves out the shape of the result and any indication of how long a large crawl takes or how failures are surfaced.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so url, maxPages, and ignoreRobots are already documented in the schema; a 3 is the baseline. The description only restates the page cap already expressed by maxPages' maximum and never mentions ignoreRobots, so it adds little parameter-level meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ('Crawl a website live') with concrete scope qualifiers (same-origin, breadth-first, robots.txt respected) and the side effect (result becomes the active dataset for the other tools). It also names the sibling load_export as the route for a different job, so an agent can distinguish it from siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the full-site-audit-with-history/Search-Console case to the desktop app plus load_export, which is a clear when-to-use-this-instead signal. It does not address the smaller-scale alternative (e.g., using get_page for a single URL), so the guidance is strong but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_issuesList SEO issuesA

List the SEO issues found in the active dataset, optionally one type only. Types: broken-page (The page returned a 4xx/5xx status or could not be fetched); missing-title (No ); long-title ( longer than 60 characters (may be truncated in search results)); duplicate-title (The same is used on more than one page); missing-meta-description (No ); long-meta-description (Meta description longer than 160 characters); missing-h1 (No on a 200 HTML page); multiple-h1 (More than one ); noindex (The page carries a robots noindex directive); redirect-chain (The URL went through 2 or more redirects); missing-canonical (No on an indexable page).

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNoRestrict to one issue type
limitNoMax issues to return (default 50)

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that results come from the 'active dataset', which is useful context, but says nothing about ordering, pagination beyond the limit parameter, permissions, or what a returned issue record looks like. A safe read operation, but behavioral coverage is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose in the first clause, then the type glossary. The long enumeration is justified because the input schema does not supply these definitions. Slightly dense, but every element earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no annotations, so the description must stand alone. It covers scope, the optional filter, and all filter values, but it omits the shape/order of returned results. Adequate for a list tool, with one gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline would be 3, but the description goes well beyond the schema by defining the meaning of every one of the 11 issue-type enum values (e.g. long-title = over 60 characters, missing-canonical = no canonical on an indexable page). That materially improves the agent's ability to pick the right filter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List the SEO issues') and scopes it to the active dataset, which an agent can parse immediately. It does not explicitly differentiate itself from the nearby list_broken_links or list_pages siblings, though the SEO-issue framing makes the boundary reasonably inferable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'optionally one type only' implies the filtering use case, but there is no explicit when/when-not guidance, no mention of prerequisites, and no routing to alternatives such as list_broken_links for link-specific work. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pageGet one pageA

Return everything the crawl recorded about one URL (status, title, meta description, H1 count, canonical, robots, redirects, broken links out).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe page URL as crawled; a missing trailing slash or scheme is tolerated

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose the payload contents (status, title, meta description, H1 count, canonical, robots, redirects, broken links out), which is useful, but says nothing about read-only/immutability, permissions, rate limits, or behavior for a URL the crawl never saw.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states the verb and scope first, then enumerates the returned fields. No filler, nothing redundant, and the parenthetical earns its place by preempting a follow-up schema read.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description admirably fills the return-value gap by listing what comes back, and the single required param is trivial. The remaining shortfall is the absence of error/empty-result behavior for an uncrawled URL, which matters slightly for a lookup tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'url' parameter is already fully documented in the schema (including trailing-slash/scheme tolerance). The description adds no syntax, format, or constraint beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Return') and resource ('everything the crawl recorded about one URL') and enumerates the returned fields, which cleanly separates it from the plural-list sibling list_pages. It stops short of explicitly naming a sibling or stating the contrast, so it stays at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied ('one URL') – an agent can infer this is the per-page drill-down versus list_pages/crawl_site, but there is no explicit when-to-use, when-not, or alternative routing. Adequate but with a clear gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_pagesList pagesA

List crawled pages with status and title, optionally filtered by status code, indexability, or a URL substring.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax pages (default 50)
statusNoOnly pages with this HTTP status
indexableNoOnly indexable (true) or non-indexable (false) pages
urlContainsNoOnly URLs containing this substring

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the disclosure burden. It does disclose the returned fields ('status and title') and that results are filterable, which is genuine behavioral content; however it omits pagination behavior, ordering, and confirmation that this is a safe read-only operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with zero filler, front-loading the resource and return content before the optional filters. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with no output schema and fully documented parameters, the description covers purpose, return content, and filter dimensions. Only minor gaps remain, such as pagination/ordering behavior and explicit read-only framing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (limit, status, indexable, urlContains) are already documented in the schema. The description's filter summary mirrors the schema without adding format or syntax detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clear verb+resource: 'List crawled pages with status and title' tells the agent exactly what is returned. It distinguishes itself from get_page (single page) and list_broken_links by naming the resource, though it does not explicitly position itself against those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'optionally filtered by status code, indexability, or a URL substring' implies when the tool is useful (filtered enumeration of crawled pages), but there is no explicit when/when-not guidance or mention of alternatives such as get_page for a specific URL. Usage is inferable but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

load_exportLoad a crawl exportA

Load a Crawl Cove desktop app JSON export (Reports → Export → JSON) or a crawlcove-cli JSON result from disk and make it the active dataset. Desktop exports carry more per-page data (depth, word count, content type) and have no page cap.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesAbsolute path to the .json file

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose that loading makes the file the active dataset and that desktop exports differ in richness (depth, word count, content type) and have no page cap, which is useful behavioral context. It omits whether the load replaces existing data, what happens on malformed/unsupported JSON, and whether any auth or file-access requirements apply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences, front-loaded with the action and source formats, followed by the one differentiating detail (desktop export richness). No filler or restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter ingestion tool with no output schema and no annotations, the description covers the accepted inputs and a meaningful difference between them. It is close to complete, though it could note replacement semantics or failure behavior on invalid files.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter and schema coverage is 100%, so the schema already documents 'path' as an absolute path to the .json file. The description adds source-format context that helps the agent pick the right file but adds no syntax beyond the schema, consistent with the baseline for a fully documented single parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (load) and resource (crawl export JSON from disk) and names the two accepted source formats with their producer tools. An agent can immediately tell this ingests an existing file rather than performing a crawl, distinguishing it from siblings like crawl_site.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It identifies when the tool applies (when you have a desktop export or a crawlcove-cli JSON) and even gives a selection criterion between the two source formats (desktop exports have more per-page data, no page cap). But it never states when not to use it or explicitly contrasts it with crawl_site for agents who lack an export file.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv1.0.0
    • First observedcrawl_site
    • First observedget_issues
    • First observedget_page
    • First observedlist_broken_links
    • First observedlist_pages
    • First observedload_export

TDQS

A4.1/5.0

Scored across 6 tools

Disambiguation5/5

crawl_site and load_export both establish the active dataset but are clearly differentiated by source (live crawl vs. file import). The four query tools target distinct resources (issues, page details, broken link pairs, page list) with no meaningful overlap.

Naming Consistency5/5

All tools use snake_case verb_noun, with a small set of verbs (crawl, load, get, list) that accurately reflect actions. No inconsistent naming conventions.

Tool Count5/5

Six tools is well-scoped for a crawl-audit MCP: two ingestion paths, two listing/query tools, and two issue-focused tools. No obvious bloating or missing essential utility tools.

Completeness4/5

Core lifecycle is covered: ingest (live or export), list pages, inspect a page, list issues, and trace broken links. A site-level summary/statistics tool or a way to clear/switch datasets would round it out, but agents can work around these gaps.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

  • Audit any site's AI visibility from your assistant: crawler access, rendering, and schema.

  • Your agent needs to crawl a site and say what is wrong with it — broken tags, duplicate content, pages nothing can index, resources that never load. **What you can ask for** • "Crawl this site and list every page with a duplicate title or missing description." • "Which pages are non-indexable, and why?" • "Run Lighthouse on these URLs and give me the failing audits." • "Show the internal link graph and the orphan pages." • "Give me this page's raw HTML and its microdata." **How to use it** Point any MCP client at https://mcp.aisa.one/seo-onpage/mcp and sign in with OAuth — there is no key to create or paste. 20 tools: submit a crawl and read its summary, pages, resources, links and waterfall; duplicate content and duplicate tags; keyword density; non-indexable and uncrawlable resources; parsed content, raw HTML, microdata, screenshots and Lighthouse. **Why this rather than the source** A crawler you drive from the agent, with the audit results as structured data rather than a PDF. **It is also a door to the rest** The same login reaches 26 sources and 580+ operations. Find the broken pages here, then ask the same agent what those pages used to rank for — without adding a second server. **What it costs** Finding and inspecting an operation is free. Running one is billed per call at API prices, with no seat and no monthly minimum, and every call takes max_price_usd so an agent cannot overspend by accident. **Where else it reaches** https://mcp.aisa.one/seo/mcp for all of it at once — rankings, keywords, backlinks, site health and AI-answer visibility across DataForSEO, Semrush and Ahrefs.

  • Query your SEO data in plain language: rankings, audits, backlinks, competitors and AI visibility.

  • Real SEO data for AI assistants: page audits, Keyword Planner volumes, Search Console history.

Related MCP Servers