crawlcove-mcp
OfficialClick on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@crawlcove-mcpcrawl https://example.com and list every SEO issue you find"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
crawlcove-mcp
An MCP server that gives Claude, Cursor, Claude Code and any other MCP client your site's SEO crawl data — ask "which pages are missing meta descriptions?", "what links to the 404s?", or "crawl staging and tell me what's wrong" and get answers from a real crawl, not a guess.
Setup
Requires Node 18+. Nothing to install: MCP clients run the server with npx straight from GitHub (the npm package is coming).
Claude Desktop — claude_desktop_config.json:
{
"mcpServers": {
"crawlcove": {
"command": "npx",
"args": ["-y", "github:CrawlCove/crawlcove-mcp"]
}
}
}Claude Code:
claude mcp add crawlcove -- npx -y github:CrawlCove/crawlcove-mcpCursor — .cursor/mcp.json in your project (or the global one):
{
"mcpServers": {
"crawlcove": {
"command": "npx",
"args": ["-y", "github:CrawlCove/crawlcove-mcp"]
}
}
}The first start takes a few seconds while npx fetches the package; after that it is cached.
Related MCP server: Screaming Frog MCP Server
Tools
Tool | What it does |
| Live crawl of a site (same-origin, breadth-first, robots.txt respected), up to 200 pages. Becomes the active dataset. |
| Load a Crawl Cove desktop app JSON export (Reports → Export → JSON) or a |
| The SEO issues in the active dataset. Types: |
| Everything recorded about one URL: status, title, meta description, H1 count, canonical, robots, redirect hops, broken links out. |
| Broken internal links as source page → broken target, so you know which page to fix. |
| Pages with status and title, filtered. |
Things to ask once it is connected:
"Crawl https://staging.example.com and list every issue."
"Which pages are missing meta descriptions? Draft one for each."
"What links to the 404s?"
"Load ~/Downloads/crawl-export.json and show me the noindex pages."
Where the data comes from
Live crawls use the same crawler as crawlcove-cli: same-origin links only, robots.txt honoured (a
User-agent: crawlcove-cligroup is respected over*), a hard cap of 200 pages so an assistant cannot accidentally hammer a site. Broken-link sources are available because the crawler keeps the link graph.Desktop exports follow crawlcove-export-spec and carry more per page (depth, word count, content type, X-Robots-Tag-aware indexability) with no page cap — but no link graph, so
list_broken_linkslists the broken pages and points you at the app for inlinks.
Works with CrawlCove
This server is the assistant-facing half of Crawl Cove, a desktop SEO crawler for Windows and Mac. For anything past 200 pages, for history over time, for Search Console data next to the crawl, or to fix findings in bulk, run the crawl in the desktop app and hand its export to load_export.
This repo has its own page on crawlcove.com: Crawl Cove MCP server.
Development
npm ci && npm run build && npm testdist/ is committed (it is what npx github:… runs); CI fails if it is stale. The server logic is in src/server.ts; the pure dataset layer (src/dataset.ts) has no MCP dependency and is exported for scripting.
Related tools
crawlcove-js —
crawlcove-export, a typed JavaScript/TypeScript library to load, query and convert Crawl Cove exports.crawlcove-sheets — Google Sheets add-on that turns a Crawl Cove export into an audit workbook (issues by type, pages by status, title/meta flags).
crawlcove-sf-import — convert a Screaming Frog export into the Crawl Cove export format, with a report of what carried over.
crawlcove-schema-validator — validate a page's JSON-LD against Google's required and recommended rich-result properties.
crawlcove-hreflang-checker — check a page's or a sitemap's hreflang tags: codes, self-reference, x-default and return tags.
crawlcove-cli — the command line crawler this server uses for live crawls.
crawlcove-action — the same checks as a GitHub Action on every PR.
crawlcove-export-spec — the JSON Schema for the desktop export
load_exportreads.crawl-cove-connector — WordPress plugin that applies Crawl Cove's approved fixes to Yoast, Rank Math, SEOPress or AIOSEO.
crawlcove-redirect-chain-checker — follow every hop of a URL’s redirects; flags chains, loops, HTTPS downgrades and meta refreshes.
crawlcove-sitemap-validator — validate an XML sitemap or sitemap index against the protocol and search-engine limits.
crawlcove-robots-txt-tester — lint a robots.txt and test which URLs each crawler may fetch, with the deciding line.
License
MIT — see LICENSE.
Available Tools
6 toolscrawl_siteCrawl a siteA
Crawl a website live (same-origin, breadth-first, robots.txt respected) and make the result the active dataset for the other tools. Capped at 200 pages; for a full-site audit with history and Search Console data, use the Crawl Cove desktop app and load its export with load_export.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Start URL, e.g. https://example.com/ | |
| maxPages | No | Page cap (default 50, max 200) | |
| ignoreRobots | No | Crawl URLs robots.txt disallows. Only for sites you own. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that the crawl is live, same-origin, breadth-first, robots.txt-respecting, capped at 200 pages, and that it mutates shared state by making the result the active dataset. It omits things like expected runtime, rate limiting, and failure behavior on unparseable or unreachable URLs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with no filler: the core behavior lands first, then the cap and the escalation path. Every clause carries information an agent needs (scope constraints, side effect, limit, alternative).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-annotation, no-output-schema mutation tool, the description covers what the crawl does, its constraints, its limit, and its effect on the rest of the toolset. It leaves out the shape of the result and any indication of how long a large crawl takes or how failures are surfaced.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so url, maxPages, and ignoreRobots are already documented in the schema; a 3 is the baseline. The description only restates the page cap already expressed by maxPages' maximum and never mentions ignoreRobots, so it adds little parameter-level meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('Crawl a website live') with concrete scope qualifiers (same-origin, breadth-first, robots.txt respected) and the side effect (result becomes the active dataset for the other tools). It also names the sibling load_export as the route for a different job, so an agent can distinguish it from siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the full-site-audit-with-history/Search-Console case to the desktop app plus load_export, which is a clear when-to-use-this-instead signal. It does not address the smaller-scale alternative (e.g., using get_page for a single URL), so the guidance is strong but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_issuesList SEO issuesA
List the SEO issues found in the active dataset, optionally one type only. Types: broken-page (The page returned a 4xx/5xx status or could not be fetched); missing-title (No ); long-title ( longer than 60 characters (may be truncated in search results)); duplicate-title (The same is used on more than one page); missing-meta-description (No ); long-meta-description (Meta description longer than 160 characters); missing-h1 (No on a 200 HTML page); multiple-h1 (More than one ); noindex (The page carries a robots noindex directive); redirect-chain (The URL went through 2 or more redirects); missing-canonical (No on an indexable page).
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Restrict to one issue type | |
| limit | No | Max issues to return (default 50) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that results come from the 'active dataset', which is useful context, but says nothing about ordering, pagination beyond the limit parameter, permissions, or what a returned issue record looks like. A safe read operation, but behavioral coverage is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose in the first clause, then the type glossary. The long enumeration is justified because the input schema does not supply these definitions. Slightly dense, but every element earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description must stand alone. It covers scope, the optional filter, and all filter values, but it omits the shape/order of returned results. Adequate for a list tool, with one gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline would be 3, but the description goes well beyond the schema by defining the meaning of every one of the 11 issue-type enum values (e.g. long-title = over 60 characters, missing-canonical = no canonical on an indexable page). That materially improves the agent's ability to pick the right filter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the SEO issues') and scopes it to the active dataset, which an agent can parse immediately. It does not explicitly differentiate itself from the nearby list_broken_links or list_pages siblings, though the SEO-issue framing makes the boundary reasonably inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'optionally one type only' implies the filtering use case, but there is no explicit when/when-not guidance, no mention of prerequisites, and no routing to alternatives such as list_broken_links for link-specific work. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_pageGet one pageA
Return everything the crawl recorded about one URL (status, title, meta description, H1 count, canonical, robots, redirects, broken links out).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The page URL as crawled; a missing trailing slash or scheme is tolerated |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose the payload contents (status, title, meta description, H1 count, canonical, robots, redirects, broken links out), which is useful, but says nothing about read-only/immutability, permissions, rate limits, or behavior for a URL the crawl never saw.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the verb and scope first, then enumerates the returned fields. No filler, nothing redundant, and the parenthetical earns its place by preempting a follow-up schema read.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description admirably fills the return-value gap by listing what comes back, and the single required param is trivial. The remaining shortfall is the absence of error/empty-result behavior for an uncrawled URL, which matters slightly for a lookup tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'url' parameter is already fully documented in the schema (including trailing-slash/scheme tolerance). The description adds no syntax, format, or constraint beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Return') and resource ('everything the crawl recorded about one URL') and enumerates the returned fields, which cleanly separates it from the plural-list sibling list_pages. It stops short of explicitly naming a sibling or stating the contrast, so it stays at 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied ('one URL') – an agent can infer this is the per-page drill-down versus list_pages/crawl_site, but there is no explicit when-to-use, when-not, or alternative routing. Adequate but with a clear gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_broken_linksList broken internal linksA
Broken internal links as source page → broken target pairs, so you know which page to fix. Needs a live crawl or crawlcove-cli dataset; a desktop export lists the broken pages instead (its link graph lives in the app).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max entries (default 50) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden and delivers the key operational constraint — the required data source (live crawl or CLI dataset) and its behavior difference on desktop exports. It also clarifies the returned pair structure. It does not mention pagination behavior or ordering, but the essentials are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with no filler; the output shape and the fix-oriented purpose are front-loaded before the data-source caveat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity, single-parameter tool with no output schema, the description supplies the return shape and the critical data-source prerequisite. Nothing essential is missing, though it could briefly state the default/limit behavior to be fully self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one optional parameter (limit) with 100% schema description coverage, so the schema already documents it fully. The description adds nothing about limit semantics, which is acceptable given the high coverage, making 3 the correct baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('broken internal links') and goes further by naming the exact output shape (source page → broken target pairs) and the practical intent ('which page to fix'). The crawl/export distinction also separates it from siblings like load_export and get_issues.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear prerequisite: it needs a live crawl or crawlcove-cli dataset, and notes that a desktop export instead lists broken pages because its link graph lives in the app. This is real when-to-use context, though the alternative tool to reach for in the export case is not named explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_pagesList pagesA
List crawled pages with status and title, optionally filtered by status code, indexability, or a URL substring.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max pages (default 50) | |
| status | No | Only pages with this HTTP status | |
| indexable | No | Only indexable (true) or non-indexable (false) pages | |
| urlContains | No | Only URLs containing this substring |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the disclosure burden. It does disclose the returned fields ('status and title') and that results are filterable, which is genuine behavioral content; however it omits pagination behavior, ordering, and confirmation that this is a safe read-only operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with zero filler, front-loading the resource and return content before the optional filters. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with no output schema and fully documented parameters, the description covers purpose, return content, and filter dimensions. Only minor gaps remain, such as pagination/ordering behavior and explicit read-only framing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (limit, status, indexable, urlContains) are already documented in the schema. The description's filter summary mirrors the schema without adding format or syntax detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb+resource: 'List crawled pages with status and title' tells the agent exactly what is returned. It distinguishes itself from get_page (single page) and list_broken_links by naming the resource, though it does not explicitly position itself against those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'optionally filtered by status code, indexability, or a URL substring' implies when the tool is useful (filtered enumeration of crawled pages), but there is no explicit when/when-not guidance or mention of alternatives such as get_page for a specific URL. Usage is inferable but not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_exportLoad a crawl exportA
Load a Crawl Cove desktop app JSON export (Reports → Export → JSON) or a crawlcove-cli JSON result from disk and make it the active dataset. Desktop exports carry more per-page data (depth, word count, content type) and have no page cap.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Absolute path to the .json file |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose that loading makes the file the active dataset and that desktop exports differ in richness (depth, word count, content type) and have no page cap, which is useful behavioral context. It omits whether the load replaces existing data, what happens on malformed/unsupported JSON, and whether any auth or file-access requirements apply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences, front-loaded with the action and source formats, followed by the one differentiating detail (desktop export richness). No filler or restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter ingestion tool with no output schema and no annotations, the description covers the accepted inputs and a meaningful difference between them. It is close to complete, though it could note replacement semantics or failure behavior on invalid files.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter and schema coverage is 100%, so the schema already documents 'path' as an absolute path to the .json file. The description adds source-format context that helps the agent pick the right file but adds no syntax beyond the schema, consistent with the baseline for a fully documented single parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (load) and resource (crawl export JSON from disk) and names the two accepted source formats with their producer tools. An agent can immediately tell this ingests an existing file rather than performing a crawl, distinguishing it from siblings like crawl_site.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It identifies when the tool applies (when you have a desktop export or a crawlcove-cli JSON) and even gives a selection criterion between the two source formats (desktop exports have more per-page data, no page cap). But it never states when not to use it or explicitly contrasts it with crawl_site for agents who lack an export file.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v1.0.0- First observed
crawl_site - First observed
get_issues - First observed
get_page - First observed
list_broken_links - First observed
list_pages - First observed
load_export
TDQS
Scored across 6 tools
crawl_site and load_export both establish the active dataset but are clearly differentiated by source (live crawl vs. file import). The four query tools target distinct resources (issues, page details, broken link pairs, page list) with no meaningful overlap.
All tools use snake_case verb_noun, with a small set of verbs (crawl, load, get, list) that accurately reflect actions. No inconsistent naming conventions.
Six tools is well-scoped for a crawl-audit MCP: two ingestion paths, two listing/query tools, and two issue-focused tools. No obvious bloating or missing essential utility tools.
Core lifecycle is covered: ingest (live or export), list pages, inspect a page, list issues, and trace broken links. A site-level summary/statistics tool or a way to clear/switch datasets would round it out, but agents can work around these gaps.
Maintenance
Related MCP Connectors
Audit any site's AI visibility from your assistant: crawler access, rendering, and schema.
Your agent needs to crawl a site and say what is wrong with it — broken tags, duplicate content, pages nothing can index, resources that never load. **What you can ask for** • "Crawl this site and list every page with a duplicate title or missing description." • "Which pages are non-indexable, and why?" • "Run Lighthouse on these URLs and give me the failing audits." • "Show the internal link graph and the orphan pages." • "Give me this page's raw HTML and its microdata." **How to use it** Point any MCP client at https://mcp.aisa.one/seo-onpage/mcp and sign in with OAuth — there is no key to create or paste. 20 tools: submit a crawl and read its summary, pages, resources, links and waterfall; duplicate content and duplicate tags; keyword density; non-indexable and uncrawlable resources; parsed content, raw HTML, microdata, screenshots and Lighthouse. **Why this rather than the source** A crawler you drive from the agent, with the audit results as structured data rather than a PDF. **It is also a door to the rest** The same login reaches 26 sources and 580+ operations. Find the broken pages here, then ask the same agent what those pages used to rank for — without adding a second server. **What it costs** Finding and inspecting an operation is free. Running one is billed per call at API prices, with no seat and no monthly minimum, and every call takes max_price_usd so an agent cannot overspend by accident. **Where else it reaches** https://mcp.aisa.one/seo/mcp for all of it at once — rankings, keywords, backlinks, site health and AI-answer visibility across DataForSEO, Semrush and Ahrefs.
Query your SEO data in plain language: rankings, audits, backlinks, competitors and AI visibility.
Real SEO data for AI assistants: page audits, Keyword Planner volumes, Search Console history.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables SEO auditing and site analysis by crawling websites, identifying issues, and generating reports like sitemaps and markdown exports.59 npm4MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to analyze Screaming Frog SEO Spider crawl data through tool calls, supporting audits for broken links, redirects, indexability, and more.1MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to crawl and audit websites for SEO issues, returning structured JSON reports with errors, warnings, and key statistics.MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to perform instant SEO audits, check robots.txt, sitemaps, and AI crawler access for any URL without API keys.MIT