crawlgraph-mcp
This server is an MCP integration for the CrawlGraph backlink-intelligence API, giving any MCP client backlink lookups and competitor analysis built on Common Crawl data.
backlinks– Look up every referring domain for a target domain, with host counts and CrawlGraph authority/rank; supports limits, sorting, and querying a specific Common Crawl release.gap_analysis– Find domains that link to your competitors but not to you, including which competitors each domain links to (found_on).gap_outreach_targets– The warm-outreach play: prioritize domains linking to all (or 2+) competitors, filter platform/CDN noise, and optionally authority-score the top targets.backlink_changes– Compare two Common Crawl snapshots to see added, removed, and authority-moved referring domains over time.releases– List the Common Crawl snapshots the API can query, free of charge.
crawlgraph-mcp
MCP server for the CrawlGraph backlink-intelligence API. Gives any MCP client — Claude Desktop, Claude Code, Cursor, Cline, Zed, Windsurf — backlink lookups and competitor gap analysis built on the public Common Crawl webgraph (4.4B edges, 120M domains).
Backlink data without the $129/month subscription. CrawlGraph is $99 lifetime; API access is included on the lifetime tier.
What you can do
backlinks— every referring domain for a target, with authority scoresgap_analysis— domains linking to your competitors but not to yougap_outreach_targets— the warm-outreach play: the domains that link to all of your competitors but not to you, ranked and de-noised. These are publishers who cover your whole space and have simply never heard of you — the warmest backlink targets you will ever pitch.backlink_changes— additions, observed absences, and authority movement between Common Crawl snapshotsreleases— list the Common Crawl snapshots you can query
Related MCP server: oncrawl-mcp-server
Install
You need a CrawlGraph API key (cg_live_...). Free tier: 15 backlink calls/month, no card - get a key emailed to you at crawlgraph.com/docs/api. The gap_analysis and gap_outreach_targets tools need the $99 lifetime tier (1,000 calls + 50 gap analyses/month, no subscription).
Claude Desktop / Claude Code
Add to your MCP config (claude_desktop_config.json, or .mcp.json for Claude Code):
{
"mcpServers": {
"crawlgraph": {
"command": "npx",
"args": ["-y", "crawlgraph-mcp"],
"env": {
"CRAWLGRAPH_API_KEY": "cg_live_your_key_here"
}
}
}
}Cursor / Windsurf / Cline / Zed
Same shape — point the client's MCP config at npx -y crawlgraph-mcp with CRAWLGRAPH_API_KEY in the env. Restart the client and the five tools appear.
Hosted endpoint
Clients that support Streamable HTTP can use the zero-install hosted endpoint
at https://crawlgraph.com/mcp with the same bearer key. This package's local
0.3.0 server exposes the five tools above; hosted package versions are released
separately and must be verified with tools/list before relying on the new
backlink_changes tool. See the hosted MCP smoke runbook
for the operator-owned release and verification process.
The outreach play, in one prompt
Once it's connected, you don't call the tools by hand — you describe the goal:
"Use gap_outreach_targets for mydomain.com against competitor-a.com and competitor-b.com, then draft a short, specific outreach email to each priority target."
Behind the scenes the server submits the gap job, polls until it completes, filters the results down to the domains that link to every competitor but not to you, strips out platform/CDN noise (amazonaws, github, facebook, ...), and hands your agent a clean ranked list to write outreach against.
Why 2-3 competitors, not one: a site linking to one competitor might be a fluke or a paid placement. A site linking to three of your competitors is a publisher who covers your whole category. That overlap is the qualifier.
Tools reference
Tool | Arguments | Quota cost |
|
| 1 backlinks call |
|
| 1 backlinks call |
|
| 1 gap job |
|
| 1 gap job |
| — | free |
Lifetime quota: 1,000 backlinks calls + 50 gap jobs per calendar month. Full API reference: crawlgraph.com/docs/api.
backlink_changes uses the newest queryable release pair when release ids are
omitted. Its removed list means a referring domain was not observed in the
newer Common Crawl snapshot, not that a live link was proven deleted. If two
queryable snapshots do not exist, it returns a successful
comparison_available: false response with the reason instead of inventing a
comparison.
Example input:
{
"domain": "example.com",
"from_release": "cc-main-2025-50",
"to_release": "cc-main-2026-04"
}Example output (abbreviated):
{
"domain": "example.com",
"comparison_available": true,
"from_release": { "id": "cc-main-2025-50", "label": "Dec 2025" },
"to_release": { "id": "cc-main-2026-04", "label": "Apr 2026" },
"counts": { "from_snapshot": 4821, "to_snapshot": 4890, "added": 92, "removed": 23, "authority_moved": 17 },
"added": [],
"removed": [],
"authority_moved": [],
"truncated": false,
"cap": 100000,
"snapshot_caveat": "Common Crawl snapshots are periodic observations, not live link monitoring."
}Configuration
Env var | Required | Default |
| yes | — |
| no |
|
Limitations
CrawlGraph is a quarterly Common Crawl snapshot, not a live crawler. It's built for one-off competitor prospecting and release-to-release comparison, not live backlink monitoring — for change-tracking within days, a continuous-crawl tool like Ahrefs is the right choice. The backlink_changes tool reports observations across indexed snapshots; an absent domain is not proof that a live link was deleted. The gap result carries which competitors each domain links to (found_on) but not per-domain authority; use the backlinks tool if you need to score an individual target.
Develop
npm install
npm run build
CRAWLGRAPH_API_KEY=cg_live_... npm startLicense
MIT
Available Tools
4 toolsbacklinksBacklink lookupARead-onlyIdempotent
Look up referring domains (backlinks) for a single target domain from the Common Crawl webgraph. Returns each linking domain with host count and CrawlGraph authority score, plus the target's own authority/rank. Costs one backlinks call against the monthly quota (1,000/mo on lifetime).
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Target domain, e.g. 'stripe.com'. | |
| limit | No | Max rows (1..10000, default 1000). | |
| sort | No | 'authority' (default) or 'hosts'. | |
| release_id | No | Common Crawl release id (defaults to latest; see the releases tool). |
Output Schema
| Name | Required | Description |
|---|---|---|
| domain | Yes | |
| release_id | Yes | |
| release_label | Yes | |
| total_linking_domains | Yes | |
| returned | Yes | |
| cg_authority | Yes | |
| cg_rank | Yes | |
| results | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive. The description adds the critical quota usage detail (1,000/mo) and data source (Common Crawl webgraph), providing extra behavioral context beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise: two sentences. First sentence front-loads purpose and output, second adds quota info. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With good annotations, full schema, and output schema, the description covers purpose, source, output summary, and quota. Missing a note on default sort or pagination, but otherwise very complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so baseline is 3. The description does not add any parameter-specific meaning beyond what the schema already provides; it mainly describes output behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool looks up referring domains for a single target domain from Common Crawl, specifying the output (linking domain, host count, authority score, target authority/rank). This distinguishes it from siblings like gap_analysis or releases.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions quota cost but does not explicitly state when to use this versus alternatives like gap_analysis or releases. It implies use for backlink data but lacks exclusion criteria or comparative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gap_analysisCompetitor backlink gap analysisARead-onlyIdempotent
Run a competitor backlink gap analysis: find domains that link to one or more of your competitors but NOT to you. Submits an async job and polls until done (usually 5-30s). Returns every gap with found_on listing which competitors each domain links to. Costs one gap job against the monthly quota (50/mo on lifetime).
| Name | Required | Description | Default |
|---|---|---|---|
| my_domain | Yes | Your domain. | |
| competitor_domains | Yes | 1 to 5 competitor domains. |
Output Schema
| Name | Required | Description |
|---|---|---|
| my_domain | Yes | |
| competitor_domains | Yes | |
| total_gaps | Yes | |
| gaps | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds significant behavioral context beyond annotations: async job with polling (5-30s), return format with 'found_on' listing, and quota limits (50/month). Annotations already indicate readOnly, openWorld, idempotent, and non-destructive, and the description aligns without contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with three sentences. The first sentence front-loads the purpose, followed by operational details and return info. No unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema (context signals indicate true), the description adequately covers return format and execution behavior. Quota and async details are included. Slight improvement could mention output schema existence or pagination, but overall complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both parameters (my_domain, competitor_domains) well-described in the schema. The description does not add extra parameter-specific semantics, so baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool performs a competitor backlink gap analysis, specifically finding domains linking to competitors but not to the user's domain. It distinguishes itself from siblings like 'backlinks' and 'gap_outreach_targets' by specifying the unique gap analysis functionality.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use the tool (to find linking domains) and provides context on async job execution and quota costs. However, it does not explicitly state when not to use it or contrast with alternatives, though the sibling list implies distinct use cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gap_outreach_targetsOutreach target finderARead-onlyIdempotent
The warm-outreach play. Runs a gap analysis, then ranks results: PRIORITY = domains linking to ALL your competitors but not you (publishers who cover your whole space and have never heard of you), SECONDARY = domains linking to 2+ competitors. Platform/CDN noise is filtered, top N priority targets are scored by authority. Use 2-3 competitors. Costs one gap job + one backlinks call per enriched target.
| Name | Required | Description | Default |
|---|---|---|---|
| my_domain | Yes | Your domain. | |
| competitor_domains | Yes | 2 to 5 competitor domains (2-3 recommended). | |
| include_platforms | No | Keep platform/CDN/social domains in the list. Default false. | |
| enrich_top | No | Authority-score the top N priority targets. Default 10; each costs one backlinks call. 0 disables. |
Output Schema
| Name | Required | Description |
|---|---|---|
| my_domain | Yes | |
| competitor_domains | Yes | |
| priority_targets | Yes | |
| secondary_targets | Yes | |
| total_gaps | Yes | |
| platforms_filtered | Yes | |
| authority_enriched | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses costs ('one gap job + one backlinks call per enriched target'), noise filtering, and ranking behavior. Adds value beyond annotations (readOnlyHint, etc.) by detailing operational impact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Concise 5 sentences, front-loaded with key purpose, each sentence adds unique value. No redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, ranking logic, costs, filtering, and parameter specifics. With output schema existing, no need to detail return values. Complete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers all parameters with descriptions (100% coverage). Description adds minor extra context (e.g., cost per enrich_top, default for include_platforms), but not substantial beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it finds outreach targets based on gap analysis, ranking priority and secondary domains. Distinct from sibling tools (backlinks, gap_analysis, releases) by combining both analyses.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use 2-3 competitors' and describes the ranking logic, giving clear context. Does not include explicit when-not-to-use, but the purpose is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
releasesList Common Crawl releasesARead-onlyIdempotent
List the Common Crawl releases the API can query. Does not count against any quota. Use a release id with the backlinks tool to query a specific snapshot.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| releases | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint, idempotentHint, destructiveHint. The description adds that it does not count against quota, which is valuable behavioral context beyond annotations. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, front-loaded with purpose, no superfluous words. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool is simple with no parameters and an output schema. The description fully explains purpose, quota impact, and how to use the output with a sibling tool. Complete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist, so baseline is 4. The description adds no parameter info, but none is needed since there are no parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists Common Crawl releases the API can query, with a specific verb and resource. It distinguishes from siblings by mentioning using a release id with the backlinks tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states it does not count against quota, and advises to use a release id with the backlinks tool to query a specific snapshot, providing clear context for when to use this tool versus alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.2.2- First observed
backlinks - First observed
gap_analysis - First observed
gap_outreach_targets - First observed
releases
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: backlinks for single domain lookup, gap_analysis for competitor gap detection, gap_outreach_targets for ranked outreach targets, and releases for listing data snapshots. No overlap.
All tool names follow a consistent pattern of lowercase with underscores, using descriptive noun phrases (e.g., gap_analysis, gap_outreach_targets). No mixing of conventions.
With 4 tools, the set is well-scoped for a specialized backlink analysis server. Each tool is necessary and the count is appropriate for the domain.
Covers essential workflows: single domain lookup, competitor gap analysis, and enriched outreach targeting. Missing bulk queries or historical comparisons, but the releases tool enables snapshot selection, mitigating gaps.
Maintenance
Related MCP Connectors
CrawlGraph MCP — backlink intelligence on the public Common Crawl webgraph
Your agent needs the link graph — who points at a competitor, what anchor text they use, what was won or lost last month, and which sites link to all of your rivals but not to you. **What you can ask for** • "Who links to my competitor and not to me?" • "What anchors point at this domain, and how spammy are the sources?" • "Which links did this site win and lose in the last 30 days?" • "Compare the referring domains of these five competitors." • "Summarise the backlink profile of these 100 domains in bulk." **How to use it** Point any MCP client at https://mcp.aisa.one/seo-backlinks/mcp and sign in with OAuth — there is no key to create or paste. 27 tools: backlinks and referring domains, anchors, new and lost links over time, domain and page intersections, referring networks, spam scores, bulk summaries for many domains at once, plus Semrush's own backlink and indexed-page sets. **Why this rather than the source** Two independent link indexes behind one account — coverage differs, and the gap is usually the interesting part. **It is also a door to the rest** The same login reaches 26 sources and 580+ operations. Find the link gap here, then ask the same agent who runs those sites and how to reach them — without adding a second server. **What it costs** Finding and inspecting an operation is free. Running one is billed per call at API prices, with no seat and no monthly minimum, and every call takes max_price_usd so an agent cannot overspend by accident. **Where else it reaches** https://mcp.aisa.one/seo/mcp for all of it at once — rankings, keywords, backlinks, site health and AI-answer visibility across DataForSEO, Semrush and Ahrefs.
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
- QuallaaOAuthcom.quallaa
Talk to your public-facing AI from any MCP client — Claude, ChatGPT, Cursor, Cline, Windsurf.
Related MCP Servers
- AlicenseAqualityCmaintenanceAutomateLab AI-SEO audits, scores, and rewrites web pages for AI citation eligibility, AEO and GEO. No API keys or registration. Works with Claude, Cursor, Codex, and other MCP clients. Product and documentation: https://automatelab.tech/products/mcp/ai-seo/2058 npm3MIT
- AlicenseNot gradedqualityFmaintenanceMCP server that exposes OnCrawl's API for use with Claude Code and Claude Desktop. Enables Claude to perform deep technical SEO analysis by querying crawl data, Google Search Console metrics, and crawl-over-crawl comparisons.2MIT
- AlicenseAqualityBmaintenanceCode intelligence MCP server for Claude Code providing multi-project code graph, semantic search, session history, knowledge base, and web search.154MIT
- AlicenseAqualityAmaintenanceThe MCP server for SEO. Find prospects, draft outreach, and monitor backlinks from your AI agent.14MIT