mcp-common-crawl
# mcp-common-crawl
> Built by **[Artur Ferreira](https://github.com/arturseo-geo)** @ **The GEO Lab** ยท [๐ @TheGEO\_Lab](https://x.com/TheGEO_Lab) ยท [LinkedIn](https://linkedin.com/in/arturgeo) ยท [Reddit](https://www.reddit.com/user/Alternative_Teach_74/)



MCP server for Common Crawl CDX โ backlink discovery, expired domain finder, competitor gap analysis. Free alternative to Ahrefs/Semrush backlink APIs ($100+/month).
## Tools
| Tool | Description |
|------|-------------|
| `discover_backlinks` | Find backlinks to any domain across 3 CC indexes |
| `find_expired` | Search for expired/parked domains in a niche via CC CDX |
| `check_domain` | Deep single domain check โ live/expired/parked + CC page count |
| `competitor_gap` | Find domains linking to competitors but not to you |
## Features
โ
**Production-tested** โ patterns used in production at [**TheGEOLab**](https://thegeolab.net)
## Install
```bash
# Claude Code
claude mcp add common-crawl -- npx mcp-common-crawl
# Or in .mcp.json
{
"mcpServers": {
"common-crawl": {
"command": "npx",
"args": ["mcp-common-crawl"]
}
}
}
```
## No API Keys Required
Common Crawl is a free, open web archive. No API keys, no rate limits, no paid tiers.
## Usage
```
> find backlinks to thegeolab.net using Common Crawl
> search for expired domains in the "seo tools" niche
> check if example.com is expired or parked
> find link gap between my site and competitors
```
## Important Notes
- Uses native `fetch()` for CC CDX (axios returns 404 on CC CDX โ known issue)
- Queries the 3 most recent CC indexes for best coverage
- Expired domain detection: ECONNREFUSED/ENOTFOUND = expired, parked page pattern matching for parked domains
---
## Attributions & Licence
Built and maintained by **[Artur Ferreira](https://github.com/arturseo-geo)** @ **[TheGEOLab](https://thegeolab.net)**.
Email: artur@thegeolab.net
### Best Practice Attribution
This MCP server was built following the open source Best Practice Approach โ
reading community work for inspiration, then writing original content,
and crediting every source.
**Based on:**
- [Model Context Protocol specification](https://modelcontextprotocol.io) by Anthropic
- [MCP SDK](https://github.com/modelcontextprotocol/sdk) (MIT)
**Data source:**
- [Common Crawl](https://commoncrawl.org/) โ free, open web archive (non-profit)
- [Common Crawl CDX API](https://index.commoncrawl.org/) โ index search endpoint
**Backlink analysis concepts inspired by:**
- [Ahrefs](https://ahrefs.com/) โ backlink discovery and competitor gap methodology
- [Semrush](https://semrush.com/) โ backlink analytics and domain comparison
- [Majestic](https://majestic.com/) โ historic backlink index concepts
**Technical decisions:**
- Native `fetch()` used instead of axios for CC CDX queries (axios returns 404 on CC CDX from inside Express โ persistent debugging issue documented in geolab-backlinks)
All server code is original writing. No files were copied or adapted from any source. MIT licence.
---
Found this useful? โญ Star the repo and connect:
[๐ thegeolab.net](https://thegeolab.net) ยท [๐ @TheGEO_Lab](https://x.com/TheGEO_Lab) ยท [LinkedIn](https://linkedin.com/in/arturgeo) ยท [Reddit](https://www.reddit.com/user/Alternative_Teach_74/)
## Related Repos
- [claude-code-mcps](https://github.com/arturseo-geo/claude-code-mcps) โ All 5 MCP servers in one collection
- [mcp-seo-auditor](https://github.com/arturseo-geo/mcp-seo-auditor) โ On-page SEO audit + JSON-LD validation
- [mcp-serp-intel](https://github.com/arturseo-geo/mcp-serp-intel) โ SERP weak spots, PAA trees, intent comparison
- [mcp-common-crawl](https://github.com/arturseo-geo/mcp-common-crawl) โ Free backlink discovery via Common Crawl
- [mcp-gsc-advanced](https://github.com/arturseo-geo/mcp-gsc-advanced) โ GSC cannibalization, rank changes
- [mcp-wordpress-setup](https://github.com/arturseo-geo/mcp-wordpress-setup) โ WordPress MCP server setup guide
## Licence
MIT โ see LICENSE
---
Built and maintained by **[Artur Ferreira](https://github.com/arturseo-geo)** @ **[TheGEOLab](https://thegeolab.net)** ยท [MIT License](LICENSE)
TDQS
Scored across 4 tools
Tools have distinct purposes, but discover_backlinks and competitor_gap both involve backlink queries, while find_expired and check_domain both assess domain status. Descriptions clarify the differences, so ambiguity is low but not zero.
Naming is inconsistent: discover_backlinks and check_domain use verb_noun, find_expired uses verb_adjective, and competitor_gap is a noun phrase. All use lowercase underscores, but the patterns are mixed.
Four tools is well-scoped for a Common Crawl domain analysis server. Each tool addresses a distinct task without redundancy, and the count fits the niche purpose perfectly.
The surface covers core domain-backlink workflows: backlink discovery, expired domain hunting, domain status checks, and competitor gap analysis. A general URL search or content fetch tool is missing, but agents can work around it.