Skip to main content
Glama
arturseo-geo

mcp-common-crawl

by arturseo-geo
README.md
# mcp-common-crawl

> Built by **[Artur Ferreira](https://github.com/arturseo-geo)** @ **The GEO Lab** ยท [๐• @TheGEO\_Lab](https://x.com/TheGEO_Lab) ยท [LinkedIn](https://linkedin.com/in/arturgeo) ยท [Reddit](https://www.reddit.com/user/Alternative_Teach_74/)

![Version](https://img.shields.io/badge/version-1.0.0-blue)
![Licence](https://img.shields.io/badge/licence-MIT-green)
![Claude Code](https://img.shields.io/badge/Claude_Code-MCP_Server-blueviolet)

MCP server for Common Crawl CDX โ€” backlink discovery, expired domain finder, competitor gap analysis. Free alternative to Ahrefs/Semrush backlink APIs ($100+/month).

## Tools

| Tool | Description |
|------|-------------|
| `discover_backlinks` | Find backlinks to any domain across 3 CC indexes |
| `find_expired` | Search for expired/parked domains in a niche via CC CDX |
| `check_domain` | Deep single domain check โ€” live/expired/parked + CC page count |
| `competitor_gap` | Find domains linking to competitors but not to you |

## Features

โœ… **Production-tested** โ€” patterns used in production at [**TheGEOLab**](https://thegeolab.net)

## Install

```bash
# Claude Code
claude mcp add common-crawl -- npx mcp-common-crawl

# Or in .mcp.json
{
  "mcpServers": {
    "common-crawl": {
      "command": "npx",
      "args": ["mcp-common-crawl"]
    }
  }
}
```

## No API Keys Required

Common Crawl is a free, open web archive. No API keys, no rate limits, no paid tiers.

## Usage

```
> find backlinks to thegeolab.net using Common Crawl
> search for expired domains in the "seo tools" niche
> check if example.com is expired or parked
> find link gap between my site and competitors
```

## Important Notes

- Uses native `fetch()` for CC CDX (axios returns 404 on CC CDX โ€” known issue)
- Queries the 3 most recent CC indexes for best coverage
- Expired domain detection: ECONNREFUSED/ENOTFOUND = expired, parked page pattern matching for parked domains

---

## Attributions & Licence

Built and maintained by **[Artur Ferreira](https://github.com/arturseo-geo)** @ **[TheGEOLab](https://thegeolab.net)**.

Email: artur@thegeolab.net

### Best Practice Attribution

This MCP server was built following the open source Best Practice Approach โ€”
reading community work for inspiration, then writing original content,
and crediting every source.

**Based on:**
- [Model Context Protocol specification](https://modelcontextprotocol.io) by Anthropic
- [MCP SDK](https://github.com/modelcontextprotocol/sdk) (MIT)

**Data source:**
- [Common Crawl](https://commoncrawl.org/) โ€” free, open web archive (non-profit)
- [Common Crawl CDX API](https://index.commoncrawl.org/) โ€” index search endpoint

**Backlink analysis concepts inspired by:**
- [Ahrefs](https://ahrefs.com/) โ€” backlink discovery and competitor gap methodology
- [Semrush](https://semrush.com/) โ€” backlink analytics and domain comparison
- [Majestic](https://majestic.com/) โ€” historic backlink index concepts

**Technical decisions:**
- Native `fetch()` used instead of axios for CC CDX queries (axios returns 404 on CC CDX from inside Express โ€” persistent debugging issue documented in geolab-backlinks)

All server code is original writing. No files were copied or adapted from any source. MIT licence.

---

Found this useful? โญ Star the repo and connect:
[๐ŸŒ thegeolab.net](https://thegeolab.net) ยท [๐• @TheGEO_Lab](https://x.com/TheGEO_Lab) ยท [LinkedIn](https://linkedin.com/in/arturgeo) ยท [Reddit](https://www.reddit.com/user/Alternative_Teach_74/)

## Related Repos

- [claude-code-mcps](https://github.com/arturseo-geo/claude-code-mcps) โ€” All 5 MCP servers in one collection
- [mcp-seo-auditor](https://github.com/arturseo-geo/mcp-seo-auditor) โ€” On-page SEO audit + JSON-LD validation
- [mcp-serp-intel](https://github.com/arturseo-geo/mcp-serp-intel) โ€” SERP weak spots, PAA trees, intent comparison
- [mcp-common-crawl](https://github.com/arturseo-geo/mcp-common-crawl) โ€” Free backlink discovery via Common Crawl
- [mcp-gsc-advanced](https://github.com/arturseo-geo/mcp-gsc-advanced) โ€” GSC cannibalization, rank changes
- [mcp-wordpress-setup](https://github.com/arturseo-geo/mcp-wordpress-setup) โ€” WordPress MCP server setup guide


## Licence

MIT โ€” see LICENSE

---

Built and maintained by **[Artur Ferreira](https://github.com/arturseo-geo)** @ **[TheGEOLab](https://thegeolab.net)** ยท [MIT License](LICENSE)

TDQS

A3.9/5.0

Scored across 4 tools

Disambiguation4/5

Tools have distinct purposes, but discover_backlinks and competitor_gap both involve backlink queries, while find_expired and check_domain both assess domain status. Descriptions clarify the differences, so ambiguity is low but not zero.

Naming Consistency2/5

Naming is inconsistent: discover_backlinks and check_domain use verb_noun, find_expired uses verb_adjective, and competitor_gap is a noun phrase. All use lowercase underscores, but the patterns are mixed.

Tool Count5/5

Four tools is well-scoped for a Common Crawl domain analysis server. Each tool addresses a distinct task without redundancy, and the count fits the niche purpose perfectly.

Completeness4/5

The surface covers core domain-backlink workflows: backlink discovery, expired domain hunting, domain status checks, and competitor gap analysis. A general URL search or content fetch tool is missing, but agents can work around it.

Maintenance

ActivityInactive
ResponsivenessNo issues