Scrapy MCP
by B1unt-alt
README.md
# Scrapy MCP — Universal Web Crawling & Scraping
A general-purpose MCP (Model Context Protocol) server for crawling and extracting structured data from any website. Inspired by [Scrapy](https://github.com/scrapy/scrapy)'s architecture.
## Architecture
```
User (Claude) → MCP Tools → Engine → Scheduler → Downloader → Spider → Pipeline → Result
↓ ↓ ↓
Settings Middleware Extractors
```
- **Engine** — Central coordinator (BFS crawl loop)
- **Scheduler** — FIFO queue with URL deduplication
- **Downloader** — Async HTTP via `httpx` with middleware chain (retry, UA rotation, rate limiting)
- **Spider** — Parses responses using configurable CSS/XPath/regex extractors
- **Pipeline** — Post-processing chain (clean, validate, transform, dedup)
- **Middleware** — Pluggable hooks for request/response processing
## Installation
```bash
pip install -r requirements.txt
```
## Register with Claude Code
Edit `~/.claude/mcp.json`:
```json
{
"mcpServers": {
"scrapy-mcp": {
"command": "python",
"args": ["d:/project/scrapy-mcp/server.py"]
}
}
}
```
## MCP Tools
### 1. `crawl` — Structured crawling
Crawl a website and extract structured data using CSS selectors or XPath.
**Parameters:**
| Param | Type | Required | Description |
|-------|------|----------|-------------|
| `urls` | string[] | ✅ | Starting URLs |
| `spider` | string | ❌ | Pre-configured spider name (from spiders/ directory) |
| `item_selector` | string | ❌ | CSS/XPath selector for item containers |
| `fields` | object | ❌ | Field extraction rules |
| `pagination` | object | ❌ | Pagination config |
| `max_pages` | integer | ❌ | Max pages to crawl (default: 10) |
| `max_depth` | integer | ❌ | Max link depth (default: 2) |
| `output_format` | string | ❌ | json / markdown / csv (default: json) |
| `allowed_domains` | string[] | ❌ | Restrict to domains |
**Example:**
```
crawl({
"urls": ["https://books.toscrape.com"],
"item_selector": "article.product_pod",
"fields": {
"title": {"selector": "h3 a", "attr": "title"},
"price": {"selector": ".price_color", "regex": "([\\d.]+)"}
},
"pagination": {"type": "next_link", "selector": ".next a", "max_pages": 3}
})
```
### 2. `scrape_page` — Single page extraction
Fetch and extract content from a single page. Supports multiple modes.
**Parameters:**
| Param | Type | Required | Description |
|-------|------|----------|-------------|
| `url` | string | ✅ | Target page URL |
| `mode` | string | ❌ | auto / markdown / structured / raw (default: auto) |
| `extract_schema` | boolean | ❌ | Extract Schema.org / Open Graph (default: true) |
| `extract_links` | boolean | ❌ | Extract all links (default: false) |
| `extract_images` | boolean | ❌ | Extract all images (default: false) |
| `selectors` | object | ❌ | Custom CSS selectors for specific elements |
**Modes:**
- **auto** — Smart: returns structured JSON if schema data found, otherwise markdown
- **markdown** — Clean readable text with headings, images, links
- **structured** — Full JSON with all extracted metadata + schema data
- **raw** — Original HTML source
### 3. `search_and_crawl` — Search then crawl
Search any website and crawl each result page for details.
**Parameters:**
| Param | Type | Required | Description |
|-------|------|----------|-------------|
| `search_url` | string | ✅ | URL template with `{query}` placeholder |
| `query` | string | ✅ | Search keywords |
| `spider` | string | ❌ | Spider name for result pages |
| `result_item_selector` | string | ❌ | Selector for result items on search page |
| `result_link_selector` | string | ❌ | Selector for links within results |
| `fields` | object | ❌ | Fields to extract from each result page |
| `max_results` | integer | ❌ | Max results (default: 10) |
| `crawl_each_result` | boolean | ❌ | Crawl each result page (default: true) |
### 4. `extract_schema` — Structured data extraction
Extract Schema.org JSON-LD, Open Graph, Twitter Cards, and meta tags.
**Parameters:**
| Param | Type | Required | Description |
|-------|------|----------|-------------|
| `url` | string | ❌ | Page URL |
| `html` | string | ❌ | Or provide raw HTML |
| `schema_types` | string[] | ❌ | Filter by type (e.g. ["Product", "Article"]) |
| `include_raw` | boolean | ❌ | Include raw JSON-LD objects |
### 5. `list_spiders` — Show available spiders
Lists all pre-configured spider definitions from the `spiders/` directory.
## Spider Configuration (YAML)
Create YAML files in `spiders/` to define reusable spiders without writing code:
```yaml
name: my-spider
description: "What this spider does"
start_urls:
- "https://example.com/items"
item_selector: ".item"
fields:
title:
selector: ".title"
type: text
link:
selector: "a"
type: attr
attr: href
transform: absolute_url
pagination:
type: next_link
selector: ".next a"
max_pages: 5
```
See [spiders/example.yml](spiders/example.yml) for an annotated template and [spiders/books_toscrape.yml](spiders/books_toscrape.yml) for a working example.
### Field config reference
| Key | Description |
|-----|-------------|
| `selector` | CSS or XPath expression |
| `type` | `text` (default), `attr`, `html`, `list` |
| `attr` | HTML attribute to extract (e.g. `href`, `src`) |
| `regex` | Regex to apply after extraction (first group returned) |
| `transform` | `strip`, `int`, `float`, `absolute_url`, `lower`, `upper` |
| `default` | Fallback value |
### Pagination types
| Type | Description |
|------|-------------|
| `next_link` | Follow a "next" link using a CSS selector |
| `url_pattern` | Generate URLs from a `{page}` template |
| `scroll` | Scroll-based (requires JS renderer — not yet supported) |
## Settings
Default settings (overridable via spider config):
| Setting | Default | Description |
|---------|---------|-------------|
| `DOWNLOAD_DELAY` | 0.5s | Delay between requests to same domain |
| `CONCURRENT_REQUESTS` | 8 | Max simultaneous requests |
| `TIMEOUT` | 30s | Request timeout |
| `MAX_DEPTH` | 3 | Max link depth |
| `MAX_PAGES` | 50 | Max pages to crawl |
| `MAX_ITEMS` | 500 | Max items to collect |
| `RETRY_TIMES` | 3 | Retry attempts on failure |
| `RETRY_HTTP_CODES` | [500, 502, 503, 504, 408, 429] | Codes to retry |
| `ROBOTSTXT_OBEY` | false | Respect robots.txt |
## Supported Extraction
- **CSS selectors** — `.class`, `#id`, `div > p`, `a[href]`
- **XPath** — `/html/body/div`, `//article//h2`
- **Regex** — Post-process extracted text
- **Schema.org** — JSON-LD, microdata
- **Open Graph** — og:title, og:description, og:image, etc.
- **Meta tags** — description, keywords, etc.
- **Auto-extraction** — All links, images, tables from a page
## Quick Test
```bash
cd d:/project/scrapy-mcp
# Test extraction
python -c "
from mcp_scrapy.extractors import FieldExtractor
import httpx, asyncio
async def test():
async with httpx.AsyncClient() as c:
r = await c.get('https://httpbin.org/html')
ext = FieldExtractor(r.text)
print('Title:', ext.extract_field({'selector': 'h1'}))
print('Links:', len(ext.extract_all_links()))
asyncio.run(test())
"
# Test spider loading
python -c "
from mcp_scrapy.spider import SpiderLoader
loader = SpiderLoader(['spiders'])
for s in loader.list_all():
print(f'{s[\"name\"]}: {s[\"description\"]}')
"
```
## Project Structure
```
scrapy-mcp/
├── server.py # MCP entry point
├── requirements.txt
├── mcp_scrapy/
│ ├── __init__.py
│ ├── engine.py # CrawlEngine coordinator
│ ├── scheduler.py # Request queue + dedup
│ ├── downloader.py # Async HTTP + middleware
│ ├── spider.py # Spider + YAML loader
│ ├── request.py # Request/Response/Item types
│ ├── pipelines.py # Item post-processing
│ ├── middleware.py # Downloader middleware
│ ├── extractors.py # CSS/XPath/regex field extraction
│ ├── exporters.py # JSON/Markdown/CSV output
│ ├── schema_extractor.py # Schema.org / Open Graph
│ ├── settings.py # Configuration
│ └── tools/
│ ├── __init__.py
│ ├── crawl.py # crawl tool
│ ├── scrape_page.py # scrape_page tool
│ ├── search.py # search_and_crawl tool
│ ├── extract_schema.py # extract_schema tool
│ └── list_spiders.py # list_spiders tool
├── spiders/ # Spider YAML configs
│ ├── example.yml
│ ├── hackernews.yml
│ └── books_toscrape.yml
└── tests/
├── test_extractors.py
├── test_pipelines.py
└── test_scheduler.py
```
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues