Skip to main content
Glama
MrGreenlaw

icecrawl MCP server

by MrGreenlaw
README.md
# ICEcrawl

Self-hosted Firecrawl-alternative for personal use: crawl + JS-render +
markdown-convert web pages via [Crawl4AI](https://github.com/unclecode/crawl4ai),
with optional Claude-powered structured extraction, behind a small job-queued
HTTP API, all in Docker Compose on one machine.

**Architecture:** FastAPI (`api`) accepts jobs → Redis/RQ queue → Python
`worker` containers call the Crawl4AI container's REST API for fetch/render/
markdown, then (optionally) the Claude Messages API for extraction → results
land in SQLite on a shared volume → clients poll job/crawl endpoints.

> **Status: verified working end-to-end (2026-07-22).** `docker compose up`,
> `/healthz`, a real scrape of a live site, and both extraction endpoints
> (real Claude API calls) have all been tested against a running stack. See
> `DEVIATIONS.md` for the one real bug that turned up during that
> verification and how it was fixed — it's already applied here.

## What you need before you start

- **Docker Desktop** (or another Docker Compose-compatible runtime) installed
  and running.
- Two values to put in a config file (below): any random password you make up
  yourself, and — only if you want AI-powered data extraction, not just plain
  scraping — an Anthropic API key from [console.anthropic.com](https://console.anthropic.com).
  Without a key, everything except the `extract` feature works fine.

## Quickstart

```bash
cp .env.example .env
```

Open `.env` in any text editor and fill in:
- `SCRAPER_API_KEY` — make up any long random string, e.g. `openssl rand -hex 32`. This is the password your own requests will use.
- `CRAWL4AI_API_TOKEN` — same idea, another random string. Required for the containers to talk to each other correctly.
- `ANTHROPIC_API_KEY` — optional, only needed for the `extract` feature.

Then, from this folder:

```bash
docker compose up -d --build
```

The first run pulls a few images and can take a few minutes. Once
`docker compose ps` shows everything healthy:

```bash
curl http://localhost:8080/healthz
```

should return `{"status": "ok", "redis": true, "db": true, "crawl4ai": true}`.
If `crawl4ai` shows `false`, give it another 10-20 seconds and retry — it's
usually still starting up.

That's it — the API is live at `http://localhost:8080`.

Scale workers if you're running a lot of jobs at once (each worker holds a
Chromium page in the crawl4ai container while a job is in flight):

```bash
docker compose up -d --scale worker=4
```

## API reference

All endpoints except `/healthz` require `X-API-Key: <SCRAPER_API_KEY>` (the
value you put in `.env`). All bodies are JSON. Errors are
`{"error": {"code": "<UPPER_SNAKE>", "message": "..."}}`.

### Scrape a single page

```bash
curl -X POST http://localhost:8080/v1/scrape \
  -H "X-API-Key: $SCRAPER_API_KEY" -H "Content-Type: application/json" \
  -d '{"url": "https://example.com", "wait": true}'
```

Fire-and-poll instead of waiting:

```bash
curl -X POST http://localhost:8080/v1/scrape \
  -H "X-API-Key: $SCRAPER_API_KEY" -H "Content-Type: application/json" \
  -d '{"url": "https://example.com"}'
# -> {"job_id": "...", "status": "queued", "status_url": "/v1/jobs/..."}

curl http://localhost:8080/v1/jobs/<job_id> -H "X-API-Key: $SCRAPER_API_KEY"
```

Scrape with structured extraction (requires `ANTHROPIC_API_KEY`):

```bash
curl -X POST http://localhost:8080/v1/scrape \
  -H "X-API-Key: $SCRAPER_API_KEY" -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com",
    "wait": true,
    "extract": {
      "schema": {"type": "object", "properties": {"title": {"type": "string"}}, "required": ["title"]},
      "prompt": "Extract the page title"
    }
  }'
```

### Crawl a site (depth/page-limited BFS)

```bash
curl -X POST http://localhost:8080/v1/crawl \
  -H "X-API-Key: $SCRAPER_API_KEY" -H "Content-Type: application/json" \
  -d '{"url": "https://books.toscrape.com/", "max_depth": 1, "max_pages": 5}'
# -> {"crawl_id": "...", "status_url": "/v1/crawls/..."}

curl http://localhost:8080/v1/crawls/<crawl_id> -H "X-API-Key: $SCRAPER_API_KEY"

# with full page results, paginated:
curl "http://localhost:8080/v1/crawls/<crawl_id>?include=results&limit=50&offset=0" \
  -H "X-API-Key: $SCRAPER_API_KEY"

curl -X DELETE http://localhost:8080/v1/crawls/<crawl_id> -H "X-API-Key: $SCRAPER_API_KEY"
```

**Recommended pattern for extraction on a crawl:** crawl without `extract`
first (extraction on every page of a crawl multiplies Claude calls by page
count — the response echoes `estimated_extraction_calls` as a warning when
`extract` is present), then run `/v1/extract` selectively on the pages you
actually want structured data from.

### Extract structured data from raw content (no fetch, synchronous)

```bash
curl -X POST http://localhost:8080/v1/extract \
  -H "X-API-Key: $SCRAPER_API_KEY" -H "Content-Type: application/json" \
  -d '{
    "content": "# Hello World\n\nThis page is about widgets.",
    "schema": {"type": "object", "properties": {"title": {"type": "string"}}, "required": ["title"]}
  }'
```

Returns `{"extraction": {...}, "extraction_meta": {"model", "input_tokens",
"output_tokens", "validation_errors", "truncated"}}`. 422 if `ANTHROPIC_API_KEY`
is not configured (`EXTRACTION_UNCONFIGURED`), or if content/schema fail
validation.

## Using it from Claude (MCP server)

`mcp-server/` wraps the HTTP API as two Claude-friendly tools — `parse` (one
page) and `scrape` (a whole domain, following internal links). This is
optional; the HTTP API above works fine on its own.

```bash
cd mcp-server
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
```

Then register it with Claude Code:

```bash
claude mcp add icecrawl --scope user \
  --env ICECRAWL_BASE_URL=http://localhost:8080 \
  --env SCRAPER_API_KEY=<same key as your .env> \
  -- /full/path/to/mcp-server/.venv/bin/python /full/path/to/mcp-server/server.py
```

(For Claude Desktop, add the same command/env under `mcpServers` in its
config file instead.)

By default, every completed `parse`/`scrape` also saves a markdown note to
`~/Library/Mobile Documents/iCloud~md~obsidian/Documents/My Brain/LLM World/ICEcrawl`
— that's an Obsidian vault path specific to the original author's machine.
If you don't have (or want) that, set `ICECRAWL_OBSIDIAN_DIR` to any folder
on your own machine, or just ignore it — it'll create that folder path the
first time it saves a note.

## Security notes

- Single static API key (`X-API-Key`), constant-time compared. **No TLS** —
  this service is meant to sit on localhost/LAN, not the public internet. Put
  a reverse proxy with real TLS + auth in front before ever exposing it
  publicly.
- **SSRF guard:** every URL that reaches crawl4ai is checked first — it
  rejects loopback/private/link-local/reserved/multicast/unique-local
  addresses (including the cloud metadata IP `169.254.169.254`) and
  non-http(s) schemes. `allow_private: true` (default false) opts a specific
  request out of the private/reserved range check for deliberately scraping
  your own LAN hosts — loopback and link-local stay blocked even then.
  - **Residual risk:** this is a pre-fetch check only; a public URL that
    redirects (or DNS-rebinds) to an internal address can still be fetched.
    Compensating control: `redis` and `crawl4ai` aren't published to the
    host, and the whole stack is meant to stay off the public internet.
- `ANTHROPIC_API_KEY` is read only from `.env` at container runtime, never
  baked into the image, never logged, never returned by any endpoint.
- Request bodies are capped at 2 MB; `/v1/extract` content is separately
  capped at 200,000 characters.

## Known limitations

| Limitation | Fallback |
|---|---|
| Anti-bot walls (Cloudflare/CAPTCHAs) | Job fails with `FETCH_FAILED`. Proxy rotation is out of scope. |
| Heavy SPA / infinite scroll / login-gated content | `wait_for` + `page_timeout_ms` cover most cases; true interaction is out of scope. |
| Extraction cost on crawls | Prefer crawl-then-selective-`/v1/extract` (see above). |
| SQLite under concurrency | Fine at moderate volume; keep workers ≤ 4. |
| Worker crash mid-job | Startup reaper auto-marks stuck jobs failed on next API boot. |
| No dedupe/caching across jobs | Re-scraping a URL re-fetches every time. |

## Cost trade-offs (single glance)

- **Haiku default vs Sonnet override** — Sonnet-class is roughly 10x the
  per-token price of Haiku-class; override (`extract.model`) only when a page
  is hard enough to need it.
- **`EXTRACT_MAX_INPUT_CHARS` / `EXTRACT_MAX_TOKENS`** in `.env` are the two
  spend knobs; defaults (120,000 chars / 4096 tokens) are conservative.
- No Claude call ever happens implicitly — extraction only runs when a
  request explicitly includes `extract` or hits `/v1/extract`.

## Development

```bash
docker compose run --rm api pytest          # full suite inside the container
docker compose run --rm api pytest -m "not integration"  # skip the live crawl4ai check
```

115 tests pass (1 deselected without `-m "not integration"` — the live
crawl4ai integration test self-skips outside a running compose stack).