ScrapeForge MCP Server
README.md
# ScrapeForge
> **A web scraper that learns page layouts and remembers them.**
> One LLM call per page *template*, not per page.
[](https://python.org)
[](https://fastapi.tiangolo.com)
[](https://github.com/unclecode/crawl4ai)
[](https://www.docker.com)
[]()
---
## What is this?
ScrapeForge is a **self-learning, LLM-powered web scraper** that doesn't just extract data , it *understands* page templates and caches them for reuse.
Traditional scrapers break when a site redesigns. ScrapeForge breaks this cycle.
## The problem with every other "AI scraper"
Most of them do the same thing: send the page to an LLM, get structured data back, charge you for the token, repeat for every single URL. Scrape 1,000 products, pay for 1,000 calls. Site redesigns? Start over.
That's not intelligence. That's an API invoice with extra steps.
ScrapeForge works differently. It looks at a page once, figures out the underlying *template* (`product/iphone` and `product/samsung` are the same layout), generates a CSS extraction schema, and caches it. From then on, extraction is pure CSS selectors no LLM, no cost, no waiting.
**1,000 product pages on the same template = 1 LLM call.** That's the whole pitch.
---
## Origin & Design
I needed a solid scraping system for AI agents and RAG pipelines something reliable, fast, and cheap enough to use at scale. Sending every page to an LLM wasn’t sustainable, and maintaining brittle CSS selectors wasn’t appealing. The idea of caching page *templates* rather than page *data* stayed in my mind for months, but I never had the right foundation to build it cleanly.
That changed when I discovered [Crawl4AI](https://github.com/unclecode/crawl4ai). Its flexible extraction strategies and anti‑bot support gave me exactly the building blocks I needed. I started researching how to separate *layout understanding* from *data extraction*, and iterated on the design until the system could:
- **Learn** a template from a single page using an LLM (most time taken part).
- **Cache** the schema with a matching URL pattern.
- **Test** itself empirically before trusting a cached schema.
- **Merge** templates automatically when they share the same field structure.
The result is ScrapeForge: learn once with an LLM, extract thousands of times with pure CSS cutting cost and latency by over 90% compared to per‑page AI scrapers.
---
## How it works
```javascript
First visit to a URL:
crawl HTML -> LLM generates extraction schema -> validate it
-> test it on the live page -> cache it in SQLite -> return data
Every visit after that:
match URL against cached patterns -> run CSS extraction -> return data
```
And because websites love ruining your week, the cache isn't trusted blindly:
- **Empirical verification** before a cached schema is used for a new URL, the extractor actually *runs* and the fill-rate is measured. Too many empty fields? The schema gets flagged, not silently returned.
- **Rot detection** when a cached schema keeps failing, hit `/schemas/refresh` and a fresh one is generated and swapped in.
- **Template merging** new schemas are compared against existing ones using **Jaccard similarity** on field signatures. Same layout detected? The URL patterns get generalized (`product/iphone` + `product/samsung` -> `product/*`) instead of duplicating schemas.
---
## Architecture
```mermaid
flowchart TD
A["Client<br/>curl / Claude / Cursor / any MCP client"] --> B["FastAPI<br/>+ fastapi-mcp"]
B --> C["SecurityRateLimitMiddleware<br/>X-API-Key check + per-IP rate limit"]
C --> D["/scrape request"]
D --> D1{"crawl_type?"}
%% Markdown / HTML path – skips schema service entirely
D1 -->|"markdown / html"| R1["crawl_with_filter<br/>Crawl4AI + Playwright (stealth mode)"]
R1 --> P1["Markdown or HTML response"]
R1 -.->|"pdf + screenshot<br/>if requested"| Q[("media volume")]
%% Structured path – goes through schema service
D1 -->|"structured"| E["Schema service<br/>(per-domain asyncio lock)"]
E --> F{"Cache lookup<br/>SQLite via SQLAlchemy"}
F -->|"exact URL match"| J
F -->|"pattern candidates"| G["Empirical test:<br/>run extraction, measure fill-rate"]
G -->|"passes (>= 50% fields filled)"| J
G -->|"fails"| H
H["Crawl raw HTML<br/>Crawl4AI + Playwright (stealth mode)"] --> I["LLM generates<br/>JsonCssExtractionStrategy schema"]
%% H -.->|"pdf + screenshot<br/>if requested"| Q
I --> K["Validate schema shape"]
K --> L{"Jaccard similarity >= 0.8<br/>vs existing domain schemas?"}
L -->|"same template"| M["Merge: generalize URL pattern,<br/>append example URL"]
L -->|"new template"| N["Insert new schema row"]
M --> J["Structured extraction<br/>CSS-based, zero LLM cost"]
N --> J
J --> O["Verify output<br/>fill-rate check"]
J -.->|"pdf + screenshot<br/>if requested"| Q
O --> P2["JSON response"]
E -.->|schemas + patterns| R[("SQLite schema.db")]
style E fill:#1f2937,color:#fff
style J fill:#065f46,color:#fff
style R fill:#78350f,color:#fff
```
---
## Features
- **LLM schema generation**: Feed raw HTML to Gemini, OpenAI, Groq, Ollama, anything supported by `crawl4ai.LLMConfig`. Provider (uses `LiteLLM` as proxy) is a runtime parameter, not a hardcoded import.
- **Pattern learning**: URLs sharing a layout are automatically grouped into wildcard patterns. One schema ends up covering thousands of URLs.
- **Empirical cache validation**: cached schemas are *tested*, not assumed. A schema that extracts nothing useful gets rejected before it wastes your time.
- **Self-healing**: stale schema after a redesign? One call to `/schemas/refresh` regenerates and replaces it.
- **Three output modes**: markdown, raw HTML, or structured JSON from a single endpoint.
- **Anti-bot stealth**: Crawl4AI's `magic` mode (stealth, user simulation) on by default, with robots.txt support when you want to be polite.
- **MCP server built in**: every endpoint is automatically exposed as a Model Context Protocol tool. Codex, Claude, Cursor, and any MCP client can call the scraper natively.
- **Media capture**: full-page PDF and screenshots, saved per-request and served via `/download/{id}`.
- **Actually tested**: 25 unit tests covering the URL generalizer, schema similarity, validators, repository layer, retry logic, and rate limiter.
---
## Quickstart
### With Docker (recommended)
```bash
git clone https://github.com/ajazhussainsiddiqui/ScrapeForge.git
cd ScrapeForge
docker build -t scrapeforge .
docker run -d --name scrapeforge \
-p 8000:8000 \
-e ENDPOINT_API_KEY="pick-something-strong" \
-v scrapeforge-data:/data \
--shm-size=1g \
scrapeforge
```
### Local
```bash
python -m venv venv && source venv/bin/activate
pip install -r requirements.txt
playwright install
echo "ENDPOINT_API_KEY=pick-something-strong" > .env
echo "LLM_API_KEY=your_key_here" >> .env # only needed for structured mode
uvicorn api:app --reload
```
> Markdown/HTML scraping needs no LLM key at all. Structured extraction needs one, passed via `.env`.
---
## API
All endpoints except `/health` and `/` require the `X-API-Key` header.
| Method | Endpoint | What it does |
| --- | --- | --- |
| `POST` | `/scrape` | Scrape a URL as markdown, HTML, or structured JSON |
| `GET` | `/schemas` | List cached schemas (optional `?domain=` filter) |
| `DELETE` | `/schemas/{id}` | Delete a cached schema |
| `POST` | `/schemas/refresh` | Force schema regeneration for a URL (post-redesign rescue) |
| `GET` | `/download/{request_id}/{pdf\\|screenshot}` | Fetch saved media |
| `GET` | `/health` | Health check |
### Scrape a page
> you can test this via Swagger UI (`http://localhost:8000/docs`)
```bash
curl -X POST http://localhost:8000/scrape \
-H "X-API-Key: pick-something-strong" \
-H "Content-Type: application/json" \
-d '{
"url": "https://www.bbc.com/news/articles/cvgy5k4n07ko",
"crawl_type": "structured",
"model_provider": "gemini/gemini-2.0-flash",
"api_key": "your_llm_key"
}'
```
First call: slow (LLM is generating the schema). Second call to any URL on the same template: fast, cheap, pure CSS.
```json
{
"success": true,
"data": "{\"content\": \"[{ \\\"title\\\": \\\"...\\\", \\\"author\\\": \\\"...\\\" }]\", \"type\": \"structured\"}"
}
```
Interactive docs at `/docs`. MCP tools at the mounted MCP endpoint for Claude/Cursor.
---
## Configuration
Everything lives in `config.py`, overridable via environment variables:
| Variable | Default | Notes |
| --- | --- | --- |
| `ENDPOINT_API_KEY` | - | **Required.** No key, no server startup. |
| `DATABASE_URL` | `sqlite:///./schema.db` | Any SQLAlchemy URL. Postgres works too. |
| `RATE_LIMIT_REQUESTS` | `10` | Max requests per IP per window |
| `RATE_LIMIT_WINDOW` | `3600` | Window in seconds |
| `SIMILARITY_THRESHOLD` | `0.8` | Jaccard score for merging templates |
| `SCHEMA_FAILURE_SCORE_THRESHOLD` | `0.5` | Max empty-field ratio before a schema is rejected |
| `MAX_RETRIES` / `BACKOFF_BASE_SECONDS` | `3` / `2.0` | Exponential backoff for crawls and LLM calls |
| `CRAWL_TIMEOUT_SECONDS` | `70` | Hard timeout per crawl |
| `MEDIA_DIR` / `LOG_DIR` | `./media` / `./logs` | Overridden to `/data/...` in Docker |
---
## Project Structure (high‑level)
```javascript
ScrapeForge/
├── core/
│ ├── crawler.py # Crawl4AI wrapper: markdown / HTML / structured + media
│ ├── middleware.py # API key auth + per-IP rate limiting
│ ├── schema_gen.py # LLM schema generation (any provider)
│ └── validators.py # schema shape validation + fill-rate verification
├── services/
│ └── schema_service.py # the brain: cache lookup, empirical tests, pattern merging
├── db/
│ ├── connection.py # SQLAlchemy engine
│ ├── models.py # Schema table
│ └── repository.py # CRUD + pattern queries
├── utils/
│ ├── url.py # domain extraction, wildcard matching, URL generalization
│ ├── schema.py # field signatures + Jaccard similarity
│ └── async_helpers.py # retry decorator + RateLimiter
├── tests/ # 25 tests, no network required
├── api.py # FastAPI app + MCP server
├── main.py # orchestration entrypoint
└── Dockerfile
```
---
## Deployed on AWS
The included `Dockerfile` is production-shaped: non-root user, Chromium + system deps baked in, healthcheck on `/health`, `/data` volume for the SQLite DB and media, `shm-size` handled.
```javascript
docker build -t scrapeforge .
docker run -d -p 8000:8000 \
-e ENDPOINT_API_KEY=... \
-v scrapeforge-data:/data \
--shm-size=1g scrapeforge
```
Runs happily on a single EC2 instance, Lightsail container, or any host that can hold a Docker container. Two honest caveats for multi-instance deployments: the rate limiter and per-domain locks are in-memory (pin to one replica, or move the limiter to Redis), and SQLite is a single-writer database (swap `DATABASE_URL` to Postgres when you scale out, the models are already portable).
---
## Honest limitations
A README that only shows strengths is a marketing page. Here's where ScrapeForge struggles:
- **Detectable on some anti‑bot pages.** Heavily restricted pages (like Bloomberg) aren't always bypassed (in current version).
- **LLM output is non-deterministic.** Schema generation mostly works, occasionally produces a weird selector. The validators catch most of it; the retry wrapper catches the rest. Still, expect the occasional failed first crawl on ugly HTML.
- **Heavily personalized / infinite-scroll pages** can defeat static schema caching. `scan_full_page` and stealth mode help, but some SPAs will need `update_schema=true` more often.
- **LLM cost is per-template, not zero.** If every URL on a domain has a unique layout (some forums, some listing sites), you pay per page anyway. ScrapeForge shines when templates repeat, on the web, they almost always do.
---
## Testing
```bash
pytest tests/ -v
```
25 tests across the URL generalizer, schema similarity, validators, repository layer, retry logic, and rate limiter. All run in-memory no network, no API keys, no Playwright.
---
## Why I built this
Every scraping tools ends the same way: "now maintain your selectors forever." (or use LLM tokens for each scraping) I wanted to see what happens if you make the *maintenance* the interesting part schemas that test themselves, patterns that generalize themselves, a cache that knows when it's wrong.
---
## License
MIT. Scrape responsibly, respect `robots.txt`, don't be a jerk with other people's servers.This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues