ScrapeForge MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ScrapeForge MCP ServerExtract product names and prices from these Amazon product URLs and cache the schema."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ScrapeForge
A web scraper that learns page layouts and remembers them.
One LLM call per page template, not per page.
What is this?
ScrapeForge is a self-learning, LLM-powered web scraper that doesn't just extract data , it understands page templates and caches them for reuse.
Traditional scrapers break when a site redesigns. ScrapeForge breaks this cycle.
Related MCP server: extracto-mcp
The problem with every other "AI scraper"
Most of them do the same thing: send the page to an LLM, get structured data back, charge you for the token, repeat for every single URL. Scrape 1,000 products, pay for 1,000 calls. Site redesigns? Start over.
That's not intelligence. That's an API invoice with extra steps.
ScrapeForge works differently. It looks at a page once, figures out the underlying template (product/iphone and product/samsung are the same layout), generates a CSS extraction schema, and caches it. From then on, extraction is pure CSS selectors no LLM, no cost, no waiting.
1,000 product pages on the same template = 1 LLM call. That's the whole pitch.
Origin & Design
I needed a solid scraping system for AI agents and RAG pipelines something reliable, fast, and cheap enough to use at scale. Sending every page to an LLM wasn’t sustainable, and maintaining brittle CSS selectors wasn’t appealing. The idea of caching page templates rather than page data stayed in my mind for months, but I never had the right foundation to build it cleanly.
That changed when I discovered Crawl4AI. Its flexible extraction strategies and anti‑bot support gave me exactly the building blocks I needed. I started researching how to separate layout understanding from data extraction, and iterated on the design until the system could:
Learn a template from a single page using an LLM (most time taken part).
Cache the schema with a matching URL pattern.
Test itself empirically before trusting a cached schema.
Merge templates automatically when they share the same field structure.
The result is ScrapeForge: learn once with an LLM, extract thousands of times with pure CSS cutting cost and latency by over 90% compared to per‑page AI scrapers.
How it works
First visit to a URL:
crawl HTML -> LLM generates extraction schema -> validate it
-> test it on the live page -> cache it in SQLite -> return data
Every visit after that:
match URL against cached patterns -> run CSS extraction -> return dataAnd because websites love ruining your week, the cache isn't trusted blindly:
Empirical verification before a cached schema is used for a new URL, the extractor actually runs and the fill-rate is measured. Too many empty fields? The schema gets flagged, not silently returned.
Rot detection when a cached schema keeps failing, hit
/schemas/refreshand a fresh one is generated and swapped in.Template merging new schemas are compared against existing ones using Jaccard similarity on field signatures. Same layout detected? The URL patterns get generalized (
product/iphone+product/samsung->product/*) instead of duplicating schemas.
Architecture
flowchart TD
A["Client<br/>curl / Claude / Cursor / any MCP client"] --> B["FastAPI<br/>+ fastapi-mcp"]
B --> C["SecurityRateLimitMiddleware<br/>X-API-Key check + per-IP rate limit"]
C --> D["/scrape request"]
D --> D1{"crawl_type?"}
%% Markdown / HTML path – skips schema service entirely
D1 -->|"markdown / html"| R1["crawl_with_filter<br/>Crawl4AI + Playwright (stealth mode)"]
R1 --> P1["Markdown or HTML response"]
R1 -.->|"pdf + screenshot<br/>if requested"| Q[("media volume")]
%% Structured path – goes through schema service
D1 -->|"structured"| E["Schema service<br/>(per-domain asyncio lock)"]
E --> F{"Cache lookup<br/>SQLite via SQLAlchemy"}
F -->|"exact URL match"| J
F -->|"pattern candidates"| G["Empirical test:<br/>run extraction, measure fill-rate"]
G -->|"passes (>= 50% fields filled)"| J
G -->|"fails"| H
H["Crawl raw HTML<br/>Crawl4AI + Playwright (stealth mode)"] --> I["LLM generates<br/>JsonCssExtractionStrategy schema"]
%% H -.->|"pdf + screenshot<br/>if requested"| Q
I --> K["Validate schema shape"]
K --> L{"Jaccard similarity >= 0.8<br/>vs existing domain schemas?"}
L -->|"same template"| M["Merge: generalize URL pattern,<br/>append example URL"]
L -->|"new template"| N["Insert new schema row"]
M --> J["Structured extraction<br/>CSS-based, zero LLM cost"]
N --> J
J --> O["Verify output<br/>fill-rate check"]
J -.->|"pdf + screenshot<br/>if requested"| Q
O --> P2["JSON response"]
E -.->|schemas + patterns| R[("SQLite schema.db")]
style E fill:#1f2937,color:#fff
style J fill:#065f46,color:#fff
style R fill:#78350f,color:#fffFeatures
LLM schema generation: Feed raw HTML to Gemini, OpenAI, Groq, Ollama, anything supported by
crawl4ai.LLMConfig. Provider (usesLiteLLMas proxy) is a runtime parameter, not a hardcoded import.Pattern learning: URLs sharing a layout are automatically grouped into wildcard patterns. One schema ends up covering thousands of URLs.
Empirical cache validation: cached schemas are tested, not assumed. A schema that extracts nothing useful gets rejected before it wastes your time.
Self-healing: stale schema after a redesign? One call to
/schemas/refreshregenerates and replaces it.Three output modes: markdown, raw HTML, or structured JSON from a single endpoint.
Anti-bot stealth: Crawl4AI's
magicmode (stealth, user simulation) on by default, with robots.txt support when you want to be polite.MCP server built in: every endpoint is automatically exposed as a Model Context Protocol tool. Codex, Claude, Cursor, and any MCP client can call the scraper natively.
Media capture: full-page PDF and screenshots, saved per-request and served via
/download/{id}.Actually tested: 25 unit tests covering the URL generalizer, schema similarity, validators, repository layer, retry logic, and rate limiter.
Quickstart
With Docker (recommended)
git clone https://github.com/ajazhussainsiddiqui/ScrapeForge.git
cd ScrapeForge
docker build -t scrapeforge .
docker run -d --name scrapeforge \
-p 8000:8000 \
-e ENDPOINT_API_KEY="pick-something-strong" \
-v scrapeforge-data:/data \
--shm-size=1g \
scrapeforgeLocal
python -m venv venv && source venv/bin/activate
pip install -r requirements.txt
playwright install
echo "ENDPOINT_API_KEY=pick-something-strong" > .env
echo "LLM_API_KEY=your_key_here" >> .env # only needed for structured mode
uvicorn api:app --reloadMarkdown/HTML scraping needs no LLM key at all. Structured extraction needs one, passed via
.env.
API
All endpoints except /health and / require the X-API-Key header.
Method | Endpoint | What it does |
|
| Scrape a URL as markdown, HTML, or structured JSON |
|
| List cached schemas (optional |
|
| Delete a cached schema |
|
| Force schema regeneration for a URL (post-redesign rescue) |
| `/download/{request_id}/{pdf\ | screenshot}` |
|
| Health check |
Scrape a page
you can test this via Swagger UI (
http://localhost:8000/docs)
curl -X POST http://localhost:8000/scrape \
-H "X-API-Key: pick-something-strong" \
-H "Content-Type: application/json" \
-d '{
"url": "https://www.bbc.com/news/articles/cvgy5k4n07ko",
"crawl_type": "structured",
"model_provider": "gemini/gemini-2.0-flash",
"api_key": "your_llm_key"
}'First call: slow (LLM is generating the schema). Second call to any URL on the same template: fast, cheap, pure CSS.
{
"success": true,
"data": "{\"content\": \"[{ \\\"title\\\": \\\"...\\\", \\\"author\\\": \\\"...\\\" }]\", \"type\": \"structured\"}"
}Interactive docs at /docs. MCP tools at the mounted MCP endpoint for Claude/Cursor.
Configuration
Everything lives in config.py, overridable via environment variables:
Variable | Default | Notes |
| - | Required. No key, no server startup. |
|
| Any SQLAlchemy URL. Postgres works too. |
|
| Max requests per IP per window |
|
| Window in seconds |
|
| Jaccard score for merging templates |
|
| Max empty-field ratio before a schema is rejected |
|
| Exponential backoff for crawls and LLM calls |
|
| Hard timeout per crawl |
|
| Overridden to |
Project Structure (high‑level)
ScrapeForge/
├── core/
│ ├── crawler.py # Crawl4AI wrapper: markdown / HTML / structured + media
│ ├── middleware.py # API key auth + per-IP rate limiting
│ ├── schema_gen.py # LLM schema generation (any provider)
│ └── validators.py # schema shape validation + fill-rate verification
├── services/
│ └── schema_service.py # the brain: cache lookup, empirical tests, pattern merging
├── db/
│ ├── connection.py # SQLAlchemy engine
│ ├── models.py # Schema table
│ └── repository.py # CRUD + pattern queries
├── utils/
│ ├── url.py # domain extraction, wildcard matching, URL generalization
│ ├── schema.py # field signatures + Jaccard similarity
│ └── async_helpers.py # retry decorator + RateLimiter
├── tests/ # 25 tests, no network required
├── api.py # FastAPI app + MCP server
├── main.py # orchestration entrypoint
└── DockerfileDeployed on AWS
The included Dockerfile is production-shaped: non-root user, Chromium + system deps baked in, healthcheck on /health, /data volume for the SQLite DB and media, shm-size handled.
docker build -t scrapeforge .
docker run -d -p 8000:8000 \
-e ENDPOINT_API_KEY=... \
-v scrapeforge-data:/data \
--shm-size=1g scrapeforgeRuns happily on a single EC2 instance, Lightsail container, or any host that can hold a Docker container. Two honest caveats for multi-instance deployments: the rate limiter and per-domain locks are in-memory (pin to one replica, or move the limiter to Redis), and SQLite is a single-writer database (swap DATABASE_URL to Postgres when you scale out, the models are already portable).
Honest limitations
A README that only shows strengths is a marketing page. Here's where ScrapeForge struggles:
Detectable on some anti‑bot pages. Heavily restricted pages (like Bloomberg) aren't always bypassed (in current version).
LLM output is non-deterministic. Schema generation mostly works, occasionally produces a weird selector. The validators catch most of it; the retry wrapper catches the rest. Still, expect the occasional failed first crawl on ugly HTML.
Heavily personalized / infinite-scroll pages can defeat static schema caching.
scan_full_pageand stealth mode help, but some SPAs will needupdate_schema=truemore often.LLM cost is per-template, not zero. If every URL on a domain has a unique layout (some forums, some listing sites), you pay per page anyway. ScrapeForge shines when templates repeat, on the web, they almost always do.
Testing
pytest tests/ -v25 tests across the URL generalizer, schema similarity, validators, repository layer, retry logic, and rate limiter. All run in-memory no network, no API keys, no Playwright.
Why I built this
Every scraping tools ends the same way: "now maintain your selectors forever." (or use LLM tokens for each scraping) I wanted to see what happens if you make the maintenance the interesting part schemas that test themselves, patterns that generalize themselves, a cache that knows when it's wrong.
License
MIT. Scrape responsibly, respect robots.txt, don't be a jerk with other people's servers.
This server cannot be deployed
Maintenance
Related MCP Connectors
Cloud scraping & crawling API for AI agents. Turn any URL into clean, LLM-ready markdown.
Turn any website into structured JSON data matching your custom schema.
Enable language models to perform advanced AI-powered web scraping with enterprise-grade reliabili…
Web scraping for AI agents. Converts URLs to clean, LLM-ready Markdown with anti-bot bypass.
Related MCP Servers
- AlicenseAqualityDmaintenanceStructured web context infrastructure for AI agents. Extract reliable schema-guided JSON from websites using Claude-powered parsing, Browserless fallback rendering, and MCP-native workflows.11MIT
- AlicenseAqualityDmaintenanceEnables Claude and any MCP client to turn a URL plus a schema into validated, typed JSON without HTML parsing or hallucinated fields.431 npmMIT
- AlicenseAqualityBmaintenanceMCP server that extracts structured JSON from public URLs for AI agents using schemas like product, article, and company.644 npmMIT
- AlicenseNot gradedqualityCmaintenanceEnables lead generation and web scraping through 16 MCP tools, allowing AI clients like Claude, Cursor, and Windsurf to perform scraping tasks via natural language.25 npmMIT