web-retrieval-mcp
# Web Retrieval MCP
A small MCP server that gives an AI agent a unified web-retrieval interface while routing work to specialized backends:
- **Scrapling** for fast static retrieval, targeted extraction, dynamic pages, and stealth fallback.
- **Crawl4AI** for JavaScript rendering and bounded same-site crawling.
- **Firecrawl** for web search (optional; requires an API key).
The model sees four stable tools instead of backend-specific APIs:
| Tool | Purpose |
| --- | --- |
| `web_fetch` | Read one URL as clean Markdown with automatic escalation. |
| `web_crawl` | Crawl a bounded set of internal pages from one site. |
| `web_extract` | Extract repeated elements with a CSS selector. |
| `web_search` | Search the public web through Firecrawl. |
## Routing
`web_fetch(render="auto")` uses a cost-aware fallback chain:
```text
Scrapling static HTTP
↓ content missing / placeholder / blocked
Crawl4AI browser render
↓ still missing / blocked
Scrapling stealth browser
```
`web_extract(mode="auto")` similarly escalates from static → dynamic → stealth.
`web_crawl` uses Crawl4AI BFS with same-site-only traversal and hard page/depth caps.
`web_search` uses Firecrawl and is only enabled when `FIRECRAWL_API_KEY` is configured.
## Requirements
- Python 3.10+
- Windows, Linux, or macOS
- `uv` recommended
Current tested dependency targets for v0.1:
- MCP Python SDK 2.x
- Crawl4AI 0.9.2.x
- Scrapling 0.4.15.x
- Firecrawl Python SDK 4.40.x
## Windows quick setup
Clone the repository, then run:
```powershell
git clone https://github.com/Tyr1onX/web-retrieval-mcp.git
cd web-retrieval-mcp
powershell -ExecutionPolicy Bypass -File .\scripts\setup-windows.ps1
```
The script installs the Python environment plus Crawl4AI/Scrapling browser dependencies. At the end it prints an executable path similar to:
```text
C:\Users\YOU\web-retrieval-mcp\.venv\Scripts\web-retrieval-mcp.exe
```
### Add it to ChatGPT Desktop
In **Settings → Plugins → MCP → Add → Custom MCP**:
```text
Name: Web Retrieval
Type: STDIO
Startup command: C:\...\web-retrieval-mcp\.venv\Scripts\web-retrieval-mcp.exe
Arguments: (leave empty)
```
Environment variables are optional. To enable Firecrawl search, add:
```text
FIRECRAWL_API_KEY=fc-...
```
Do not put API keys in the repository or `.env` files that you commit.
After enabling the MCP server, a new Work/Codex chat should discover the four tools. A good first test is:
```text
Use web_fetch to read https://example.com and tell me which backend handled it.
```
Then test JavaScript rendering with a JS-heavy public page and `render="always"`.
## Manual setup
```bash
uv sync --extra all
uv run crawl4ai-setup
uv run scrapling install
uv run web-retrieval-mcp
```
The default transport is STDIO.
You can install only the backends you need:
```bash
uv sync --extra scrapling --extra crawl4ai
uv sync --extra firecrawl
```
## Tool details
### `web_fetch`
Inputs:
- `url`: public HTTP(S) URL.
- `render`: `auto`, `never`, or `always`.
- `selector`: optional CSS selector to narrow the returned content.
- `max_chars`: optional response cap.
`auto` starts cheap and escalates. `never` prevents browser rendering. `always` skips the static path and starts with Crawl4AI.
### `web_crawl`
Inputs include `max_depth`, `max_pages`, `per_page_chars`, and `respect_robots`.
The server applies its own hard limits even when a client requests larger values. External links are not followed.
### `web_extract`
Inputs:
- `selector`: CSS selector.
- `attribute`: optional attribute name; otherwise visible text is returned.
- `mode`: `auto`, `static`, `dynamic`, or `stealth`.
- `limit`: maximum returned matches (hard capped at 100).
### `web_search`
Requires `FIRECRAWL_API_KEY`.
Optional filters include `category` (`github`, `research`, `pdf`, `developer`) and a Firecrawl time filter such as `qdr:d` or `qdr:w`.
## Configuration
Copy `.env.example` for the available variables. Important defaults:
```text
WEB_RETRIEVAL_TRANSPORT=stdio
WEB_RETRIEVAL_TIMEOUT_MS=30000
WEB_RETRIEVAL_MAX_CHARS=50000
WEB_RETRIEVAL_MAX_CRAWL_PAGES=20
WEB_RETRIEVAL_MAX_CRAWL_DEPTH=3
WEB_RETRIEVAL_ALLOW_PRIVATE_NETWORKS=false
```
`WEB_RETRIEVAL_ALLOW_PRIVATE_NETWORKS=true` is intentionally opt-in. Do not enable it on a remotely reachable deployment.
## Streamable HTTP / server deployment
Set:
```text
WEB_RETRIEVAL_TRANSPORT=streamable-http
WEB_RETRIEVAL_HOST=127.0.0.1
WEB_RETRIEVAL_PORT=8765
```
Then:
```bash
uv run web-retrieval-mcp
```
The MCP endpoint is:
```text
http://127.0.0.1:8765/mcp
```
A Dockerfile and `docker-compose.yml` are included. Compose intentionally publishes the service only on loopback:
```bash
docker compose up -d --build
```
### Do not expose v0.1 directly to the public Internet
The application blocks private/loopback targets by default, but application-level URL checks cannot eliminate all DNS-rebinding/redirect/egress risks. A remote production deployment should additionally have:
- authentication/access control in front of the MCP endpoint,
- network-level egress restrictions that prevent access to metadata/private networks,
- request/concurrency/rate limits,
- resource caps for browser containers.
Keep the service loopback-only until those controls are in place.
## Security model
The MCP tools accept URLs supplied by an AI model, so all URLs and returned page content are treated as untrusted.
Current safeguards include:
- only `http://` and `https://`,
- embedded URL credentials rejected,
- loopback/private/link-local/reserved/multicast/unspecified IPs blocked by default,
- DNS resolution checked before retrieval,
- obvious final redirect destinations checked again,
- bounded crawl depth/pages,
- bounded output sizes,
- no arbitrary request headers, cookies, POST bodies, JavaScript snippets, local files, or proxies exposed through the MCP schema,
- server instructions explicitly tell the model not to obey instructions found inside retrieved webpage content.
For remote deployments, use network egress controls as the final SSRF boundary.
## Development
Base CI intentionally does not download browsers:
```bash
uv sync --dev
uv run ruff check src tests
uv run pytest -q
```
Full local integration setup:
```bash
uv sync --dev --extra all
uv run crawl4ai-setup
uv run scrapling install
```
## License
MIT. Third-party backends retain their own licenses and are installed as dependencies; their source code is not vendored into this repository.
TDQS
Scored across 4 tools
Each tool targets a clear, distinct retrieval mode: single-page fetch, multi-page crawl, structured extraction, and web search. The descriptions make the boundaries between fetch and crawl or extract easy to distinguish.
All tools follow the same web_<verb> snake_case pattern, making it predictable and easy to remember. The naming convention is consistent across the entire tool set.
Four tools is a well-scoped size for a web retrieval server. Each tool covers a distinct core operation without unnecessary redundancy or overwhelming the agent.
The tool surface covers the main web retrieval workflows: fetching a single page, crawling a site, extracting structured elements, and searching the web. No obvious dead ends or missing core operations are apparent for the stated purpose.