Skip to main content
Glama
Tyr1onX

web-retrieval-mcp

by Tyr1onX
README.md
# Web Retrieval MCP

A small MCP server that gives an AI agent a unified web-retrieval interface while routing work to specialized backends:

- **Scrapling** for fast static retrieval, targeted extraction, dynamic pages, and stealth fallback.
- **Crawl4AI** for JavaScript rendering and bounded same-site crawling.
- **Firecrawl** for web search (optional; requires an API key).

The model sees four stable tools instead of backend-specific APIs:

| Tool | Purpose |
| --- | --- |
| `web_fetch` | Read one URL as clean Markdown with automatic escalation. |
| `web_crawl` | Crawl a bounded set of internal pages from one site. |
| `web_extract` | Extract repeated elements with a CSS selector. |
| `web_search` | Search the public web through Firecrawl. |

## Routing

`web_fetch(render="auto")` uses a cost-aware fallback chain:

```text
Scrapling static HTTP
        ↓ content missing / placeholder / blocked
Crawl4AI browser render
        ↓ still missing / blocked
Scrapling stealth browser
```

`web_extract(mode="auto")` similarly escalates from static → dynamic → stealth.

`web_crawl` uses Crawl4AI BFS with same-site-only traversal and hard page/depth caps.

`web_search` uses Firecrawl and is only enabled when `FIRECRAWL_API_KEY` is configured.

## Requirements

- Python 3.10+
- Windows, Linux, or macOS
- `uv` recommended

Current tested dependency targets for v0.1:

- MCP Python SDK 2.x
- Crawl4AI 0.9.2.x
- Scrapling 0.4.15.x
- Firecrawl Python SDK 4.40.x

## Windows quick setup

Clone the repository, then run:

```powershell
git clone https://github.com/Tyr1onX/web-retrieval-mcp.git
cd web-retrieval-mcp
powershell -ExecutionPolicy Bypass -File .\scripts\setup-windows.ps1
```

The script installs the Python environment plus Crawl4AI/Scrapling browser dependencies. At the end it prints an executable path similar to:

```text
C:\Users\YOU\web-retrieval-mcp\.venv\Scripts\web-retrieval-mcp.exe
```

### Add it to ChatGPT Desktop

In **Settings → Plugins → MCP → Add → Custom MCP**:

```text
Name: Web Retrieval
Type: STDIO
Startup command: C:\...\web-retrieval-mcp\.venv\Scripts\web-retrieval-mcp.exe
Arguments: (leave empty)
```

Environment variables are optional. To enable Firecrawl search, add:

```text
FIRECRAWL_API_KEY=fc-...
```

Do not put API keys in the repository or `.env` files that you commit.

After enabling the MCP server, a new Work/Codex chat should discover the four tools. A good first test is:

```text
Use web_fetch to read https://example.com and tell me which backend handled it.
```

Then test JavaScript rendering with a JS-heavy public page and `render="always"`.

## Manual setup

```bash
uv sync --extra all
uv run crawl4ai-setup
uv run scrapling install
uv run web-retrieval-mcp
```

The default transport is STDIO.

You can install only the backends you need:

```bash
uv sync --extra scrapling --extra crawl4ai
uv sync --extra firecrawl
```

## Tool details

### `web_fetch`

Inputs:

- `url`: public HTTP(S) URL.
- `render`: `auto`, `never`, or `always`.
- `selector`: optional CSS selector to narrow the returned content.
- `max_chars`: optional response cap.

`auto` starts cheap and escalates. `never` prevents browser rendering. `always` skips the static path and starts with Crawl4AI.

### `web_crawl`

Inputs include `max_depth`, `max_pages`, `per_page_chars`, and `respect_robots`.

The server applies its own hard limits even when a client requests larger values. External links are not followed.

### `web_extract`

Inputs:

- `selector`: CSS selector.
- `attribute`: optional attribute name; otherwise visible text is returned.
- `mode`: `auto`, `static`, `dynamic`, or `stealth`.
- `limit`: maximum returned matches (hard capped at 100).

### `web_search`

Requires `FIRECRAWL_API_KEY`.

Optional filters include `category` (`github`, `research`, `pdf`, `developer`) and a Firecrawl time filter such as `qdr:d` or `qdr:w`.

## Configuration

Copy `.env.example` for the available variables. Important defaults:

```text
WEB_RETRIEVAL_TRANSPORT=stdio
WEB_RETRIEVAL_TIMEOUT_MS=30000
WEB_RETRIEVAL_MAX_CHARS=50000
WEB_RETRIEVAL_MAX_CRAWL_PAGES=20
WEB_RETRIEVAL_MAX_CRAWL_DEPTH=3
WEB_RETRIEVAL_ALLOW_PRIVATE_NETWORKS=false
```

`WEB_RETRIEVAL_ALLOW_PRIVATE_NETWORKS=true` is intentionally opt-in. Do not enable it on a remotely reachable deployment.

## Streamable HTTP / server deployment

Set:

```text
WEB_RETRIEVAL_TRANSPORT=streamable-http
WEB_RETRIEVAL_HOST=127.0.0.1
WEB_RETRIEVAL_PORT=8765
```

Then:

```bash
uv run web-retrieval-mcp
```

The MCP endpoint is:

```text
http://127.0.0.1:8765/mcp
```

A Dockerfile and `docker-compose.yml` are included. Compose intentionally publishes the service only on loopback:

```bash
docker compose up -d --build
```

### Do not expose v0.1 directly to the public Internet

The application blocks private/loopback targets by default, but application-level URL checks cannot eliminate all DNS-rebinding/redirect/egress risks. A remote production deployment should additionally have:

- authentication/access control in front of the MCP endpoint,
- network-level egress restrictions that prevent access to metadata/private networks,
- request/concurrency/rate limits,
- resource caps for browser containers.

Keep the service loopback-only until those controls are in place.

## Security model

The MCP tools accept URLs supplied by an AI model, so all URLs and returned page content are treated as untrusted.

Current safeguards include:

- only `http://` and `https://`,
- embedded URL credentials rejected,
- loopback/private/link-local/reserved/multicast/unspecified IPs blocked by default,
- DNS resolution checked before retrieval,
- obvious final redirect destinations checked again,
- bounded crawl depth/pages,
- bounded output sizes,
- no arbitrary request headers, cookies, POST bodies, JavaScript snippets, local files, or proxies exposed through the MCP schema,
- server instructions explicitly tell the model not to obey instructions found inside retrieved webpage content.

For remote deployments, use network egress controls as the final SSRF boundary.

## Development

Base CI intentionally does not download browsers:

```bash
uv sync --dev
uv run ruff check src tests
uv run pytest -q
```

Full local integration setup:

```bash
uv sync --dev --extra all
uv run crawl4ai-setup
uv run scrapling install
```

## License

MIT. Third-party backends retain their own licenses and are installed as dependencies; their source code is not vendored into this repository.

TDQS

A3.8/5.0

Scored across 4 tools

Disambiguation5/5

Each tool targets a clear, distinct retrieval mode: single-page fetch, multi-page crawl, structured extraction, and web search. The descriptions make the boundaries between fetch and crawl or extract easy to distinguish.

Naming Consistency5/5

All tools follow the same web_<verb> snake_case pattern, making it predictable and easy to remember. The naming convention is consistent across the entire tool set.

Tool Count5/5

Four tools is a well-scoped size for a web retrieval server. Each tool covers a distinct core operation without unnecessary redundancy or overwhelming the agent.

Completeness5/5

The tool surface covers the main web retrieval workflows: fetching a single page, crawling a site, extracting structured elements, and searching the web. No obvious dead ends or missing core operations are apparent for the stated purpose.

Maintenance

ActivityMaintained
ResponsivenessSyncing