Skip to main content
Glama
redup-ai

redup.mcp-web-parser

Official
by redup-ai
README.md
# redup.mcp-web-parser

![Docker test](https://github.com/redup-ai/redup.mcp-web-parser/actions/workflows/docker-test.yml/badge.svg?branch=master)
![Python test](https://github.com/redup-ai/redup.mcp-web-parser/actions/workflows/python-test.yml/badge.svg?branch=master)

MCP Streamable HTTP service that parses web pages into cleaned markdown via
[Crawl4AI](https://github.com/unclecode/crawl4ai) `POST /crawl` (0.8.x).

## Model

- Thin MCP façade over Crawl4AI HTTP API — **no browser** in this image.
- Crawl requests set ``exclude_all_images`` (and disable screenshot/pdf) so
  upstream payloads stay small; markdown and links are unchanged for tools.
- **`upstream_base_url` is required** at runtime (config or
  `McpWebParser___upstream_base_url`). Defaults ship **empty** (OSS-safe: no
  cluster hostnames or internal proxies in the repo).
- Optional **egress proxy** for Crawl4AI IP substitution via server config
  `default_proxy` → `crawler_config.proxy_config.server`. Empty = direct fetch.
  Proxy is **not** a tool argument (deploy/runtime only).
- Tool results are **JSON** (`success` / `markdown` / `status_code` / …),
  not a concatenated text dump.
- Targeted at Crawl4AI **0.8.x** (per-request `proxy_config` works). On 0.9+
  Docker API may reject `proxy` / `proxy_config` in the request body.

Contract: MCP tools `parse_page`, `fetch_binary`.
Endpoint: `POST http://<host>:8000/mcp` (stateless Streamable HTTP, JSON).
Metrics: `GET http://<host>:9999/metrics` (Prometheus via `redup-servicekit`).

**Tool args:**
- `parse_page` — HTML pages → markdown JSON. Not for PDF/DOCX/ZIP/images
  (`is_binary=true` → switch to `fetch_binary`).
- `fetch_binary` — download binary files (pdf/docx/zip/images/…). Download only
  (no OCR/unzip). Bytes only in JSON `content_base64` (no shared disk path).
  Downstream tools must accept those bytes via their own input contract.
  Not a fallback when HTML `parse_page` fails.

**Agent registration example:** `{"id":"web-parser","url":"http://…:8000/mcp"}`
→ LLM names `mcp__web-parser__parse_page` / `mcp__web-parser__fetch_binary`.

## Configuration

`config/config.yaml`:

```yaml
service:
  console_log_level: INFO
  host: "0.0.0.0"
  port: 8000
  path: /mcp
  max_workers: 4
  hpa_max_workers: 2

McpWebParser:
  upstream_base_url: ""
  upstream_token: ""
  default_proxy: ""
  request_timeout_seconds: 120
  max_timeout_seconds: 300
  max_markdown_chars: 100000
  max_binary_bytes: 15728640
  delay_before_return_html: 2.5
  json_response: true
  stateless_http: true
```

Override via servicekit env substitution (`section___key`):

```bash
export McpWebParser___upstream_base_url=https://crawl4ai.example.com
export McpWebParser___default_proxy=http://user:pass@proxy.example:3128
export McpWebParser___upstream_token=
export service___port=8000
```

Startup fails fast if `upstream_base_url` is empty.

## Run with Docker

```bash
docker run --rm -p 8000:8000 -p 9999:9999 \
  -e McpWebParser___upstream_base_url=https://crawl4ai.example.com \
  -e McpWebParser___default_proxy=http://proxy.example:3128 \
  redup4ai/redup.mcp-web-parser:0.1.0-3.13-slim
```

MCP URL: `http://127.0.0.1:8000/mcp`. Metrics: `http://127.0.0.1:9999/metrics`.

GitHub Release publishes `{VERSION}-3.13-slim` to Docker Hub (`DOCKERHUB_USER` /
`DOCKERHUB_PASSWORD` secrets).

## Run locally without Docker

Requires Python 3.13+ and a reachable Crawl4AI base URL:

```bash
export McpWebParser___upstream_base_url=https://crawl4ai.example.com
uv sync
uv run python -m redup_mcp_web_parser.service config/config.yaml
```

Desktop MCP clients (stdio):

```bash
uv run redup-mcp-web-parser \
  --transport stdio \
  --upstream-base-url https://crawl4ai.example.com
```

## Tests

```bash
uv sync --dev
uv run pytest tests -q -m "not live"
```

Optional live smoke (needs a real Crawl4AI):

```bash
export McpWebParser___upstream_base_url=https://crawl4ai.example.com
uv run pytest tests -m live -q
```

## License

MIT — see `LICENSE` and `NOTICE`.

TDQS

A4.8/5.0

Scored across 2 tools

Disambiguation5/5

The two tools are cleanly separated: parse_page handles HTML pages and returns markdown, while fetch_binary handles file downloads and returns metadata plus bytes. They also include when-to-use and when-not-to-use guidance, plus a clear fallback path when parse_page detects binary content.

Naming Consistency5/5

Both tools use the verb_noun pattern: parse_page and fetch_binary. Naming is regular, consistent, and immediately communicates what each tool does.

Tool Count4/5

Two tools is lower than the typical 3-15 range, but it maps exactly to the server's intended scope: HTML parsing and binary fetching. The set is slightly thin but not excessive or insufficient for such a narrow purpose.

Completeness5/5

The tool set covers the main web resource categories: HTML pages are parsed to markdown, and binaries are downloaded with metadata and bytes. The is_binary fallback closes the biggest edge case, so there are no obvious dead ends.

Maintenance

ActivitySlowing
ResponsivenessNo issues