Skip to main content
Glama
Tomjuerui

WebIntel-MCP

by Tomjuerui
README.md
# WebIntel-MCP

[![CI](https://github.com/jianx/webintel-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/jianx/webintel-mcp/actions/workflows/ci.yml)

A read-only web-intelligence MCP server with **exactly three tools** — `mcp_browser_navigate`, `mcp_extract_table`, `mcp_take_screenshot`. It fetches pages through a security gate, purifies them down to the article body, and hands an agent Markdown or structured rows instead of a raw DOM dump.

---

## Why

Two well-known servers already drive a browser over MCP. Neither is a drop-in for an agent whose job is *reading pages*, not *operating sites*.

| | `@playwright/mcp` | `mcp-server-fetch`-style fetchers | WebIntel-MCP |
|---|---|---|---|
| Tool surface | 72 tools (click / fill / drag / network / storage …) | 1 tool, no browser | **3 tools**, frozen |
| Security boundary | Self-described *"not a security boundary"*; `--allowed-origins` is documented as **not** covering redirects | none | Hard gate before any request + post-navigation re-check |
| Content | Raw DOM / accessibility tree | Raw HTML → naive text | Readability → Markdown, `<details>` expanded, boundary-safe truncation |
| Context cost | High (agent must navigate a tree) | High | Bounded by `WEBINTEL_MAX_CHARS` |
| Code execution | `page.evaluate` reachable via tool surface | — | Not exposed; no code-execution tool exists by design |

Three concrete reasons this server exists:

1. **Fewer tools, fewer ways to go wrong.** A 72-tool surface lets a model improvise. Reading a page needs exactly three capabilities, so the surface is frozen at three and a fourth tool is not accepted.
2. **The gate is the product.** Domain allowlisting that only inspects the pre-navigation URL is trivially bypassable — the browser follows the redirect itself. Here the post-redirect final URL is re-checked against the same verdicts.
3. **Purification is not optional.** A raw DOM or a full-HTML text dump costs context and buries the answer. Readability plus a container fallback turns a page into a bounded Markdown document.

There is deliberately **no** `evaluate` / `run_code` tool. An RCE-shaped tool would poison any client that mounts this server, so the capability is absent rather than gated.

---

## Tools

Three tools. The surface is frozen — adding a fourth is out of scope by design.

All three return a JSON-encoded string. Guard refusals are returned as MCP tool errors (`isError: true`) whose message is a JSON payload, so the agent can branch on `error` instead of parsing prose.

### `mcp_browser_navigate`

Navigate to a URL and return the purified page body as Markdown.

```jsonc
// input
{
  "type": "object",
  "properties": {
    "url":        { "type": "string", "description": "Absolute http/https URL." },
    "wait_until": { "type": "string", "default": "networkidle",
                    "description": "Playwright wait state. Use \"domcontentloaded\" for pages holding long-poll connections that keep networkidle from firing." },
    "timeout_ms": { "type": "integer", "default": 15000 }
  },
  "required": ["url"]
}
```

```jsonc
// output (JSON string)
{
  "url": "https://github.com/langchain-ai/langgraph/releases",
  "final_url": "https://github.com/langchain-ai/langgraph/releases",  // post-redirect; was re-checked
  "title": "Releases · langchain-ai/langgraph",
  "elapsed_ms": 1843,
  "extracted_chars": 12044,
  "truncated": false,
  "markdown": "# Releases ..."   // last, so a client that clips long results keeps the metadata above
}
```

### `mcp_extract_table`

Extract every table matching a selector as structured rows.

```jsonc
// input
{
  "type": "object",
  "properties": {
    "url":            { "type": "string" },
    "table_selector": { "type": "string", "default": "table" },
    "max_rows":       { "type": "integer", "default": 200 }
  },
  "required": ["url"]
}
```

```jsonc
// output (JSON string)
{
  "url": "https://github.com/...",
  "final_url": "https://github.com/...",
  "row_counts": [5],
  "tables": [
    [ { "Version": "0.2.60", "Release Date": "2024-05-01", "Highlights": "..." } ]
  ]
}
```

Tables are a separate tool rather than a `mcp_browser_navigate` variant on purpose: Readability routinely drops or mangles interleaved `<table>` / `<details>` markup, so narrative text and tabular data are extracted by different code paths.

### `mcp_take_screenshot`

Screenshot a URL — optionally element-scoped or full-page — into the shared artifacts directory.

```jsonc
// input
{
  "type": "object",
  "properties": {
    "url":       { "type": "string" },
    "selector":  { "type": "string", "default": "", "description": "CSS selector; empty = viewport (or whole page when full_page)." },
    "full_page": { "type": "boolean", "default": false }
  },
  "required": ["url"]
}
```

```jsonc
// output (JSON string)
{ "url": "https://...", "final_url": "https://...", "path": "github.com/20260930T101533-1-000042.png", "bytes": 84591 }
```

`path` is relative to `WEBINTEL_ARTIFACTS_DIR`. Callers cannot influence it beyond the host segment, which is sanitized.

### Error payloads

Guard refusals use stable `error` codes:

| code | meaning |
|---|---|
| `scheme_not_allowed` | scheme is not `http` / `https` (`file://`, `data:`, `ftp://` …) |
| `blocked_target` | host resolves to a private / loopback / link-local / reserved / multicast address |
| `domain_not_allowed` | host is not in `CRAWL_ALLOW_DOMAINS`; payload lists `allowed_domains` |
| `rate_limited` | per-host token bucket empty; payload carries `retry_after_ms` |
| `navigation_timeout` | `page.goto` exceeded `timeout_ms` |
| `extraction_failed` | purification produced nothing usable |

> `error` codes are stable and English. The human-readable `message` / `hint` fields in the payloads are currently written in Chinese.

---

## Security Model

Every request passes one gate; navigation never starts on a refused URL.

```
requested URL
   │
   ├─ 1. scheme          http / https only              → SchemeNotAllowedError
   ├─ 2. DNS resolve     every address family, before any request
   ├─ 3. SSRF verdict    private / loopback / link-local /
   │                     reserved / multicast / unspecified → BlockedTargetError
   ├─ 4. allowlist       exact host match (no wildcards)  → DomainNotAllowedError
   ├─ 5. rate limit      per-host token bucket, fail fast → RateLimitedError
   └─ 6. demo rewrite    DEMO_MODE=true → host becomes mock-web
                         │
                         ▼
                    page.goto(url)
                         │
                         ▼
   ┌─ final-URL re-check: page.url re-enters steps 1 + 2 + 3 + 4
   └─                          → DomainNotAllowedError / BlockedTargetError
```

**The re-check is the point.** An origin allowlist that inspects only the URL you asked for is bypassable: your URL is allowed, the server answers `302` to somewhere else, and the browser follows it. Step 3 of the pipeline runs again on the *final* URL, so a redirect into an unlisted host or an internal address fails the call rather than silently fetching it. Only the rate-limit step is skipped on the re-check — the request already happened, so that is a verdict, not admission control.

**Isolation.** Each call gets a throwaway browser context: cookies, `localStorage` and sessions never survive a call, and no login state is ever carried. Content extraction runs on the fetched HTML; `page.evaluate` is not exposed through any tool. Screenshot and HTML paths are built from a sanitized host segment and a generated filename — never from caller input — so they cannot escape `WEBINTEL_ARTIFACTS_DIR`.

**What "security" means here.** It means the gate above and nothing more. It is not a sandbox, not a hardened egress proxy, and not a claim about the pages you fetch. The known gaps are listed in the next section rather than hidden.

---

## Configuration

All configuration is environment variables, read once at process start (a container restart refreshes them; there is no hot reload by design).

| Variable | Default | Effect |
|---|---|---|
| `CRAWL_ALLOW_DOMAINS` | `github.com,news.ycombinator.com,arxiv.org` | Comma-separated allowlist, **exact host match** — `*.github.com` is not supported. Empty string disables the allowlist (local debugging only; the private-IP block still applies). |
| `DEMO_MODE` | `false` | When `true`, the host of every request is rewritten to the in-network mock host `mock-web` for offline demos. |
| `WEBINTEL_ARTIFACTS_DIR` | `/artifacts` | Directory for screenshots and raw HTML. |
| `WEBINTEL_MAX_CHARS` | `30000` | Character budget for `mcp_browser_navigate` Markdown. Truncation retreats to a paragraph boundary and sets `truncated`. |
| `WEBINTEL_RATE_PER_SEC` | `1` | Per-host token-bucket refill rate (burst = capacity). Over-limit calls raise immediately; they do not queue. |

Transport is selected by CLI flags, not environment: `--transport stdio|http|sse`, plus `--host` and `--port` (default `0.0.0.0:9002`) for network transports.

---

## Known Limitations

- **DNS rebinding (TOCTOU).** Hostnames are resolved and judged *before* the request, but Playwright resolves the name again when it actually connects. A name that answers with a public address during the check and a private one microseconds later can slip through. Closing this needs IP pinning plus network-layer enforcement, which is out of scope here — treat the SSRF verdict as a strong filter, not a proof.
- **Rate limiting is in-process and per-container.** The token buckets live in a module-level dict guarded by an `asyncio.Lock`. Two replicas of this server do not share state, so `WEBINTEL_RATE_PER_SEC` bounds the rate of *one* process, not of your whole deployment. Deliberate: no Redis, no database.
- **Purification depends on documented preprocessing.** Collapsed `<details>` blocks (`summary` → `h3`, wrapper → `div`) are expanded before Readability runs, because Readability mis-scores them badly on GitHub Releases. Sites that hide content behind JS interactions rather than `<details>` may under-extract; a container fallback (`main` / `[role=main]` / `article` / `body`) covers the near-total-loss case.
- **`wait_until` is a heuristic.** `networkidle` never fires on pages holding long-lived connections; use `domcontentloaded` there.
- **Single shared browser process.** Chromium is a process-wide singleton with disposable contexts per call. That is cheap, but a crashed browser affects concurrent calls until it is relaunched.
- **Error payload prose is Chinese.** Codes are stable ASCII; `message` / `hint` are not yet localized.

---

## Install

Requires Python ≥ 3.10 and a Chromium for Playwright. **Any local (non-Docker) install needs a browser binary:** `uvx` installs the Python package but not Chromium, so run `playwright install chromium` once, or use the Docker path which bundles it.

### 1. stdio via `uvx` (git — works before any PyPI release)

```json
{
  "mcpServers": {
    "webintel-mcp": {
      "command": "uvx",
      "args": ["--from", "git+https://github.com/jianx/webintel-mcp", "webintel-mcp"],
      "env": {
        "CRAWL_ALLOW_DOMAINS": "github.com,news.ycombinator.com,arxiv.org"
      }
    }
  }
}
```

### 2. stdio via `uvx` (PyPI)

Same config, shorter args — available once `webintel-mcp` is published to PyPI:

```json
{
  "mcpServers": {
    "webintel-mcp": {
      "command": "uvx",
      "args": ["webintel-mcp"],
      "env": {
        "CRAWL_ALLOW_DOMAINS": "github.com,news.ycombinator.com,arxiv.org"
      }
    }
  }
}
```

### 3. stdio via Docker (no local Python, no `playwright install`)

The image's default `CMD` serves HTTP, so stdio needs an explicit command override, and environment variables must be forwarded with `-e` (the client's `env` block reaches the `docker` CLI, not the container):

```json
{
  "mcpServers": {
    "webintel-mcp": {
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "CRAWL_ALLOW_DOMAINS=github.com,news.ycombinator.com,arxiv.org",
        "webintel-mcp:0.1.0",
        "webintel-mcp", "--transport", "stdio"
      ]
    }
  }
}
```

Build that image from the repo:

```bash
docker build -f docker/Dockerfile -t webintel-mcp:0.1.0 .
```

The `Dockerfile` builds on the official Playwright Python image so Chromium and its OS libraries come with the base.

### 4. streamable-http (long-running container, e.g. for a backend service)

```json
{
  "mcpServers": {
    "webintel-mcp": {
      "url": "http://localhost:9002/mcp"
    }
  }
}
```

```bash
docker run -d --rm -p 9002:9002 \
  -v webintel-artifacts:/artifacts \
  -e CRAWL_ALLOW_DOMAINS=github.com,news.ycombinator.com,arxiv.org \
  webintel-mcp:0.1.0
```

The image's default `CMD` already serves streamable-http on `0.0.0.0:9002`. `sse` is also available via `--transport sse`.

### Verify

```bash
npx @modelcontextprotocol/inspector uvx --from git+https://github.com/jianx/webintel-mcp webintel-mcp
```

Connect, run `tools/list` — it must show exactly the three tools above — then call `mcp_browser_navigate` with `https://github.com/langchain-ai/langgraph/releases` and check that you get Markdown back.

### Publishing (maintainers)

`server.json` is the official MCP Registry manifest (`io.github.jianx/webintel-mcp`). The Registry listing and the `pypi` package entry it declares require the package to exist first:

```bash
python -m build && twine upload dist/*          # PyPI, so `uvx webintel-mcp` resolves
git tag v0.1.0 && git push --tags               # Registry entries resolve against a tagged repo
```

Until then, use snippet 1 or 3 above.

## License

MIT — see [LICENSE](LICENSE).

TDQS

A3.7/5.0

Scored across 3 tools

Disambiguation5/5

Each tool targets a distinct output mode: purified Markdown body (navigate), structured JSON rows (extract_table), and a rendered image (screenshot). Descriptions explicitly delineate boundaries, e.g. telling the agent that tables must go through extract_table rather than navigate, removing any misselection risk.

Naming Consistency4/5

All tools share the mcp_ prefix, but the verb patterns are slightly mixed: 'browser_navigate' (noun_verb) versus 'extract_table' and 'take_screenshot' (verb_noun). Still readable and predictable enough to infer purpose at a glance.

Tool Count4/5

Three tools is lean but each earns its place by covering a distinct retrieval modality (text, tabular, visual). It is on the thin side for a server branded 'WebIntel', but not problematic.

Completeness3/5

Fetching, table extraction, and screenshots are covered, but the surface lacks discovery/search operations and link or metadata extraction that a web-intelligence server would typically expose. Agents can work around this by navigating to known URLs, but it is a notable gap.

Maintenance

ActivityMaintained
ResponsivenessNo issues