WebIntel-MCP
# WebIntel-MCP
[](https://github.com/jianx/webintel-mcp/actions/workflows/ci.yml)
A read-only web-intelligence MCP server with **exactly three tools** — `mcp_browser_navigate`, `mcp_extract_table`, `mcp_take_screenshot`. It fetches pages through a security gate, purifies them down to the article body, and hands an agent Markdown or structured rows instead of a raw DOM dump.
---
## Why
Two well-known servers already drive a browser over MCP. Neither is a drop-in for an agent whose job is *reading pages*, not *operating sites*.
| | `@playwright/mcp` | `mcp-server-fetch`-style fetchers | WebIntel-MCP |
|---|---|---|---|
| Tool surface | 72 tools (click / fill / drag / network / storage …) | 1 tool, no browser | **3 tools**, frozen |
| Security boundary | Self-described *"not a security boundary"*; `--allowed-origins` is documented as **not** covering redirects | none | Hard gate before any request + post-navigation re-check |
| Content | Raw DOM / accessibility tree | Raw HTML → naive text | Readability → Markdown, `<details>` expanded, boundary-safe truncation |
| Context cost | High (agent must navigate a tree) | High | Bounded by `WEBINTEL_MAX_CHARS` |
| Code execution | `page.evaluate` reachable via tool surface | — | Not exposed; no code-execution tool exists by design |
Three concrete reasons this server exists:
1. **Fewer tools, fewer ways to go wrong.** A 72-tool surface lets a model improvise. Reading a page needs exactly three capabilities, so the surface is frozen at three and a fourth tool is not accepted.
2. **The gate is the product.** Domain allowlisting that only inspects the pre-navigation URL is trivially bypassable — the browser follows the redirect itself. Here the post-redirect final URL is re-checked against the same verdicts.
3. **Purification is not optional.** A raw DOM or a full-HTML text dump costs context and buries the answer. Readability plus a container fallback turns a page into a bounded Markdown document.
There is deliberately **no** `evaluate` / `run_code` tool. An RCE-shaped tool would poison any client that mounts this server, so the capability is absent rather than gated.
---
## Tools
Three tools. The surface is frozen — adding a fourth is out of scope by design.
All three return a JSON-encoded string. Guard refusals are returned as MCP tool errors (`isError: true`) whose message is a JSON payload, so the agent can branch on `error` instead of parsing prose.
### `mcp_browser_navigate`
Navigate to a URL and return the purified page body as Markdown.
```jsonc
// input
{
"type": "object",
"properties": {
"url": { "type": "string", "description": "Absolute http/https URL." },
"wait_until": { "type": "string", "default": "networkidle",
"description": "Playwright wait state. Use \"domcontentloaded\" for pages holding long-poll connections that keep networkidle from firing." },
"timeout_ms": { "type": "integer", "default": 15000 }
},
"required": ["url"]
}
```
```jsonc
// output (JSON string)
{
"url": "https://github.com/langchain-ai/langgraph/releases",
"final_url": "https://github.com/langchain-ai/langgraph/releases", // post-redirect; was re-checked
"title": "Releases · langchain-ai/langgraph",
"elapsed_ms": 1843,
"extracted_chars": 12044,
"truncated": false,
"markdown": "# Releases ..." // last, so a client that clips long results keeps the metadata above
}
```
### `mcp_extract_table`
Extract every table matching a selector as structured rows.
```jsonc
// input
{
"type": "object",
"properties": {
"url": { "type": "string" },
"table_selector": { "type": "string", "default": "table" },
"max_rows": { "type": "integer", "default": 200 }
},
"required": ["url"]
}
```
```jsonc
// output (JSON string)
{
"url": "https://github.com/...",
"final_url": "https://github.com/...",
"row_counts": [5],
"tables": [
[ { "Version": "0.2.60", "Release Date": "2024-05-01", "Highlights": "..." } ]
]
}
```
Tables are a separate tool rather than a `mcp_browser_navigate` variant on purpose: Readability routinely drops or mangles interleaved `<table>` / `<details>` markup, so narrative text and tabular data are extracted by different code paths.
### `mcp_take_screenshot`
Screenshot a URL — optionally element-scoped or full-page — into the shared artifacts directory.
```jsonc
// input
{
"type": "object",
"properties": {
"url": { "type": "string" },
"selector": { "type": "string", "default": "", "description": "CSS selector; empty = viewport (or whole page when full_page)." },
"full_page": { "type": "boolean", "default": false }
},
"required": ["url"]
}
```
```jsonc
// output (JSON string)
{ "url": "https://...", "final_url": "https://...", "path": "github.com/20260930T101533-1-000042.png", "bytes": 84591 }
```
`path` is relative to `WEBINTEL_ARTIFACTS_DIR`. Callers cannot influence it beyond the host segment, which is sanitized.
### Error payloads
Guard refusals use stable `error` codes:
| code | meaning |
|---|---|
| `scheme_not_allowed` | scheme is not `http` / `https` (`file://`, `data:`, `ftp://` …) |
| `blocked_target` | host resolves to a private / loopback / link-local / reserved / multicast address |
| `domain_not_allowed` | host is not in `CRAWL_ALLOW_DOMAINS`; payload lists `allowed_domains` |
| `rate_limited` | per-host token bucket empty; payload carries `retry_after_ms` |
| `navigation_timeout` | `page.goto` exceeded `timeout_ms` |
| `extraction_failed` | purification produced nothing usable |
> `error` codes are stable and English. The human-readable `message` / `hint` fields in the payloads are currently written in Chinese.
---
## Security Model
Every request passes one gate; navigation never starts on a refused URL.
```
requested URL
│
├─ 1. scheme http / https only → SchemeNotAllowedError
├─ 2. DNS resolve every address family, before any request
├─ 3. SSRF verdict private / loopback / link-local /
│ reserved / multicast / unspecified → BlockedTargetError
├─ 4. allowlist exact host match (no wildcards) → DomainNotAllowedError
├─ 5. rate limit per-host token bucket, fail fast → RateLimitedError
└─ 6. demo rewrite DEMO_MODE=true → host becomes mock-web
│
▼
page.goto(url)
│
▼
┌─ final-URL re-check: page.url re-enters steps 1 + 2 + 3 + 4
└─ → DomainNotAllowedError / BlockedTargetError
```
**The re-check is the point.** An origin allowlist that inspects only the URL you asked for is bypassable: your URL is allowed, the server answers `302` to somewhere else, and the browser follows it. Step 3 of the pipeline runs again on the *final* URL, so a redirect into an unlisted host or an internal address fails the call rather than silently fetching it. Only the rate-limit step is skipped on the re-check — the request already happened, so that is a verdict, not admission control.
**Isolation.** Each call gets a throwaway browser context: cookies, `localStorage` and sessions never survive a call, and no login state is ever carried. Content extraction runs on the fetched HTML; `page.evaluate` is not exposed through any tool. Screenshot and HTML paths are built from a sanitized host segment and a generated filename — never from caller input — so they cannot escape `WEBINTEL_ARTIFACTS_DIR`.
**What "security" means here.** It means the gate above and nothing more. It is not a sandbox, not a hardened egress proxy, and not a claim about the pages you fetch. The known gaps are listed in the next section rather than hidden.
---
## Configuration
All configuration is environment variables, read once at process start (a container restart refreshes them; there is no hot reload by design).
| Variable | Default | Effect |
|---|---|---|
| `CRAWL_ALLOW_DOMAINS` | `github.com,news.ycombinator.com,arxiv.org` | Comma-separated allowlist, **exact host match** — `*.github.com` is not supported. Empty string disables the allowlist (local debugging only; the private-IP block still applies). |
| `DEMO_MODE` | `false` | When `true`, the host of every request is rewritten to the in-network mock host `mock-web` for offline demos. |
| `WEBINTEL_ARTIFACTS_DIR` | `/artifacts` | Directory for screenshots and raw HTML. |
| `WEBINTEL_MAX_CHARS` | `30000` | Character budget for `mcp_browser_navigate` Markdown. Truncation retreats to a paragraph boundary and sets `truncated`. |
| `WEBINTEL_RATE_PER_SEC` | `1` | Per-host token-bucket refill rate (burst = capacity). Over-limit calls raise immediately; they do not queue. |
Transport is selected by CLI flags, not environment: `--transport stdio|http|sse`, plus `--host` and `--port` (default `0.0.0.0:9002`) for network transports.
---
## Known Limitations
- **DNS rebinding (TOCTOU).** Hostnames are resolved and judged *before* the request, but Playwright resolves the name again when it actually connects. A name that answers with a public address during the check and a private one microseconds later can slip through. Closing this needs IP pinning plus network-layer enforcement, which is out of scope here — treat the SSRF verdict as a strong filter, not a proof.
- **Rate limiting is in-process and per-container.** The token buckets live in a module-level dict guarded by an `asyncio.Lock`. Two replicas of this server do not share state, so `WEBINTEL_RATE_PER_SEC` bounds the rate of *one* process, not of your whole deployment. Deliberate: no Redis, no database.
- **Purification depends on documented preprocessing.** Collapsed `<details>` blocks (`summary` → `h3`, wrapper → `div`) are expanded before Readability runs, because Readability mis-scores them badly on GitHub Releases. Sites that hide content behind JS interactions rather than `<details>` may under-extract; a container fallback (`main` / `[role=main]` / `article` / `body`) covers the near-total-loss case.
- **`wait_until` is a heuristic.** `networkidle` never fires on pages holding long-lived connections; use `domcontentloaded` there.
- **Single shared browser process.** Chromium is a process-wide singleton with disposable contexts per call. That is cheap, but a crashed browser affects concurrent calls until it is relaunched.
- **Error payload prose is Chinese.** Codes are stable ASCII; `message` / `hint` are not yet localized.
---
## Install
Requires Python ≥ 3.10 and a Chromium for Playwright. **Any local (non-Docker) install needs a browser binary:** `uvx` installs the Python package but not Chromium, so run `playwright install chromium` once, or use the Docker path which bundles it.
### 1. stdio via `uvx` (git — works before any PyPI release)
```json
{
"mcpServers": {
"webintel-mcp": {
"command": "uvx",
"args": ["--from", "git+https://github.com/jianx/webintel-mcp", "webintel-mcp"],
"env": {
"CRAWL_ALLOW_DOMAINS": "github.com,news.ycombinator.com,arxiv.org"
}
}
}
}
```
### 2. stdio via `uvx` (PyPI)
Same config, shorter args — available once `webintel-mcp` is published to PyPI:
```json
{
"mcpServers": {
"webintel-mcp": {
"command": "uvx",
"args": ["webintel-mcp"],
"env": {
"CRAWL_ALLOW_DOMAINS": "github.com,news.ycombinator.com,arxiv.org"
}
}
}
}
```
### 3. stdio via Docker (no local Python, no `playwright install`)
The image's default `CMD` serves HTTP, so stdio needs an explicit command override, and environment variables must be forwarded with `-e` (the client's `env` block reaches the `docker` CLI, not the container):
```json
{
"mcpServers": {
"webintel-mcp": {
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "CRAWL_ALLOW_DOMAINS=github.com,news.ycombinator.com,arxiv.org",
"webintel-mcp:0.1.0",
"webintel-mcp", "--transport", "stdio"
]
}
}
}
```
Build that image from the repo:
```bash
docker build -f docker/Dockerfile -t webintel-mcp:0.1.0 .
```
The `Dockerfile` builds on the official Playwright Python image so Chromium and its OS libraries come with the base.
### 4. streamable-http (long-running container, e.g. for a backend service)
```json
{
"mcpServers": {
"webintel-mcp": {
"url": "http://localhost:9002/mcp"
}
}
}
```
```bash
docker run -d --rm -p 9002:9002 \
-v webintel-artifacts:/artifacts \
-e CRAWL_ALLOW_DOMAINS=github.com,news.ycombinator.com,arxiv.org \
webintel-mcp:0.1.0
```
The image's default `CMD` already serves streamable-http on `0.0.0.0:9002`. `sse` is also available via `--transport sse`.
### Verify
```bash
npx @modelcontextprotocol/inspector uvx --from git+https://github.com/jianx/webintel-mcp webintel-mcp
```
Connect, run `tools/list` — it must show exactly the three tools above — then call `mcp_browser_navigate` with `https://github.com/langchain-ai/langgraph/releases` and check that you get Markdown back.
### Publishing (maintainers)
`server.json` is the official MCP Registry manifest (`io.github.jianx/webintel-mcp`). The Registry listing and the `pypi` package entry it declares require the package to exist first:
```bash
python -m build && twine upload dist/* # PyPI, so `uvx webintel-mcp` resolves
git tag v0.1.0 && git push --tags # Registry entries resolve against a tagged repo
```
Until then, use snippet 1 or 3 above.
## License
MIT — see [LICENSE](LICENSE).
TDQS
Scored across 3 tools
Each tool targets a distinct output mode: purified Markdown body (navigate), structured JSON rows (extract_table), and a rendered image (screenshot). Descriptions explicitly delineate boundaries, e.g. telling the agent that tables must go through extract_table rather than navigate, removing any misselection risk.
All tools share the mcp_ prefix, but the verb patterns are slightly mixed: 'browser_navigate' (noun_verb) versus 'extract_table' and 'take_screenshot' (verb_noun). Still readable and predictable enough to infer purpose at a glance.
Three tools is lean but each earns its place by covering a distinct retrieval modality (text, tabular, visual). It is on the thin side for a server branded 'WebIntel', but not problematic.
Fetching, table extraction, and screenshots are covered, but the surface lacks discovery/search operations and link or metadata extraction that a web-intelligence server would typically expose. Agents can work around this by navigating to known URLs, but it is a notable gap.