Skip to main content
Glama
Tomjuerui

WebIntel-MCP

by Tomjuerui

WebIntel-MCP

CI

A read-only web-intelligence MCP server with exactly three tools — mcp_browser_navigate, mcp_extract_table, mcp_take_screenshot. It fetches pages through a security gate, purifies them down to the article body, and hands an agent Markdown or structured rows instead of a raw DOM dump.


Why

Two well-known servers already drive a browser over MCP. Neither is a drop-in for an agent whose job is reading pages, not operating sites.

@playwright/mcp

mcp-server-fetch-style fetchers

WebIntel-MCP

Tool surface

72 tools (click / fill / drag / network / storage …)

1 tool, no browser

3 tools, frozen

Security boundary

Self-described "not a security boundary"; --allowed-origins is documented as not covering redirects

none

Hard gate before any request + post-navigation re-check

Content

Raw DOM / accessibility tree

Raw HTML → naive text

Readability → Markdown, <details> expanded, boundary-safe truncation

Context cost

High (agent must navigate a tree)

High

Bounded by WEBINTEL_MAX_CHARS

Code execution

page.evaluate reachable via tool surface

—

Not exposed; no code-execution tool exists by design

Three concrete reasons this server exists:

  1. Fewer tools, fewer ways to go wrong. A 72-tool surface lets a model improvise. Reading a page needs exactly three capabilities, so the surface is frozen at three and a fourth tool is not accepted.

  2. The gate is the product. Domain allowlisting that only inspects the pre-navigation URL is trivially bypassable — the browser follows the redirect itself. Here the post-redirect final URL is re-checked against the same verdicts.

  3. Purification is not optional. A raw DOM or a full-HTML text dump costs context and buries the answer. Readability plus a container fallback turns a page into a bounded Markdown document.

There is deliberately no evaluate / run_code tool. An RCE-shaped tool would poison any client that mounts this server, so the capability is absent rather than gated.


Related MCP server: webfetch-mcp

Tools

Three tools. The surface is frozen — adding a fourth is out of scope by design.

All three return a JSON-encoded string. Guard refusals are returned as MCP tool errors (isError: true) whose message is a JSON payload, so the agent can branch on error instead of parsing prose.

mcp_browser_navigate

Navigate to a URL and return the purified page body as Markdown.

// input
{
  "type": "object",
  "properties": {
    "url":        { "type": "string", "description": "Absolute http/https URL." },
    "wait_until": { "type": "string", "default": "networkidle",
                    "description": "Playwright wait state. Use \"domcontentloaded\" for pages holding long-poll connections that keep networkidle from firing." },
    "timeout_ms": { "type": "integer", "default": 15000 }
  },
  "required": ["url"]
}
// output (JSON string)
{
  "url": "https://github.com/langchain-ai/langgraph/releases",
  "final_url": "https://github.com/langchain-ai/langgraph/releases",  // post-redirect; was re-checked
  "title": "Releases · langchain-ai/langgraph",
  "elapsed_ms": 1843,
  "extracted_chars": 12044,
  "truncated": false,
  "markdown": "# Releases ..."   // last, so a client that clips long results keeps the metadata above
}

mcp_extract_table

Extract every table matching a selector as structured rows.

// input
{
  "type": "object",
  "properties": {
    "url":            { "type": "string" },
    "table_selector": { "type": "string", "default": "table" },
    "max_rows":       { "type": "integer", "default": 200 }
  },
  "required": ["url"]
}
// output (JSON string)
{
  "url": "https://github.com/...",
  "final_url": "https://github.com/...",
  "row_counts": [5],
  "tables": [
    [ { "Version": "0.2.60", "Release Date": "2024-05-01", "Highlights": "..." } ]
  ]
}

Tables are a separate tool rather than a mcp_browser_navigate variant on purpose: Readability routinely drops or mangles interleaved <table> / <details> markup, so narrative text and tabular data are extracted by different code paths.

mcp_take_screenshot

Screenshot a URL — optionally element-scoped or full-page — into the shared artifacts directory.

// input
{
  "type": "object",
  "properties": {
    "url":       { "type": "string" },
    "selector":  { "type": "string", "default": "", "description": "CSS selector; empty = viewport (or whole page when full_page)." },
    "full_page": { "type": "boolean", "default": false }
  },
  "required": ["url"]
}
// output (JSON string)
{ "url": "https://...", "final_url": "https://...", "path": "github.com/20260930T101533-1-000042.png", "bytes": 84591 }

path is relative to WEBINTEL_ARTIFACTS_DIR. Callers cannot influence it beyond the host segment, which is sanitized.

Error payloads

Guard refusals use stable error codes:

code

meaning

scheme_not_allowed

scheme is not http / https (file://, data:, ftp:// …)

blocked_target

host resolves to a private / loopback / link-local / reserved / multicast address

domain_not_allowed

host is not in CRAWL_ALLOW_DOMAINS; payload lists allowed_domains

rate_limited

per-host token bucket empty; payload carries retry_after_ms

navigation_timeout

page.goto exceeded timeout_ms

extraction_failed

purification produced nothing usable

error codes are stable and English. The human-readable message / hint fields in the payloads are currently written in Chinese.


Security Model

Every request passes one gate; navigation never starts on a refused URL.

requested URL
   │
   ├─ 1. scheme          http / https only              → SchemeNotAllowedError
   ├─ 2. DNS resolve     every address family, before any request
   ├─ 3. SSRF verdict    private / loopback / link-local /
   │                     reserved / multicast / unspecified → BlockedTargetError
   ├─ 4. allowlist       exact host match (no wildcards)  → DomainNotAllowedError
   ├─ 5. rate limit      per-host token bucket, fail fast → RateLimitedError
   └─ 6. demo rewrite    DEMO_MODE=true → host becomes mock-web
                         │
                         ▼
                    page.goto(url)
                         │
                         ▼
   ┌─ final-URL re-check: page.url re-enters steps 1 + 2 + 3 + 4
   └─                          → DomainNotAllowedError / BlockedTargetError

The re-check is the point. An origin allowlist that inspects only the URL you asked for is bypassable: your URL is allowed, the server answers 302 to somewhere else, and the browser follows it. Step 3 of the pipeline runs again on the final URL, so a redirect into an unlisted host or an internal address fails the call rather than silently fetching it. Only the rate-limit step is skipped on the re-check — the request already happened, so that is a verdict, not admission control.

Isolation. Each call gets a throwaway browser context: cookies, localStorage and sessions never survive a call, and no login state is ever carried. Content extraction runs on the fetched HTML; page.evaluate is not exposed through any tool. Screenshot and HTML paths are built from a sanitized host segment and a generated filename — never from caller input — so they cannot escape WEBINTEL_ARTIFACTS_DIR.

What "security" means here. It means the gate above and nothing more. It is not a sandbox, not a hardened egress proxy, and not a claim about the pages you fetch. The known gaps are listed in the next section rather than hidden.


Configuration

All configuration is environment variables, read once at process start (a container restart refreshes them; there is no hot reload by design).

Variable

Default

Effect

CRAWL_ALLOW_DOMAINS

github.com,news.ycombinator.com,arxiv.org

Comma-separated allowlist, exact host match — *.github.com is not supported. Empty string disables the allowlist (local debugging only; the private-IP block still applies).

DEMO_MODE

false

When true, the host of every request is rewritten to the in-network mock host mock-web for offline demos.

WEBINTEL_ARTIFACTS_DIR

/artifacts

Directory for screenshots and raw HTML.

WEBINTEL_MAX_CHARS

30000

Character budget for mcp_browser_navigate Markdown. Truncation retreats to a paragraph boundary and sets truncated.

WEBINTEL_RATE_PER_SEC

1

Per-host token-bucket refill rate (burst = capacity). Over-limit calls raise immediately; they do not queue.

Transport is selected by CLI flags, not environment: --transport stdio|http|sse, plus --host and --port (default 0.0.0.0:9002) for network transports.


Known Limitations

  • DNS rebinding (TOCTOU). Hostnames are resolved and judged before the request, but Playwright resolves the name again when it actually connects. A name that answers with a public address during the check and a private one microseconds later can slip through. Closing this needs IP pinning plus network-layer enforcement, which is out of scope here — treat the SSRF verdict as a strong filter, not a proof.

  • Rate limiting is in-process and per-container. The token buckets live in a module-level dict guarded by an asyncio.Lock. Two replicas of this server do not share state, so WEBINTEL_RATE_PER_SEC bounds the rate of one process, not of your whole deployment. Deliberate: no Redis, no database.

  • Purification depends on documented preprocessing. Collapsed <details> blocks (summary → h3, wrapper → div) are expanded before Readability runs, because Readability mis-scores them badly on GitHub Releases. Sites that hide content behind JS interactions rather than <details> may under-extract; a container fallback (main / [role=main] / article / body) covers the near-total-loss case.

  • wait_until is a heuristic. networkidle never fires on pages holding long-lived connections; use domcontentloaded there.

  • Single shared browser process. Chromium is a process-wide singleton with disposable contexts per call. That is cheap, but a crashed browser affects concurrent calls until it is relaunched.

  • Error payload prose is Chinese. Codes are stable ASCII; message / hint are not yet localized.


Install

Requires Python ≥ 3.10 and a Chromium for Playwright. Any local (non-Docker) install needs a browser binary: uvx installs the Python package but not Chromium, so run playwright install chromium once, or use the Docker path which bundles it.

1. stdio via uvx (git — works before any PyPI release)

{
  "mcpServers": {
    "webintel-mcp": {
      "command": "uvx",
      "args": ["--from", "git+https://github.com/jianx/webintel-mcp", "webintel-mcp"],
      "env": {
        "CRAWL_ALLOW_DOMAINS": "github.com,news.ycombinator.com,arxiv.org"
      }
    }
  }
}

2. stdio via uvx (PyPI)

Same config, shorter args — available once webintel-mcp is published to PyPI:

{
  "mcpServers": {
    "webintel-mcp": {
      "command": "uvx",
      "args": ["webintel-mcp"],
      "env": {
        "CRAWL_ALLOW_DOMAINS": "github.com,news.ycombinator.com,arxiv.org"
      }
    }
  }
}

3. stdio via Docker (no local Python, no playwright install)

The image's default CMD serves HTTP, so stdio needs an explicit command override, and environment variables must be forwarded with -e (the client's env block reaches the docker CLI, not the container):

{
  "mcpServers": {
    "webintel-mcp": {
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "CRAWL_ALLOW_DOMAINS=github.com,news.ycombinator.com,arxiv.org",
        "webintel-mcp:0.1.0",
        "webintel-mcp", "--transport", "stdio"
      ]
    }
  }
}

Build that image from the repo:

docker build -f docker/Dockerfile -t webintel-mcp:0.1.0 .

The Dockerfile builds on the official Playwright Python image so Chromium and its OS libraries come with the base.

4. streamable-http (long-running container, e.g. for a backend service)

{
  "mcpServers": {
    "webintel-mcp": {
      "url": "http://localhost:9002/mcp"
    }
  }
}
docker run -d --rm -p 9002:9002 \
  -v webintel-artifacts:/artifacts \
  -e CRAWL_ALLOW_DOMAINS=github.com,news.ycombinator.com,arxiv.org \
  webintel-mcp:0.1.0

The image's default CMD already serves streamable-http on 0.0.0.0:9002. sse is also available via --transport sse.

Verify

npx @modelcontextprotocol/inspector uvx --from git+https://github.com/jianx/webintel-mcp webintel-mcp

Connect, run tools/list — it must show exactly the three tools above — then call mcp_browser_navigate with https://github.com/langchain-ai/langgraph/releases and check that you get Markdown back.

Publishing (maintainers)

server.json is the official MCP Registry manifest (io.github.jianx/webintel-mcp). The Registry listing and the pypi package entry it declares require the package to exist first:

python -m build && twine upload dist/*          # PyPI, so `uvx webintel-mcp` resolves
git tag v0.1.0 && git push --tags               # Registry entries resolve against a tagged repo

Until then, use snippet 1 or 3 above.

License

MIT — see LICENSE.

Available Tools

3 tools
mcp_browser_navigateMcp Browser NavigateA

Navigate to a URL and return the purified page body as Markdown.

Runs the full security pipeline first (scheme / DNS / private-IP block / domain allowlist / rate limit); navigation never starts on a refused URL. The final (post-redirect) URL is re-checked against the same verdicts. Use wait_until="domcontentloaded" when the page holds long-poll connections that prevent networkidle from ever firing.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
timeout_msNo
wait_untilNonetworkidle

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses the security pipeline (scheme/DNS/private-IP block/domain allowlist/rate limit), that navigation never starts on a refused URL, and that post-redirect URLs are re-validated. It omits error/timeout behavior and rate-limit specifics, but the refusal and redirect semantics are exactly the kind of operational detail an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and return value, then security behavior, then the wait_until tip. Every sentence adds information and none is filler, though the final tip is narrow and could be folded into parameter guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out, and the description supplies the behavioral context that missing annotations would otherwise require. It is still thin on parameter semantics and failure/timeout outcomes for a 3-parameter network tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain all three parameters, yet it only addresses wait_until (and only for one scenario). url and timeout_ms are left entirely undocumented, so an agent gets no guidance on timeout units/limits or accepted URL forms.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Navigate to a URL") and even the return form ("purified page body as Markdown"). It does not explicitly contrast itself with siblings like mcp_take_screenshot or mcp_extract_table, but the navigation purpose is unambiguous on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides one concrete conditional: use wait_until="domcontentloaded" when long-poll connections prevent networkidle from firing. That is parameter-choice guidance rather than when-to-use-this-tool-vs-alternatives, and it never says when a screenshot or table extraction would be the better call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mcp_extract_tableMcp Extract TableA

Extract all tables from a URL as structured JSON rows.

Tables must go through this tool, not mcp_browser_navigate: Readability routinely drops or mangles interleaved / markup (GitHub Releases pages), so narrative text and tabular data are separated by design.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
max_rowsNo
table_selectorNotable

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It usefully discloses a non-obvious behavioral fact - narrative text and tabular data are separated by design, so this tool returns tables only. It omits operational traits like pagination limits, auth, or rate behavior, but the core surprise an agent would hit is covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short blocks with the core action front-loaded and the routing caveat immediately after. No filler sentences; every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers purpose and routing well. The remaining gap is parameter behavior, which the zero-coverage schema leaves entirely to the description and the description does not supply.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 3 parameters, and the description mentions none of them. max_rows (default 200) and table_selector (default 'table') materially affect results, yet neither their meaning nor their defaults are explained anywhere in prose.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb, resource, and output form: 'Extract all tables from a URL as structured JSON rows.' It also names the sibling it must not be confused with (mcp_browser_navigate), so an agent can discriminate without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing rule: tables must go through this tool, not mcp_browser_navigate, with a concrete justification (Readability drops/mangles interleaved table/details markup on GitHub Releases pages). This is exactly the when-to-use/when-not guidance the dimension asks for.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mcp_take_screenshotMcp Take ScreenshotB

Screenshot a URL (optionally element-scoped or full-page) into the shared artifacts volume.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
selectorNo
full_pageNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the genuinely useful fact that output lands in the shared artifacts volume rather than being returned inline. However, it says nothing about whether the tool navigates the page itself, whether it requires an already-loaded browser session, or what side effects the capture has.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the core action first and the optional modifiers parenthesized. No filler, no repetition of schema or title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the destination volume is noted. For a three-parameter capture tool this is close to sufficient, with only selector semantics and session prerequisites left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does map all three parameters conceptually: url, "element-scoped" (selector), and "full-page" (full_page). It stops short of clarifying selector syntax or how selector and full_page interact when both are set, so the compensation is only partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Screenshot a URL" gives a specific verb and resource, and the parenthetical scope modifiers clarify the operation. It is clearly distinguishable from mcp_browser_navigate and mcp_extract_table by the capture semantics, though it never explicitly contrasts itself with those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this versus mcp_browser_navigate (which presumably must run first to load a page) or mcp_extract_table. Usage is only implied by the verb; no prerequisites, exclusions, or alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedmcp_browser_navigate
    • First observedmcp_extract_table
    • First observedmcp_take_screenshot

TDQS

A3.7/5.0

Scored across 3 tools

Disambiguation5/5

Each tool targets a distinct output mode: purified Markdown body (navigate), structured JSON rows (extract_table), and a rendered image (screenshot). Descriptions explicitly delineate boundaries, e.g. telling the agent that tables must go through extract_table rather than navigate, removing any misselection risk.

Naming Consistency4/5

All tools share the mcp_ prefix, but the verb patterns are slightly mixed: 'browser_navigate' (noun_verb) versus 'extract_table' and 'take_screenshot' (verb_noun). Still readable and predictable enough to infer purpose at a glance.

Tool Count4/5

Three tools is lean but each earns its place by covering a distinct retrieval modality (text, tabular, visual). It is on the thin side for a server branded 'WebIntel', but not problematic.

Completeness3/5

Fetching, table extraction, and screenshots are covered, but the surface lacks discovery/search operations and link or metadata extraction that a web-intelligence server would typically expose. Agents can work around this by navigating to known URLs, but it is a notable gap.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    C
    maintenance
    A read-only MCP server that fetches hard-to-scrape web pages via an anti-detection browser and returns clean, token-efficient markdown. It exposes a single fetch tool that converts pages to markdown without requiring approval prompts.
    -
  • A
    license
    A
    quality
    A
    maintenance
    A token-efficient web browsing, scraping, and crawling MCP server that provides tools for fetching pages as clean markdown or schema-based JSON, verifying content, monitoring changes, and interacting with web pages via a persistent browser pool, complete with SSRF protection and rate limiting.
    7
    7 npm
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Enables AI assistants to fetch and process web content securely, including HTML-to-markdown conversion, reader mode, metadata extraction, RSS/sitemap parsing, and robots.txt-aware requests with SSRF protection.
    15
    506 npm
    2
    MIT