WebIntel-MCP
Allows fetching, purifying, and extracting content from arXiv pages as Markdown or structured tables and screenshots; arxiv.org is included in the default allowlist.
Allows fetching, purifying, and extracting content from GitHub web pages, such as release pages, as Markdown or structured tables and screenshots; github.com is included in the default allowlist.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@WebIntel-MCPfetch the LangGraph releases page as markdown"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
WebIntel-MCP
A read-only web-intelligence MCP server with exactly three tools — mcp_browser_navigate, mcp_extract_table, mcp_take_screenshot. It fetches pages through a security gate, purifies them down to the article body, and hands an agent Markdown or structured rows instead of a raw DOM dump.
Why
Two well-known servers already drive a browser over MCP. Neither is a drop-in for an agent whose job is reading pages, not operating sites.
|
| WebIntel-MCP | |
Tool surface | 72 tools (click / fill / drag / network / storage …) | 1 tool, no browser | 3 tools, frozen |
Security boundary | Self-described "not a security boundary"; | none | Hard gate before any request + post-navigation re-check |
Content | Raw DOM / accessibility tree | Raw HTML → naive text | Readability → Markdown, |
Context cost | High (agent must navigate a tree) | High | Bounded by |
Code execution |
| — | Not exposed; no code-execution tool exists by design |
Three concrete reasons this server exists:
Fewer tools, fewer ways to go wrong. A 72-tool surface lets a model improvise. Reading a page needs exactly three capabilities, so the surface is frozen at three and a fourth tool is not accepted.
The gate is the product. Domain allowlisting that only inspects the pre-navigation URL is trivially bypassable — the browser follows the redirect itself. Here the post-redirect final URL is re-checked against the same verdicts.
Purification is not optional. A raw DOM or a full-HTML text dump costs context and buries the answer. Readability plus a container fallback turns a page into a bounded Markdown document.
There is deliberately no evaluate / run_code tool. An RCE-shaped tool would poison any client that mounts this server, so the capability is absent rather than gated.
Related MCP server: webfetch-mcp
Tools
Three tools. The surface is frozen — adding a fourth is out of scope by design.
All three return a JSON-encoded string. Guard refusals are returned as MCP tool errors (isError: true) whose message is a JSON payload, so the agent can branch on error instead of parsing prose.
mcp_browser_navigate
Navigate to a URL and return the purified page body as Markdown.
// input
{
"type": "object",
"properties": {
"url": { "type": "string", "description": "Absolute http/https URL." },
"wait_until": { "type": "string", "default": "networkidle",
"description": "Playwright wait state. Use \"domcontentloaded\" for pages holding long-poll connections that keep networkidle from firing." },
"timeout_ms": { "type": "integer", "default": 15000 }
},
"required": ["url"]
}// output (JSON string)
{
"url": "https://github.com/langchain-ai/langgraph/releases",
"final_url": "https://github.com/langchain-ai/langgraph/releases", // post-redirect; was re-checked
"title": "Releases · langchain-ai/langgraph",
"elapsed_ms": 1843,
"extracted_chars": 12044,
"truncated": false,
"markdown": "# Releases ..." // last, so a client that clips long results keeps the metadata above
}mcp_extract_table
Extract every table matching a selector as structured rows.
// input
{
"type": "object",
"properties": {
"url": { "type": "string" },
"table_selector": { "type": "string", "default": "table" },
"max_rows": { "type": "integer", "default": 200 }
},
"required": ["url"]
}// output (JSON string)
{
"url": "https://github.com/...",
"final_url": "https://github.com/...",
"row_counts": [5],
"tables": [
[ { "Version": "0.2.60", "Release Date": "2024-05-01", "Highlights": "..." } ]
]
}Tables are a separate tool rather than a mcp_browser_navigate variant on purpose: Readability routinely drops or mangles interleaved <table> / <details> markup, so narrative text and tabular data are extracted by different code paths.
mcp_take_screenshot
Screenshot a URL — optionally element-scoped or full-page — into the shared artifacts directory.
// input
{
"type": "object",
"properties": {
"url": { "type": "string" },
"selector": { "type": "string", "default": "", "description": "CSS selector; empty = viewport (or whole page when full_page)." },
"full_page": { "type": "boolean", "default": false }
},
"required": ["url"]
}// output (JSON string)
{ "url": "https://...", "final_url": "https://...", "path": "github.com/20260930T101533-1-000042.png", "bytes": 84591 }path is relative to WEBINTEL_ARTIFACTS_DIR. Callers cannot influence it beyond the host segment, which is sanitized.
Error payloads
Guard refusals use stable error codes:
code | meaning |
| scheme is not |
| host resolves to a private / loopback / link-local / reserved / multicast address |
| host is not in |
| per-host token bucket empty; payload carries |
|
|
| purification produced nothing usable |
errorcodes are stable and English. The human-readablemessage/hintfields in the payloads are currently written in Chinese.
Security Model
Every request passes one gate; navigation never starts on a refused URL.
requested URL
│
├─ 1. scheme http / https only → SchemeNotAllowedError
├─ 2. DNS resolve every address family, before any request
├─ 3. SSRF verdict private / loopback / link-local /
│ reserved / multicast / unspecified → BlockedTargetError
├─ 4. allowlist exact host match (no wildcards) → DomainNotAllowedError
├─ 5. rate limit per-host token bucket, fail fast → RateLimitedError
└─ 6. demo rewrite DEMO_MODE=true → host becomes mock-web
│
▼
page.goto(url)
│
▼
┌─ final-URL re-check: page.url re-enters steps 1 + 2 + 3 + 4
└─ → DomainNotAllowedError / BlockedTargetErrorThe re-check is the point. An origin allowlist that inspects only the URL you asked for is bypassable: your URL is allowed, the server answers 302 to somewhere else, and the browser follows it. Step 3 of the pipeline runs again on the final URL, so a redirect into an unlisted host or an internal address fails the call rather than silently fetching it. Only the rate-limit step is skipped on the re-check — the request already happened, so that is a verdict, not admission control.
Isolation. Each call gets a throwaway browser context: cookies, localStorage and sessions never survive a call, and no login state is ever carried. Content extraction runs on the fetched HTML; page.evaluate is not exposed through any tool. Screenshot and HTML paths are built from a sanitized host segment and a generated filename — never from caller input — so they cannot escape WEBINTEL_ARTIFACTS_DIR.
What "security" means here. It means the gate above and nothing more. It is not a sandbox, not a hardened egress proxy, and not a claim about the pages you fetch. The known gaps are listed in the next section rather than hidden.
Configuration
All configuration is environment variables, read once at process start (a container restart refreshes them; there is no hot reload by design).
Variable | Default | Effect |
|
| Comma-separated allowlist, exact host match — |
|
| When |
|
| Directory for screenshots and raw HTML. |
|
| Character budget for |
|
| Per-host token-bucket refill rate (burst = capacity). Over-limit calls raise immediately; they do not queue. |
Transport is selected by CLI flags, not environment: --transport stdio|http|sse, plus --host and --port (default 0.0.0.0:9002) for network transports.
Known Limitations
DNS rebinding (TOCTOU). Hostnames are resolved and judged before the request, but Playwright resolves the name again when it actually connects. A name that answers with a public address during the check and a private one microseconds later can slip through. Closing this needs IP pinning plus network-layer enforcement, which is out of scope here — treat the SSRF verdict as a strong filter, not a proof.
Rate limiting is in-process and per-container. The token buckets live in a module-level dict guarded by an
asyncio.Lock. Two replicas of this server do not share state, soWEBINTEL_RATE_PER_SECbounds the rate of one process, not of your whole deployment. Deliberate: no Redis, no database.Purification depends on documented preprocessing. Collapsed
<details>blocks (summary→h3, wrapper →div) are expanded before Readability runs, because Readability mis-scores them badly on GitHub Releases. Sites that hide content behind JS interactions rather than<details>may under-extract; a container fallback (main/[role=main]/article/body) covers the near-total-loss case.wait_untilis a heuristic.networkidlenever fires on pages holding long-lived connections; usedomcontentloadedthere.Single shared browser process. Chromium is a process-wide singleton with disposable contexts per call. That is cheap, but a crashed browser affects concurrent calls until it is relaunched.
Error payload prose is Chinese. Codes are stable ASCII;
message/hintare not yet localized.
Install
Requires Python ≥ 3.10 and a Chromium for Playwright. Any local (non-Docker) install needs a browser binary: uvx installs the Python package but not Chromium, so run playwright install chromium once, or use the Docker path which bundles it.
1. stdio via uvx (git — works before any PyPI release)
{
"mcpServers": {
"webintel-mcp": {
"command": "uvx",
"args": ["--from", "git+https://github.com/jianx/webintel-mcp", "webintel-mcp"],
"env": {
"CRAWL_ALLOW_DOMAINS": "github.com,news.ycombinator.com,arxiv.org"
}
}
}
}2. stdio via uvx (PyPI)
Same config, shorter args — available once webintel-mcp is published to PyPI:
{
"mcpServers": {
"webintel-mcp": {
"command": "uvx",
"args": ["webintel-mcp"],
"env": {
"CRAWL_ALLOW_DOMAINS": "github.com,news.ycombinator.com,arxiv.org"
}
}
}
}3. stdio via Docker (no local Python, no playwright install)
The image's default CMD serves HTTP, so stdio needs an explicit command override, and environment variables must be forwarded with -e (the client's env block reaches the docker CLI, not the container):
{
"mcpServers": {
"webintel-mcp": {
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "CRAWL_ALLOW_DOMAINS=github.com,news.ycombinator.com,arxiv.org",
"webintel-mcp:0.1.0",
"webintel-mcp", "--transport", "stdio"
]
}
}
}Build that image from the repo:
docker build -f docker/Dockerfile -t webintel-mcp:0.1.0 .The Dockerfile builds on the official Playwright Python image so Chromium and its OS libraries come with the base.
4. streamable-http (long-running container, e.g. for a backend service)
{
"mcpServers": {
"webintel-mcp": {
"url": "http://localhost:9002/mcp"
}
}
}docker run -d --rm -p 9002:9002 \
-v webintel-artifacts:/artifacts \
-e CRAWL_ALLOW_DOMAINS=github.com,news.ycombinator.com,arxiv.org \
webintel-mcp:0.1.0The image's default CMD already serves streamable-http on 0.0.0.0:9002. sse is also available via --transport sse.
Verify
npx @modelcontextprotocol/inspector uvx --from git+https://github.com/jianx/webintel-mcp webintel-mcpConnect, run tools/list — it must show exactly the three tools above — then call mcp_browser_navigate with https://github.com/langchain-ai/langgraph/releases and check that you get Markdown back.
Publishing (maintainers)
server.json is the official MCP Registry manifest (io.github.jianx/webintel-mcp). The Registry listing and the pypi package entry it declares require the package to exist first:
python -m build && twine upload dist/* # PyPI, so `uvx webintel-mcp` resolves
git tag v0.1.0 && git push --tags # Registry entries resolve against a tagged repoUntil then, use snippet 1 or 3 above.
License
MIT — see LICENSE.
Available Tools
3 toolsmcp_extract_tableMcp Extract TableA
Extract all tables from a URL as structured JSON rows.
Tables must go through this tool, not mcp_browser_navigate: Readability routinely drops or mangles interleaved / markup (GitHub Releases pages), so narrative text and tabular data are separated by design.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| max_rows | No | ||
| table_selector | No | table |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It usefully discloses a non-obvious behavioral fact - narrative text and tabular data are separated by design, so this tool returns tables only. It omits operational traits like pagination limits, auth, or rate behavior, but the core surprise an agent would hit is covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short blocks with the core action front-loaded and the routing caveat immediately after. No filler sentences; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers purpose and routing well. The remaining gap is parameter behavior, which the zero-coverage schema leaves entirely to the description and the description does not supply.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 3 parameters, and the description mentions none of them. max_rows (default 200) and table_selector (default 'table') materially affect results, yet neither their meaning nor their defaults are explained anywhere in prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb, resource, and output form: 'Extract all tables from a URL as structured JSON rows.' It also names the sibling it must not be confused with (mcp_browser_navigate), so an agent can discriminate without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing rule: tables must go through this tool, not mcp_browser_navigate, with a concrete justification (Readability drops/mangles interleaved table/details markup on GitHub Releases pages). This is exactly the when-to-use/when-not guidance the dimension asks for.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcp_take_screenshotMcp Take ScreenshotB
Screenshot a URL (optionally element-scoped or full-page) into the shared artifacts volume.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| selector | No | ||
| full_page | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the genuinely useful fact that output lands in the shared artifacts volume rather than being returned inline. However, it says nothing about whether the tool navigates the page itself, whether it requires an already-loaded browser session, or what side effects the capture has.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the core action first and the optional modifiers parenthesized. No filler, no repetition of schema or title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the destination volume is noted. For a three-parameter capture tool this is close to sufficient, with only selector semantics and session prerequisites left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does map all three parameters conceptually: url, "element-scoped" (selector), and "full-page" (full_page). It stops short of clarifying selector syntax or how selector and full_page interact when both are set, so the compensation is only partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Screenshot a URL" gives a specific verb and resource, and the parenthetical scope modifiers clarify the operation. It is clearly distinguishable from mcp_browser_navigate and mcp_extract_table by the capture semantics, though it never explicitly contrasts itself with those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this versus mcp_browser_navigate (which presumably must run first to load a page) or mcp_extract_table. Usage is only implied by the verb; no prerequisites, exclusions, or alternatives are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
mcp_browser_navigate - First observed
mcp_extract_table - First observed
mcp_take_screenshot
TDQS
Scored across 3 tools
Each tool targets a distinct output mode: purified Markdown body (navigate), structured JSON rows (extract_table), and a rendered image (screenshot). Descriptions explicitly delineate boundaries, e.g. telling the agent that tables must go through extract_table rather than navigate, removing any misselection risk.
All tools share the mcp_ prefix, but the verb patterns are slightly mixed: 'browser_navigate' (noun_verb) versus 'extract_table' and 'take_screenshot' (verb_noun). Still readable and predictable enough to infer purpose at a glance.
Three tools is lean but each earns its place by covering a distinct retrieval modality (text, tabular, visual). It is on the thin side for a server branded 'WebIntel', but not problematic.
Fetching, table extraction, and screenshots are covered, but the surface lacks discovery/search operations and link or metadata extraction that a web-intelligence server would typically expose. Agents can work around this by navigating to known URLs, but it is a notable gap.
Maintenance
Related MCP Connectors
Read a URL as clean markdown, screenshot a website, url to PDF. Web access for agents, no signup.
Public web tools for agents: product extraction, claim checks, webpage QA and ranked audits.
Web search and clean-text page fetch for AI agents, with SSRF protection.
Prompt-injection scanning and safe webpage fetching for AI agents reading untrusted content.
Related MCP Servers
- AlicenseCqualityCmaintenanceMCP server for safely reading public URLs for AI agents, providing tools to fetch, extract, cache, and inspect web content as evidence.15MIT
- FlicenseNot gradedqualityCmaintenanceA read-only MCP server that fetches hard-to-scrape web pages via an anti-detection browser and returns clean, token-efficient markdown. It exposes a single fetch tool that converts pages to markdown without requiring approval prompts.-
- AlicenseAqualityAmaintenanceA token-efficient web browsing, scraping, and crawling MCP server that provides tools for fetching pages as clean markdown or schema-based JSON, verifying content, monitoring changes, and interacting with web pages via a persistent browser pool, complete with SSRF protection and rate limiting.77 npmMIT
- AlicenseAqualityAmaintenanceEnables AI assistants to fetch and process web content securely, including HTML-to-markdown conversion, reader mode, metadata extraction, RSS/sitemap parsing, and robots.txt-aware requests with SSRF protection.15506 npm2MIT