scrape-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| SCRAPE_MCP_PROXY | No | 代理,L1/L2 共用 | |
| SCRAPE_MCP_DATA_DIR | No | 登录态/缓存放盘目录 | ~/.scrape_mcp |
| SCRAPE_MCP_CACHE_TTL | No | L1 缓存 TTL(秒),`0` 关闭 | 300 |
| SCRAPE_MCP_HTTP_PORT | No | HTTP 端口 | |
| SCRAPE_MCP_TRANSPORT | No | 传输协议,stdio 或 streamable-http | stdio |
| SCRAPE_MCP_L2_CHANNEL | No | 关联浏览器:空=自带 Chromium,`chrome`/`msedge`=系统 | |
| SCRAPE_MCP_L2_ENABLED | No | 是否启用 Playwright 升级链 | true |
| SCRAPE_MCP_IMPERSONATE | No | curl_cffi 指纹模板 | chrome |
| SCRAPE_MCP_L2_TIMEZONE | No | 隐身加固时区 | Asia/Shanghai |
| SCRAPE_MCP_BATCH_MAX_URLS | No | `web_batch` 单批上限 | 20 |
| SCRAPE_MCP_REQUEST_TIMEOUT | No | L1 超时(秒) | 20 |
| SCRAPE_MCP_L2_SETTLE_TIMEOUT | No | 挑战页等待窗口(秒) | 8 |
| SCRAPE_MCP_BATCH_MAX_CONCURRENCY | No | `web_batch` 并发上限 | 4 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| web_fetchA | 抓取网页并返回极省 token 的紧凑正文(自动辨别反爬拦截页并如实上报)。 Args: url: 目标网页完整 URL,需带 http/https 协议头。 max_tokens: 正文 token 预算上限,超出时保留头尾、省略中段。 link_policy: 链接处理策略。internal=站内链接保留为相对路径、站外降级为纯文本(默认,最省); all=站内外链接全保留;none=全部降级为纯文本。 |
| web_extractA | 抓取网页并按字段 schema 抽取结构化 JSON(字段级数据,可入库)。 适合需要"数据而非整页正文"的场景:给 URL 和字段规则,直接返回字段值, 而不是一段 compact 文本。抓取链路(分级反爬、登录态、合规)与 web_fetch 一致。 字段规则示例: {"fields": { "title": {"selector": "h1", "type": "text"}, "first_heading": {"selector": "h1", "type": "text"}, "main_link": {"selector": "a", "type": "attr", "attr": "href"}, "link_count": {"selector": "a", "type": "count"}, "tags": {"selector": ".tag", "type": "list", "list_key": "text"} }} 字段类型:text(默认,节点归一化文本)/ attr(需配 attr,取属性值)/ count(匹配节点数)/ list(取所有匹配节点,list_key 决定取值方式: text / text_trimmed / attr / html)。 Args: url: 目标网页完整 URL。 schema: 字段抽取规则字典,见上方示例。 |
| web_batchA | 并发抓取一批 URL,各自走完整分级链路,返回逐个结果。 Args: urls: 目标 URL 列表(最多 batch_max_urls 个,默认 20)。 max_tokens: 每个 URL 的正文 token 预算上限。 link_policy: 同 web_fetch。 |
| loginA | 打开带界面的浏览器手动登录,并把该站登录态持久化到磁盘。 用于 web_fetch 报告 login_required=true 的站点:浏览器会以非无头方式打开目标页, 请在弹出的窗口里完成登录;检测到登录成功或超时后,登录态(cookie 等)会被保存, 之后该站的 web_fetch 会自动带上登录态。 Args: url: 目标站点任一页面 URL(按域名区分登录态)。 timeout: 等待手动登录完成的秒数,超时也会保存当前状态便于下次继续。 |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: single-page text fetch, structured field extraction, manual login for auth persistence, and batch fetching. The overlap between web_fetch and web_batch is limited to shared parameters, but their single vs. multi-URL scopes are unambiguous.
Three tools follow the web_verb pattern (web_fetch, web_extract, web_batch), which is predictable and readable. The lone 'login' tool breaks the pattern, but it represents a different kind of operation (authentication rather than fetching), so the deviation is minor.
With four tools, the server is well-scoped for scraping: a basic fetch, a structured extraction, a batch operation, and an auth helper. Each tool fills a distinct role without unnecessary bloat or obvious missing core functionality.
The tool surface covers the main scraping lifecycle: fetching text, extracting structured data, handling multiple URLs, and dealing with login-protected pages. Minor gaps exist, such as no raw HTML output or explicit pagination/crawling support, but these can be worked around with the existing extraction and batch tools.