ScrAPI MCP Server
![]()
ScrAPI MCP 服务器
MCP 服务器使用ScrAPI来抓取网页。
ScrAPI 是您的终极网络抓取解决方案,提供强大、可靠且易于使用的功能,可毫不费力地从任何网站提取数据。
工具
scrape_url_html使用 ScrAPI 服务,通过 URL 抓取网站内容,并以 HTML 格式获取结果。此功能适用于抓取因机器人检测、验证码甚至地理位置限制而难以访问的网站内容。结果将以 HTML 格式呈现,如果需要高级解析,则更适合使用 HTML 格式。
输入:
url(字符串)返回:URL 的 HTML 内容
scrape_url_markdown使用 ScrAPI 服务,通过 URL 抓取网站内容,并以 Markdown 格式获取结果。此功能适用于抓取因机器人检测、验证码甚至地理位置限制而难以访问的网站内容。如果网页的文本内容而非结构信息很重要,则结果将以 Markdown 格式呈现,因此更适合使用 Markdown 格式。
输入:
url(字符串)返回:URL 的 Markdown 内容
Related MCP server: web-scrapper-stdio
设置
API 密钥(可选)
可以选择从ScrAPI 网站获取 API 密钥。
如果没有 API 密钥,您将只能进行一次并发呼叫,每天只能进行二十次免费呼叫,并且排队功能也非常有限。
云服务器
ScrAPI MCP 服务器也可通过 SSE 在云端使用,网址为https://api.scrapi.dev/sse
云 MCP 服务器尚未得到广泛支持,但您可以直接从自定义客户端访问,或使用MCP 检查器进行测试。目前,连接到云 MCP 服务器时无法传递您的 API 密钥。

与 Claude Desktop 一起使用
将以下内容添加到您的claude_desktop_config.json中:
Docker
{
"mcpServers": {
"scrapi": {
"command": "docker",
"args": [
"run",
"-i",
"--rm",
"-e",
"SCRAPI_API_KEY",
"deventerprisesoftware/scrapi-mcp"
],
"env": {
"SCRAPI_API_KEY": "<YOUR_API_KEY>"
}
}
}
}NPX
{
"mcpServers": {
"scrapi": {
"command": "npx",
"args": [
"-y",
"@deventerprisesoftware/scrapi-mcp"
],
"env": {
"SCRAPI_API_KEY": "<YOUR_API_KEY>"
}
}
}
}
建造
Docker 构建:
docker build -t deventerprisesoftware/scrapi-mcp -f Dockerfile .执照
此 MCP 服务器采用 MIT 许可证。这意味着您可以自由使用、修改和分发该软件,但须遵守 MIT 许可证的条款和条件。更多详情,请参阅项目仓库中的 LICENSE 文件。
Available Tools
2 toolsscrape_url_htmlScrape URL and respond with HTMLA
Use a URL to scrape a website using the ScrAPI service and retrieve the result as HTML. Use this for scraping website content that is difficult to access because of bot detection, captchas or even geolocation restrictions. The result will be in HTML which is preferable if advanced parsing is required.
BROWSER COMMANDS: You can optionally provide browser commands to interact with the page before scraping (e.g., clicking buttons, filling forms, scrolling). Provide commands as a JSON array string. Available commands:
Click: {"click": "#buttonId"} - Click an element using CSS selector
Input: {"input": {"input[name='email']": "value"}} - Fill an input field
Select: {"select": {"select[name='country']": "USA"}} - Select from dropdown
Scroll: {"scroll": 1000} - Scroll down (negative values scroll up)
Wait: {"wait": 5000} - Wait milliseconds (max 15000)
WaitFor: {"waitfor": "#elementId"} - Wait for element to appear
JavaScript: {"javascript": "console.log('test')"} - Execute custom JS Example: [{"click": "#accept-cookies"}, {"wait": 2000}, {"input": {"input[name='search']": "query"}}]
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to scrape | |
| browserCommands | No | Optional JSON array of browser commands to execute before scraping. See tool description for available commands and format. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, the description details the ScrAPI service, browser commands with constraints (e.g., max wait time), and interaction capabilities. Does not mention failure modes or rate limits but covers key behaviors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-structured with clear separation of overview and command details. Slightly lengthy due to command examples, but each part adds value and is front-loaded with purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Lacks details on output structure (e.g., response format, error handling) and authentication requirements. For a scraping tool without an output schema, more on return values would be beneficial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description adds extensive semantics for browserCommands (list of commands, parameters, examples), which goes well beyond the schema's brief description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool scrapes a URL and retrieves the result as HTML, explicitly distinguishing it from the sibling tool by mentioning 'advanced parsing' for HTML output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit use cases (scraping content with bot detection, captchas, geolocation) and implies when to prefer HTML over markdown. Lacks explicit 'when not to use' but sufficient context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrape_url_markdownScrape URL and respond with MarkdownA
Use a URL to scrape a website using the ScrAPI service and retrieve the result as Markdown. Use this for scraping website content that is difficult to access because of bot detection, captchas or even geolocation restrictions. The result will be in Markdown which is preferable if the text content of the webpage is important and not the structural information of the page.
BROWSER COMMANDS: You can optionally provide browser commands to interact with the page before scraping (e.g., clicking buttons, filling forms, scrolling). Provide commands as a JSON array string. Available commands:
Click: {"click": "#buttonId"} - Click an element using CSS selector
Input: {"input": {"input[name='email']": "value"}} - Fill an input field
Select: {"select": {"select[name='country']": "USA"}} - Select from dropdown
Scroll: {"scroll": 1000} - Scroll down (negative values scroll up)
Wait: {"wait": 5000} - Wait milliseconds (max 15000)
WaitFor: {"waitfor": "#elementId"} - Wait for element to appear
JavaScript: {"javascript": "console.log('test')"} - Execute custom JS Example: [{"click": "#accept-cookies"}, {"wait": 2000}, {"input": {"input[name='search']": "query"}}]
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to scrape | |
| browserCommands | No | Optional JSON array of browser commands to execute before scraping. See tool description for available commands and format. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses using ScrAPI service, handling bot detection, and details browser command behavior. However, it does not mention failure modes or rate limits, which are relevant for scraping tools.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is somewhat lengthy due to the browser commands section, but it is clearly structured with a heading and bullet-like list. It could be more concise by trimming redundant phrases.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description states result is Markdown but lacks specifics on structure. It covers main usage and browser interaction well. For a scraping tool, this is mostly sufficient, though more detail on output format would help.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds significant value by explaining the browserCommands parameter format with a list of available commands and an example, which is not fully captured in the schema description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it scrapes a URL using ScrAPI and returns Markdown. It distinguishes from sibling tool 'scrape_url_html' by specifying Markdown output and use cases like bot detection, captchas, and geolocation restrictions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description advises when to use: when text content is important and not structural information. It implies when not to use (if structural info is needed, use HTML version) but does not explicitly state alternatives. The context includes a sibling tool, which helps.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.4.0- Changed
scrape_url_html2 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / browserCommandsAdded value: +{ + "description": "Optional JSON array of browser commands to execute before scraping. See tool description for available commands and format.", + "type": "string" +}
- Changed
scrape_url_markdown2 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / browserCommandsAdded value: +{ + "description": "Optional JSON array of browser commands to execute before scraping. See tool description for available commands and format.", + "type": "string" +}
2 tool updates
v1.0.0- Changed
scrape_url_html1 field changed- added
Input schema / properties / url / descriptionAdded value: +"The URL to scrape"
- Changed
scrape_url_markdown1 field changed- added
Input schema / properties / url / descriptionAdded value: +"The URL to scrape"
2 tool updates
- First observed
scrape_url_html - First observed
scrape_url_markdown
TDQS
Scored across 2 tools
The two tools are clearly distinguished by their output format (HTML vs Markdown), leaving no ambiguity about which to use based on desired result type.
Both tools follow a consistent 'scrape_url_{format}' pattern with identical prefix and clear format suffix, making naming predictable.
With exactly 2 tools covering the two primary output formats (HTML and Markdown), the count is minimal but complete for the server's core purpose.
The tools cover the essential use cases of scraping with browser interaction and returning structured output. Minor gaps like raw text or JSON output exist but are not critical.
Maintenance
Related MCP Connectors
Web scraping for AI agents. Converts URLs to clean, LLM-ready Markdown with anti-bot bypass.
Cloud scraping & crawling API for AI agents. Turn any URL into clean, LLM-ready markdown.
Scrape to markdown, map and crawl sites, drive a real browser. Pay per call, no API key, no signup.
Fetch, crawl, and browse protected pages with anti-bot handling - renders in a real browser and
Related MCP Servers
- AlicenseAqualityAmaintenanceThis server enables LLMs to retrieve and process content from web pages, converting HTML to markdown for easier consumption.21156,058 PyPI91,120MIT
- AlicenseNot gradedqualityCmaintenanceA headless web scraping server that extracts main content from web pages into Markdown, text, or HTML for AI and automation integration. It features per-domain rate limiting and robust error handling using Playwright and BeautifulSoup.MIT
- AlicenseAqualityDmaintenanceMCP server for fetching web content with browser fingerprint camouflage, converting HTML to clean Markdown to bypass bot detection.11MIT
- AlicenseNot gradedqualityBmaintenanceRemote MCP server for web scraping with anti-bot evasion. Provides stealth HTTP fetching, headless browser with Cloudflare bypass, CSS selectors, YouTube transcripts, and Markdown conversion.1MIT