ScrAPI MCP Server
![]()
ScrAPI MCP 서버
ScrAPI를 사용하여 웹 페이지를 스크래핑하는 MCP 서버입니다.
ScrAPI는 모든 웹사이트에서 손쉽게 데이터를 추출할 수 있는 강력하고 안정적이며 사용하기 쉬운 기능을 제공하는 최고의 웹 스크래핑 솔루션입니다.
도구
scrape_url_htmlScrAPI 서비스를 사용하여 URL을 사용하여 웹사이트를 스크래핑하고 결과를 HTML로 가져옵니다. 봇 탐지, 캡차 또는 지리적 위치 제한으로 인해 접근이 어려운 웹사이트 콘텐츠를 스크래핑하는 데 이 기능을 사용하세요. 결과는 HTML로 제공되며, 고급 파싱이 필요한 경우 더욱 유용합니다.
입력:
url(문자열)반환: URL의 HTML 콘텐츠
scrape_url_markdownScrAPI 서비스를 사용하여 URL을 사용하여 웹사이트를 스크래핑하고 결과를 마크다운 형식으로 가져오세요. 봇 탐지, 캡차 또는 지리적 위치 제한으로 인해 접근이 어려운 웹사이트 콘텐츠를 스크래핑하는 데 이 기능을 사용하세요. 결과는 마크다운 형식으로 제공되며, 웹페이지의 텍스트 콘텐츠가 중요하고 페이지의 구조적 정보가 중요하지 않은 경우 더욱 효과적입니다.
입력:
url(문자열)반환: URL의 마크다운 콘텐츠
Related MCP server: web-scrapper-stdio
설정
API 키(선택 사항)
선택적으로 ScrAPI 웹사이트 에서 API 키를 받으세요.
API 키가 없으면 동시 통화는 1회로 제한되고, 대기 기능은 최소화되어 하루 무료 통화는 20회로 제한됩니다.
클라우드 서버
ScrAPI MCP 서버는 https://api.scrapi.dev/sse 에서 SSE를 통해 클라우드에서도 사용할 수 있습니다.
클라우드 MCP 서버는 아직 널리 지원되지 않지만, 사용자 지정 클라이언트에서 직접 액세스하거나 MCP Inspector를 사용하여 테스트할 수 있습니다. 현재 클라우드 MCP 서버에 연결할 때 API 키를 전달하는 기능은 없습니다.

Claude Desktop과 함께 사용
claude_desktop_config.json 에 다음을 추가하세요.
도커
지엑스피1
엔피엑스
{
"mcpServers": {
"scrapi": {
"command": "npx",
"args": [
"-y",
"@deventerprisesoftware/scrapi-mcp"
],
"env": {
"SCRAPI_API_KEY": "<YOUR_API_KEY>"
}
}
}
}
짓다
Docker 빌드:
docker build -t deventerprisesoftware/scrapi-mcp -f Dockerfile .특허
이 MCP 서버는 MIT 라이선스에 따라 라이선스가 부여됩니다. 즉, MIT 라이선스의 약관에 따라 소프트웨어를 자유롭게 사용, 수정 및 배포할 수 있습니다. 자세한 내용은 프로젝트 저장소의 LICENSE 파일을 참조하세요.
Available Tools
2 toolsscrape_url_htmlScrape URL and respond with HTMLA
Use a URL to scrape a website using the ScrAPI service and retrieve the result as HTML. Use this for scraping website content that is difficult to access because of bot detection, captchas or even geolocation restrictions. The result will be in HTML which is preferable if advanced parsing is required.
BROWSER COMMANDS: You can optionally provide browser commands to interact with the page before scraping (e.g., clicking buttons, filling forms, scrolling). Provide commands as a JSON array string. Available commands:
Click: {"click": "#buttonId"} - Click an element using CSS selector
Input: {"input": {"input[name='email']": "value"}} - Fill an input field
Select: {"select": {"select[name='country']": "USA"}} - Select from dropdown
Scroll: {"scroll": 1000} - Scroll down (negative values scroll up)
Wait: {"wait": 5000} - Wait milliseconds (max 15000)
WaitFor: {"waitfor": "#elementId"} - Wait for element to appear
JavaScript: {"javascript": "console.log('test')"} - Execute custom JS Example: [{"click": "#accept-cookies"}, {"wait": 2000}, {"input": {"input[name='search']": "query"}}]
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to scrape | |
| browserCommands | No | Optional JSON array of browser commands to execute before scraping. See tool description for available commands and format. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, the description details the ScrAPI service, browser commands with constraints (e.g., max wait time), and interaction capabilities. Does not mention failure modes or rate limits but covers key behaviors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-structured with clear separation of overview and command details. Slightly lengthy due to command examples, but each part adds value and is front-loaded with purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Lacks details on output structure (e.g., response format, error handling) and authentication requirements. For a scraping tool without an output schema, more on return values would be beneficial.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description adds extensive semantics for browserCommands (list of commands, parameters, examples), which goes well beyond the schema's brief description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool scrapes a URL and retrieves the result as HTML, explicitly distinguishing it from the sibling tool by mentioning 'advanced parsing' for HTML output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit use cases (scraping content with bot detection, captchas, geolocation) and implies when to prefer HTML over markdown. Lacks explicit 'when not to use' but sufficient context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrape_url_markdownScrape URL and respond with MarkdownA
Use a URL to scrape a website using the ScrAPI service and retrieve the result as Markdown. Use this for scraping website content that is difficult to access because of bot detection, captchas or even geolocation restrictions. The result will be in Markdown which is preferable if the text content of the webpage is important and not the structural information of the page.
BROWSER COMMANDS: You can optionally provide browser commands to interact with the page before scraping (e.g., clicking buttons, filling forms, scrolling). Provide commands as a JSON array string. Available commands:
Click: {"click": "#buttonId"} - Click an element using CSS selector
Input: {"input": {"input[name='email']": "value"}} - Fill an input field
Select: {"select": {"select[name='country']": "USA"}} - Select from dropdown
Scroll: {"scroll": 1000} - Scroll down (negative values scroll up)
Wait: {"wait": 5000} - Wait milliseconds (max 15000)
WaitFor: {"waitfor": "#elementId"} - Wait for element to appear
JavaScript: {"javascript": "console.log('test')"} - Execute custom JS Example: [{"click": "#accept-cookies"}, {"wait": 2000}, {"input": {"input[name='search']": "query"}}]
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to scrape | |
| browserCommands | No | Optional JSON array of browser commands to execute before scraping. See tool description for available commands and format. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses using ScrAPI service, handling bot detection, and details browser command behavior. However, it does not mention failure modes or rate limits, which are relevant for scraping tools.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is somewhat lengthy due to the browser commands section, but it is clearly structured with a heading and bullet-like list. It could be more concise by trimming redundant phrases.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description states result is Markdown but lacks specifics on structure. It covers main usage and browser interaction well. For a scraping tool, this is mostly sufficient, though more detail on output format would help.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description adds significant value by explaining the browserCommands parameter format with a list of available commands and an example, which is not fully captured in the schema description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it scrapes a URL using ScrAPI and returns Markdown. It distinguishes from sibling tool 'scrape_url_html' by specifying Markdown output and use cases like bot detection, captchas, and geolocation restrictions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description advises when to use: when text content is important and not structural information. It implies when not to use (if structural info is needed, use HTML version) but does not explicitly state alternatives. The context includes a sibling tool, which helps.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.4.0- Changed
scrape_url_html2 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / browserCommandsAdded value: +{ + "description": "Optional JSON array of browser commands to execute before scraping. See tool description for available commands and format.", + "type": "string" +}
- Changed
scrape_url_markdown2 fields changed- removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / browserCommandsAdded value: +{ + "description": "Optional JSON array of browser commands to execute before scraping. See tool description for available commands and format.", + "type": "string" +}
2 tool updates
v1.0.0- Changed
scrape_url_html1 field changed- added
Input schema / properties / url / descriptionAdded value: +"The URL to scrape"
- Changed
scrape_url_markdown1 field changed- added
Input schema / properties / url / descriptionAdded value: +"The URL to scrape"
2 tool updates
- First observed
scrape_url_html - First observed
scrape_url_markdown
TDQS
Scored across 2 tools
The two tools are clearly distinguished by their output format (HTML vs Markdown), leaving no ambiguity about which to use based on desired result type.
Both tools follow a consistent 'scrape_url_{format}' pattern with identical prefix and clear format suffix, making naming predictable.
With exactly 2 tools covering the two primary output formats (HTML and Markdown), the count is minimal but complete for the server's core purpose.
The tools cover the essential use cases of scraping with browser interaction and returning structured output. Minor gaps like raw text or JSON output exist but are not critical.
Maintenance
Related MCP Connectors
Web scraping for AI agents. Converts URLs to clean, LLM-ready Markdown with anti-bot bypass.
Cloud scraping & crawling API for AI agents. Turn any URL into clean, LLM-ready markdown.
Scrape to markdown, map and crawl sites, drive a real browser. Pay per call, no API key, no signup.
Fetch, crawl, and browse protected pages with anti-bot handling - renders in a real browser and
Related MCP Servers
- AlicenseAqualityAmaintenanceThis server enables LLMs to retrieve and process content from web pages, converting HTML to markdown for easier consumption.21156,058 PyPI91,120MIT
- AlicenseNot gradedqualityCmaintenanceA headless web scraping server that extracts main content from web pages into Markdown, text, or HTML for AI and automation integration. It features per-domain rate limiting and robust error handling using Playwright and BeautifulSoup.MIT
- AlicenseAqualityDmaintenanceMCP server for fetching web content with browser fingerprint camouflage, converting HTML to clean Markdown to bypass bot detection.11MIT
- AlicenseNot gradedqualityBmaintenanceRemote MCP server for web scraping with anti-bot evasion. Provides stealth HTTP fetching, headless browser with Cloudflare bypass, CSS selectors, YouTube transcripts, and Markdown conversion.1MIT