Skip to main content
Glama
DevEnterpriseSoftware

ScrAPI MCP Server

ScrAPI 로고

ScrAPI MCP 서버

라이센스: MIT NPM 다운로드 도커 풀

ScrAPI를 사용하여 웹 페이지를 스크래핑하는 MCP 서버입니다.

ScrAPI는 모든 웹사이트에서 손쉽게 데이터를 추출할 수 있는 강력하고 안정적이며 사용하기 쉬운 기능을 제공하는 최고의 웹 스크래핑 솔루션입니다.

도구

  1. scrape_url_html

    • ScrAPI 서비스를 사용하여 URL을 사용하여 웹사이트를 스크래핑하고 결과를 HTML로 가져옵니다. 봇 탐지, 캡차 또는 지리적 위치 제한으로 인해 접근이 어려운 웹사이트 콘텐츠를 스크래핑하는 데 이 기능을 사용하세요. 결과는 HTML로 제공되며, 고급 파싱이 필요한 경우 더욱 유용합니다.

    • 입력: url (문자열)

    • 반환: URL의 HTML 콘텐츠

  2. scrape_url_markdown

    • ScrAPI 서비스를 사용하여 URL을 사용하여 웹사이트를 스크래핑하고 결과를 마크다운 형식으로 가져오세요. 봇 탐지, 캡차 또는 지리적 위치 제한으로 인해 접근이 어려운 웹사이트 콘텐츠를 스크래핑하는 데 이 기능을 사용하세요. 결과는 마크다운 형식으로 제공되며, 웹페이지의 텍스트 콘텐츠가 중요하고 페이지의 구조적 정보가 중요하지 않은 경우 더욱 효과적입니다.

    • 입력: url (문자열)

    • 반환: URL의 마크다운 콘텐츠

Related MCP server: web-scrapper-stdio

설정

API 키(선택 사항)

선택적으로 ScrAPI 웹사이트 에서 API 키를 받으세요.

API 키가 없으면 동시 통화는 1회로 제한되고, 대기 기능은 최소화되어 하루 무료 통화는 20회로 제한됩니다.

클라우드 서버

ScrAPI MCP 서버는 https://api.scrapi.dev/sse 에서 SSE를 통해 클라우드에서도 사용할 수 있습니다.

클라우드 MCP 서버는 아직 널리 지원되지 않지만, 사용자 지정 클라이언트에서 직접 액세스하거나 MCP Inspector를 사용하여 테스트할 수 있습니다. 현재 클라우드 MCP 서버에 연결할 때 API 키를 전달하는 기능은 없습니다.

MCP-검사관

Claude Desktop과 함께 사용

claude_desktop_config.json 에 다음을 추가하세요.

도커

지엑스피1

엔피엑스

{
  "mcpServers": {
    "scrapi": {
      "command": "npx",
      "args": [
        "-y",
        "@deventerprisesoftware/scrapi-mcp"
      ],
      "env": {
        "SCRAPI_API_KEY": "<YOUR_API_KEY>"
      }
    }
  }
}

클로드-데스크탑

짓다

Docker 빌드:

docker build -t deventerprisesoftware/scrapi-mcp -f Dockerfile .

특허

이 MCP 서버는 MIT 라이선스에 따라 라이선스가 부여됩니다. 즉, MIT 라이선스의 약관에 따라 소프트웨어를 자유롭게 사용, 수정 및 배포할 수 있습니다. 자세한 내용은 프로젝트 저장소의 LICENSE 파일을 참조하세요.

Available Tools

2 tools
scrape_url_htmlScrape URL and respond with HTMLA

Use a URL to scrape a website using the ScrAPI service and retrieve the result as HTML. Use this for scraping website content that is difficult to access because of bot detection, captchas or even geolocation restrictions. The result will be in HTML which is preferable if advanced parsing is required.

BROWSER COMMANDS: You can optionally provide browser commands to interact with the page before scraping (e.g., clicking buttons, filling forms, scrolling). Provide commands as a JSON array string. Available commands:

  • Click: {"click": "#buttonId"} - Click an element using CSS selector

  • Input: {"input": {"input[name='email']": "value"}} - Fill an input field

  • Select: {"select": {"select[name='country']": "USA"}} - Select from dropdown

  • Scroll: {"scroll": 1000} - Scroll down (negative values scroll up)

  • Wait: {"wait": 5000} - Wait milliseconds (max 15000)

  • WaitFor: {"waitfor": "#elementId"} - Wait for element to appear

  • JavaScript: {"javascript": "console.log('test')"} - Execute custom JS Example: [{"click": "#accept-cookies"}, {"wait": 2000}, {"input": {"input[name='search']": "query"}}]

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to scrape
browserCommandsNoOptional JSON array of browser commands to execute before scraping. See tool description for available commands and format.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite no annotations, the description details the ScrAPI service, browser commands with constraints (e.g., max wait time), and interaction capabilities. Does not mention failure modes or rate limits but covers key behaviors.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-structured with clear separation of overview and command details. Slightly lengthy due to command examples, but each part adds value and is front-loaded with purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Lacks details on output structure (e.g., response format, error handling) and authentication requirements. For a scraping tool without an output schema, more on return values would be beneficial.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the description adds extensive semantics for browserCommands (list of commands, parameters, examples), which goes well beyond the schema's brief description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool scrapes a URL and retrieves the result as HTML, explicitly distinguishing it from the sibling tool by mentioning 'advanced parsing' for HTML output.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit use cases (scraping content with bot detection, captchas, geolocation) and implies when to prefer HTML over markdown. Lacks explicit 'when not to use' but sufficient context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scrape_url_markdownScrape URL and respond with MarkdownA

Use a URL to scrape a website using the ScrAPI service and retrieve the result as Markdown. Use this for scraping website content that is difficult to access because of bot detection, captchas or even geolocation restrictions. The result will be in Markdown which is preferable if the text content of the webpage is important and not the structural information of the page.

BROWSER COMMANDS: You can optionally provide browser commands to interact with the page before scraping (e.g., clicking buttons, filling forms, scrolling). Provide commands as a JSON array string. Available commands:

  • Click: {"click": "#buttonId"} - Click an element using CSS selector

  • Input: {"input": {"input[name='email']": "value"}} - Fill an input field

  • Select: {"select": {"select[name='country']": "USA"}} - Select from dropdown

  • Scroll: {"scroll": 1000} - Scroll down (negative values scroll up)

  • Wait: {"wait": 5000} - Wait milliseconds (max 15000)

  • WaitFor: {"waitfor": "#elementId"} - Wait for element to appear

  • JavaScript: {"javascript": "console.log('test')"} - Execute custom JS Example: [{"click": "#accept-cookies"}, {"wait": 2000}, {"input": {"input[name='search']": "query"}}]

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to scrape
browserCommandsNoOptional JSON array of browser commands to execute before scraping. See tool description for available commands and format.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses using ScrAPI service, handling bot detection, and details browser command behavior. However, it does not mention failure modes or rate limits, which are relevant for scraping tools.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is somewhat lengthy due to the browser commands section, but it is clearly structured with a heading and bullet-like list. It could be more concise by trimming redundant phrases.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description states result is Markdown but lacks specifics on structure. It covers main usage and browser interaction well. For a scraping tool, this is mostly sufficient, though more detail on output format would help.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds significant value by explaining the browserCommands parameter format with a list of available commands and an example, which is not fully captured in the schema description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it scrapes a URL using ScrAPI and returns Markdown. It distinguishes from sibling tool 'scrape_url_html' by specifying Markdown output and use cases like bot detection, captchas, and geolocation restrictions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description advises when to use: when text content is important and not structural information. It implies when not to use (if structural info is needed, use HTML version) but does not explicitly state alternatives. The context includes a sibling tool, which helps.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.4.0
    • Changedscrape_url_html2 fields changed
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / browserCommands
        Added value: +{
        +  "description": "Optional JSON array of browser commands to execute before scraping. See tool description for available commands and format.",
        +  "type": "string"
        +}
    • Changedscrape_url_markdown2 fields changed
      • removedInput schema / additionalProperties
        Removed value: -false
      • addedInput schema / properties / browserCommands
        Added value: +{
        +  "description": "Optional JSON array of browser commands to execute before scraping. See tool description for available commands and format.",
        +  "type": "string"
        +}
  2. 2 tool updatesv1.0.0
    • Changedscrape_url_html1 field changed
      • addedInput schema / properties / url / description
        Added value: +"The URL to scrape"
    • Changedscrape_url_markdown1 field changed
      • addedInput schema / properties / url / description
        Added value: +"The URL to scrape"
  3. 2 tool updates
    • First observedscrape_url_html
    • First observedscrape_url_markdown

TDQS

A4.4/5.0

Scored across 2 tools

Disambiguation5/5

The two tools are clearly distinguished by their output format (HTML vs Markdown), leaving no ambiguity about which to use based on desired result type.

Naming Consistency5/5

Both tools follow a consistent 'scrape_url_{format}' pattern with identical prefix and clear format suffix, making naming predictable.

Tool Count5/5

With exactly 2 tools covering the two primary output formats (HTML and Markdown), the count is minimal but complete for the server's core purpose.

Completeness4/5

The tools cover the essential use cases of scraping with browser interaction and returning structured output. Minor gaps like raw text or JSON output exist but are not critical.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    This server enables LLMs to retrieve and process content from web pages, converting HTML to markdown for easier consumption.
    2
    1
    156,058 PyPI
    91,120
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    A headless web scraping server that extracts main content from web pages into Markdown, text, or HTML for AI and automation integration. It features per-domain rate limiting and robust error handling using Playwright and BeautifulSoup.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Remote MCP server for web scraping with anti-bot evasion. Provides stealth HTTP fetching, headless browser with Cloudflare bypass, CSS selectors, YouTube transcripts, and Markdown conversion.
    1
    MIT