Skip to main content
Glama
dylangroos

Patchright Lite MCP Server

by dylangroos

Patchright Lite MCP 서버

Patchright Node.js SDK를 래핑하여 AI 모델에 은밀한 브라우저 자동화 기능을 제공하는 간소화된 모델 컨텍스트 프로토콜(MCP) 서버입니다. 이 가벼운 서버는 간단한 AI 모델의 사용을 용이하게 하는 필수 기능에 중점을 둡니다.

Patchright란 무엇인가요?

Patchright는 Playwright 테스트 및 자동화 프레임워크의 탐지되지 않는 버전입니다. Playwright를 즉시 대체하도록 설계되었지만, 안티 봇 시스템의 탐지를 피할 수 있는 고급 스텔스 기능을 갖추고 있습니다. Patchright는 다음을 포함한 다양한 탐지 기술을 패치합니다.

  • 런타임.enable 누수

  • Console.enable 누수

  • 명령 플래그 누출

  • 일반 감지 지점

  • 닫힌 Shadow Root 상호 작용

이 MCP 서버는 Patchright의 Node.js 버전을 래핑하여 간단하고 표준화된 프로토콜을 통해 AI 모델에서 해당 기능을 사용할 수 있도록 합니다.

Related MCP server: Puppeteer-Extra MCP Server

특징

  • 간단한 인터페이스 : 4가지 필수 도구만으로 핵심 기능에 집중

  • 스텔스 자동화 : Patchright의 스텔스 모드를 사용하여 감지를 방지합니다.

  • MCP 표준 : AI 통합을 용이하게 하기 위한 모델 컨텍스트 프로토콜 구현

  • Stdio Transport : 원활한 통합을 위해 표준 입력/출력을 사용합니다.

필수 조건

  • 노드.js 18+

  • npm 또는 yarn

설치

  1. 이 저장소를 복제하세요:

    지엑스피1

  2. 종속성 설치:

    npm install
  3. TypeScript 코드를 작성합니다.

    npm run build

용법

다음을 사용하여 서버를 실행합니다.

npm start

이렇게 하면 stdio 전송으로 서버가 시작되어 MCP를 지원하는 AI 도구와 통합할 준비가 됩니다.

AI 모델과 통합

클로드 데스크탑

claude-desktop-config.json 파일에 다음을 추가하세요.

{
  "mcpServers": {
    "patchright": {
      "command": "node",
      "args": ["path/to/patchright-lite-mcp-server/dist/index.js"]
    }
  }
}

GitHub Copilot을 사용한 VS 코드

VS Code CLI를 사용하여 MCP 서버를 추가합니다.

code --add-mcp '{"name":"patchright","command":"node","args":["path/to/patchright-lite-mcp-server/dist/index.js"]}'

사용 가능한 도구

서버는 4가지 필수 도구만 제공합니다.

1. 찾아보기

브라우저를 실행하고 URL로 이동하여 콘텐츠를 추출합니다.

Tool: browse
Parameters: {
  "url": "https://example.com",
  "headless": true,
  "waitFor": 1000
}

보고:

  • 페이지 제목

  • 보이는 텍스트 미리보기

  • 브라우저 ID(후속 작업용)

  • 페이지 ID(후속 작업용)

  • 스크린샷 경로

2. 상호 작용하다

페이지에서 간단한 상호작용을 수행합니다.

Tool: interact
Parameters: {
  "browserId": "browser-id-from-browse",
  "pageId": "page-id-from-browse",
  "action": "click", // can be "click", "fill", or "select"
  "selector": "#submit-button",
  "value": "Hello World" // only needed for fill and select
}

보고:

  • 작업 결과

  • 현재 URL

  • 스크린샷 경로

3. 추출하다

현재 페이지에서 특정 콘텐츠를 추출합니다.

Tool: extract
Parameters: {
  "browserId": "browser-id-from-browse",
  "pageId": "page-id-from-browse",
  "type": "text" // can be "text", "html", or "screenshot"
}

보고:

  • 요청된 유형에 따라 추출된 콘텐츠

4. 닫다

리소스를 확보하기 위해 브라우저를 닫습니다.

Tool: close
Parameters: {
  "browserId": "browser-id-from-browse"
}

사용 흐름 예시

  1. 브라우저를 실행하고 사이트로 이동합니다.

    Tool: browse
    Parameters: {
      "url": "https://example.com/login",
      "headless": false
    }
  2. 로그인 양식을 작성하세요:

    Tool: interact
    Parameters: {
      "browserId": "browser-id-from-step-1",
      "pageId": "page-id-from-step-1",
      "action": "fill",
      "selector": "#username",
      "value": "user@example.com"
    }
  3. 비밀번호를 입력하세요:

    Tool: interact
    Parameters: {
      "browserId": "browser-id-from-step-1",
      "pageId": "page-id-from-step-1",
      "action": "fill",
      "selector": "#password",
      "value": "password123"
    }
  4. 로그인 버튼을 클릭하세요:

    Tool: interact
    Parameters: {
      "browserId": "browser-id-from-step-1",
      "pageId": "page-id-from-step-1",
      "action": "click",
      "selector": "#login-button"
    }
  5. 로그인 확인을 위한 텍스트 추출:

    Tool: extract
    Parameters: {
      "browserId": "browser-id-from-step-1",
      "pageId": "page-id-from-step-1",
      "type": "text"
    }
  6. 브라우저를 닫습니다:

    Tool: close
    Parameters: {
      "browserId": "browser-id-from-step-1"
    }

보안 고려 사항

  • 이 서버는 강력한 자동화 기능을 제공합니다. 책임감 있고 윤리적으로 사용하세요.

  • 웹사이트의 서비스 약관을 위반하는 작업은 자동화하지 마세요.

  • 속도 제한을 염두에 두고 웹사이트에 과도한 요청을 보내지 마세요.

특허

이 프로젝트는 MIT 라이선스에 따라 라이선스가 부여되었습니다. 자세한 내용은 라이선스 파일을 참조하세요.

감사의 말

  • Kaliiiiiiiii-Vinyzu의 Patchright-nodejs

  • modelcontextprotocol의 모델 컨텍스트 프로토콜

Docker 사용법

Docker를 사용하여 이 서버를 실행할 수 있습니다.

docker run -it --rm dylangroos/patchright-mcp

Docker 이미지를 로컬로 빌드하기

Docker 이미지를 빌드합니다.

docker build -t patchright-mcp .

컨테이너를 실행합니다.

docker run -it --rm patchright-mcp

도커 허브

변경 사항이 메인 브랜치에 병합되면 이미지가 Docker Hub에 자동으로 게시됩니다. 최신 이미지는 dylangroos/patchright-mcp 에서 확인할 수 있습니다.

Available Tools

4 tools
browseB

Browse to a URL and return the page title and visible text

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to navigate to
headlessNoWhether to run the browser in headless mode
waitForNoTime to wait after page load (milliseconds)

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the return values (page title and visible text) but lacks critical details such as error handling (e.g., for invalid URLs), performance implications (e.g., timeouts), authentication needs, or rate limits, which are important for a browsing tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the core purpose and outcome with zero wasted words. It is appropriately sized for the tool's complexity and gets straight to the point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (3 parameters, no output schema, no annotations), the description is minimally adequate. It covers the basic purpose and return values but lacks details on behavioral traits and usage guidelines, leaving gaps that could hinder an AI agent's effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, so the schema already documents all parameters (url, headless, waitFor) thoroughly. The description adds no additional meaning beyond what the schema provides, such as explaining parameter interactions or usage nuances, but this is acceptable given the high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('browse to a URL') and the outcome ('return the page title and visible text'), using specific verbs and resources. It distinguishes itself from sibling tools like 'close', 'extract', and 'interact' by focusing on navigation and content retrieval rather than closing, extraction, or interaction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'extract' or 'interact', nor does it mention any prerequisites, exclusions, or specific contexts for usage. It states what the tool does but not when it's appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

closeC

Close browser to free resources

ParametersJSON Schema
NameRequiredDescriptionDefault
browserIdYesBrowser ID to close

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It states the tool closes a browser to free resources, implying a destructive action that terminates a session, but doesn't disclose behavioral traits like whether it's reversible, requires specific permissions, affects other tools, or has side effects (e.g., losing unsaved data). The mention of 'free resources' adds some context but is minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero waste, front-loading the key action and purpose. It's appropriately sized for a simple tool with one parameter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and a destructive action implied by 'close', the description is incomplete. It lacks details on what happens after closing (e.g., return values, error conditions), prerequisites, or integration with sibling tools. For a tool that likely terminates a resource, more context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the parameter 'browserId' fully documented in the schema. The description adds no meaning beyond the schema, as it doesn't explain what a 'browserId' is, how to obtain it, or its format. Baseline 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Close') and the resource ('browser'), specifying it's to 'free resources'. It distinguishes from sibling tools like 'browse' (open/access) and 'interact' (use while open), but doesn't explicitly contrast with 'extract' (which might operate on a closed or open browser).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when resources need freeing, but provides no explicit guidance on when to use this tool versus alternatives (e.g., whether to close after 'browse' or 'extract'), prerequisites, or exclusions. It lacks context for decision-making relative to siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extractC

Extract information from the current page as text, html, or screenshot

ParametersJSON Schema
NameRequiredDescriptionDefault
browserIdYesBrowser ID from a previous browse operation
pageIdYesPage ID from a previous browse operation
typeYesType of content to extract

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but lacks behavioral details. It doesn't disclose whether extraction is read-only (implied but not stated), if it requires specific permissions, rate limits, or what happens on failure (e.g., invalid IDs). The description only states what it does, not how it behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero wasted words. It front-loads the core action ('extract information') and specifies key details (source: current page; formats: text, html, screenshot). Every element earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 3 required parameters, no annotations, and no output schema, the description is incomplete. It doesn't explain the relationship to 'browse' (source of browserId/pageId), what 'extract' returns (e.g., raw text, file path), or error handling. The agent lacks context for effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so parameters are well-documented in the schema. The description adds minimal value by mentioning the 'type' enum options (text, html, screenshot), but doesn't explain semantics beyond what the schema already provides (e.g., what 'text' extraction includes vs. 'html'). Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'extract' and the resource 'information from the current page', specifying the output formats (text, html, or screenshot). It distinguishes from sibling tools like 'browse' (which likely navigates) and 'interact' (which likely performs actions), but doesn't explicitly differentiate from 'close' (which likely terminates a session).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing a browser/page ID from 'browse'), exclusions, or contextual cues for choosing between extraction types. The agent must infer usage from parameter names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

interactC

Perform simple interactions on a page

ParametersJSON Schema
NameRequiredDescriptionDefault
browserIdYesBrowser ID from a previous browse operation
pageIdYesPage ID from a previous browse operation
actionYesThe type of interaction to perform
selectorYesCSS selector for the element to interact with
valueNoValue for fill/select actions

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'simple interactions' but doesn't specify what happens (e.g., page changes, errors, side effects), whether it's read-only or mutative, or any constraints like rate limits or authentication needs. This leaves significant gaps for a tool that performs actions on a page.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no wasted words. It's front-loaded and clear in its brevity, though it could benefit from more detail to improve other dimensions. The structure is straightforward and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of performing interactions on a page (likely involving mutations), lack of annotations, and no output schema, the description is incomplete. It doesn't cover behavioral aspects, return values, or usage context, making it inadequate for safe and effective tool invocation by an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds no additional meaning beyond what the schema provides, such as explaining how actions like 'click', 'fill', or 'select' work in context. Baseline 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool 'Perform simple interactions on a page', which provides a basic verb+resource combination but lacks specificity. It doesn't clarify what types of interactions beyond the generic term 'simple', nor does it distinguish this tool from potential siblings like 'extract' or 'browse'. The purpose is understandable but vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing a browserId and pageId from a previous browse operation), exclusions, or comparisons to sibling tools like 'extract' or 'close'. Usage is implied through parameter names but not explicitly stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updates
    • First observedbrowse
    • First observedclose
    • First observedextract
    • First observedinteract

TDQS

B3.3/5.0

Scored across 4 tools

Disambiguation4/5

The tools have mostly distinct purposes with clear boundaries: browse for navigation, extract for content retrieval, interact for actions, and close for cleanup. However, 'extract' and 'interact' could potentially overlap in some use cases (e.g., extracting after an interaction), but their descriptions help differentiate them.

Naming Consistency5/5

All tool names follow a consistent, simple verb-based pattern (browse, close, extract, interact) without any mixing of conventions. This makes the set predictable and easy to understand at a glance.

Tool Count5/5

With 4 tools, this server is well-scoped for its apparent purpose of web browsing and interaction. Each tool serves a clear, essential function, and there are no extraneous or redundant tools, making the count appropriate for the domain.

Completeness4/5

The toolset covers the core web browsing lifecycle: navigate (browse), retrieve content (extract), perform actions (interact), and clean up (close). A minor gap is the lack of explicit tools for handling multiple tabs or sessions, but agents can likely work around this with the provided tools.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    D
    quality
    D
    maintenance
    AI-driven browser automation server that implements the Model Context Protocol to enable natural language control of web browsers for tasks like navigation, form filling, and visual interaction.
    1
    2
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    A Model Context Protocol server that provides enhanced browser automation capabilities using Puppeteer-Extra with Stealth Plugin, enabling LLMs to interact with web pages in a way that better emulates human behavior and avoids detection as automation.
    3
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    A Model Context Protocol server that enables AI assistants to interact with web pages through browser automation, supporting web scraping, form filling, navigation, and other browser-based tasks using Playwright.
    1
    MIT