Patchright Lite MCP Server
Patchright Lite MCP サーバー
Patchright Node.js SDKをラップし、AIモデルにステルスブラウザ自動化機能を提供する、合理化されたモデルコンテキストプロトコル(MCP)サーバーです。この軽量サーバーは、よりシンプルなAIモデルを容易に利用できるよう、必須機能に特化しています。
Patchright とは何ですか?
Patchrightは、Playwrightテストおよび自動化フレームワークの検出されないバージョンです。Playwrightの代替として設計されていますが、アンチボットシステムによる検出を回避するための高度なステルス機能を備えています。Patchrightは、以下を含む様々な検出手法にパッチを適用します。
Runtime.enable リーク
Console.enable リーク
コマンドフラグの漏洩
一般的な検出ポイント
クローズドシャドウルートインタラクション
この MCP サーバーは、Patchright の Node.js バージョンをラップし、シンプルで標準化されたプロトコルを通じて AI モデルでその機能を利用できるようにします。
Related MCP server: Puppeteer-Extra MCP Server
特徴
シンプルなインターフェース: 4つの必須ツールのみでコア機能に焦点を合わせています
ステルス自動化: Patchrightのステルスモードを使用して検出を回避します
MCP 標準: AI 統合を容易にするモデル コンテキスト プロトコルを実装
Stdioトランスポート: シームレスな統合のために標準入出力を使用します
前提条件
Node.js 18歳以上
npmまたはyarn
インストール
このリポジトリをクローンします:
git clone https://github.com/yourusername/patchright-lite-mcp-server.git cd patchright-lite-mcp-server依存関係をインストールします:
npm installTypeScript コードをビルドします。
npm run build
使用法
次のコマンドでサーバーを実行します。
npm startこれにより、stdio トランスポートを使用してサーバーが起動し、MCP をサポートする AI ツールと統合できるようになります。
AIモデルとの統合
クロードデスクトップ
claude-desktop-config.jsonファイルに以下を追加します。
{
"mcpServers": {
"patchright": {
"command": "node",
"args": ["path/to/patchright-lite-mcp-server/dist/index.js"]
}
}
}GitHub Copilot を使用した VS Code
VS Code CLI を使用して MCP サーバーを追加します。
code --add-mcp '{"name":"patchright","command":"node","args":["path/to/patchright-lite-mcp-server/dist/index.js"]}'利用可能なツール
サーバーは 4 つの必須ツールのみを提供します。
1. 閲覧
ブラウザを起動し、URL に移動してコンテンツを抽出します。
Tool: browse
Parameters: {
"url": "https://example.com",
"headless": true,
"waitFor": 1000
}戻り値:
ページタイトル
表示されるテキストのプレビュー
ブラウザID(後続の操作用)
ページID(後続の操作用)
スクリーンショットのパス
2. 交流する
ページ上で簡単な操作を実行します。
Tool: interact
Parameters: {
"browserId": "browser-id-from-browse",
"pageId": "page-id-from-browse",
"action": "click", // can be "click", "fill", or "select"
"selector": "#submit-button",
"value": "Hello World" // only needed for fill and select
}戻り値:
アクション結果
現在のURL
スクリーンショットのパス
3. 抽出
現在のページから特定のコンテンツを抽出します。
Tool: extract
Parameters: {
"browserId": "browser-id-from-browse",
"pageId": "page-id-from-browse",
"type": "text" // can be "text", "html", or "screenshot"
}戻り値:
要求されたタイプに基づいて抽出されたコンテンツ
4. 閉じる
ブラウザを閉じてリソースを解放します。
Tool: close
Parameters: {
"browserId": "browser-id-from-browse"
}使用フローの例
ブラウザを起動してサイトに移動します。
Tool: browse Parameters: { "url": "https://example.com/login", "headless": false }ログインフォームに記入してください:
Tool: interact Parameters: { "browserId": "browser-id-from-step-1", "pageId": "page-id-from-step-1", "action": "fill", "selector": "#username", "value": "user@example.com" }パスワードを入力してください:
Tool: interact Parameters: { "browserId": "browser-id-from-step-1", "pageId": "page-id-from-step-1", "action": "fill", "selector": "#password", "value": "password123" }ログインボタンをクリックします:
Tool: interact Parameters: { "browserId": "browser-id-from-step-1", "pageId": "page-id-from-step-1", "action": "click", "selector": "#login-button" }ログインを確認するためのテキストを抽出します:
Tool: extract Parameters: { "browserId": "browser-id-from-step-1", "pageId": "page-id-from-step-1", "type": "text" }ブラウザを閉じます:
Tool: close Parameters: { "browserId": "browser-id-from-step-1" }
セキュリティに関する考慮事項
このサーバーは強力な自動化機能を提供します。責任を持って倫理的にご利用ください。
ウェブサイトの利用規約に違反するアクションの自動化は避けてください。
レート制限に留意し、リクエストでウェブサイトに過負荷をかけないようにしてください。
ライセンス
このプロジェクトは MIT ライセンスに基づいてライセンスされています - 詳細については LICENSE ファイルを参照してください。
謝辞
Kaliiiiiiiiii-Vinyzu による Patchright-nodejs
modelcontextprotocol によるモデルコンテキストプロトコル
Dockerの使用
このサーバーは Docker を使用して実行できます。
docker run -it --rm dylangroos/patchright-mcpローカルでDockerイメージを構築する
Docker イメージをビルドします。
docker build -t patchright-mcp .コンテナを実行します。
docker run -it --rm patchright-mcpDockerハブ
変更がメインブランチにマージされると、イメージは自動的にDocker Hubに公開されます。最新のイメージはdylangroos/patchright-mcpで入手できます。
Available Tools
4 toolsbrowseB
Browse to a URL and return the page title and visible text
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to navigate to | |
| headless | No | Whether to run the browser in headless mode | |
| waitFor | No | Time to wait after page load (milliseconds) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the return values (page title and visible text) but lacks critical details such as error handling (e.g., for invalid URLs), performance implications (e.g., timeouts), authentication needs, or rate limits, which are important for a browsing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose and outcome with zero wasted words. It is appropriately sized for the tool's complexity and gets straight to the point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (3 parameters, no output schema, no annotations), the description is minimally adequate. It covers the basic purpose and return values but lacks details on behavioral traits and usage guidelines, leaving gaps that could hinder an AI agent's effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, so the schema already documents all parameters (url, headless, waitFor) thoroughly. The description adds no additional meaning beyond what the schema provides, such as explaining parameter interactions or usage nuances, but this is acceptable given the high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('browse to a URL') and the outcome ('return the page title and visible text'), using specific verbs and resources. It distinguishes itself from sibling tools like 'close', 'extract', and 'interact' by focusing on navigation and content retrieval rather than closing, extraction, or interaction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'extract' or 'interact', nor does it mention any prerequisites, exclusions, or specific contexts for usage. It states what the tool does but not when it's appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
closeC
Close browser to free resources
| Name | Required | Description | Default |
|---|---|---|---|
| browserId | Yes | Browser ID to close |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It states the tool closes a browser to free resources, implying a destructive action that terminates a session, but doesn't disclose behavioral traits like whether it's reversible, requires specific permissions, affects other tools, or has side effects (e.g., losing unsaved data). The mention of 'free resources' adds some context but is minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste, front-loading the key action and purpose. It's appropriately sized for a simple tool with one parameter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and a destructive action implied by 'close', the description is incomplete. It lacks details on what happens after closing (e.g., return values, error conditions), prerequisites, or integration with sibling tools. For a tool that likely terminates a resource, more context is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the parameter 'browserId' fully documented in the schema. The description adds no meaning beyond the schema, as it doesn't explain what a 'browserId' is, how to obtain it, or its format. Baseline 3 is appropriate since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Close') and the resource ('browser'), specifying it's to 'free resources'. It distinguishes from sibling tools like 'browse' (open/access) and 'interact' (use while open), but doesn't explicitly contrast with 'extract' (which might operate on a closed or open browser).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when resources need freeing, but provides no explicit guidance on when to use this tool versus alternatives (e.g., whether to close after 'browse' or 'extract'), prerequisites, or exclusions. It lacks context for decision-making relative to siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extractC
Extract information from the current page as text, html, or screenshot
| Name | Required | Description | Default |
|---|---|---|---|
| browserId | Yes | Browser ID from a previous browse operation | |
| pageId | Yes | Page ID from a previous browse operation | |
| type | Yes | Type of content to extract |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but lacks behavioral details. It doesn't disclose whether extraction is read-only (implied but not stated), if it requires specific permissions, rate limits, or what happens on failure (e.g., invalid IDs). The description only states what it does, not how it behaves.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero wasted words. It front-loads the core action ('extract information') and specifies key details (source: current page; formats: text, html, screenshot). Every element earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 3 required parameters, no annotations, and no output schema, the description is incomplete. It doesn't explain the relationship to 'browse' (source of browserId/pageId), what 'extract' returns (e.g., raw text, file path), or error handling. The agent lacks context for effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are well-documented in the schema. The description adds minimal value by mentioning the 'type' enum options (text, html, screenshot), but doesn't explain semantics beyond what the schema already provides (e.g., what 'text' extraction includes vs. 'html'). Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'extract' and the resource 'information from the current page', specifying the output formats (text, html, or screenshot). It distinguishes from sibling tools like 'browse' (which likely navigates) and 'interact' (which likely performs actions), but doesn't explicitly differentiate from 'close' (which likely terminates a session).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing a browser/page ID from 'browse'), exclusions, or contextual cues for choosing between extraction types. The agent must infer usage from parameter names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
interactC
Perform simple interactions on a page
| Name | Required | Description | Default |
|---|---|---|---|
| browserId | Yes | Browser ID from a previous browse operation | |
| pageId | Yes | Page ID from a previous browse operation | |
| action | Yes | The type of interaction to perform | |
| selector | Yes | CSS selector for the element to interact with | |
| value | No | Value for fill/select actions |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'simple interactions' but doesn't specify what happens (e.g., page changes, errors, side effects), whether it's read-only or mutative, or any constraints like rate limits or authentication needs. This leaves significant gaps for a tool that performs actions on a page.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no wasted words. It's front-loaded and clear in its brevity, though it could benefit from more detail to improve other dimensions. The structure is straightforward and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of performing interactions on a page (likely involving mutations), lack of annotations, and no output schema, the description is incomplete. It doesn't cover behavioral aspects, return values, or usage context, making it inadequate for safe and effective tool invocation by an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds no additional meaning beyond what the schema provides, such as explaining how actions like 'click', 'fill', or 'select' work in context. Baseline 3 is appropriate since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool 'Perform simple interactions on a page', which provides a basic verb+resource combination but lacks specificity. It doesn't clarify what types of interactions beyond the generic term 'simple', nor does it distinguish this tool from potential siblings like 'extract' or 'browse'. The purpose is understandable but vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing a browserId and pageId from a previous browse operation), exclusions, or comparisons to sibling tools like 'extract' or 'close'. Usage is implied through parameter names but not explicitly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
- First observed
browse - First observed
close - First observed
extract - First observed
interact
TDQS
Scored across 4 tools
The tools have mostly distinct purposes with clear boundaries: browse for navigation, extract for content retrieval, interact for actions, and close for cleanup. However, 'extract' and 'interact' could potentially overlap in some use cases (e.g., extracting after an interaction), but their descriptions help differentiate them.
All tool names follow a consistent, simple verb-based pattern (browse, close, extract, interact) without any mixing of conventions. This makes the set predictable and easy to understand at a glance.
With 4 tools, this server is well-scoped for its apparent purpose of web browsing and interaction. Each tool serves a clear, essential function, and there are no extraneous or redundant tools, making the count appropriate for the domain.
The toolset covers the core web browsing lifecycle: navigate (browse), retrieve content (extract), perform actions (interact), and clean up (close). A minor gap is the lack of explicit tools for handling multiple tabs or sessions, but agents can likely work around this with the provided tools.
Maintenance
Related MCP Connectors
A Model Context Protocol server for Wix AI tools
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
Stealth web browser for agents: search, fetch, click, download and type in persistent MCP sessions.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Related MCP Servers
- AlicenseDqualityDmaintenanceAI-driven browser automation server that implements the Model Context Protocol to enable natural language control of web browsers for tasks like navigation, form filling, and visual interaction.12MIT
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that provides enhanced browser automation capabilities using Puppeteer-Extra with Stealth Plugin, enabling LLMs to interact with web pages in a way that better emulates human behavior and avoids detection as automation.3MIT
- AlicenseBqualityDmaintenanceA Model Context Protocol server that enables AI assistants to control a real web browser with stealth capabilities, avoiding bot detection while performing tasks like clicking, filling forms, taking screenshots, and extracting data.1158 npm26MIT
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that enables AI assistants to interact with web pages through browser automation, supporting web scraping, form filling, navigation, and other browser-based tasks using Playwright.1MIT