Puppeteer MCP Server
Puppeteer MCP 服务器
该 MCP 服务器通过 Puppeteer 提供浏览器自动化功能,允许与新的浏览器实例和现有的 Chrome 窗口进行交互。
致谢
该项目是受@modelcontextprotocol/server-puppeteer启发的实验性实现。虽然它们有着相似的目标和概念,但它探索了通过模型上下文协议 (MCP) 实现浏览器自动化的替代方法。
Related MCP server: Puppeteer MCP Server
特征
浏览网页
截取屏幕截图
点击元素
填写表格
选择选项
悬停元素
执行 JavaScript
智能 Chrome 标签管理:
连接到活动的 Chrome 标签页
保留现有的 Chrome 实例
智能连接处理
项目结构
/
├── src/
│ ├── config/ # Configuration modules
│ ├── tools/ # Tool definitions and handlers
│ ├── browser/ # Browser connection management
│ ├── types/ # TypeScript type definitions
│ ├── resources/ # Resource handlers
│ └── server.ts # Server initialization
├── index.ts # Entry point
└── README.md # Documentation安装
选项 1:从 npm 安装
npm install -g puppeteer-mcp-server您也可以使用 npx 直接运行它而无需安装:
npx puppeteer-mcp-server选项 2:从源安装
克隆此存储库或下载源代码
安装依赖项:
npm install构建项目:
npm run build运行服务器:
npm startMCP 服务器配置
要将此工具与 Claude 一起使用,您需要将其添加到您的 MCP 设置配置文件中。
对于克劳德桌面应用程序
将以下内容添加到您的 Claude Desktop 配置文件(位于 Windows 上的%APPDATA%\Claude\claude_desktop_config.json或 macOS 上的~/Library/Application Support/Claude/claude_desktop_config.json ):
如果通过 npm 全局安装:
{
"mcpServers": {
"puppeteer": {
"command": "puppeteer-mcp-server",
"args": [],
"env": {}
}
}
}使用 npx(无需安装):
{
"mcpServers": {
"puppeteer": {
"command": "npx",
"args": ["-y", "puppeteer-mcp-server"],
"env": {}
}
}
}如果从源安装:
{
"mcpServers": {
"puppeteer": {
"command": "node",
"args": ["path/to/puppeteer-mcp-server/dist/index.js"],
"env": {
"NODE_OPTIONS": "--experimental-modules"
}
}
}
}对于 Claude VSCode 扩展
将以下内容添加到您的 Claude VSCode 扩展 MCP 设置文件(位于 Windows 上的%APPDATA%\Code\User\globalStorage\saoudrizwan.claude-dev\settings\cline_mcp_settings.json或 macOS 上的~/Library/Application Support/Code/User/globalStorage/saoudrizwan.claude-dev/settings/cline_mcp_settings.json ):
如果通过 npm 全局安装:
{
"mcpServers": {
"puppeteer": {
"command": "puppeteer-mcp-server",
"args": [],
"env": {}
}
}
}使用 npx(无需安装):
{
"mcpServers": {
"puppeteer": {
"command": "npx",
"args": ["-y", "puppeteer-mcp-server"],
"env": {}
}
}
}如果从源安装:
{
"mcpServers": {
"puppeteer": {
"command": "node",
"args": ["path/to/puppeteer-mcp-server/dist/index.js"],
"env": {
"NODE_OPTIONS": "--experimental-modules"
}
}
}
}对于源安装,将path/to/puppeteer-mcp-server替换为您安装此工具的实际路径。
用法
标准模式
服务器将默认启动一个新的浏览器实例。
活动标签模式
要连接到现有的 Chrome 窗口:
完全关闭所有现有的 Chrome 实例
启动启用远程调试的 Chrome:
# Windows "C:\Program Files\Google\Chrome\Application\chrome.exe" --remote-debugging-port=9222 # macOS /Applications/Google\ Chrome.app/Contents/MacOS/Google\ Chrome --remote-debugging-port=9222 # Linux google-chrome --remote-debugging-port=9222在 Chrome 中导航到您想要的网页
使用
puppeteer_connect_active_tab工具进行连接:{ "targetUrl": "https://example.com", // Optional: specific tab URL "debugPort": 9222 // Optional: defaults to 9222 }
服务器将:
检测并连接到启用远程调试的 Chrome 实例
保留您的 Chrome 实例(不会关闭它)
查找并连接到非扩展选项卡
如果连接失败,提供清晰的错误消息
可用工具
puppeteer_connect_active_tab
连接到启用远程调试的现有 Chrome 实例。
选修的:
targetUrl- 要连接的特定选项卡的 URLdebugPort- Chrome 调试端口(默认值:9222)
puppeteer_navigate
导航到某个 URL。
必需:
url- 要导航到的 URL
puppeteer_screenshot
截取当前页面或特定元素的屏幕截图。
必填:
name- 屏幕截图的名称选修的:
selector- 用于截图元素的 CSS 选择器width- 宽度(以像素为单位)(默认值:800)height- 高度(以像素为单位)(默认值:600)
puppeteer_click
单击页面上的某个元素。
必需:
selector- 要点击元素的 CSS 选择器
puppeteer_fill
填写输入字段。
必需的:
selector- 输入字段的 CSS 选择器value- 要输入的文本
puppeteer_select
使用下拉菜单。
必需的:
selector- 选择元素的 CSS 选择器value- 要选择的选项值
puppeteer_hover
将鼠标悬停在元素上。
必需:
selector- 用于悬停元素的 CSS 选择器
puppeteer_evaluate
在浏览器控制台中执行 JavaScript。
必需:
script- 要执行的 JavaScript 代码
安全注意事项
使用远程调试时:
仅在受信任的网络上启用
使用唯一的调试端口
不使用时关闭调试端口
切勿将调试端口暴露给公共网络
日志记录和调试
基于文件的日志记录
服务器使用 Winston 实现全面的日志记录:
位置:
logs/目录文件模式:
mcp-puppeteer-YYYY-MM-DD.log日志旋转:
每日轮换
最大尺寸:每个文件 20MB
保留:14天
自动压缩旧日志
日志级别
DEBUG:详细的调试信息
信息:一般操作信息
WARN:警告信息
错误:错误事件和异常
记录的信息
服务器启动/关闭事件
浏览器操作(启动、连接、关闭)
导航尝试和结果
工具执行和结果
带有堆栈跟踪的错误详细信息
浏览器控制台输出
资源使用情况(屏幕截图、控制台日志)
错误处理
服务器提供以下详细的错误消息:
连接失败
缺失元素
无效的选择器
JavaScript 执行错误
截图失败
每次工具调用都会返回:
成功/失败状态
如果失败则显示详细的错误信息
操作成功时的结果数据
所有错误也会记录到日志文件中:
时间戳
错误信息
堆栈跟踪(可用时)
上下文信息
贡献
欢迎贡献!请阅读我们的贡献指南,了解如何提交拉取请求、报告问题以及为项目做出贡献。
执照
该项目根据 MIT 许可证获得许可 - 有关详细信息,请参阅LICENSE文件。
Available Tools
8 toolspuppeteer_clickC
Click an element on the page
| Name | Required | Description | Default |
|---|---|---|---|
| selector | Yes | CSS selector for element to click |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is a tautology of the name and adds no behavioral details beyond 'click'. It does not specify side effects (e.g., navigation trigger), waiting behavior, or error conditions. With no annotations, this is insufficient for safe invocation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise (one sentence) but lacks necessary detail. It is front-loaded but does not earn its place by providing substantive information beyond the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, no output schema), the description fails to cover essential context like post-click behavior, return value, or error handling. For a click action, crucial details (e.g., navigation wait) are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already describes the selector parameter (CSS selector for element to click). The description adds no further meaning, but schema coverage is 100%, meeting the baseline. No extra value is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('click') and the target ('an element on the page'), making the purpose obvious. It distinguishes from sibling tools like fill, hover, or navigate, which have different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives (e.g., puppeteer_hover or puppeteer_fill). It does not mention prerequisites or context, leaving the agent to infer usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
puppeteer_connect_active_tabB
Connect to an existing Chrome instance with remote debugging enabled
| Name | Required | Description | Default |
|---|---|---|---|
| targetUrl | No | Optional URL of the target tab to connect to. If not provided, connects to the first available tab. | |
| debugPort | No | Optional Chrome debugging port (default: 9222) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description does not disclose what happens if connection fails, whether state changes, or prerequisites (e.g., Chrome must already be running with remote debugging).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with 10 words, efficiently conveying the core purpose without unnecessary details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description is too brief for a connection tool. It lacks information about return values, success/failure indicators, and prerequisites for using the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema coverage is 100% and already includes parameter descriptions. The tool description adds no additional behavioral details beyond what is in the schema, so baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Connect') and resource ('existing Chrome instance with remote debugging enabled'), distinguishing it from sibling tools that perform page manipulations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, such as launching a new browser. It does not indicate it is a prerequisite for other puppeteer actions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
puppeteer_evaluateB
Execute JavaScript in the browser console
| Name | Required | Description | Default |
|---|---|---|---|
| script | Yes | JavaScript code to execute |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears full responsibility. It mentions execution in the browser console but does not specify whether the script runs synchronously, how return values are handled, or potential side effects on the page state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, direct sentence that immediately conveys the tool's purpose. It is front-loaded and contains no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is minimal given the tool's potential complexity. It does not explain what the execution returns, how to handle asynchronous code, or any security considerations. No output schema exists to compensate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'script' has a clear description in the schema ('JavaScript code to execute'), and the tool description adds no extra meaning. With 100% schema coverage, baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Execute') and the target ('JavaScript in the browser console'), distinguishing it from sibling tools that perform specific UI interactions like clicking, filling, or navigating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives such as puppeteer_click or puppeteer_fill. The description lacks context about typical use cases for executing arbitrary JavaScript.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
puppeteer_fillC
Fill out an input field
| Name | Required | Description | Default |
|---|---|---|---|
| selector | Yes | CSS selector for input field | |
| value | Yes | Value to fill |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description does not disclose behavioral traits such as whether the field is cleared before filling, whether events are triggered, or what happens if the selector is invalid. With no annotations, the description carries the full burden but fails to provide any behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is front-loaded and concise. However, it could be slightly improved by adding more context without losing brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity of the tool (fill input field), the description is minimally adequate. However, the lack of output schema and behavioral details (e.g., does it clear existing values?) leaves gaps for the agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both 'selector' and 'value' having descriptions. The description adds no additional meaning beyond what the schema provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states 'Fill out an input field' which clearly identifies the action (fill) and resource (input field). However, it does not differentiate from sibling tools like puppeteer_select or puppeteer_click, which could also interact with input fields.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like puppeteer_select or puppeteer_click. The description does not mention prerequisites, exclusions, or typical use cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
puppeteer_hoverC
Hover an element on the page
| Name | Required | Description | Default |
|---|---|---|---|
| selector | Yes | CSS selector for element to hover |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It only states the action without disclosing behavior like asynchronous execution, event side effects, or whether the hover is a one-time action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence. It is front-loaded and wastes no words, but could benefit from slightly more detail without losing conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity (1 param, no output schema), the description omits important context like return behavior (likely void) or side effects (e.g., triggering event listeners). Not complete for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (selector described), so baseline is 3. The description adds no new meaning beyond the schema; it merely restates the action.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the specific verb 'Hover' and resource 'an element on the page', clearly distinguishing it from sibling tools like click, fill, or navigate. However, it lacks nuance about what 'hover' entails, such as triggering mouseover events.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like puppeteer_click for interactions or puppeteer_evaluate for custom actions. There is no mention of prerequisites or context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
puppeteer_screenshotB
Take a screenshot of the current page or a specific element
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Name for the screenshot | |
| selector | No | CSS selector for element to screenshot | |
| width | No | Width in pixels (default: 800) | |
| height | No | Height in pixels (default: 600) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It does not disclose what happens to the screenshot (e.g., saved to disk, returned as base64), nor any side effects or permissions needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, which is concise, but it lacks essential context such as output format or usage details. It is under-specified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of output schema and annotations, the description should explain what happens with the screenshot. It fails to mention return value or storage, making it incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the description adds no additional meaning beyond the schema. Baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool takes a screenshot of the current page or a specific element, with a specific verb and resource. It differentiates from sibling tools like puppeteer_navigate or puppeteer_click.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, no preconditions, and no mention of when to capture a full page versus an element.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
puppeteer_selectC
Select an element on the page with Select tag
| Name | Required | Description | Default |
|---|---|---|---|
| selector | Yes | CSS selector for element to select | |
| value | Yes | Value to select |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the burden of behavioral disclosure. It only states the intended action but does not mention side effects (e.g., page state change), error conditions, or if it waits for the element.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence. However, it is so brief that it sacrifices clarity; slightly more detail would be beneficial without losing conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with well-described parameters, the description is adequate but incomplete. It does not explain change events, support for multiple select, or error handling, which would be useful for full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with clear parameter descriptions. The tool description adds the context that it works with a 'Select tag', which clarifies the element type beyond the schema, but this is minimal additional value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Select an element on the page with Select tag', which identifies the target element type (select tag) and action, but it is vague about selecting an option from a dropdown. It does not clearly distinguish from sibling tools like puppeteer_click or puppeteer_fill.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives (e.g., puppeteer_click for generic clicks, puppeteer_fill for input fields). There is no mention of use cases or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
- First observed
puppeteer_click - First observed
puppeteer_connect_active_tab - First observed
puppeteer_evaluate - First observed
puppeteer_fill - First observed
puppeteer_hover - First observed
puppeteer_navigate - First observed
puppeteer_screenshot - First observed
puppeteer_select
TDQS
Scored across 8 tools
Each tool targets a distinct browser action (click, navigate, fill, hover, etc.) with no overlap, making selection unambiguous.
All tools follow a consistent pattern: 'puppeteer_' prefix followed by a clear action verb, ensuring predictability.
With 8 tools covering essential browser interactions, the count is well-scoped for a focused automation server.
Core operations are covered, but missing explicit waiting or content extraction tools; however, puppeteer_evaluate can compensate.
Maintenance
Related MCP Connectors
AI-powered browser automation — navigate, click, fill forms, and extract data from any website.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Provides cloud browser automation capabilities using Stagehand and Browserbase, enabling LLMs to i…
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI coding assistants to control and inspect a live Chrome browser through DevTools for automated testing, performance analysis, debugging, and web scraping. Provides reliable browser automation using Puppeteer with comprehensive DevTools access.2,051,149 npm3Apache 2.0
- AlicenseBqualityDmaintenanceEnables AI agents to automate browser interactions including navigation, content extraction, form filling, screenshots, and JavaScript execution across multiple tabs using Puppeteer.2728 npm1MIT
- AlicenseAqualityDmaintenanceEnables browser automation with concurrent tab pool management using Puppeteer. Supports navigation, content extraction, screenshots, element interaction, and JavaScript execution across multiple browser tabs with auto-recovery and idle timeout features.117 npm1MIT
- AlicenseNot gradedqualityDmaintenanceEnables LLMs to perform browser automation including web navigation, element interaction, and screenshot capture using Puppeteer. It provides capabilities for executing JavaScript in the browser and monitoring console logs for debugging and data extraction.36,709 npmMIT