Skip to main content
Glama

MCP 浏览器代理

铁匠徽章

特征

  • 高级浏览器自动化

    • 使用可自定义的加载策略导航到任何 URL

    • 捕获整页或特定元素的屏幕截图

    • 执行精确的 DOM 交互(单击、填充、选择、悬停)

    • 使用控制台日志捕获在浏览器上下文中执行任意 JavaScript

  • 强大的API客户端

    • 执行 HTTP 请求(GET、POST、PUT、PATCH、DELETE)

    • 配置请求标头和正文内容

    • 使用 JSON 格式处理响应数据

    • 带有详细反馈的错误处理

  • MCP资源管理

    • 将浏览器控制台日志作为资源进行访问

    • 通过MCP资源接口获取截图

    • 与 headful 浏览器实例的持久会话

  • AI代理功能

    • 链接多个浏览器操作以完成复杂任务

    • 遵循多步骤指令,智能错误恢复

    • 通过自然语言指令实现技术任务自动化

Related MCP server: Browserbeam MCP Server

演示

点击任意时间戳即可跳转到视频的相应部分

00:00 - Google 搜索 MCP
导航到 Google 主页并搜索“模型上下文协议”。演示如何使用 MCP 集成执行基本的网页搜索并处理结果。

00:33 -截图
使用自定义文件名截取搜索结果的屏幕截图,并在 Finder 中展示。展示了 Claude 如何在浏览器自动化过程中捕获并保存网页中的视觉内容。

01:00 -维基百科搜索
导航到 Wikipedia.org 并搜索“模型上下文协议”。演示了 Claude 通过 MCP 集成与不同网站及其搜索功能进行交互的能力。

01:38 -下拉菜单交互 I
导航到测试网站 (the-internet.herokuapp.com/dropdown) 并从下拉菜单中选择“选项 1”。演示了 Claude 与表单元素交互并进行选择的能力。

01:56 -下拉菜单交互 II
从同一个下拉菜单中将选项更改为“选项 2”。这展现了 Claude 能够多次操作同一表单元素并做出不同选择的能力。

02:09 -登录表单完成
导航到登录页面 (the-internet.herokuapp.com/login),并在用户名字段中填写“tomsmith”,在密码字段中填写“SuperSecretPassword!”。演示了表单填写的自动化。

02:28 -登录提交
提交登录凭证并完成身份验证过程。展示了 Claude 触发表单提交和浏览多步骤流程的能力。

02:36 - API 请求执行
向 JSONPlaceholder API 端点执行 GET 请求。演示了 Claude 能够直接调用 API 并通过 MCP 集成处理返回的数据。

要求

  • Node.js 16 或更高版本

  • 克劳德桌面

  • 剧作家依赖关系

浏览器支持

npm init playwright@latest

此软件包包含 Playwright 以及运行浏览器自动化所需的依赖项。运行npm install时,将安装所需的 Playwright 依赖项。此软件包支持以下浏览器:

  • Chrome(默认)

  • 火狐

  • 微软 Edge

  • WebKit(Safari 引擎)

首次使用某种浏览器类型时,Playwright 会根据需要自动安装相应的浏览器驱动程序。您也可以使用以下命令手动安装:

npx playwright install chrome
npx playwright install firefox
npx playwright install webkit
npx playwright install msedge

关于 Safari 的说明:Playwright 不直接支持 Safari 浏览器。相反,它使用 WebKit,它是驱动 Safari 的浏览器引擎。

关于 Edge 的注意事项:选择 Edge 作为浏览器类型时,代理实际上会启动 Microsoft Edge(而非 Chromium)。从技术上讲,在 Playwright 中,Edge 是使用 Chromium 浏览器实例和“msedge”通道参数启动的,因为 Microsoft Edge 基于 Chromium。

安装

手动安装

  1. 克隆或下载此存储库:

git clone https://github.com/imprvhub/mcp-browser-agent
cd mcp-browser-agent
  1. 安装依赖项:

npm install
  1. 构建项目:

npm run build

运行 MCP 服务器

运行 MCP 服务器有两种方式:

选项 1:手动运行

  1. 打开终端或命令提示符

  2. 导航到项目目录

  3. 直接运行服务器:

node dist/index.js

使用 Claude Desktop 时,请保持此终端窗口打开。服务器将一直运行,直到您关闭终端。

选项 2:使用 Claude Desktop 自动启动(建议定期使用)

Claude Desktop 可以在需要时自动启动 MCP 服务器。设置方法如下:

配置

Claude Desktop 配置文件位于:

  • macOS : ~/Library/Application Support/Claude/claude_desktop_config.json

  • Windows : %APPDATA%\Claude\claude_desktop_config.json

  • Linux : ~/.config/Claude/claude_desktop_config.json

编辑此文件以添加浏览器代理 MCP 配置。如果该文件不存在,请创建它:

{
  "mcpServers": {
    "browserAgent": {
      "command": "node",
      "args": ["ABSOLUTE_PATH_TO_DIRECTORY/mcp-browser-agent/dist/index.js",
      "--browser",
      "chrome"
    ]
    }
  }
}

重要提示:将ABSOLUTE_PATH_TO_DIRECTORY替换为您安装 MCP 的完整绝对路径

  • macOS/Linux 示例: /Users/username/mcp-browser-agent

  • Windows 示例: C:\\Users\\username\\mcp-browser-agent

如果您已配置其他 MCP,只需在“mcpServers”对象中添加“browserAgent”部分即可。以下是包含多个 MCP 的配置示例:

{
  "mcpServers": {
    "otherMcp1": {
      "command": "...",
      "args": ["..."]
    },
    "otherMcp2": {
      "command": "...",
      "args": ["..."]
    },
    "browserAgent": {
      "command": "node",
      "args": [
        "ABSOLUTE_PATH_TO_DIRECTORY/mcp-browser-agent/dist/index.js",
      "--browser",
      "chrome"
    ]
    }
  }
}

浏览器选择

MCP 浏览器代理支持多种浏览器类型。默认情况下,它使用 Chrome,但您可以通过多种方式指定其他浏览器:

选项 1:配置文件

在您的主目录中创建或编辑文件.mcp_browser_agent_config.json :

{
  "browserType": "chrome"
}

browserType支持的值为:

  • chrome - 使用已安装的 Chrome(默认)

  • firefox - 使用 Firefox Nightly 浏览器

  • webkit - 使用 WebKit 引擎(注意:这不是 Safari 本身,而是为 Safari 提供支持的 WebKit 渲染引擎)

  • edge - 使用 Microsoft Edge

关于 Safari 的说明:Playwright 不直接支持 Safari 浏览器。它使用的是 WebKit,它是驱动 Safari 的浏览器引擎。Playwright 中的 WebKit 实现提供了类似的功能,但体验与 Safari 浏览器并不完全相同。

选项 2:命令行参数

手动启动 MCP 服务器时,可以指定浏览器类型:

node dist/index.js --browser firefox

选项 3:环境变量

设置MCP_BROWSER_TYPE环境变量:

MCP_BROWSER_TYPE=firefox node dist/index.js

选项 4:Claude 桌面配置

在 Claude Desktop 的claude_desktop_config.json中配置 MCP 时,可以指定浏览器类型:

{
  "mcpServers": {
    "browserAgent": {
      "command": "node",
      "args": [
        "ABSOLUTE_PATH_TO_DIRECTORY/mcp-browser-agent/dist/index.js",
        "--browser",
        "chrome"
      ]
    }
  }
}

技术实现

MCP 浏览器代理基于模型上下文协议 (MCP Protocol) 构建,使 Claude 能够通过 Playwright 与 Headful 浏览器进行交互。该实现包含四个主要组件:

  1. 服务器(index.ts)

    • 使用模型上下文协议标准协议初始化 MCP 服务器

    • 配置工具和资源的服务器功能

    • 通过 stdio 传输与 Claude 建立通信

  2. 工具注册表(tools.ts)

    • 定义浏览器和 API 工具模式

    • 指定参数、验证规则和说明

    • 向 MCP 服务器注册工具,用于 Claude 的发现

  3. 请求处理程序(handlers.ts)

    • 管理工具和资源的 MCP 协议请求

    • 将浏览器日志和屏幕截图公开为可查询的资源

    • 将工具执行请求路由到适当的处理程序

  4. 执行器(executor.ts)

    • 管理浏览器和 API 客户端生命周期

    • 使用 Playwright 实现浏览器自动化功能

    • 使用适当的错误处理和响应解析来处理 API 请求

    • 在命令之间维护有状态的浏览器会话

代理能力

与基本集成不同,MCP 浏览器代理通过以下方式充当真正的 AI 代理:

  • 在多个命令之间保持持久浏览器状态

  • 捕获详细的控制台日志以进行调试

  • 保存屏幕截图以供参考和审查

  • 管理复杂的交互序列

  • 提供详细的错误信息以便恢复

  • 支持复杂工作流程的链式操作

可用工具

浏览器工具

工具名称

描述

参数

browser_navigate

导航到 URL

url (必需)、 timeout 、 waitUntil

browser_screenshot

截取屏幕截图

name (必需)、 selector 、 fullPage 、 mask 、 savePath

browser_click

单击元素

selector (必需)

browser_fill

填写表单输入

selector (必需), value (必需)

browser_select

选择下拉选项

selector (必需), value (必需)

browser_hover

将鼠标悬停在元素上

selector (必需)

browser_evaluate

执行 JavaScript

script (必需)

API 工具

工具名称

描述

参数

api_get

GET 请求

url (必需)、 headers

api_post

POST 请求

url (必需)、 data (必需)、 headers

api_put

PUT 请求

url (必需)、 data (必需)、 headers

api_patch

PATCH 请求

url (必需)、 data (必需)、 headers

api_delete

删除请求

url (必需)、 headers

资源访问

MCP 浏览器代理公开以下资源:

  • browser://logs - 访问浏览器控制台日志

  • screenshot://[name] - 通过名称访问截图

示例用法

以下是一些如何将 MCP 浏览器代理与 Claude 结合使用的实际示例:

基本浏览器导航

Navigate to the Google homepage at https://www.google.com
Take a screenshot of the current page and name it "google-homepage"
Type "weather forecast" in the search box

简单交互

Navigate to https://www.wikipedia.org and search for "Model Context Protocol"
Go to https://the-internet.herokuapp.com/dropdown and select the option "Option 1" from the dropdown

基本表格填写

Navigate to https://the-internet.herokuapp.com/login and fill in the username field with "tomsmith" and the password field with "SuperSecretPassword!"
Go to https://the-internet.herokuapp.com/login, fill in the username and password fields, then click the login button

简单的 JavaScript 执行

Go to https://example.com and execute a JavaScript script to return the page title
Navigate to https://www.google.com and execute a JavaScript script to count the number of links on the page

基本 API 请求

Perform a GET request to https://jsonplaceholder.typicode.com/todos/1
Make a POST request to https://jsonplaceholder.typicode.com/posts with appropriate JSON data

这些示例代表了 MCP 浏览器代理的实际功能,并且更现实地说明了它在当前状态下可以完成的任务。

故障排除

“服务器断开连接”错误

如果您在 Claude Desktop 中看到错误“MCP 浏览器代理:服务器已断开连接”:

  1. 验证服务器正在运行:

    • 打开终端并从项目目录手动运行node dist/index.js

    • 如果服务器启动成功,则使用 Claude 并保持此终端打开

  2. 检查您的配置:

    • 确保claude_desktop_config.json中的绝对路径对于您的系统来说是正确的

    • 仔细检查 Windows 路径是否使用了双反斜杠 ( \\ )

    • 验证您使用的文件系统根目录的完整路径

浏览器未显示

如果浏览器没有启动或者您没有看到它:

  1. 检查指定的浏览器是否已安装

    • 验证您的系统上是否安装了浏览器(Chrome、Firefox、Edge 或 Safari/WebKit)

    • 浏览器驱动程序由 Playwright 自动处理

  2. 重启服务器和 Claude Desktop

    • 终止可能正在运行服务器的任何现有节点进程

    • 重新启动 Claude Desktop 以建立新的连接

浏览器进程未正确关闭

Chromium 和 Chrome 浏览器存在已知问题,在使用后进程有时无法正常终止。如果您遇到此问题:

  1. 手动关闭浏览器进程:

    • Windows :按 Ctrl+Shift+Esc 打开任务管理器,找到 Chrome/Chromium 进程并结束它

    • macOS :打开活动监视器(应用程序 > 实用程序 > 活动监视器),找到 Chrome/Chromium 进程并单击 X 将其终止

    • Linux :运行ps aux | grep chrome或ps aux | grep chromium来查找进程,然后kill <PID>终止它

  2. 关于浏览器兼容性的注意事项:

    • 此问题主要出现在 Chromium 和 Chrome 中

    • Firefox 和 Playwright 的内置浏览器通常不会遇到此问题

注意:此 MCP 集成基于 Playwright 构建,Playwright 存在已知问题和错误,可能会影响其运行。请将您在浏览器自动化过程中遇到的任何问题报告至Playwright 的 GitHub 问题。Playwright 团队正在持续努力解决这些问题,但尽管存在这些限制,此代理仍为 Claude Desktop 的浏览器自动化功能奠定了基础。

发展

项目结构

  • src/index.ts :主入口点和 MCP 服务器初始化

  • src/tools.ts :工具模式和注册

  • src/handlers.ts :工具和资源的 MCP 请求处理程序

  • src/executor.ts :使用 Playwright 的工具实现逻辑

建筑

npm run build

观察变化

npm run watch

测试

该项目包括验证核心功能和浏览器处理的测试。

npm test               # Run tests
npm run test:watch     # Watch mode
npm run test:coverage  # Coverage report

测试验证配置完整性、浏览器自动化功能、错误处理和进程清理。由于 Chrome/Chromium 终止存在已知问题,该测试套件尤其侧重于确保浏览器进程的正确处理。

安全注意事项

重要提示:此 MCP 集成为 Claude 提供了自主浏览器控制功能。请查看我们的安全政策,了解有关禁止使用、安全隐患和最佳实践的重要信息。

MCP 浏览器代理旨在执行合法的自动化任务,但也可能被滥用。用户有责任确保其使用符合所有适用法律、服务条款和道德准则。请参阅我们详细的安全政策,了解更多信息。

贡献

欢迎为 MCP 浏览器代理做出贡献!以下是您可以提供帮助的领域:

  • 添加新的浏览器自动化功能

  • 改进错误处理和恢复

  • 增强屏幕截图和资源管理

  • 创建有用的工作流程和示例

  • 优化复杂操作的性能

执照

该项目根据 Mozilla 公共许可证 2.0 获得许可 - 有关详细信息,请参阅LICENSE文件。

相关链接

Available Tools

13 tools
api_deleteB

Perform a DELETE request to an API endpoint

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesAPI endpoint URL
headersNoRequest headers

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden but only states the action without disclosing side effects, idempotency, authentication needs, rate limits, or return format. This is minimal for an agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It is appropriately concise for a simple tool, though slightly more context could be included without losing efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (2 params, no output schema), the description covers the basic action but omits expected return values, error conditions, or usage scope, leaving the description marginally adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptive parameter names (url, headers). The description adds no additional meaning beyond the schema, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool performs a DELETE request to an API endpoint, using a specific verb and resource that distinguishes it from sibling tools (api_get, api_patch, etc.) which handle other HTTP methods.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like api_get or api_post. The description merely repeats the method, missing explicit when/when-not context or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

api_getC

Perform a GET request to an API endpoint

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesAPI endpoint URL
headersNoRequest headers

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It fails to mention that GET is typically safe and idempotent, how errors are handled, or whether redirects are followed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single clear sentence with no unnecessary words. It efficiently conveys the core action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple but lacks information about return values, error handling, authentication requirements, or default behavior. The description is too minimal for an agent to use confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions for both parameters. The tool description adds no extra meaning beyond the schema, so baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Perform a GET request to an API endpoint', which identifies the HTTP method and the action. It distinguishes from sibling tools like api_post and api_delete by the method name, but lacks mention of read-only nature.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus siblings. The description does not specify that GET should be used for retrieving data, nor does it mention alternatives for modifying or deleting resources.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

api_patchB

Perform a PATCH request to an API endpoint

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesAPI endpoint URL
dataYesRequest body data (JSON string)
headersNoRequest headers

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It only states it performs a PATCH request, implying mutation, but does not disclose side effects, authentication needs, rate limits, or response behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no unnecessary words. However, it is very brief and could benefit from additional context without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is incomplete given the tool's complexity. No output schema exists, and the description does not explain return values or error handling. Sibling tools suggest it is part of an HTTP client set, but the description lacks depth.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description does not add any additional meaning beyond what is in the input schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it performs a PATCH request to an API endpoint, which is a specific verb and resource. This distinguishes it from sibling tools like api_get, api_post, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus other HTTP methods (e.g., POST, PUT) or alternatives. The context signals show siblings, but the description offers no differentiation advice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

api_postB

Perform a POST request to an API endpoint

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesAPI endpoint URL
dataYesRequest body data (JSON string)
headersNoRequest headers

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description must disclose behavior. It fails to mention typical POST behavior (resource creation), data validation, or error handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence is concise and front-loaded, but it is too brief and lacks substance for a practical tool description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple HTTP tool, description should mention typical use (e.g., 'sends data to URL'). Schema covers parameters but context is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (all parameters described), so baseline is 3. Description adds no additional meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states 'POST request' verb and 'API endpoint' resource, clearly distinguishing from sibling tools like api_get, api_put, etc.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, no mention of prerequisites or context such as authentication or data format.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

api_putC

Perform a PUT request to an API endpoint

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesAPI endpoint URL
dataYesRequest body data (JSON string)
headersNoRequest headers

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided; description only states 'PUT request' implying mutation, but omits critical details like idempotency, side effects, authentication needs, or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence is concise and front-loaded, but does not add meaningful content beyond the tool name; minimal but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no behavioral details; for a tool with 3 parameters and nested objects, the description is insufficient to fully understand usage and return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers all 3 parameters with descriptions; description adds no extra meaning beyond what schema already provides, meeting the baseline for high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the action ('Perform a PUT request') and resource ('API endpoint'), but lacks differentiation from sibling tools like api_patch or api_post, which perform similar HTTP methods.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use PUT versus other HTTP methods (e.g., PATCH for partial updates, POST for creation). Does not mention idempotency or replacement semantics.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_clickC

Click an element on the page

ParametersJSON Schema
NameRequiredDescriptionDefault
selectorYesCSS selector for element to click

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description does not disclose behavioral traits beyond the basic action, such as potential side effects (e.g., navigation, page changes) or element visibility requirements. No annotations exist to compensate for this lack of detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no unnecessary words. It is appropriately sized for a simple action, though it could benefit from slight elaboration without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the sibling tools and the single parameter, the description is incomplete. It fails to mention crucial context like element visibility, clicks causing navigation, or waiting behavior, leaving the agent with insufficient information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully describes the 'selector' parameter with a clear definition. The description adds no extra meaning beyond the schema, which is acceptable given 100% coverage, but it does not enhance understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Click an element') and the resource ('on the page'), distinguishing it from sibling tools like browser_fill or browser_hover. However, it could benefit from specifying that it operates within the current page context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives, such as when to click vs. hover, or prerequisites like page navigation. The description lacks context for proper selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_evaluateB

Execute JavaScript in the browser context

ParametersJSON Schema
NameRequiredDescriptionDefault
scriptYesJavaScript code to execute

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must fully disclose behavior. It mentions script execution but omits critical details: return value, side effects, permissions, sandboxing, or error handling. This is insufficient for a potentially powerful tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with no unnecessary words. It is optimally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema, so the description should explain return values or behavior. It does not. Additionally, it lacks details on execution context (e.g., async support, timeout). This leaves the agent underinformed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single required parameter 'script' has a description in the schema. The tool description adds no additional meaning beyond what the schema already provides, making it adequate but not additive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Execute JavaScript in the browser context' clearly states the action and resource. It is specific and distinct from sibling tools like browser_click or api_get.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives (e.g., browser_navigate or API calls). An agent receives no context for decision-making.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_fillB

Fill a form input with text

ParametersJSON Schema
NameRequiredDescriptionDefault
valueYesText to enter in the field
selectorYesCSS selector for input field

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description should disclose behavioral details, but it only says 'fill' without specifying whether it overwrites existing text, waits for elements, or handles disabled fields. This leaves significant behavioral ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence, front-loading the action. However, it could be slightly expanded with useful context while remaining brief.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of the tool (2 params, no output schema), the description is adequate but lacks details on behavior like clearing the field or submission. Sibling tools exist for other form actions, but no comparative guidance is provided.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema covers 100% of parameters with descriptions (selector and value), so baseline is 3. The description adds no additional semantic information beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Fill a form input with text' clearly states the action (fill) and the target (form input), distinguishing it from sibling tools like browser_click or browser_select which perform different actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as browser_evaluate for setting values or browser_click for activation. There is no mention of prerequisites like element visibility or state.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_hoverC

Hover over an element on the page

ParametersJSON Schema
NameRequiredDescriptionDefault
selectorYesCSS selector for element to hover over

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavioral traits. It only states 'hover over an element' without explaining whether it triggers JavaScript events, waits for any transitions, or is safe. Essential behavioral context is missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise at 6 words and front-loaded. It is not verbose, but the brevity may sacrifice necessary detail. It earns its place but could be slightly more informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (1 parameter, no output schema), the description is incomplete. It does not mention the return value (likely void or success), side effects, or behavior after hovering. Sibling tools suggest a sequence of actions, but this tool's role is under-described.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a single parameter 'selector' described as 'CSS selector for element to hover over'. The description adds no additional meaning beyond what the schema already provides, so it meets the baseline but does not enhance understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Hover over an element on the page' clearly states the action (hover) and target (element on page). It differentiates from siblings like browser_click and browser_fill. However, it could be more specific about the effect (e.g., triggering hover state) but is not a tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like browser_click or browser_evaluate. The description lacks any context about prerequisites, typical scenarios, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_navigateC

Navigate to a specific URL

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to navigate to
timeoutNoNavigation timeout in milliseconds
waitUntilNoNavigation wait criteria

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description does not disclose behavioral traits such as whether navigation replaces the current page, how timeouts affect behavior, or what happens on failure. The parameters timeout and waitUntil are defined in the schema but not mentioned in the description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with no unnecessary words. However, it could be slightly expanded to include key behavioral details without losing conciseness, but as is, it is efficient and to the point.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 3 parameters (one with enum), sibling tools, and no output schema, the description is too brief. It does not cover return values, error handling, or when to use timeout/waitUntil. The agent receiving this description would need to rely entirely on the schema for context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the description does not need to add much. The description itself adds no parameter information beyond the schema, but given full coverage, a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action of navigating to a URL, which is the primary purpose. However, it does not differentiate from sibling browser tools like browser_click or browser_fill, but the verb 'navigate' and parameter 'url' make it unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus other browser actions such as browser_click or browser_fill. The description lacks context for appropriate usage scenarios or prerequisites like requiring a current page.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_screenshotB

Capture a screenshot of the current page or a specific element

ParametersJSON Schema
NameRequiredDescriptionDefault
maskNoSelectors for elements to mask
nameYesIdentifier for the screenshot
fullPageNoCapture full page height
savePathNoPath to save screenshot (default: user's Downloads folder)
selectorNoCSS selector for element to capture

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description does not disclose behavioral traits such as whether the capture affects the page state, any authorization requirements, or rate limits. It only states the action, leaving the agent without important context about side effects or constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single well-formed sentence that directly states the tool's purpose. It wastes no words, though it could benefit from slightly more detail without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is brief and lacks context about return values (e.g., image path), default behavior, or how the tool interacts with other browser tools. Given the five parameters and no output schema, the description should provide more operational context to be fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Since schema description coverage is 100%, the baseline is 3. The description does not add additional meaning beyond the parameter names and schema descriptions; for example, it doesn't explain when to use fullPage or mask options more concretely.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool captures a screenshot of the current page or a specific element, which is a specific verb-resource combination. It distinguishes itself from sibling tools like browser_navigate or browser_click by focusing on capture rather than navigation or element interaction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used when a screenshot is needed, but it provides no explicit guidance on when to use it versus alternative methods (e.g., browser_evaluate for custom captures) or any exclusions (e.g., not for video capture).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_selectC

Select an option from a dropdown menu

ParametersJSON Schema
NameRequiredDescriptionDefault
valueYesValue or label to select
selectorYesCSS selector for select element

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description does not disclose behavioral traits such as whether it waits for options to load, supports custom dropdowns, triggers events, or requires scrolling. The description is too minimal to inform the agent of important behaviors.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that is front-loaded. However, it sacrifices completeness for brevity. It earns a 4 for being efficient, but could include more key details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that no annotations or output schema exist, the description is insufficiently complete. It does not explain return values, constraints, or behavior for complex dropdowns, leaving gaps for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema describes both parameters (selector and value) with 100% coverage. The description adds no additional meaning beyond the schema, so baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it selects an option from a dropdown menu, using a specific verb and resource. It distinguishes from sibling tools like browser_fill (text input) and browser_click (clicking), though it could be more precise by specifying HTML <select> elements.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives (e.g., browser_click for custom dropdowns) or any prerequisites. No exclusions or context are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_set_viewportA

Change the browser's viewport size and scale factor

ParametersJSON Schema
NameRequiredDescriptionDefault
widthNoViewport width in pixels
heightNoViewport height in pixels
deviceScaleFactorNoDevice scale factor (affects how content is scaled)

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description should disclose behavioral traits. It only states what the tool changes, but not side effects (e.g., impact on screenshots, persistence across navigation). Lacks important context for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, no fluff. Every word is relevant. Excellent conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Adequately describes the core function, but lacks details on required fields, defaults, or behavioral context. Acceptable for a simple tool, but could be more complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with meaningful descriptions for each parameter. The description adds no extra meaning beyond the schema, justifying the baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the action 'Change' and the target 'browser's viewport size and scale factor'. It is specific and distinct from sibling tools like browser_navigate or browser_click.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use guidance is provided. The description implies usage for adjusting viewport, but does not mention alternatives or exclusions. Minimal viable score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updatesv1.0.0
    • First observedapi_delete
    • First observedapi_get
    • First observedapi_patch
    • First observedapi_post
    • First observedapi_put
    • First observedbrowser_click
    • First observedbrowser_evaluate
    • First observedbrowser_fill
    • First observedbrowser_hover
    • First observedbrowser_navigate
    • First observedbrowser_screenshot
    • First observedbrowser_select
    • First observedbrowser_set_viewport

TDQS

A3.5/5.0

Scored across 13 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: API tools are differentiated by HTTP method, and browser tools cover unique interactions like clicking, hovering, filling, etc. No two tools overlap in function.

Naming Consistency5/5

All tools follow a consistent '<domain>_<action>' pattern, with 'api_' prefix for HTTP methods and 'browser_' prefix for browser actions. Naming is unambiguous and predictable.

Tool Count5/5

13 tools is well-scoped for a browser automation and API testing server. Each tool covers a fundamental operation without unnecessary bloat or gaps.

Completeness4/5

Core browser interactions (navigation, clicking, form filling, selecting, screenshot) and all major HTTP methods are covered. Minor omissions like file upload or wait-for-element are acceptable for this scope.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    A browser automation agent that enables Claude to interact with web browsers through the Model Context Protocol, allowing for actions like navigating websites, manipulating elements, and managing browser state.
    2
    9
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables real browser automation as tools in Cursor, Claude Desktop, Windsurf, and any MCP-compatible client, allowing AI agents to interact with web pages through natural language.
    17 npm
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables Claude Code to control a real browser using AI for web scraping, competitive intelligence, and UX auditing through the MCP protocol.
    -
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables browser automation through the Claude Chrome Extension, allowing agents to navigate websites, fill forms, take screenshots, and debug web apps via standard MCP protocols.
    1
    MIT