Skip to main content
Glama
ztobs

Browser Use Server

by ztobs

浏览器使用服务器

铁匠徽章

一个使用 Python 脚本实现浏览器自动化的模型上下文协议服务器。可与 Cline 配合使用

特征

浏览器操作

  • screenshot :捕获网页截图(整页或视口)

  • get_html :检索网页的 HTML 内容

  • execute_js网页上的 JavaScript

  • get_console_logs :从网页获取控制台日志

所有操作都支持页面加载后自定义交互步骤(例如点击元素、滚动)。

Related MCP server: Playwright MCP Server for Security

先决条件

  1. (可选但推荐)安装 Xvfb 以实现无头浏览器自动化:

# Ubuntu/Debian
sudo apt-get install xvfb

# CentOS/RHEL
sudo yum install xorg-x11-server-Xvfb

# Arch Linux
sudo pacman -S xorg-server-xvfb

Xvfb(X 虚拟帧缓冲区)创建虚拟显示器,允许浏览器自动化运行,而不会被检测为机器人程序。点击此处了解更多关于 Xvfb 的信息。

  1. 安装 Miniconda 或 Anaconda

  2. 创建 Conda 环境:

conda create -n browser-use python=3.11
conda activate browser-use
pip install -r requirements.txt
  1. 设置 LLM 配置:

该服务器支持多个 LLM 提供程序。您可以使用以下任意 API 密钥:

# Required: Set at least one of these API keys
export GLHF_API_KEY=your_api_key
export GROQ_API_KEY=your_api_key
export OPENAI_API_KEY=your_api_key
export OPENROUTER_API_KEY=your_api_key
export GITHUB_API_KEY=your_api_key
export DEEPSEEK_API_KEY=your_api_key
export GEMINI_API_KEY=your_api_key
export OLLAMA_API_KEY=your_api_key

# Optional: Override default configuration
export MODEL=your_preferred_model  # Override the default model
export BASE_URL=your_custom_url    # Override the default API endpoint
export USE_VISION=false  # Enable/disable vision capabilities (default: false)

服务器将自动使用找到的第一个可用的 API 密钥。您可以选择使用环境变量为任何提供程序自定义模型和基本 URL。

安装

通过 Smithery 安装

要通过Smithery自动安装 Claude Desktop 的浏览器使用服务器:

npx -y @smithery/cli install @ztobs/cline-browser-use-mcp --client claude
  1. 将此存储库克隆到/home/YOUR_HOME/Documents/Cline/目录

  2. 安装依赖项:

npm install
  1. 构建服务器:

npm run build

MCP 配置

将以下配置添加到您的 Cline MCP 设置:

"browser-use": {
  "command": "node",
  "args": [
    "/home/YOUR_HOME/Documents/Cline/MCP/browser-use-server/build/index.js"
  ],
  "env": {
    // Required: Set at least one API key
    "GLHF_API_KEY": "your_api_key",
    "GROQ_API_KEY": "your_api_key",
    "OPENAI_API_KEY": "your_api_key",
    "OPENROUTER_API_KEY": "your_api_key",
    "GITHUB_API_KEY": "your_api_key",
    "DEEPSEEK_API_KEY": "your_api_key",
    "GEMINI_API_KEY": "your_api_key",
    "OLLAMA_API_KEY": "your_api_key",
    // Optional: Configuration overrides
    "MODEL": "your_preferred_model",
    "BASE_URL": "your_custom_url",
    "USE_VISION": "false"
  },
  "disabled": false,
  "autoApprove": []
}

代替:

  • YOUR_HOME替换为您的实际主目录名称

  • your_api_key替换为您的实际 API 密钥

用法

运行服务器:

node build/index.js

该服务器将在 stdio 上可用并支持以下操作:

截屏

参数:

  • url:网页URL(必填)

  • full_page:是否捕获整个页面或仅捕获视口(可选,默认值:false)

  • 步骤:以逗号分隔的操作或句子,描述页面加载后要采取的步骤(可选)

获取 HTML

参数:

  • url:网页URL(必填)

  • 步骤:以逗号分隔的操作或句子,描述页面加载后要采取的步骤(可选)

执行 JavaScript

参数:

  • url:网页URL(必填)

  • script:要执行的 JavaScript 代码(必需)

  • 步骤:以逗号分隔的操作或句子,描述页面加载后要采取的步骤(可选)

获取控制台日志

参数:

  • url:网页URL(必填)

  • 步骤:以逗号分隔的操作或句子,描述页面加载后要采取的步骤(可选)

Cline 使用示例

以下是使用 Cline 的浏览器服务器可以完成的一些示例任务:

在开发过程中修改网页元素

要更改需要身份验证的页面上的标题颜色:

Change the colour of the headline with the text "Alle Foren im Überblick." to deep blue on https://localhost:3000/foren/ page

To check/see the page, use browser-use MCP server to:
Open https://localhost:3000/auth,
Login with ztobs:Password123,
Navigate to https://localhost:3000/foren/,
Accept cookies if required

hint: execute all browser actions in one command with multiple comma-separated steps

此任务演示:

  • 使用逗号分隔的步骤实现多步骤浏览器自动化

  • 身份验证处理

  • 接受 Cookie

  • DOM 操作

  • CSS 样式更改

服务器将按顺序执行这些步骤,并处理过程中所需的任何交互。

配置

LLM 配置

该服务器支持多个 LLM 提供程序及其默认配置:

  • GLHF:使用 deepseek-ai/DeepSeek-V3 模型

  • Ollama:使用 qwen2.5:32b-instruct-q4_K_M 模型和 32k 上下文窗口

  • Groq:使用 deepseek-r1-distill-llama-70b 模型

  • OpenAI:使用gpt-4o-mini模型

  • Openrouter:使用 deepseek/deepseek-chat 模型

  • Github:使用 gpt-4o-mini 模型

  • DeepSeek:使用 deepseek-chat 模型

  • Gemini:使用 gemini-2.0-flash-exp 模型

您可以使用环境变量覆盖这些默认值:

  • MODEL :为任何提供商设置自定义模型名称

  • BASE_URL :设置自定义 API 端点 URL(如果提供商支持)

视觉支持

服务器通过 USE_VISION 环境变量支持视觉功能:

  • 设置 USE_VISION=true 以启用浏览器操作的视觉功能

  • 默认值为 false,以便在不需要视觉时优化性能

  • 对于需要视觉理解网页内容的任务很有用

Xvfb 支持

服务器会自动检测 Xvfb 是否已安装并且:

  • 在可用时使用 xvfb-run,实现更好的浏览器自动化,无需机器人检测

  • 当未安装 Xvfb 时,回退到直接执行

  • 相应地设置 RUNNING_UNDER_XVFB 环境变量

暂停

默认超时时间为 5 分钟(300000 毫秒)。修改build/index.js中的 TIMEOUT 常量即可更改此设置。

错误处理

服务器提供以下详细的错误消息:

  • Python 脚本执行失败

  • 浏览器操作超时

  • 参数无效

调试

使用 MCP Inspector 进行调试:

npm run inspector

用途

浏览器使用

执照

麻省理工学院

Available Tools

4 tools
execute_jsC

Execute JavaScript code on a webpage

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to navigate to
scriptYesThe JavaScript code to execute
stepsNoComma-separated actions or sentences describing steps to take after page load (e.g., "click #submit, scroll down" or "Fill the login form and submit")

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions execution but lacks details on permissions needed, potential side effects (e.g., page modifications), error handling, or execution environment. This is inadequate for a tool that performs code execution.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with no wasted words. It's front-loaded and efficiently communicates the core function without unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of executing JavaScript on a webpage, the lack of annotations, and no output schema, the description is insufficient. It doesn't cover behavioral aspects like safety, return values, or error conditions, leaving significant gaps for an agent to use this tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters (url, script, steps). The description adds no additional meaning or context beyond what's in the schema, such as examples or constraints, but doesn't contradict it either.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Execute JavaScript code') and target ('on a webpage'), which is specific and unambiguous. However, it doesn't explicitly differentiate from sibling tools like get_console_logs or get_html, which also interact with webpages but for different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like get_html or screenshot, nor does it mention prerequisites or constraints. It simply states what the tool does without context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_console_logsC

Get the console logs of a webpage

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to navigate to
stepsNoComma-separated actions or sentences describing steps to take after page load (e.g., "click #submit, scroll down" or "Fill the login form and submit")

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states what the tool does but lacks critical details: it doesn't specify if this requires browser automation, what types of console logs are captured (e.g., errors, warnings), whether it's a read-only operation, or any limitations like timeouts or authentication needs. This leaves significant gaps in understanding how the tool behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with zero waste—it directly states the tool's purpose without unnecessary words. It's appropriately sized and front-loaded, making it easy to grasp immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of interacting with webpages and the lack of annotations and output schema, the description is incomplete. It doesn't address key contextual aspects like what the tool returns (e.g., log format, error handling), behavioral constraints, or how it differs from siblings. For a tool with two parameters and no structured safety hints, more detail is needed to be fully helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, clearly documenting both parameters ('url' and 'steps'). The description adds no additional meaning beyond what the schema provides, such as explaining the format of console logs or how steps interact with log capture. Since the schema does the heavy lifting, the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Get') and resource ('console logs of a webpage'), making it immediately understandable. However, it doesn't differentiate from sibling tools like 'execute_js' or 'get_html', which might also interact with webpage content, so it doesn't reach the highest score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'execute_js' (which might execute JavaScript and potentially capture logs) or 'get_html' (which retrieves HTML content). There's no mention of prerequisites, such as whether the webpage needs to be loaded first, or any exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_htmlC

Get the HTML content of a webpage

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to navigate to
stepsNoComma-separated actions or sentences describing steps to take after page load (e.g., "click #submit, scroll down" or "Fill the login form and submit")

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states what the tool does but doesn't describe how it behaves—e.g., whether it follows redirects, handles authentication, respects rate limits, or returns errors. This leaves critical operational details unspecified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with zero wasted words. It's front-loaded and efficiently communicates the core function without unnecessary elaboration, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete for a tool with two parameters and potential behavioral complexity. It doesn't address what the tool returns (e.g., raw HTML, status codes), error handling, or dependencies, leaving significant gaps for an AI agent to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters ('url' and 'steps') thoroughly. The description doesn't add any meaning beyond what the schema provides, such as clarifying the interaction between parameters or providing examples of 'steps' usage, resulting in a baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and resource ('HTML content of a webpage'), making the purpose immediately understandable. It doesn't specifically differentiate from sibling tools like 'execute_js' or 'screenshot', but the core function is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'execute_js' or 'screenshot'. It doesn't mention prerequisites, limitations, or scenarios where this tool is preferred over others, leaving usage context entirely implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshotC

Take a screenshot of a webpage

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to navigate to
full_pageNoWhether to capture the full page or just the viewport
stepsNoComma-separated actions or sentences describing steps to take after page load (e.g., "click #submit, scroll down" or "Fill the login form and submit")

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but only states the basic action without disclosing behavioral traits. It lacks details on permissions needed, potential rate limits, output format (e.g., image type), error handling, or whether it's a read-only or mutative operation, leaving significant gaps for an agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero waste, front-loading the core purpose. Every word earns its place, making it highly concise and well-structured for quick comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (involving webpage interaction and screenshot capture), no annotations, and no output schema, the description is incomplete. It fails to address critical context like what the output returns (e.g., image data or file path), error conditions, or behavioral nuances, leaving the agent under-informed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters (url, full_page, steps). The description adds no additional meaning beyond implying webpage capture, which is redundant with the schema's details. Baseline 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('Take') and resource ('screenshot of a webpage'), making the purpose immediately understandable. It distinguishes from siblings like execute_js or get_html by focusing on visual capture rather than code execution or HTML retrieval, though it doesn't explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like get_html for content extraction or execute_js for interactive actions. The description implies usage for webpage capture but offers no context about prerequisites, limitations, or comparative scenarios with sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updates
    • First observedexecute_js
    • First observedget_console_logs
    • First observedget_html
    • First observedscreenshot

TDQS

A3.5/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: execute_js runs code, get_console_logs retrieves logs, get_html fetches content, and screenshot captures visual output. There is no overlap or ambiguity between these functions.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern with snake_case (e.g., execute_js, get_console_logs, get_html, screenshot). The naming is predictable and readable throughout.

Tool Count5/5

With 4 tools, this server is well-scoped for browser automation, covering key operations like executing scripts, retrieving logs, getting content, and taking screenshots. Each tool earns its place without being excessive or insufficient.

Completeness4/5

The tool set covers essential browser interactions for the domain, including execution, logging, content retrieval, and visualization. A minor gap exists in navigation or page manipulation tools (e.g., navigate, click), but core workflows are well-supported.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers