Skip to main content
Glama
ztobs

Browser Use Server

by ztobs

브라우저 사용 서버

대장간 배지

Python 스크립트를 사용한 브라우저 자동화를 위한 모델 컨텍스트 프로토콜 서버입니다. Cline과 함께 사용 가능

특징

브라우저 작업

  • screenshot : 웹페이지(전체 페이지 또는 뷰포트)의 스크린샷을 캡처합니다.

  • get_html : 웹페이지의 HTML 콘텐츠를 검색합니다.

  • execute_js : 웹페이지에서 JavaScript 실행

  • get_console_logs : 웹페이지에서 콘솔 로그를 가져옵니다.

모든 작업은 페이지 로드 후 사용자 지정 상호작용 단계(예: 요소 클릭, 스크롤)를 지원합니다.

Related MCP server: Playwright MCP Server for Security

필수 조건

  1. (선택 사항이지만 권장됨) 헤드리스 브라우저 자동화를 위해 Xvfb를 설치하세요.

지엑스피1

Xvfb(X Virtual Frame Buffer)는 가상 디스플레이를 생성하여 봇으로 감지되지 않고 브라우저 자동화를 가능하게 합니다. Xvfb에 대한 자세한 내용은 여기를 참조하세요.

  1. Miniconda 또는 Anaconda 설치

  2. Conda 환경을 만듭니다.

conda create -n browser-use python=3.11
conda activate browser-use
pip install -r requirements.txt
  1. LLM 구성 설정:

이 서버는 여러 LLM 공급자를 지원합니다. 다음 API 키를 사용할 수 있습니다.

# Required: Set at least one of these API keys
export GLHF_API_KEY=your_api_key
export GROQ_API_KEY=your_api_key
export OPENAI_API_KEY=your_api_key
export OPENROUTER_API_KEY=your_api_key
export GITHUB_API_KEY=your_api_key
export DEEPSEEK_API_KEY=your_api_key
export GEMINI_API_KEY=your_api_key
export OLLAMA_API_KEY=your_api_key

# Optional: Override default configuration
export MODEL=your_preferred_model  # Override the default model
export BASE_URL=your_custom_url    # Override the default API endpoint
export USE_VISION=false  # Enable/disable vision capabilities (default: false)

서버는 자동으로 찾은 첫 번째 사용 가능한 API 키를 사용합니다. 환경 변수를 사용하여 모든 공급자의 모델과 기본 URL을 사용자 지정할 수 있습니다.

설치

Smithery를 통해 설치

Smithery를 통해 Claude Desktop용 Browser Use Server를 자동으로 설치하려면:

npx -y @smithery/cli install @ztobs/cline-browser-use-mcp --client claude
  1. 이 저장소를 /home/YOUR_HOME/Documents/Cline/ 디렉토리로 복제합니다.

  2. 종속성 설치:

npm install
  1. 서버를 빌드하세요:

npm run build

MCP 구성

Cline MCP 설정에 다음 구성을 추가하세요.

"browser-use": {
  "command": "node",
  "args": [
    "/home/YOUR_HOME/Documents/Cline/MCP/browser-use-server/build/index.js"
  ],
  "env": {
    // Required: Set at least one API key
    "GLHF_API_KEY": "your_api_key",
    "GROQ_API_KEY": "your_api_key",
    "OPENAI_API_KEY": "your_api_key",
    "OPENROUTER_API_KEY": "your_api_key",
    "GITHUB_API_KEY": "your_api_key",
    "DEEPSEEK_API_KEY": "your_api_key",
    "GEMINI_API_KEY": "your_api_key",
    "OLLAMA_API_KEY": "your_api_key",
    // Optional: Configuration overrides
    "MODEL": "your_preferred_model",
    "BASE_URL": "your_custom_url",
    "USE_VISION": "false"
  },
  "disabled": false,
  "autoApprove": []
}

바꾸다:

  • YOUR_HOME 실제 홈 디렉토리 이름으로 변경

  • 실제 API 키와 your_api_key 함께 사용하세요

용법

서버를 실행합니다:

node build/index.js

서버는 stdio에서 사용할 수 있으며 다음 작업을 지원합니다.

스크린샷

매개변수:

  • url: 웹페이지 URL(필수)

  • full_page: 전체 페이지를 캡처할지 아니면 뷰포트만 캡처할지(선택 사항, 기본값: false)

  • steps: 페이지 로드 후 수행해야 할 단계를 설명하는 쉼표로 구분된 작업 또는 문장(선택 사항)

HTML 가져오기

매개변수:

  • url: 웹페이지 URL(필수)

  • steps: 페이지 로드 후 수행해야 할 단계를 설명하는 쉼표로 구분된 작업 또는 문장(선택 사항)

JavaScript 실행

매개변수:

  • url: 웹페이지 URL(필수)

  • 스크립트: 실행할 JavaScript 코드(필수)

  • steps: 페이지 로드 후 수행해야 할 단계를 설명하는 쉼표로 구분된 작업 또는 문장(선택 사항)

콘솔 로그 가져오기

매개변수:

  • url: 웹페이지 URL(필수)

  • steps: 페이지 로드 후 수행해야 할 단계를 설명하는 쉼표로 구분된 작업 또는 문장(선택 사항)

클라인 사용 예시

다음은 Cline과 함께 브라우저 사용 서버를 사용하여 수행할 수 있는 몇 가지 작업의 예입니다.

개발 중 웹 페이지 요소 수정

인증이 필요한 페이지에서 제목의 색상을 변경하려면:

Change the colour of the headline with the text "Alle Foren im Überblick." to deep blue on https://localhost:3000/foren/ page

To check/see the page, use browser-use MCP server to:
Open https://localhost:3000/auth,
Login with ztobs:Password123,
Navigate to https://localhost:3000/foren/,
Accept cookies if required

hint: execute all browser actions in one command with multiple comma-separated steps

이 작업에서는 다음 사항을 보여줍니다.

  • 쉼표로 구분된 단계를 사용한 다단계 브라우저 자동화

  • 인증 처리

  • 쿠키 수락

  • DOM 조작

  • CSS 스타일 변경

서버는 이러한 단계를 순차적으로 실행하면서 그 과정에서 필요한 상호작용을 처리합니다.

구성

LLM 구성

서버는 기본 구성을 사용하여 여러 LLM 공급자를 지원합니다.

  • GLHF: deepseek-ai/DeepSeek-V3 모델을 사용합니다.

  • Ollama: 32k 컨텍스트 창을 사용하는 qwen2.5:32b-instruct-q4_K_M 모델을 사용합니다.

  • Groq: deepseek-r1-distill-llama-70b 모델을 사용합니다.

  • OpenAI: gpt-4o-mini 모델을 사용합니다.

  • Openrouter: deepseek/deepseek-chat 모델을 사용합니다.

  • Github: gpt-4o-mini 모델 사용

  • DeepSeek: deepseek-chat 모델을 사용합니다

  • Gemini: gemini-2.0-flash-exp 모델을 사용합니다.

환경 변수를 사용하여 이러한 기본값을 재정의할 수 있습니다.

  • MODEL : 모든 공급자에 대한 사용자 정의 모델 이름 설정

  • BASE_URL : 사용자 정의 API 엔드포인트 URL을 설정합니다(공급자가 지원하는 경우)

비전 지원

서버는 USE_VISION 환경 변수를 통해 비전 기능을 지원합니다.

  • 브라우저 작업에 대한 비전 기능을 활성화하려면 USE_VISION=true를 설정합니다.

  • 비전이 필요하지 않을 때 성능을 최적화하기 위해 기본값은 false입니다.

  • 웹 페이지 콘텐츠의 시각적 이해가 필요한 작업에 유용합니다.

Xvfb 지원

서버는 Xvfb가 설치되어 있는지 자동으로 감지합니다.

  • 사용 가능한 경우 xvfb-run을 사용하여 봇 감지 없이 더 나은 브라우저 자동화를 활성화합니다.

  • Xvfb가 설치되지 않은 경우 직접 실행으로 돌아갑니다.

  • RUNNING_UNDER_XVFB 환경 변수를 적절히 설정합니다.

타임아웃

기본 제한 시간은 5분(300000ms)입니다. build/index.js 파일의 TIMEOUT 상수를 수정하여 이 값을 변경하세요.

오류 처리

서버는 다음에 대한 자세한 오류 메시지를 제공합니다.

  • Python 스크립트 실행 실패

  • 브라우저 작업 시간 초과

  • 잘못된 매개변수

디버깅

디버깅을 위해 MCP Inspector를 사용하세요.

npm run inspector

용도

브라우저 사용

특허

MIT

Available Tools

4 tools
execute_jsC

Execute JavaScript code on a webpage

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to navigate to
scriptYesThe JavaScript code to execute
stepsNoComma-separated actions or sentences describing steps to take after page load (e.g., "click #submit, scroll down" or "Fill the login form and submit")

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions execution but lacks details on permissions needed, potential side effects (e.g., page modifications), error handling, or execution environment. This is inadequate for a tool that performs code execution.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with no wasted words. It's front-loaded and efficiently communicates the core function without unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of executing JavaScript on a webpage, the lack of annotations, and no output schema, the description is insufficient. It doesn't cover behavioral aspects like safety, return values, or error conditions, leaving significant gaps for an agent to use this tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters (url, script, steps). The description adds no additional meaning or context beyond what's in the schema, such as examples or constraints, but doesn't contradict it either.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Execute JavaScript code') and target ('on a webpage'), which is specific and unambiguous. However, it doesn't explicitly differentiate from sibling tools like get_console_logs or get_html, which also interact with webpages but for different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like get_html or screenshot, nor does it mention prerequisites or constraints. It simply states what the tool does without context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_console_logsC

Get the console logs of a webpage

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to navigate to
stepsNoComma-separated actions or sentences describing steps to take after page load (e.g., "click #submit, scroll down" or "Fill the login form and submit")

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states what the tool does but lacks critical details: it doesn't specify if this requires browser automation, what types of console logs are captured (e.g., errors, warnings), whether it's a read-only operation, or any limitations like timeouts or authentication needs. This leaves significant gaps in understanding how the tool behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with zero waste—it directly states the tool's purpose without unnecessary words. It's appropriately sized and front-loaded, making it easy to grasp immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of interacting with webpages and the lack of annotations and output schema, the description is incomplete. It doesn't address key contextual aspects like what the tool returns (e.g., log format, error handling), behavioral constraints, or how it differs from siblings. For a tool with two parameters and no structured safety hints, more detail is needed to be fully helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, clearly documenting both parameters ('url' and 'steps'). The description adds no additional meaning beyond what the schema provides, such as explaining the format of console logs or how steps interact with log capture. Since the schema does the heavy lifting, the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Get') and resource ('console logs of a webpage'), making it immediately understandable. However, it doesn't differentiate from sibling tools like 'execute_js' or 'get_html', which might also interact with webpage content, so it doesn't reach the highest score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'execute_js' (which might execute JavaScript and potentially capture logs) or 'get_html' (which retrieves HTML content). There's no mention of prerequisites, such as whether the webpage needs to be loaded first, or any exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_htmlC

Get the HTML content of a webpage

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to navigate to
stepsNoComma-separated actions or sentences describing steps to take after page load (e.g., "click #submit, scroll down" or "Fill the login form and submit")

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states what the tool does but doesn't describe how it behaves—e.g., whether it follows redirects, handles authentication, respects rate limits, or returns errors. This leaves critical operational details unspecified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with zero wasted words. It's front-loaded and efficiently communicates the core function without unnecessary elaboration, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete for a tool with two parameters and potential behavioral complexity. It doesn't address what the tool returns (e.g., raw HTML, status codes), error handling, or dependencies, leaving significant gaps for an AI agent to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters ('url' and 'steps') thoroughly. The description doesn't add any meaning beyond what the schema provides, such as clarifying the interaction between parameters or providing examples of 'steps' usage, resulting in a baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and resource ('HTML content of a webpage'), making the purpose immediately understandable. It doesn't specifically differentiate from sibling tools like 'execute_js' or 'screenshot', but the core function is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'execute_js' or 'screenshot'. It doesn't mention prerequisites, limitations, or scenarios where this tool is preferred over others, leaving usage context entirely implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshotC

Take a screenshot of a webpage

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to navigate to
full_pageNoWhether to capture the full page or just the viewport
stepsNoComma-separated actions or sentences describing steps to take after page load (e.g., "click #submit, scroll down" or "Fill the login form and submit")

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but only states the basic action without disclosing behavioral traits. It lacks details on permissions needed, potential rate limits, output format (e.g., image type), error handling, or whether it's a read-only or mutative operation, leaving significant gaps for an agent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero waste, front-loading the core purpose. Every word earns its place, making it highly concise and well-structured for quick comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (involving webpage interaction and screenshot capture), no annotations, and no output schema, the description is incomplete. It fails to address critical context like what the output returns (e.g., image data or file path), error conditions, or behavioral nuances, leaving the agent under-informed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters (url, full_page, steps). The description adds no additional meaning beyond implying webpage capture, which is redundant with the schema's details. Baseline 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('Take') and resource ('screenshot of a webpage'), making the purpose immediately understandable. It distinguishes from siblings like execute_js or get_html by focusing on visual capture rather than code execution or HTML retrieval, though it doesn't explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like get_html for content extraction or execute_js for interactive actions. The description implies usage for webpage capture but offers no context about prerequisites, limitations, or comparative scenarios with sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updates
    • First observedexecute_js
    • First observedget_console_logs
    • First observedget_html
    • First observedscreenshot

TDQS

A3.5/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: execute_js runs code, get_console_logs retrieves logs, get_html fetches content, and screenshot captures visual output. There is no overlap or ambiguity between these functions.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern with snake_case (e.g., execute_js, get_console_logs, get_html, screenshot). The naming is predictable and readable throughout.

Tool Count5/5

With 4 tools, this server is well-scoped for browser automation, covering key operations like executing scripts, retrieving logs, getting content, and taking screenshots. Each tool earns its place without being excessive or insufficient.

Completeness4/5

The tool set covers essential browser interactions for the domain, including execution, logging, content retrieval, and visualization. A minor gap exists in navigation or page manipulation tools (e.g., navigate, click), but core workflows are well-supported.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers