MCP Browser Agent
MCP 브라우저 에이전트
AGI House MCP 해커톤에서 제작
개요
이 프로젝트는 모델 컨텍스트 프로토콜(MCP)을 사용하여 브라우저 상호작용을 지원하는 브라우저 자동화 에이전트입니다. MCP 서버를 통해 Claude와 브라우저 자동화 기능을 원활하게 통합합니다.
MCP 서버를 구동하는 데 도움이 되는 브라우저 에이전트 기능을 제공해 준 Browser-Use에 감사드립니다!
Related MCP server: selenium-mcp
시스템 요구 사항
macOS(다윈 24.2.0)
Python 3.12 이상
uv패키지 관리자Google Chrome 브라우저(작업을 실행하기 전에 브라우저를 닫아두세요.)
설치
Smithery를 통해 설치
Smithery를 통해 Claude Desktop용 브라우저 자동화 에이전트를 자동으로 설치하려면:
지엑스피1
수동 설치
저장소를 복제합니다.
git clone <repository-url>
cd mcpuv사용하여 Python 환경을 설정합니다.
uv venv
source .venv/bin/activate
uv sync구성
클로드 데스크톱 구성
Claude Desktop 구성 파일을 만들거나 수정하세요.
{
"mcpServers": {
"browser-use": {
"command": "uv",
"args": [
"--directory",
"/ABSOLUTE/PATH/TO/mcp",
"run",
"browser-use.py"
]
}
}
}/ABSOLUTE/PATH/TO/browser-use 프로젝트 디렉토리의 절대 경로로 바꾸세요.
브라우저 구성
에이전트는 다음 기본 설정으로 Google Chrome을 사용하도록 구성되어 있습니다.
개발을 위한 비헤드리스 모드
창 크기: 1280x1100
테스트를 위한 보안 기능 비활성화
녹음 경로: ./tmp/recordings
특징
MCP 도구를 통한 브라우저 자동화
국가 관리 및 계획 역량
대화형 요소 감지 및 조작
구성 가능한 브라우저 컨텍스트
로깅 및 디버깅 지원
용법
에이전트는 두 가지 주요 도구를 제공합니다.
get_planner_state: 현재 브라우저 상태 및 계획 컨텍스트를 검색합니다.execute_actions: 브라우저에서 계획된 작업을 실행합니다.
개발
벌채 반출
이 프로젝트는 다음 구성을 사용하여 Python의 내장 로깅을 사용합니다.
모든 로그는 stderr로 전송됩니다.
사용자 지정 서식:
%(levelname)-8s [%(name)s] %(message)s루트 로거 레벨: INFO
타사 로거 수준: 경고
프로젝트 구조
browser-use.py: 주요 진입점 및 서버 구현tmp/recordings: 브라우저 세션 녹음을 위한 디렉토리uv통해 관리되는 종속성
기여하다
이 프로젝트는 AGI House MCP 해커톤 기간 동안 진행되었습니다. 여러분의 참여를 환영합니다!
특허
이 프로젝트는 MIT 라이선스에 따라 라이선스가 부여되었습니다. 자세한 내용은 라이선스 파일을 참조하세요.
Copyright (c) 2025 하재윤, 하애슐리
본 소프트웨어 및 관련 문서 파일(이하 "소프트웨어")의 사본을 취득한 모든 사람에게 소프트웨어를 제한 없이 거래할 수 있는 권한을 무상으로 부여합니다. 여기에는 소프트웨어 사본을 사용, 복사, 수정, 병합, 게시, 배포, 하위 라이선스 및/또는 판매할 수 있는 권한이 포함되나 이에 국한되지 않으며, 소프트웨어가 제공된 사람에게도 이러한 권한을 부여합니다. 단, 다음 조건에 따라야 합니다.
위의 저작권 고지와 본 허가 고지는 소프트웨어의 모든 사본 또는 실질적인 부분에 포함되어야 합니다.
본 소프트웨어는 상품성, 특정 목적 적합성 및 비침해에 대한 보증을 포함하되 이에 국한되지 않는 명시적 또는 묵시적 보증 없이 "있는 그대로" 제공됩니다. 어떠한 경우에도 저작자 또는 저작권자는 본 소프트웨어 또는 본 소프트웨어의 사용 또는 기타 거래와 관련하여 발생하는 계약, 불법 행위 또는 기타 소송을 포함한 모든 청구, 손해 또는 기타 책임에 대해 책임을 지지 않습니다.
Available Tools
2 toolsexecute_actionsB
Execute actions from the planner state.
Args:
actions: A dictionary containing the planner state and actions in format:
{
"current_state": {
"evaluation_previous_goal": str,
"memory": str,
"next_goal": str
},
"action": [
{"action_name": {"param1": "value1"}},
...
]
}
Note: If the page state changes (new elements appear) during action execution,
the sequence will be interrupted and you'll need to get a new planner state.
| Name | Required | Description | Default |
|---|---|---|---|
| actions | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that execution can be interrupted by page state changes, which is a key behavioral trait, but doesn't cover other aspects like error handling, side effects, or response format. It adds some context but is incomplete for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded with the main purpose, followed by an 'Args' section and a note. The structure is clear, but the note could be more integrated; overall, it's efficient with minimal waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (1 parameter with nested objects, no annotations, no output schema), the description covers the parameter structure well and includes a behavioral note. However, it lacks details on return values, error cases, and full usage context, making it adequate but with gaps for a tool that likely performs mutations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, so the description must compensate. It provides a detailed example of the 'actions' parameter structure, including nested objects and keys like 'current_state' and 'action', which adds significant meaning beyond the schema's generic 'object' type. This effectively documents the parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool 'Execute actions from the planner state', which provides a verb ('Execute') and resource ('actions from the planner state'), but it's vague about what 'actions' specifically entail (e.g., UI interactions, API calls) and doesn't clearly distinguish from the sibling tool 'get_planner_state'. It's not tautological but lacks specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description includes a note about interruption when 'the page state changes', which implies a usage context (e.g., web automation), but it doesn't explicitly state when to use this tool versus alternatives like 'get_planner_state' or provide prerequisites. The guidance is minimal and not comprehensive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_planner_stateA
Get the current browser state and planning context. This tool must be executed before execute_actions tool.
Must return a JSON string in the format:
{
"current_state": {
"evaluation_previous_goal": "Success|Failed|Unknown - Analysis of previous actions",
"memory": "Description of what has been done and what to remember",
"next_goal": "What needs to be done with the next immediate action"
},
"action": [
{"action_name": {"param1": "value1", ...}},
...
]
}
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It effectively describes key behavioral traits: it's a read operation ('Get'), it returns specific structured data (a JSON string with defined format), and it has a prerequisite relationship with another tool. It doesn't cover aspects like error handling or performance, but for a zero-parameter tool with no annotations, this is reasonably comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and usage guideline in the first two sentences, which is good. However, it includes a detailed JSON format specification that might be better suited for an output schema. While this adds value, it makes the description longer than necessary for conciseness, as the output details could be separated into structured data.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 0 parameters, no annotations, and no output schema, the description provides good contextual completeness. It explains the purpose, usage guidelines, and output format in detail. The only gap is the lack of an output schema, but the description compensates by specifying the return format explicitly, making it sufficient for the agent to understand how to use the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters with 100% schema description coverage, so the baseline is 4. The description doesn't need to add parameter information, and it doesn't attempt to, which is appropriate. No parameters are present to document, so this score reflects that the description doesn't introduce confusion or redundancy regarding inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Get the current browser state and planning context.' It specifies the verb ('Get') and resource ('browser state and planning context'), making it easy to understand what the tool does. However, it doesn't explicitly differentiate from its sibling tool 'execute_actions' beyond stating a prerequisite relationship, which is more about usage than purpose distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidance: 'This tool must be executed before execute_actions tool.' It clearly states when to use this tool (as a prerequisite for 'execute_actions') and implies an alternative (use 'execute_actions' after this). This is a strong, directive guideline that helps the agent understand the tool's role in the workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- First observed
execute_actions - First observed
get_planner_state
TDQS
Scored across 2 tools
The two tools have completely distinct purposes: get_planner_state retrieves browser state and planning context, while execute_actions performs actions based on that state. There is no overlap or ambiguity between these functions.
Both tools follow a consistent verb_noun pattern with clear action-oriented names (get_planner_state, execute_actions). The naming convention is uniform and predictable throughout the set.
With only 2 tools for a browser automation server, the surface feels severely limited. While the tools cover a basic planning-execution loop, typical browser automation requires more granular operations like navigation, element interaction, or content extraction.
The toolset provides only a high-level planning/execution abstraction without direct browser manipulation capabilities. There are significant gaps for common browser tasks like navigating to URLs, clicking elements, extracting text, or handling dialogs, which agents would need for robust automation.
Maintenance
Related MCP Connectors
AI-powered web automation. Navigate websites using AI agents for one page or a thousand
AI-powered web automation. Navigate websites using AI agents for one page or a thousand
- openhelmOAuthai.openhelm
Autonomous cloud agent tasks: real browser + your tools, structured evidence-backed results.
AI-powered browser automation — navigate, click, fill forms, and extract data from any website.
Related MCP Servers
- AlicenseAqualityAmaintenanceA Model Context Protocol (MCP) integration that provides Claude Desktop with autonomous browser automation capabilities. This agent enables Claude to interact with web content, manipulate DOM elements, execute JavaScript, and perform API requests.135 npm41TypeScriptMozilla Public 2.0
- AlicenseAqualityAmaintenanceEnables browser automation through the Model Context Protocol, allowing AI agents to control Chrome, Firefox, or Edge for tasks like navigation, clicking, typing, and screenshots.4192 npmMIT
- AlicenseNot gradedqualityDmaintenanceEnables browser automation through the Claude Chrome Extension, allowing agents to navigate websites, fill forms, take screenshots, and debug web apps via standard MCP protocols.1MIT

Browseagent MCPofficial
AlicenseAqualityDmaintenanceEnables AI agents to control web browsers through the Model Context Protocol, supporting navigation, clicking, typing, and screenshots.128 npm1MIT