Browser Agent MCP
MCP 브라우저 에이전트
특징
고급 브라우저 자동화
사용자 정의 가능한 로드 전략을 사용하여 모든 URL로 이동
전체 페이지 또는 요소별 스크린샷 캡처
정확한 DOM 상호 작용 수행(클릭, 채우기, 선택, 호버)
콘솔 로그 캡처를 사용하여 브라우저 컨텍스트에서 임의의 JavaScript 실행
강력한 API 클라이언트
HTTP 요청(GET, POST, PUT, PATCH, DELETE) 실행
요청 헤더 및 본문 콘텐츠 구성
JSON 포맷으로 응답 데이터 처리
자세한 피드백을 통한 오류 처리
MCP 리소스 관리
리소스로 브라우저 콘솔 로그에 액세스
MCP 리소스 인터페이스를 통해 스크린샷 검색
headful 브라우저 인스턴스와의 지속적인 세션
AI 에이전트 기능
복잡한 작업을 위해 여러 브라우저 작업을 연결합니다.
지능형 오류 복구를 통해 여러 단계의 지침을 따르세요
자연어 지침을 통한 기술 작업 자동화
Related MCP server: Browserbeam MCP Server
데모
비디오의 해당 섹션으로 이동하려면 타임스탬프를 클릭하세요.
00:00 - MCP에 대한 Google 검색
Google 홈페이지로 이동하여 "Model Context Protocol"을 검색합니다. MCP 통합을 사용하여 Claude Desktop에서 기본 웹 검색을 수행하고 결과를 처리하는 방법을 보여줍니다.
00:33 - 스크린샷 캡처
사용자 지정 파일 이름으로 검색 결과의 스크린샷을 찍어 Finder에 표시하는 방법입니다. Claude가 브라우저 자동화 중에 웹 페이지의 시각적 콘텐츠를 캡처하고 저장하는 방법을 보여줍니다.
01:00 - 위키피디아 검색
Wikipedia.org로 이동하여 "모델 컨텍스트 프로토콜"을 검색합니다. MCP 통합을 통해 Claude가 다양한 웹사이트 및 검색 기능과 상호 작용하는 능력을 보여줍니다.
01:38 - 드롭다운 메뉴 상호작용 I
테스트 웹사이트(the-internet.herokuapp.com/dropdown)로 이동하여 드롭다운 메뉴에서 "옵션 1"을 선택합니다. Claude가 양식 요소와 상호 작용하고 항목을 선택하는 능력을 보여줍니다.
01:56 - 드롭다운 메뉴 상호작용 II
같은 드롭다운 메뉴에서 선택 항목을 "옵션 2"로 변경합니다. 클로드가 같은 양식 요소를 여러 번 조작하고 다른 선택을 할 수 있는 능력을 보여줍니다.
02:09 - 로그인 양식 완료
로그인 페이지(the-internet.herokuapp.com/login)로 이동하여 사용자 이름 필드에 "tomsmith"를, 비밀번호 필드에 "SuperSecretPassword!"를 입력합니다. 양식 작성 자동화를 보여줍니다.
02:28 - 로그인 제출
로그인 자격 증명을 제출하고 인증 절차를 완료합니다. Claude가 양식 제출을 트리거하고 여러 단계의 절차를 탐색하는 능력을 보여줍니다.
02:36 - API 요청 실행
JSONPlaceholder API 엔드포인트에 GET 요청을 수행합니다. Claude가 직접 API 호출을 수행하고 MCP 통합을 통해 반환된 데이터를 처리하는 능력을 보여줍니다.
요구 사항
Node.js 16 이상
클로드 데스크탑
극작가 종속성
브라우저 지원
지엑스피1
이 패키지에는 Playwright와 브라우저 자동화 실행에 필요한 종속성이 포함되어 있습니다. npm install 실행하면 필요한 Playwright 종속성이 설치됩니다. 이 패키지는 다음 브라우저를 지원합니다.
크롬(기본)
파이어폭스
마이크로소프트 엣지
WebKit(Safari 엔진)
브라우저 유형을 처음 사용하면 Playwright가 필요에 따라 해당 브라우저 드라이버를 자동으로 설치합니다. 다음 명령을 사용하여 수동으로 설치할 수도 있습니다.
npx playwright install chrome
npx playwright install firefox
npx playwright install webkit
npx playwright install msedgeSafari 관련 참고 사항 : Playwright는 Safari 브라우저를 직접 지원하지 않습니다. 대신 Safari를 구동하는 브라우저 엔진인 WebKit을 사용합니다.
Edge 관련 참고 사항 : 브라우저 유형으로 Edge를 선택하면 에이전트가 실제로 Microsoft Edge(Chromium이 아님)를 실행합니다. 기술적으로 Playwright에서 Edge는 Chromium 기반 브라우저 인스턴스에 'msedge' 채널 매개변수를 사용하여 실행됩니다.
설치
수동 설치
이 저장소를 복제하거나 다운로드하세요:
git clone https://github.com/imprvhub/mcp-browser-agent
cd mcp-browser-agent종속성 설치:
npm install프로젝트를 빌드하세요:
npm run buildMCP 서버 실행
MCP 서버를 실행하는 방법은 두 가지가 있습니다.
옵션 1: 수동 실행
터미널이나 명령 프롬프트를 엽니다
프로젝트 디렉토리로 이동합니다
서버를 직접 실행합니다.
node dist/index.jsClaude Desktop을 사용하는 동안 이 터미널 창을 열어 두세요. 터미널을 닫을 때까지 서버가 실행됩니다.
옵션 2: Claude Desktop으로 자동 시작(일반 사용 시 권장)
Claude Desktop은 필요 시 MCP 서버를 자동으로 시작할 수 있습니다. 설정 방법은 다음과 같습니다.
구성
Claude Desktop 구성 파일은 다음 위치에 있습니다.
macOS :
~/Library/Application Support/Claude/claude_desktop_config.json윈도우 :
%APPDATA%\Claude\claude_desktop_config.json리눅스 :
~/.config/Claude/claude_desktop_config.json
이 파일을 편집하여 브라우저 에이전트 MCP 구성을 추가하세요. 파일이 없으면 새로 만드세요.
{
"mcpServers": {
"browserAgent": {
"command": "node",
"args": ["ABSOLUTE_PATH_TO_DIRECTORY/mcp-browser-agent/dist/index.js",
"--browser",
"chrome"
]
}
}
}중요 : ABSOLUTE_PATH_TO_DIRECTORY MCP를 설치한 전체 절대 경로 로 바꾸세요.
macOS/Linux 예:
/Users/username/mcp-browser-agentWindows 예:
C:\\Users\\username\\mcp-browser-agent
이미 다른 MCP를 구성했다면 "mcpServers" 객체 안에 "browserAgent" 섹션을 추가하기만 하면 됩니다. 다음은 여러 MCP를 사용한 구성의 예입니다.
{
"mcpServers": {
"otherMcp1": {
"command": "...",
"args": ["..."]
},
"otherMcp2": {
"command": "...",
"args": ["..."]
},
"browserAgent": {
"command": "node",
"args": [
"ABSOLUTE_PATH_TO_DIRECTORY/mcp-browser-agent/dist/index.js",
"--browser",
"chrome"
]
}
}
}브라우저 선택
MCP 브라우저 에이전트는 여러 브라우저 유형을 지원합니다. 기본적으로 Chrome을 사용하지만, 여러 가지 방법으로 다른 브라우저를 지정할 수 있습니다.
옵션 1: 구성 파일
홈 디렉토리에서 .mcp_browser_agent_config.json 파일을 만들거나 편집합니다.
{
"browserType": "chrome"
}browserType 에 지원되는 값은 다음과 같습니다.
chrome- 설치된 Chrome을 사용합니다(기본값)firefox- Firefox 'Nightly' 브라우저를 사용합니다webkit- WebKit 엔진을 사용합니다(참고: 이것은 Safari 자체가 아니라 Safari를 구동하는 WebKit 렌더링 엔진입니다)edge- Microsoft Edge를 사용합니다
Safari 관련 참고 사항 : Playwright는 Safari 브라우저를 직접 지원하지 않습니다. 대신 Safari를 구동하는 브라우저 엔진인 WebKit을 사용합니다. Playwright의 WebKit 구현은 유사한 기능을 제공하지만 Safari 브라우저 환경과 동일하지는 않습니다.
옵션 2: 명령줄 인수
MCP 서버를 수동으로 시작할 때 브라우저 유형을 지정할 수 있습니다.
node dist/index.js --browser firefox옵션 3: 환경 변수
MCP_BROWSER_TYPE 환경 변수를 설정합니다.
MCP_BROWSER_TYPE=firefox node dist/index.js옵션 4: Claude 데스크톱 구성
Claude Desktop의 claude_desktop_config.json 에서 MCP를 구성할 때 브라우저 유형을 지정할 수 있습니다.
{
"mcpServers": {
"browserAgent": {
"command": "node",
"args": [
"ABSOLUTE_PATH_TO_DIRECTORY/mcp-browser-agent/dist/index.js",
"--browser",
"chrome"
]
}
}
}기술 구현
MCP 브라우저 에이전트는 모델 컨텍스트 프로토콜(Model Context Protocol)을 기반으로 구축되어 Claude가 Playwright를 통해 헤드풀 브라우저와 상호 작용할 수 있도록 합니다. 구현은 네 가지 주요 구성 요소로 구성됩니다.
서버(index.ts)
Model Context Protocol 표준 프로토콜을 사용하여 MCP 서버를 초기화합니다.
도구 및 리소스에 대한 서버 기능을 구성합니다.
stdio 전송을 통해 Claude와 통신을 설정합니다.
도구 레지스트리(tools.ts)
브라우저 및 API 도구 스키마를 정의합니다.
매개변수, 검증 규칙 및 설명을 지정합니다.
Claude의 발견을 위해 MCP 서버에 도구를 등록합니다.
요청 핸들러(handlers.ts)
도구 및 리소스에 대한 MCP 프로토콜 요청을 관리합니다.
브라우저 로그와 스크린샷을 쿼리 가능한 리소스로 노출합니다.
적절한 핸들러에 도구 실행 요청을 라우팅합니다.
실행자(executor.ts)
브라우저 및 API 클라이언트 수명 주기를 관리합니다.
Playwright를 사용하여 브라우저 자동화 기능을 구현합니다.
적절한 오류 처리 및 응답 구문 분석을 통해 API 요청을 처리합니다.
명령 간에 상태 저장 브라우저 세션을 유지합니다.
에이전트 기능
기본 통합과 달리 MCP Browser Agent는 다음과 같은 기능을 통해 진정한 AI 에이전트로 기능합니다.
여러 명령에 걸쳐 지속적인 브라우저 상태 유지
디버깅을 위한 자세한 콘솔 로그 캡처
참고 및 검토를 위해 스크린샷 저장
복잡한 상호작용 시퀀스 관리
복구를 위한 자세한 오류 정보 제공
복잡한 워크플로에 대한 체인 작업 지원
사용 가능한 도구
브라우저 도구
도구 이름 | 설명 | 매개변수 |
| URL로 이동 |
|
| 스크린샷 캡처 |
|
| 클릭 요소 |
|
| 양식 입력 |
|
| 드롭다운 옵션 선택 |
|
| 요소 위에 마우스를 올려 놓으세요 |
|
| JavaScript 실행 |
|
API 도구
도구 이름 | 설명 | 매개변수 |
| GET 요청 |
|
| POST 요청 |
|
| PUT 요청 |
|
| 패치 요청 |
|
| 삭제 요청 |
|
리소스 액세스
MCP 브라우저 에이전트는 다음 리소스를 제공합니다.
browser://logs- 브라우저 콘솔 로그에 액세스합니다.screenshot://[name]- 이름으로 스크린샷에 접근합니다.
사용 예
다음은 Claude와 함께 MCP 브라우저 에이전트를 사용하는 방법에 대한 몇 가지 현실적인 예입니다.
기본 브라우저 탐색
Navigate to the Google homepage at https://www.google.comTake a screenshot of the current page and name it "google-homepage"Type "weather forecast" in the search box간단한 상호작용
Navigate to https://www.wikipedia.org and search for "Model Context Protocol"Go to https://the-internet.herokuapp.com/dropdown and select the option "Option 1" from the dropdown기본 양식 작성
Navigate to https://the-internet.herokuapp.com/login and fill in the username field with "tomsmith" and the password field with "SuperSecretPassword!"Go to https://the-internet.herokuapp.com/login, fill in the username and password fields, then click the login button간단한 JavaScript 실행
Go to https://example.com and execute a JavaScript script to return the page titleNavigate to https://www.google.com and execute a JavaScript script to count the number of links on the page기본 API 요청
Perform a GET request to https://jsonplaceholder.typicode.com/todos/1Make a POST request to https://jsonplaceholder.typicode.com/posts with appropriate JSON data이러한 예는 MCP 브라우저 에이전트의 실제 기능을 나타내며 현재 상태에서 무엇을 달성할 수 있는지에 대해 보다 현실적인 정보를 제공합니다.
문제 해결
"서버 연결 끊김" 오류
Claude Desktop에서 "MCP 브라우저 에이전트: 서버 연결 끊김" 오류가 표시되는 경우:
서버가 실행 중인지 확인하세요 .
터미널을 열고 프로젝트 디렉토리에서
node dist/index.js수동으로 실행합니다.서버가 성공적으로 시작되면 이 터미널을 열어둔 채로 Claude를 사용하세요.
구성을 확인하세요 :
claude_desktop_config.json의 절대 경로가 시스템에 맞는지 확인하세요.Windows 경로에 이중 백슬래시(
\\)를 사용했는지 다시 한 번 확인하세요.파일 시스템의 루트에서 전체 경로를 사용하고 있는지 확인하세요.
브라우저가 나타나지 않습니다
브라우저가 실행되지 않거나 보이지 않는 경우:
지정된 브라우저가 설치되어 있는지 확인하세요
시스템에 브라우저(Chrome, Firefox, Edge 또는 Safari/WebKit)가 설치되어 있는지 확인하세요.
브라우저 드라이버는 Playwright에 의해 자동으로 처리됩니다.
서버와 Claude Desktop을 다시 시작하세요
서버를 실행 중일 수 있는 기존 노드 프로세스를 모두 종료합니다.
새로운 연결을 설정하려면 Claude Desktop을 다시 시작하세요.
브라우저 프로세스가 제대로 닫히지 않습니다
Chromium 및 Chrome 브라우저에는 사용 후 프로세스가 제대로 종료되지 않는 알려진 문제가 있습니다. 이 문제가 발생하는 경우:
브라우저 프로세스를 수동으로 닫습니다 .
Windows : Ctrl+Shift+Esc를 눌러 작업 관리자를 열고 Chrome/Chromium 프로세스를 찾아 종료합니다.
macOS : 활동 모니터를 엽니다(응용 프로그램 > 유틸리티 > 활동 모니터). Chrome/Chromium 프로세스를 찾은 다음 X를 클릭하여 종료합니다.
Linux :
ps aux | grep chrome또는ps aux | grep chromium실행하여 프로세스를 찾은 다음kill <PID>실행하여 프로세스를 종료합니다.
브라우저 호환성에 대한 참고사항 :
이 문제는 주로 Chromium 및 Chrome에서 관찰되었습니다.
Firefox와 Playwright의 내장 브라우저에서는 일반적으로 이 문제가 발생하지 않습니다.
[!주의] 이 MCP 통합은 Playwright를 기반으로 구축되었으며, Playwright의 작동에 영향을 미칠 수 있는 알려진 문제와 버그가 있습니다. 브라우저 자동화 관련 문제가 발생하면 Playwright의 GitHub 문제 게시판에 보고해 주세요. Playwright 팀은 이러한 문제를 해결하기 위해 끊임없이 노력하고 있지만, 이 에이전트는 이러한 제한에도 불구하고 Claude Desktop에서 브라우저 자동화 기능을 위한 기반을 제공합니다.
개발
프로젝트 구조
src/index.ts: 메인 진입점 및 MCP 서버 초기화src/tools.ts: 도구 스키마 및 등록src/handlers.ts: 도구 및 리소스에 대한 MCP 요청 핸들러src/executor.ts: Playwright를 이용한 도구 구현 로직
건물
npm run build변화를 지켜보다
npm run watch테스트
이 프로젝트에는 핵심 기능과 브라우저 처리를 검증하는 테스트가 포함되어 있습니다.
npm test # Run tests
npm run test:watch # Watch mode
npm run test:coverage # Coverage report테스트는 구성 무결성, 브라우저 자동화 기능, 오류 처리 및 프로세스 정리를 검증합니다. 테스트 모음은 특히 Chrome/Chromium 종료와 관련된 알려진 문제로 인해 발생하는 브라우저 프로세스의 적절한 처리에 중점을 둡니다.
보안 고려 사항
[!중요] 이 MCP 통합은 Claude에게 자율적인 브라우저 제어 기능을 제공합니다. 금지된 사용, 보안 관련 영향 및 모범 사례에 대한 중요 정보는 보안 정책을 검토하십시오.
MCP 브라우저 에이전트는 합법적인 자동화 작업을 위해 설계되었지만, 잠재적으로 오용될 수 있습니다. 사용자는 모든 관련 법률, 서비스 약관 및 윤리 지침을 준수하여 사용할 책임이 있습니다. 자세한 내용은 자세한 보안 정책을 참조하십시오.
기여하다
MCP 브라우저 에이전트에 대한 기여를 환영합니다! 다음은 여러분이 도움을 줄 수 있는 몇 가지 분야입니다.
새로운 브라우저 자동화 기능 추가
오류 처리 및 복구 개선
스크린샷 및 리소스 관리 강화
유용한 워크플로 및 예제 만들기
복잡한 작업에 대한 성능 최적화
특허
이 프로젝트는 Mozilla Public License 2.0에 따라 라이선스가 부여되었습니다. 자세한 내용은 LICENSE 파일을 참조하세요.
관련 링크
Available Tools
13 toolsapi_deleteB
Perform a DELETE request to an API endpoint
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | API endpoint URL | |
| headers | No | Request headers |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden but only states the action without disclosing side effects, idempotency, authentication needs, rate limits, or return format. This is minimal for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It is appropriately concise for a simple tool, though slightly more context could be included without losing efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 params, no output schema), the description covers the basic action but omits expected return values, error conditions, or usage scope, leaving the description marginally adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptive parameter names (url, headers). The description adds no additional meaning beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool performs a DELETE request to an API endpoint, using a specific verb and resource that distinguishes it from sibling tools (api_get, api_patch, etc.) which handle other HTTP methods.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like api_get or api_post. The description merely repeats the method, missing explicit when/when-not context or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
api_getC
Perform a GET request to an API endpoint
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | API endpoint URL | |
| headers | No | Request headers |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It fails to mention that GET is typically safe and idempotent, how errors are handled, or whether redirects are followed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single clear sentence with no unnecessary words. It efficiently conveys the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple but lacks information about return values, error handling, authentication requirements, or default behavior. The description is too minimal for an agent to use confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions for both parameters. The tool description adds no extra meaning beyond the schema, so baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Perform a GET request to an API endpoint', which identifies the HTTP method and the action. It distinguishes from sibling tools like api_post and api_delete by the method name, but lacks mention of read-only nature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus siblings. The description does not specify that GET should be used for retrieving data, nor does it mention alternatives for modifying or deleting resources.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
api_patchB
Perform a PATCH request to an API endpoint
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | API endpoint URL | |
| data | Yes | Request body data (JSON string) | |
| headers | No | Request headers |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It only states it performs a PATCH request, implying mutation, but does not disclose side effects, authentication needs, rate limits, or response behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no unnecessary words. However, it is very brief and could benefit from additional context without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is incomplete given the tool's complexity. No output schema exists, and the description does not explain return values or error handling. Sibling tools suggest it is part of an HTTP client set, but the description lacks depth.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description does not add any additional meaning beyond what is in the input schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it performs a PATCH request to an API endpoint, which is a specific verb and resource. This distinguishes it from sibling tools like api_get, api_post, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus other HTTP methods (e.g., POST, PUT) or alternatives. The context signals show siblings, but the description offers no differentiation advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
api_postB
Perform a POST request to an API endpoint
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | API endpoint URL | |
| data | Yes | Request body data (JSON string) | |
| headers | No | Request headers |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description must disclose behavior. It fails to mention typical POST behavior (resource creation), data validation, or error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence is concise and front-loaded, but it is too brief and lacks substance for a practical tool description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple HTTP tool, description should mention typical use (e.g., 'sends data to URL'). Schema covers parameters but context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (all parameters described), so baseline is 3. Description adds no additional meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states 'POST request' verb and 'API endpoint' resource, clearly distinguishing from sibling tools like api_get, api_put, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, no mention of prerequisites or context such as authentication or data format.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
api_putC
Perform a PUT request to an API endpoint
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | API endpoint URL | |
| data | Yes | Request body data (JSON string) | |
| headers | No | Request headers |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description only states 'PUT request' implying mutation, but omits critical details like idempotency, side effects, authentication needs, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence is concise and front-loaded, but does not add meaningful content beyond the tool name; minimal but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no behavioral details; for a tool with 3 parameters and nested objects, the description is insufficient to fully understand usage and return values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers all 3 parameters with descriptions; description adds no extra meaning beyond what schema already provides, meeting the baseline for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the action ('Perform a PUT request') and resource ('API endpoint'), but lacks differentiation from sibling tools like api_patch or api_post, which perform similar HTTP methods.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use PUT versus other HTTP methods (e.g., PATCH for partial updates, POST for creation). Does not mention idempotency or replacement semantics.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_clickC
Click an element on the page
| Name | Required | Description | Default |
|---|---|---|---|
| selector | Yes | CSS selector for element to click |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description does not disclose behavioral traits beyond the basic action, such as potential side effects (e.g., navigation, page changes) or element visibility requirements. No annotations exist to compensate for this lack of detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no unnecessary words. It is appropriately sized for a simple action, though it could benefit from slight elaboration without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the sibling tools and the single parameter, the description is incomplete. It fails to mention crucial context like element visibility, clicks causing navigation, or waiting behavior, leaving the agent with insufficient information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema fully describes the 'selector' parameter with a clear definition. The description adds no extra meaning beyond the schema, which is acceptable given 100% coverage, but it does not enhance understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Click an element') and the resource ('on the page'), distinguishing it from sibling tools like browser_fill or browser_hover. However, it could benefit from specifying that it operates within the current page context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives, such as when to click vs. hover, or prerequisites like page navigation. The description lacks context for proper selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_evaluateB
Execute JavaScript in the browser context
| Name | Required | Description | Default |
|---|---|---|---|
| script | Yes | JavaScript code to execute |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must fully disclose behavior. It mentions script execution but omits critical details: return value, side effects, permissions, sandboxing, or error handling. This is insufficient for a potentially powerful tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with no unnecessary words. It is optimally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema, so the description should explain return values or behavior. It does not. Additionally, it lacks details on execution context (e.g., async support, timeout). This leaves the agent underinformed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required parameter 'script' has a description in the schema. The tool description adds no additional meaning beyond what the schema already provides, making it adequate but not additive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Execute JavaScript in the browser context' clearly states the action and resource. It is specific and distinct from sibling tools like browser_click or api_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives (e.g., browser_navigate or API calls). An agent receives no context for decision-making.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_fillB
Fill a form input with text
| Name | Required | Description | Default |
|---|---|---|---|
| value | Yes | Text to enter in the field | |
| selector | Yes | CSS selector for input field |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description should disclose behavioral details, but it only says 'fill' without specifying whether it overwrites existing text, waits for elements, or handles disabled fields. This leaves significant behavioral ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence, front-loading the action. However, it could be slightly expanded with useful context while remaining brief.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity of the tool (2 params, no output schema), the description is adequate but lacks details on behavior like clearing the field or submission. Sibling tools exist for other form actions, but no comparative guidance is provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema covers 100% of parameters with descriptions (selector and value), so baseline is 3. The description adds no additional semantic information beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Fill a form input with text' clearly states the action (fill) and the target (form input), distinguishing it from sibling tools like browser_click or browser_select which perform different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as browser_evaluate for setting values or browser_click for activation. There is no mention of prerequisites like element visibility or state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_hoverC
Hover over an element on the page
| Name | Required | Description | Default |
|---|---|---|---|
| selector | Yes | CSS selector for element to hover over |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits. It only states 'hover over an element' without explaining whether it triggers JavaScript events, waits for any transitions, or is safe. Essential behavioral context is missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise at 6 words and front-loaded. It is not verbose, but the brevity may sacrifice necessary detail. It earns its place but could be slightly more informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (1 parameter, no output schema), the description is incomplete. It does not mention the return value (likely void or success), side effects, or behavior after hovering. Sibling tools suggest a sequence of actions, but this tool's role is under-described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single parameter 'selector' described as 'CSS selector for element to hover over'. The description adds no additional meaning beyond what the schema already provides, so it meets the baseline but does not enhance understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Hover over an element on the page' clearly states the action (hover) and target (element on page). It differentiates from siblings like browser_click and browser_fill. However, it could be more specific about the effect (e.g., triggering hover state) but is not a tautology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like browser_click or browser_evaluate. The description lacks any context about prerequisites, typical scenarios, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_screenshotB
Capture a screenshot of the current page or a specific element
| Name | Required | Description | Default |
|---|---|---|---|
| mask | No | Selectors for elements to mask | |
| name | Yes | Identifier for the screenshot | |
| fullPage | No | Capture full page height | |
| savePath | No | Path to save screenshot (default: user's Downloads folder) | |
| selector | No | CSS selector for element to capture |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description does not disclose behavioral traits such as whether the capture affects the page state, any authorization requirements, or rate limits. It only states the action, leaving the agent without important context about side effects or constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single well-formed sentence that directly states the tool's purpose. It wastes no words, though it could benefit from slightly more detail without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is brief and lacks context about return values (e.g., image path), default behavior, or how the tool interacts with other browser tools. Given the five parameters and no output schema, the description should provide more operational context to be fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Since schema description coverage is 100%, the baseline is 3. The description does not add additional meaning beyond the parameter names and schema descriptions; for example, it doesn't explain when to use fullPage or mask options more concretely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool captures a screenshot of the current page or a specific element, which is a specific verb-resource combination. It distinguishes itself from sibling tools like browser_navigate or browser_click by focusing on capture rather than navigation or element interaction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used when a screenshot is needed, but it provides no explicit guidance on when to use it versus alternative methods (e.g., browser_evaluate for custom captures) or any exclusions (e.g., not for video capture).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_selectC
Select an option from a dropdown menu
| Name | Required | Description | Default |
|---|---|---|---|
| value | Yes | Value or label to select | |
| selector | Yes | CSS selector for select element |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description does not disclose behavioral traits such as whether it waits for options to load, supports custom dropdowns, triggers events, or requires scrolling. The description is too minimal to inform the agent of important behaviors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that is front-loaded. However, it sacrifices completeness for brevity. It earns a 4 for being efficient, but could include more key details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that no annotations or output schema exist, the description is insufficiently complete. It does not explain return values, constraints, or behavior for complex dropdowns, leaving gaps for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema describes both parameters (selector and value) with 100% coverage. The description adds no additional meaning beyond the schema, so baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it selects an option from a dropdown menu, using a specific verb and resource. It distinguishes from sibling tools like browser_fill (text input) and browser_click (clicking), though it could be more precise by specifying HTML <select> elements.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives (e.g., browser_click for custom dropdowns) or any prerequisites. No exclusions or context are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_set_viewportA
Change the browser's viewport size and scale factor
| Name | Required | Description | Default |
|---|---|---|---|
| width | No | Viewport width in pixels | |
| height | No | Viewport height in pixels | |
| deviceScaleFactor | No | Device scale factor (affects how content is scaled) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description should disclose behavioral traits. It only states what the tool changes, but not side effects (e.g., impact on screenshots, persistence across navigation). Lacks important context for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, no fluff. Every word is relevant. Excellent conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequately describes the core function, but lacks details on required fields, defaults, or behavioral context. Acceptable for a simple tool, but could be more complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with meaningful descriptions for each parameter. The description adds no extra meaning beyond the schema, justifying the baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the action 'Change' and the target 'browser's viewport size and scale factor'. It is specific and distinct from sibling tools like browser_navigate or browser_click.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance is provided. The description implies usage for adjusting viewport, but does not mention alternatives or exclusions. Minimal viable score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v1.0.0- First observed
api_delete - First observed
api_get - First observed
api_patch - First observed
api_post - First observed
api_put - First observed
browser_click - First observed
browser_evaluate - First observed
browser_fill - First observed
browser_hover - First observed
browser_navigate - First observed
browser_screenshot - First observed
browser_select - First observed
browser_set_viewport
TDQS
Scored across 13 tools
Each tool has a clearly distinct purpose: API tools are differentiated by HTTP method, and browser tools cover unique interactions like clicking, hovering, filling, etc. No two tools overlap in function.
All tools follow a consistent '<domain>_<action>' pattern, with 'api_' prefix for HTTP methods and 'browser_' prefix for browser actions. Naming is unambiguous and predictable.
13 tools is well-scoped for a browser automation and API testing server. Each tool covers a fundamental operation without unnecessary bloat or gaps.
Core browser interactions (navigation, clicking, form filling, selecting, screenshot) and all major HTTP methods are covered. Minor omissions like file upload or wait-for-element are acceptable for this scope.
Maintenance
Related MCP Connectors
AI-powered web automation. Navigate websites using AI agents for one page or a thousand
AI-powered web automation. Navigate websites using AI agents for one page or a thousand
AI-powered browser automation — navigate, click, fill forms, and extract data from any website.
An agent-first office suite Claude & ChatGPT read and write over one MCP URL.
Related MCP Servers
- AlicenseBqualityDmaintenanceA browser automation agent that enables Claude to interact with web browsers through the Model Context Protocol, allowing for actions like navigating websites, manipulating elements, and managing browser state.29MIT
- AlicenseNot gradedqualityDmaintenanceEnables real browser automation as tools in Cursor, Claude Desktop, Windsurf, and any MCP-compatible client, allowing AI agents to interact with web pages through natural language.17 npmMIT
- FlicenseNot gradedqualityBmaintenanceEnables Claude Code to control a real browser using AI for web scraping, competitive intelligence, and UX auditing through the MCP protocol.-
- AlicenseNot gradedqualityDmaintenanceEnables browser automation through the Claude Chrome Extension, allowing agents to navigate websites, fill forms, take screenshots, and debug web apps via standard MCP protocols.1MIT