Browser Agent MCP
MCP ブラウザエージェント
特徴
高度なブラウザ自動化
カスタマイズ可能なロード戦略で任意の URL にナビゲートします
全ページまたは特定の要素のスクリーンショットをキャプチャする
正確な DOM インタラクション (クリック、塗りつぶし、選択、ホバー) を実行します
コンソールログをキャプチャしてブラウザコンテキストで任意のJavaScriptを実行する
強力なAPIクライアント
HTTP リクエスト (GET、POST、PUT、PATCH、DELETE) を実行する
リクエストヘッダーと本文の内容を構成する
JSON形式で応答データを処理する
詳細なフィードバックによるエラー処理
MCP リソース管理
ブラウザコンソールのログをリソースとしてアクセスする
MCP リソース インターフェースを通じてスクリーンショットを取得する
ヘッドフルブラウザインスタンスによる永続セッション
AIエージェントの機能
複雑なタスクのために複数のブラウザ操作を連鎖させる
インテリジェントなエラー回復機能を使用して、複数の手順の指示に従います
自然言語指示による技術タスクの自動化
Related MCP server: Browserbeam MCP Server
デモ
タイムスタンプをクリックすると、ビデオのそのセクションにジャンプします。
00:00 - MCPをGoogleで検索
Googleホームページにアクセスし、「Model Context Protocol」を検索します。Claude DesktopでMCP統合を使用して基本的なウェブ検索を実行し、結果を処理するデモです。
00:33 -スクリーンショットキャプチャ
検索結果のスクリーンショットをカスタムファイル名で撮影し、Finderで表示します。ブラウザ自動化中に、ClaudeがWebページからビジュアルコンテンツをキャプチャして保存する方法を紹介します。
01:00 - Wikipedia検索
Wikipedia.org にアクセスし、「Model Context Protocol」を検索します。Claude が MCP 統合を通じてさまざまなウェブサイトやその検索機能とやり取りする様子を示します。
01:38 -ドロップダウンメニューインタラクション I
テストウェブサイト(the-internet.herokuapp.com/dropdown)に移動し、ドロップダウンメニューから「オプション1」を選択します。Claudeがフォーム要素を操作して選択する能力を示します。
01:56 -ドロップダウンメニューインタラクション II
同じドロップダウンメニューから「オプション2」を選択に変更します。Claude が同じフォーム要素を複数回操作し、異なる選択を行う能力を示しています。
02:09 -ログインフォームの完了
ログインページ(the-internet.herokuapp.com/login)に移動し、ユーザー名フィールドに「tomsmith」、パスワードフィールドに「SuperSecretPassword!」を入力します。フォーム入力の自動化のデモです。
02:28 -ログイン送信
ログイン資格情報を送信し、認証プロセスを完了します。Claude がフォームの送信をトリガーし、複数のステップからなるプロセスをナビゲートする能力を示します。
02:36 - APIリクエストの実行
JSONPlaceholder APIエンドポイントへのGETリクエストを実行します。Claudeが直接API呼び出しを行い、MCP統合を通じて返されたデータを処理できる能力を示します。
要件
Node.js 16以上
クロードデスクトップ
Playwrightの依存関係
ブラウザのサポート
npm init playwright@latestこのパッケージには、Playwrightとブラウザ自動化の実行に必要な依存関係が含まれています。npm npm install実行すると、必要なPlaywrightの依存関係がインストールされます。このパッケージは以下のブラウザをサポートしています。
Chrome(デフォルト)
ファイアフォックス
マイクロソフトエッジ
WebKit(Safariエンジン)
Playwright は、ブラウザの種類を初めて使用する際に、必要に応じて対応するブラウザドライバを自動的にインストールします。また、以下のコマンドで手動でインストールすることもできます。
npx playwright install chrome
npx playwright install firefox
npx playwright install webkit
npx playwright install msedgeSafariに関する注意:PlaywrightはSafariブラウザを直接サポートしていません。代わりに、Safariを動かすブラウザエンジンであるWebKitを使用しています。
Edgeに関する注意:ブラウザの種類としてEdgeを選択した場合、エージェントは実際にはChromiumではなくMicrosoft Edgeを起動します。技術的には、Playwrightでは、Microsoft EdgeはChromiumをベースにしているため、Chromiumブラウザインスタンスと「msedge」チャネルパラメータを使用してEdgeが起動されます。
インストール
手動でインストールする
このリポジトリをクローンまたはダウンロードします:
git clone https://github.com/imprvhub/mcp-browser-agent
cd mcp-browser-agent依存関係をインストールします:
npm installプロジェクトをビルドします。
npm run buildMCPサーバーの実行
MCP サーバーを実行するには 2 つの方法があります。
オプション1: 手動で実行する
ターミナルまたはコマンドプロンプトを開きます
プロジェクトディレクトリに移動する
サーバーを直接実行します。
node dist/index.jsClaude Desktopの使用中は、このターミナルウィンドウを開いたままにしてください。ターミナルを閉じるまでサーバーは稼働し続けます。
オプション 2: Claude Desktop で自動起動 (通常の使用に推奨)
Claude Desktopは、必要に応じてMCPサーバーを自動的に起動できます。設定手順は次のとおりです。
構成
Claude Desktop の構成ファイルは次の場所にあります。
macOS :
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows :
%APPDATA%\Claude\claude_desktop_config.jsonLinux :
~/.config/Claude/claude_desktop_config.json
このファイルを編集して、ブラウザエージェントMCP設定を追加します。ファイルが存在しない場合は作成してください。
{
"mcpServers": {
"browserAgent": {
"command": "node",
"args": ["ABSOLUTE_PATH_TO_DIRECTORY/mcp-browser-agent/dist/index.js",
"--browser",
"chrome"
]
}
}
}重要: ABSOLUTE_PATH_TO_DIRECTORY 、MCP をインストールした完全な絶対パスに置き換えてください。
macOS/Linuxの例:
/Users/username/mcp-browser-agentWindows の例:
C:\\Users\\username\\mcp-browser-agent
既に他のMCPを設定している場合は、「mcpServers」オブジェクト内に「browserAgent」セクションを追加するだけです。複数のMCPを設定する例を以下に示します。
{
"mcpServers": {
"otherMcp1": {
"command": "...",
"args": ["..."]
},
"otherMcp2": {
"command": "...",
"args": ["..."]
},
"browserAgent": {
"command": "node",
"args": [
"ABSOLUTE_PATH_TO_DIRECTORY/mcp-browser-agent/dist/index.js",
"--browser",
"chrome"
]
}
}
}ブラウザの選択
MCPブラウザエージェントは複数のブラウザタイプをサポートしています。デフォルトではChromeが使用されますが、以下の方法で別のブラウザを指定することもできます。
オプション1: 構成ファイル
ホームディレクトリに.mcp_browser_agent_config.jsonファイルを作成または編集します。
{
"browserType": "chrome"
}browserTypeでサポートされている値は次のとおりです。
chrome- インストールされているChromeを使用します(デフォルト)firefox- Firefox 'Nightly' ブラウザを使用webkit- WebKit エンジンを使用します (注: これは Safari 自体ではなく、Safari を動かす WebKit レンダリング エンジンです)edge- Microsoft Edge を使用
Safariに関する注意:PlaywrightはSafariブラウザを直接サポートしていません。代わりに、Safariを動かすブラウザエンジンであるWebKitを使用しています。PlaywrightのWebKit実装はSafariブラウザと同様の機能を提供しますが、Safariブラウザのエクスペリエンスと完全に同一ではありません。
オプション2: コマンドライン引数
MCP サーバーを手動で起動する場合は、ブラウザの種類を指定できます。
node dist/index.js --browser firefoxオプション3: 環境変数
MCP_BROWSER_TYPE環境変数を設定します。
MCP_BROWSER_TYPE=firefox node dist/index.jsオプション4: クロードデスクトップ構成
Claude Desktop のclaude_desktop_config.jsonで MCP を構成するときに、ブラウザの種類を指定できます。
{
"mcpServers": {
"browserAgent": {
"command": "node",
"args": [
"ABSOLUTE_PATH_TO_DIRECTORY/mcp-browser-agent/dist/index.js",
"--browser",
"chrome"
]
}
}
}技術的実装
MCPブラウザエージェントはモデルコンテキストプロトコル(MCP)上に構築されており、ClaudeがPlaywrightを介してヘッドフルブラウザと対話することを可能にします。実装は4つの主要コンポーネントで構成されています。
サーバー (index.ts)
モデルコンテキストプロトコル標準プロトコルを使用してMCPサーバーを初期化します
ツールとリソースのサーバー機能を構成します
stdioトランスポートを介してClaudeとの通信を確立する
ツールレジストリ (tools.ts)
ブラウザとAPIツールのスキーマを定義します
パラメータ、検証ルール、説明を指定します
クロードの発見のためにMCPサーバーにツールを登録する
リクエストハンドラー (handlers.ts)
ツールとリソースのMCPプロトコル要求を管理します
ブラウザのログとスクリーンショットをクエリ可能なリソースとして公開します
ツール実行リクエストを適切なハンドラーにルーティングします
エグゼキュータ (executor.ts)
ブラウザとAPIクライアントのライフサイクルを管理します
Playwrightを使用してブラウザ自動化機能を実装します
適切なエラー処理とレスポンス解析でAPIリクエストを処理する
コマンド間のステートフルなブラウザセッションを維持する
エージェントの機能
基本的な統合とは異なり、MCP ブラウザ エージェントは次の方法で真の AI エージェントとして機能します。
複数のコマンドにわたってブラウザの状態を永続的に維持する
デバッグのための詳細なコンソールログのキャプチャ
参照およびレビュー用にスクリーンショットを保存する
複雑なインタラクションシーケンスの管理
回復のための詳細なエラー情報の提供
複雑なワークフローの連鎖操作をサポート
利用可能なツール
ブラウザツール
ツール名 | 説明 | パラメータ |
| URLに移動する |
|
| スクリーンショットをキャプチャする |
|
| 要素をクリック |
|
| フォーム入力 |
|
| ドロップダウンオプションを選択 |
|
| 要素の上にマウスを移動 |
|
| JavaScriptを実行する |
|
APIツール
ツール名 | 説明 | パラメータ |
| GETリクエスト |
|
| POSTリクエスト |
|
| PUTリクエスト |
|
| PATCHリクエスト |
|
| 削除リクエスト |
|
リソースアクセス
MCP ブラウザ エージェントは次のリソースを公開します。
browser://logs- ブラウザコンソールのログにアクセスするscreenshot://[name]- 名前でスクリーンショットにアクセスします
使用例
Claude で MCP ブラウザ エージェントを使用する実際の例をいくつか示します。
基本的なブラウザナビゲーション
Navigate to the Google homepage at https://www.google.comTake a screenshot of the current page and name it "google-homepage"Type "weather forecast" in the search boxシンプルなインタラクション
Navigate to https://www.wikipedia.org and search for "Model Context Protocol"Go to https://the-internet.herokuapp.com/dropdown and select the option "Option 1" from the dropdown基本的なフォーム入力
Navigate to https://the-internet.herokuapp.com/login and fill in the username field with "tomsmith" and the password field with "SuperSecretPassword!"Go to https://the-internet.herokuapp.com/login, fill in the username and password fields, then click the login buttonシンプルなJavaScript実行
Go to https://example.com and execute a JavaScript script to return the page titleNavigate to https://www.google.com and execute a JavaScript script to count the number of links on the page基本的なAPIリクエスト
Perform a GET request to https://jsonplaceholder.typicode.com/todos/1Make a POST request to https://jsonplaceholder.typicode.com/posts with appropriate JSON dataこれらの例は、MCP ブラウザ エージェントの実際の機能を表し、現在の状態で何が達成できるかをより現実的に示しています。
トラブルシューティング
「サーバーが切断されました」エラー
Claude Desktop で「MCP ブラウザ エージェント: サーバーが切断されました」というエラーが表示される場合:
サーバーが実行中であることを確認します:
ターミナルを開き、プロジェクトディレクトリから
node dist/index.js手動で実行します。サーバーが正常に起動したら、このターミナルを開いたままClaudeを使用します。
設定を確認してください:
claude_desktop_config.jsonの絶対パスがシステムに合っていることを確認してくださいWindowsのパスに二重のバックスラッシュ(
\\)を使用していることを確認してくださいファイルシステムのルートからの完全なパスを使用していることを確認してください
ブラウザが表示されない
ブラウザが起動しない、または表示されない場合は、次の手順に従ってください。
指定されたブラウザがインストールされているかどうかを確認します
システムにブラウザ(Chrome、Firefox、Edge、またはSafari/WebKit)がインストールされていることを確認します。
ブラウザドライバはPlaywrightによって自動的に処理されます
サーバーとClaude Desktopを再起動します
サーバーを実行している可能性のある既存のノードプロセスをすべて強制終了します。
Claude Desktopを再起動して新しい接続を確立してください
ブラウザのプロセスが正しく終了しない
ChromiumおよびChromeブラウザでは、使用後にプロセスが正常に終了しないことがあるという既知の問題があります。この問題が発生した場合は、以下の手順に従ってください。
ブラウザのプロセスを手動で閉じます:
Windows : Ctrl+Shift+Escを押してタスクマネージャーを開き、Chrome/Chromiumプロセスを見つけて終了します。
macOS : アクティビティモニタ(アプリケーション > ユーティリティ > アクティビティモニタ)を開き、Chrome/Chromiumプロセスを見つけてXをクリックして終了します。
Linux :
ps aux | grep chromeまたはps aux | grep chromiumを実行してプロセスを見つけ、kill <PID>で終了します。
ブラウザの互換性に関する注意:
この問題は主にChromiumとChromeで確認されています。
FirefoxとPlaywrightの組み込みブラウザでは通常この問題は発生しません。
[!注意] このMCP統合はPlaywrightをベースに構築されており、 Playwrightには動作に影響を与える可能性のある既知の問題やバグがあります。ブラウザ自動化で問題が発生した場合は、PlaywrightのGitHub issuesまでご報告ください。Playwrightチームはこれらの問題への対応に継続的に取り組んでいますが、このエージェントはこれらの制限にもかかわらず、Claude Desktopのブラウザ自動化機能の基盤を提供します。
発達
プロジェクト構造
src/index.ts: メインエントリポイントとMCPサーバーの初期化src/tools.ts: ツールのスキーマと登録src/handlers.ts: ツールとリソースの MCP リクエスト ハンドラーsrc/executor.ts: Playwright を使用したツール実装ロジック
建物
npm run build変化に注意する
npm run watchテスト
このプロジェクトには、コア機能とブラウザの処理を検証するためのテストが含まれています。
npm test # Run tests
npm run test:watch # Watch mode
npm run test:coverage # Coverage reportテストでは、設定の整合性、ブラウザ自動化機能、エラー処理、プロセスのクリーンアップを検証します。テストスイートは、Chrome/Chromiumの終了に関する既知の問題を考慮し、ブラウザプロセスの適切な処理に特に重点を置いています。
セキュリティに関する考慮事項
[!重要] このMCP統合により、Claudeは自律的なブラウザ制御機能を利用できるようになります。禁止事項、セキュリティへの影響、ベストプラクティスに関する重要な情報については、セキュリティポリシーをご確認ください。
MCPブラウザエージェントは正当な自動化タスクのために設計されていますが、悪用される可能性があります。ユーザーは、適用されるすべての法律、利用規約、および倫理ガイドラインを遵守して使用する責任を負います。詳細については、詳細なセキュリティポリシーをご覧ください。
貢献
MCP ブラウザエージェントへの貢献を歓迎します!貢献できる分野は以下のとおりです。
新しいブラウザ自動化機能の追加
エラー処理と回復の改善
スクリーンショットとリソース管理の強化
便利なワークフローと例を作成する
複雑な操作のパフォーマンスの最適化
ライセンス
このプロジェクトは、Mozilla Public License 2.0 に基づいてライセンスされています。詳細については、 LICENSEファイルを参照してください。
関連リンク
Available Tools
13 toolsapi_deleteB
Perform a DELETE request to an API endpoint
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | API endpoint URL | |
| headers | No | Request headers |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden but only states the action without disclosing side effects, idempotency, authentication needs, rate limits, or return format. This is minimal for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It is appropriately concise for a simple tool, though slightly more context could be included without losing efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 params, no output schema), the description covers the basic action but omits expected return values, error conditions, or usage scope, leaving the description marginally adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptive parameter names (url, headers). The description adds no additional meaning beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool performs a DELETE request to an API endpoint, using a specific verb and resource that distinguishes it from sibling tools (api_get, api_patch, etc.) which handle other HTTP methods.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like api_get or api_post. The description merely repeats the method, missing explicit when/when-not context or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
api_getC
Perform a GET request to an API endpoint
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | API endpoint URL | |
| headers | No | Request headers |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It fails to mention that GET is typically safe and idempotent, how errors are handled, or whether redirects are followed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single clear sentence with no unnecessary words. It efficiently conveys the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple but lacks information about return values, error handling, authentication requirements, or default behavior. The description is too minimal for an agent to use confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions for both parameters. The tool description adds no extra meaning beyond the schema, so baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Perform a GET request to an API endpoint', which identifies the HTTP method and the action. It distinguishes from sibling tools like api_post and api_delete by the method name, but lacks mention of read-only nature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus siblings. The description does not specify that GET should be used for retrieving data, nor does it mention alternatives for modifying or deleting resources.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
api_patchB
Perform a PATCH request to an API endpoint
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | API endpoint URL | |
| data | Yes | Request body data (JSON string) | |
| headers | No | Request headers |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It only states it performs a PATCH request, implying mutation, but does not disclose side effects, authentication needs, rate limits, or response behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no unnecessary words. However, it is very brief and could benefit from additional context without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is incomplete given the tool's complexity. No output schema exists, and the description does not explain return values or error handling. Sibling tools suggest it is part of an HTTP client set, but the description lacks depth.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description does not add any additional meaning beyond what is in the input schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it performs a PATCH request to an API endpoint, which is a specific verb and resource. This distinguishes it from sibling tools like api_get, api_post, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus other HTTP methods (e.g., POST, PUT) or alternatives. The context signals show siblings, but the description offers no differentiation advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
api_postB
Perform a POST request to an API endpoint
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | API endpoint URL | |
| data | Yes | Request body data (JSON string) | |
| headers | No | Request headers |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description must disclose behavior. It fails to mention typical POST behavior (resource creation), data validation, or error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence is concise and front-loaded, but it is too brief and lacks substance for a practical tool description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple HTTP tool, description should mention typical use (e.g., 'sends data to URL'). Schema covers parameters but context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (all parameters described), so baseline is 3. Description adds no additional meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states 'POST request' verb and 'API endpoint' resource, clearly distinguishing from sibling tools like api_get, api_put, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, no mention of prerequisites or context such as authentication or data format.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
api_putC
Perform a PUT request to an API endpoint
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | API endpoint URL | |
| data | Yes | Request body data (JSON string) | |
| headers | No | Request headers |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description only states 'PUT request' implying mutation, but omits critical details like idempotency, side effects, authentication needs, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence is concise and front-loaded, but does not add meaningful content beyond the tool name; minimal but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no behavioral details; for a tool with 3 parameters and nested objects, the description is insufficient to fully understand usage and return values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers all 3 parameters with descriptions; description adds no extra meaning beyond what schema already provides, meeting the baseline for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the action ('Perform a PUT request') and resource ('API endpoint'), but lacks differentiation from sibling tools like api_patch or api_post, which perform similar HTTP methods.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use PUT versus other HTTP methods (e.g., PATCH for partial updates, POST for creation). Does not mention idempotency or replacement semantics.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_clickC
Click an element on the page
| Name | Required | Description | Default |
|---|---|---|---|
| selector | Yes | CSS selector for element to click |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description does not disclose behavioral traits beyond the basic action, such as potential side effects (e.g., navigation, page changes) or element visibility requirements. No annotations exist to compensate for this lack of detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no unnecessary words. It is appropriately sized for a simple action, though it could benefit from slight elaboration without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the sibling tools and the single parameter, the description is incomplete. It fails to mention crucial context like element visibility, clicks causing navigation, or waiting behavior, leaving the agent with insufficient information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema fully describes the 'selector' parameter with a clear definition. The description adds no extra meaning beyond the schema, which is acceptable given 100% coverage, but it does not enhance understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Click an element') and the resource ('on the page'), distinguishing it from sibling tools like browser_fill or browser_hover. However, it could benefit from specifying that it operates within the current page context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives, such as when to click vs. hover, or prerequisites like page navigation. The description lacks context for proper selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_evaluateB
Execute JavaScript in the browser context
| Name | Required | Description | Default |
|---|---|---|---|
| script | Yes | JavaScript code to execute |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must fully disclose behavior. It mentions script execution but omits critical details: return value, side effects, permissions, sandboxing, or error handling. This is insufficient for a potentially powerful tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with no unnecessary words. It is optimally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema, so the description should explain return values or behavior. It does not. Additionally, it lacks details on execution context (e.g., async support, timeout). This leaves the agent underinformed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required parameter 'script' has a description in the schema. The tool description adds no additional meaning beyond what the schema already provides, making it adequate but not additive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Execute JavaScript in the browser context' clearly states the action and resource. It is specific and distinct from sibling tools like browser_click or api_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives (e.g., browser_navigate or API calls). An agent receives no context for decision-making.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_fillB
Fill a form input with text
| Name | Required | Description | Default |
|---|---|---|---|
| value | Yes | Text to enter in the field | |
| selector | Yes | CSS selector for input field |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description should disclose behavioral details, but it only says 'fill' without specifying whether it overwrites existing text, waits for elements, or handles disabled fields. This leaves significant behavioral ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence, front-loading the action. However, it could be slightly expanded with useful context while remaining brief.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity of the tool (2 params, no output schema), the description is adequate but lacks details on behavior like clearing the field or submission. Sibling tools exist for other form actions, but no comparative guidance is provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema covers 100% of parameters with descriptions (selector and value), so baseline is 3. The description adds no additional semantic information beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Fill a form input with text' clearly states the action (fill) and the target (form input), distinguishing it from sibling tools like browser_click or browser_select which perform different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as browser_evaluate for setting values or browser_click for activation. There is no mention of prerequisites like element visibility or state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_hoverC
Hover over an element on the page
| Name | Required | Description | Default |
|---|---|---|---|
| selector | Yes | CSS selector for element to hover over |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits. It only states 'hover over an element' without explaining whether it triggers JavaScript events, waits for any transitions, or is safe. Essential behavioral context is missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise at 6 words and front-loaded. It is not verbose, but the brevity may sacrifice necessary detail. It earns its place but could be slightly more informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (1 parameter, no output schema), the description is incomplete. It does not mention the return value (likely void or success), side effects, or behavior after hovering. Sibling tools suggest a sequence of actions, but this tool's role is under-described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single parameter 'selector' described as 'CSS selector for element to hover over'. The description adds no additional meaning beyond what the schema already provides, so it meets the baseline but does not enhance understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Hover over an element on the page' clearly states the action (hover) and target (element on page). It differentiates from siblings like browser_click and browser_fill. However, it could be more specific about the effect (e.g., triggering hover state) but is not a tautology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like browser_click or browser_evaluate. The description lacks any context about prerequisites, typical scenarios, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_screenshotB
Capture a screenshot of the current page or a specific element
| Name | Required | Description | Default |
|---|---|---|---|
| mask | No | Selectors for elements to mask | |
| name | Yes | Identifier for the screenshot | |
| fullPage | No | Capture full page height | |
| savePath | No | Path to save screenshot (default: user's Downloads folder) | |
| selector | No | CSS selector for element to capture |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description does not disclose behavioral traits such as whether the capture affects the page state, any authorization requirements, or rate limits. It only states the action, leaving the agent without important context about side effects or constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single well-formed sentence that directly states the tool's purpose. It wastes no words, though it could benefit from slightly more detail without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is brief and lacks context about return values (e.g., image path), default behavior, or how the tool interacts with other browser tools. Given the five parameters and no output schema, the description should provide more operational context to be fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Since schema description coverage is 100%, the baseline is 3. The description does not add additional meaning beyond the parameter names and schema descriptions; for example, it doesn't explain when to use fullPage or mask options more concretely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool captures a screenshot of the current page or a specific element, which is a specific verb-resource combination. It distinguishes itself from sibling tools like browser_navigate or browser_click by focusing on capture rather than navigation or element interaction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used when a screenshot is needed, but it provides no explicit guidance on when to use it versus alternative methods (e.g., browser_evaluate for custom captures) or any exclusions (e.g., not for video capture).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_selectC
Select an option from a dropdown menu
| Name | Required | Description | Default |
|---|---|---|---|
| value | Yes | Value or label to select | |
| selector | Yes | CSS selector for select element |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description does not disclose behavioral traits such as whether it waits for options to load, supports custom dropdowns, triggers events, or requires scrolling. The description is too minimal to inform the agent of important behaviors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that is front-loaded. However, it sacrifices completeness for brevity. It earns a 4 for being efficient, but could include more key details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that no annotations or output schema exist, the description is insufficiently complete. It does not explain return values, constraints, or behavior for complex dropdowns, leaving gaps for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema describes both parameters (selector and value) with 100% coverage. The description adds no additional meaning beyond the schema, so baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it selects an option from a dropdown menu, using a specific verb and resource. It distinguishes from sibling tools like browser_fill (text input) and browser_click (clicking), though it could be more precise by specifying HTML <select> elements.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives (e.g., browser_click for custom dropdowns) or any prerequisites. No exclusions or context are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_set_viewportA
Change the browser's viewport size and scale factor
| Name | Required | Description | Default |
|---|---|---|---|
| width | No | Viewport width in pixels | |
| height | No | Viewport height in pixels | |
| deviceScaleFactor | No | Device scale factor (affects how content is scaled) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description should disclose behavioral traits. It only states what the tool changes, but not side effects (e.g., impact on screenshots, persistence across navigation). Lacks important context for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, no fluff. Every word is relevant. Excellent conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequately describes the core function, but lacks details on required fields, defaults, or behavioral context. Acceptable for a simple tool, but could be more complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with meaningful descriptions for each parameter. The description adds no extra meaning beyond the schema, justifying the baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the action 'Change' and the target 'browser's viewport size and scale factor'. It is specific and distinct from sibling tools like browser_navigate or browser_click.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance is provided. The description implies usage for adjusting viewport, but does not mention alternatives or exclusions. Minimal viable score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v1.0.0- First observed
api_delete - First observed
api_get - First observed
api_patch - First observed
api_post - First observed
api_put - First observed
browser_click - First observed
browser_evaluate - First observed
browser_fill - First observed
browser_hover - First observed
browser_navigate - First observed
browser_screenshot - First observed
browser_select - First observed
browser_set_viewport
TDQS
Scored across 13 tools
Each tool has a clearly distinct purpose: API tools are differentiated by HTTP method, and browser tools cover unique interactions like clicking, hovering, filling, etc. No two tools overlap in function.
All tools follow a consistent '<domain>_<action>' pattern, with 'api_' prefix for HTTP methods and 'browser_' prefix for browser actions. Naming is unambiguous and predictable.
13 tools is well-scoped for a browser automation and API testing server. Each tool covers a fundamental operation without unnecessary bloat or gaps.
Core browser interactions (navigation, clicking, form filling, selecting, screenshot) and all major HTTP methods are covered. Minor omissions like file upload or wait-for-element are acceptable for this scope.
Maintenance
Related MCP Connectors
AI-powered web automation. Navigate websites using AI agents for one page or a thousand
AI-powered web automation. Navigate websites using AI agents for one page or a thousand
AI-powered browser automation — navigate, click, fill forms, and extract data from any website.
An agent-first office suite Claude & ChatGPT read and write over one MCP URL.
Related MCP Servers
- AlicenseBqualityDmaintenanceA browser automation agent that enables Claude to interact with web browsers through the Model Context Protocol, allowing for actions like navigating websites, manipulating elements, and managing browser state.29MIT
- AlicenseNot gradedqualityDmaintenanceEnables real browser automation as tools in Cursor, Claude Desktop, Windsurf, and any MCP-compatible client, allowing AI agents to interact with web pages through natural language.17 npmMIT
- FlicenseNot gradedqualityBmaintenanceEnables Claude Code to control a real browser using AI for web scraping, competitive intelligence, and UX auditing through the MCP protocol.-
- AlicenseNot gradedqualityDmaintenanceEnables browser automation through the Claude Chrome Extension, allowing agents to navigate websites, fill forms, take screenshots, and debug web apps via standard MCP protocols.1MIT