Skip to main content
Glama

ie-mode-mcp

Microsoft Edge の IE モード で動作するレガシー Web アプリケーションを、AI エージェントから MCP (Model Context Protocol) 経由で操作するための MCP Server。

AI Agent ──(MCP / stdio)──> ie-mode-mcp ──> BrowserManager ──> selenium-webdriver
                                                                     │
                                                          IEDriverServer.exe
                                                                     │
                                                     Microsoft Edge (IE Mode)
                                                                     │
                                                       Legacy Web Application
  • Node.js 22 / TypeScript / selenium-webdriver のみで構成(HTTP Server・DB・DI・Logging Framework なし)

  • MCP Transport は stdio のみ

  • ブラウザセッションは 1 つのみ、WebDriver 操作は 完全逐次実行

  • HTML 全文は返さず、inspect_page が LLM 向けに要約した画面情報を返す

  • 承認フローなし。Tool を呼び出した時点で操作を実行する


目次

  1. クイックスタート

  2. 前提条件

  3. Windows 側の事前設定

  4. インストールとビルド

  5. 環境変数

  6. 起動方法

  7. AI Agent への登録

  8. Tool リファレンス

  9. 利用例

  10. エラーと対処

  11. ログ

  12. トラブルシューティング

  13. 開発

  14. 制限事項


Related MCP server: ie-mcp

1. クイックスタート

Windows 上で以下を実行する。

git clone https://github.com/sumikof/iedriver-mcp.git
cd iedriver-mcp
npm install
npm run build

# IEDriverServer.exe のパスと、遷移を許可する Origin を指定して起動
$env:IE_MCP_DRIVER_PATH = "C:\tools\IEDriverServer.exe"
$env:IE_MCP_ALLOWED_ORIGINS = "http://legacy01.local"
node dist/index.js

{"level":"info","event":"started","transport":"stdio"} が stderr に出力されれば起動成功。 通常は手動で起動せず、AI Agent 側の MCP 設定から自動起動させる。


2. 前提条件

項目

内容

OS

Windows 11 / Windows 10(ログイン済みのインタラクティブセッション)

Node.js

22 以上

ブラウザ

Microsoft Edge(IE モードが利用可能なこと)

Driver

IEDriverServer.exe(Selenium 4.x 系。32bit 版を推奨)

  • IEDriverServer.exe は Selenium のダウンロードページから取得し、 任意のフォルダ(例: C:\tools\)に配置する。 64bit 版には既知の制約があるため、Selenium 公式は 32bit 版の利用を推奨している。

  • IEDriver は GUI・ウィンドウフォーカス・ネイティブイベントの影響を受けるため、 専用の Windows VM または専用の Windows セッションでの利用を推奨する。

  • Windows Service(Session 0)上でブラウザを動作させる構成は想定していない。

  • MCP Server と IEDriver / Edge は同一 Windows 環境で動作させる。


3. Windows 側の事前設定

IEDriver は環境設定の影響を強く受ける。先に手動で設定を済ませてから MCP Server を起動する。

3.1 Edge の IE モードを利用可能にする

対象サイトが IE モードで開けることを、先に Edge の手動操作で確認しておく。IE モードは以下の いずれかのポリシーで有効化する(Software\Policies\Microsoft\Edge 配下)。

ポリシー(表示名)

レジストリ値名

Configure Internet Explorer integration

InternetExplorerIntegrationLevel

Configure the Enterprise Mode Site List

InternetExplorerIntegrationSiteList

Send all intranet sites to Internet Explorer

(Edge 77 以降のグループポリシーで設定)

具体的な構成は組織のポリシーに依存するため、詳細は Microsoft の IE モードのドキュメント と自組織の管理者に確認すること。Windows / Edge は最新の更新を適用しておく。

3.2 IEDriver の要求する設定

項目

必要な状態

本 Server での扱い

ブラウザのズーム

100%

ignoreZoomSetting(true) を設定済みのため必須ではないが、100% を推奨

保護モード(Protected Mode)

すべてのゾーンで同じ設定

未統一の場合は起動時に例外となる。Internet オプション → セキュリティ で統一する

IEDriverServer の bit 数

32bit 推奨

—

保護モードの設定が統一されていないと browser_start が失敗する。IEDriver の introduceFlakinessByIgnoringProtectedModeSettings は動作が不安定になるため使用していない。


4. インストールとビルド

npm install     # 依存パッケージの取得
npm run build   # TypeScript を dist/ へビルド

生成物は dist/index.js。ビルド後は npm start(= node dist/index.js)でも起動できる。


5. 環境変数

設定ファイル(YAML / JSON)は使用せず、環境変数のみで設定する。

環境変数

説明

既定値

IE_MCP_EDGE_PATH

msedge.exe のパス

未指定(IEDriver が自動検出)

IE_MCP_DRIVER_PATH

IEDriverServer.exe のパス

未指定(PATH から探索)

IE_MCP_ALLOWED_ORIGINS

navigate を許可する Origin のカンマ区切り。* で無制限

*

IE_MCP_TIMEOUT_MS

要素検索・待機の既定タイムアウト(ms)

10000

IE_MCP_EDGE_PATH=C:\Program Files (x86)\Microsoft\Edge\Application\msedge.exe
IE_MCP_DRIVER_PATH=C:\tools\IEDriverServer.exe
IE_MCP_ALLOWED_ORIGINS=http://legacy01.local,http://legacy02.local
IE_MCP_TIMEOUT_MS=10000
  • IE Driver 4.5.0 以降は IE 非搭載環境(Windows 11 の既定)で Edge を自動検出するため、 IE_MCP_EDGE_PATH は通常不要。自動検出に失敗する場合のみ明示指定する。

  • 運用の再現性を優先する場合は IE_MCP_DRIVER_PATH を明示指定することを推奨する。

  • IE_MCP_ALLOWED_ORIGINS は誤操作防止用の簡易的な制限であり、Origin(scheme + host + port) の完全一致で判定する。パス単位の制限は行わない。


6. 起動方法

手動起動(動作確認用)

PowerShell:

$env:IE_MCP_DRIVER_PATH = "C:\tools\IEDriverServer.exe"
$env:IE_MCP_ALLOWED_ORIGINS = "http://legacy01.local"
node dist/index.js

コマンドプロンプト:

set IE_MCP_DRIVER_PATH=C:\tools\IEDriverServer.exe
set IE_MCP_ALLOWED_ORIGINS=http://legacy01.local
node dist\index.js

stdio でクライアントからの接続を待ち受ける。標準入出力が MCP のプロトコルに使用されるため、 この状態でキーボード入力しても応答はない(正常)。ログはすべて stderr に出力される。 終了は Ctrl+C(ブラウザも自動的に閉じる)。

注意: MCP Server の起動だけではブラウザは起動しない。ブラウザは Agent が browser_start を呼び出した時点で起動する。

通常運用

AI Agent(MCP クライアント)が本 Server を子プロセスとして起動する。手動起動は不要。 次章の設定を行う。


7. AI Agent への登録

MCP クライアントの設定ファイルに以下を追加する。

{
  "mcpServers": {
    "ie-mode": {
      "command": "node",
      "args": ["C:\\ie-mode-mcp\\dist\\index.js"],
      "env": {
        "IE_MCP_DRIVER_PATH": "C:\\tools\\IEDriverServer.exe",
        "IE_MCP_EDGE_PATH": "C:\\Program Files (x86)\\Microsoft\\Edge\\Application\\msedge.exe",
        "IE_MCP_ALLOWED_ORIGINS": "http://legacy01.local,http://legacy02.local",
        "IE_MCP_TIMEOUT_MS": "10000"
      }
    }
  }
}
  • パスは JSON 内でバックスラッシュをエスケープする(C:\\...)。

  • args にはビルド後の dist/index.js の絶対パスを指定する。

  • Claude Code の場合は claude mcp add でも登録できる。

claude mcp add ie-mode --env IE_MCP_DRIVER_PATH=C:\tools\IEDriverServer.exe --env IE_MCP_ALLOWED_ORIGINS=http://legacy01.local -- node C:\ie-mode-mcp\dist\index.js

登録後、クライアント側で browser_start を含む 10 個の Tool が見えていれば接続成功。


8. Tool リファレンス

公開する Tool は 10 個。WebDriver の低レベル API(findElement / executeScript など)は公開しない。

Tool

入力

概要

browser_start

なし

Edge IE Mode を起動する。起動済みなら既存セッションを再利用する

browser_close

なし

ブラウザを終了する。何度呼んでもエラーにならない

navigate

url

URL Allowlist を確認してから遷移する

inspect_page

frame?

URL / title / 画面テキスト / 操作可能要素を返す

click

selector, frame?

表示・有効を待ってからクリックする

type

selector, frame?, text, clear?

input / textarea へ入力する

select

selector, frame?, by, value

<select> の option を選択する

wait_for

type, selector?, frame?, text?, timeoutMs?

条件が満たされるまで待機する

switch_window

target:"newest" / index, timeoutMs?

popup・別 Window へ切り替える

screenshot

なし

現在の画面を PNG(MCP image content)で返す

共通: Selector

{ "by": "id | name | css | xpath | linkText", "value": "searchButton" }

レガシー Web アプリでは name と xpath の使用頻度が高いため対応している。

共通: frame(iframe は 1 階層)

すべての要素操作 Tool は任意の frame を受け取る。指定すると defaultContent に戻してから frame に切り替え、その中で要素を検索する。

{
  "frame": { "by": "name", "value": "mainFrame" },
  "selector": { "by": "id", "value": "searchButton" }
}

browser_start

{}
{ "status": "ready", "reused": false }

reused: true は既存セッションをそのまま使ったことを示す。既存セッションが死んでいる場合は 自動的に起動し直す。

navigate

{ "url": "http://legacy01.local/customer" }
{ "url": "http://legacy01.local/customer", "title": "顧客検索" }

inspect_page

Agent が画面を理解するための主要 Tool。HTML 全文は返さず、URL / title / 表示テキスト / 操作可能要素(a button input textarea select iframe)のみを返す。 非表示の要素と type="hidden" の input は除外される。

{ "frame": { "by": "name", "value": "mainFrame" } }
{
  "url": "http://legacy01.local/customer",
  "title": "顧客検索",
  "text": "顧客検索 顧客名 支店 検索",
  "elements": [
    { "tag": "input", "id": "customerName", "name": "customerName", "type": "text" },
    { "tag": "select", "id": "branch", "name": "branch", "text": "東京支店", "optionCount": 12 },
    { "tag": "button", "id": "searchButton", "text": "検索" },
    { "tag": "iframe", "name": "mainFrame" }
  ],
  "truncated": false
}
  • truncated: true は要素が上限(300 件)で打ち切られたことを示す。

  • 要素一覧に iframe が含まれる場合、その中身を見るには frame を指定して再度呼び出す。

click

{ "selector": { "by": "id", "value": "searchButton" } }
{ "url": "http://legacy01.local/customer", "title": "顧客検索" }

表示・有効になるまで待ってからクリックする。click は自動 Retry しない(登録・更新・送信が 既に成功している状態での再クリックによる二重処理を防ぐため)。

type

{
  "selector": { "by": "id", "value": "customerName" },
  "text": "山田太郎",
  "clear": true
}

clear(既定 true)が true なら clear() 後に入力、false なら追記する。

select

{
  "selector": { "by": "id", "value": "branch" },
  "by": "text",
  "value": "東京支店"
}
{ "text": "東京支店", "value": "13", "index": 2 }

by は text / value / index(index は 0 始まり)。

wait_for

固定 sleep を使わず、明示的に待機する。

{
  "type": "visible",
  "selector": { "by": "id", "value": "resultTable" },
  "timeoutMs": 10000
}

type

必要な入力

条件

present

selector

要素が DOM に存在する

visible

selector

要素が表示されている

enabled

selector

要素が表示され、かつ操作可能

text

selector, text

要素のテキストが text を含む

url

text

現在の URL が text を含む

title

text

title が text を含む

timeoutMs 省略時は IE_MCP_TIMEOUT_MS を使用する。

switch_window

{ "target": "newest" }
{ "index": 1 }
{ "url": "http://legacy01.local/detail", "title": "顧客詳細", "index": 1, "windowCount": 2 }

newest は新しい Window Handle が現れるまで短時間ポーリングする。検出できなかった場合は 現存する最後の Window に切り替える。

screenshot

{}

PNG 画像(MCP の image content)を返す。DOM だけでは判断できないレイアウト・エラー画面の確認に使う。


9. 利用例

基本ループ

browser_start → navigate → inspect_page → click / type / select → wait_for → inspect_page

inspect_page で画面を把握 → 操作 → wait_for で結果を待つ → 再度 inspect_page、を繰り返す。

例: 顧客「山田太郎」を検索して詳細画面を開く

#

Tool

引数

1

browser_start

{}

2

navigate

{ "url": "http://legacy01.local/customer" }

3

inspect_page

{}

4

type

{ "selector": { "by": "id", "value": "customerName" }, "text": "山田太郎" }

5

select

{ "selector": { "by": "id", "value": "branch" }, "by": "text", "value": "東京支店" }

6

click

{ "selector": { "by": "id", "value": "searchButton" } }

7

wait_for

{ "type": "visible", "selector": { "by": "id", "value": "resultTable" } }

8

inspect_page

{}

9

click

{ "selector": { "by": "linkText", "value": "山田太郎" } }

10

wait_for

{ "type": "title", "text": "顧客詳細" }

11

inspect_page

{}

例: iframe 内を操作する

{"tool": "inspect_page", "args": {}}
{"tool": "inspect_page", "args": { "frame": { "by": "name", "value": "mainFrame" } }}
{"tool": "click", "args": {
  "frame": { "by": "name", "value": "mainFrame" },
  "selector": { "by": "id", "value": "searchButton" }
}}

frame の指定は操作ごとに毎回渡す(内部で毎回 defaultContent に戻してから切り替えるため、 状態は持ち越されない)。

例: popup を操作して元の Window に戻る

{"tool": "click",         "args": { "selector": { "by": "id", "value": "openPopup" } }}
{"tool": "switch_window", "args": { "target": "newest" }}
{"tool": "inspect_page",  "args": {}}
{"tool": "switch_window", "args": { "index": 0 }}

10. エラーと対処

エラーは Selenium の Stack Trace ではなく、次のコードで返る(isError: true)。

{
  "error": "ELEMENT_NOT_FOUND",
  "message": "Element was not found: id=searchButton",
  "selector": { "by": "id", "value": "searchButton" }
}

エラーコード

意味

対処

BROWSER_NOT_STARTED

ブラウザ未起動

browser_start を呼ぶ

ELEMENT_NOT_FOUND

要素・frame が見つからない

inspect_page で実際の要素を確認し、Selector を見直す

TIMEOUT

wait_for の条件が満たされなかった

条件・timeoutMs を見直す。画面が想定と異なる可能性

WINDOW_NOT_FOUND

指定 Window が存在しない

switch_window の index を見直す

NAVIGATION_FAILED

遷移に失敗

URL・ネットワーク・認証を確認

DRIVER_LOST

IEDriver / Edge が異常終了

browser_start で再起動する(下記参照)

URL_NOT_ALLOWED

Allowlist 外の Origin

IE_MCP_ALLOWED_ORIGINS を見直す

INVALID_ARGUMENT

引数不正

Tool の入力仕様を確認

INTERNAL_ERROR

その他(起動失敗を含む)

message と stderr のログを確認

DRIVER_LOST からの復旧

ブラウザまたは Driver が落ちた場合、内部の WebDriver は破棄され、以後の操作は BROWSER_NOT_STARTED になる。自動復旧・直前操作の自動再実行は行わない(二重登録などの 副作用を防ぐため)。Agent 側で browser_start を呼び直し、画面の状態を inspect_page で 確認してから操作を再開する。直前の操作が既に成立している可能性があるため、登録・更新系の 操作をそのまま再実行してはならない。


11. ログ

stdout は MCP のプロトコルが使用するため、ログはすべて stderr に JSON 1 行で出力する。

{"level":"info","event":"started","transport":"stdio"}
{"level":"info","tool":"navigate","url":"http://legacy01.local/customer","durationMs":842}
{"level":"info","tool":"type","selector":{"by":"id","value":"password"},"textLength":16,"durationMs":128}
{"level":"error","tool":"click","selector":{"by":"id","value":"x"},"error":"ELEMENT_NOT_FOUND","message":"Element was not found: id=x","durationMs":5012}

入力文字列そのもの・Cookie・認証情報・HTML 全文は記録しない(type は文字数のみ)。 ファイルに残す場合は stderr をリダイレクトする。

node dist/index.js 2>> C:\logs\ie-mode-mcp.log

12. トラブルシューティング

症状

確認すること

browser_start が INTERNAL_ERROR になる

IE_MCP_DRIVER_PATH が正しいか。IEDriverServer.exe を単体で起動できるか

保護モード関連の例外が出る

Internet オプション → セキュリティ で全ゾーンの保護モード設定を統一する

ズーム関連の例外が出る

Edge / IE のズームを 100% に戻す

Edge は起動するが IE モードにならない

IE モードのポリシー(サイトリスト等)を確認する。手動で IE モード表示できるか先に確認

操作が固まる・要素をクリックできない

ウィンドウが最小化・非アクティブになっていないか。リモートデスクトップ切断中は不安定になる

inspect_page の要素が空

frame 内の画面ではないか(frame を指定して再取得)。screenshot で実画面を確認

Agent 側に Tool が見えない

dist/index.js を絶対パスで指定しているか。npm run build 済みか

標準出力に何も出ない

正常。ログは stderr に出力される

screenshot は原因調査に有効。DOM 情報だけでは判断できない状態(モーダル、認証ダイアログ、 レンダリング崩れ)を確認できる。


13. 開発

src/
├─ index.ts      MCP Server のエントリーポイント(stdio)
├─ config.ts     環境変数と stderr ログ
├─ tools.ts      MCP Tool の Schema と Handler
├─ browser.ts    BrowserManager(Selenium / IEDriver 操作の集約)
├─ selectors.ts  Selector → Selenium の By 変換
└─ errors.ts     Selenium Error → MCP Error Code 変換
npm run build   # tsc でビルド
npm start       # node dist/index.js
  • MCP Tool は Selenium を直接触らず、必ず BrowserManager を経由する。

  • すべての WebDriver 操作は Promise Chain で逐次化されており、Tool が並列に呼ばれても IEDriver へは 1 件ずつしか送られない。

  • 副作用のない操作(要素検索・Window Handle 検出)のみ Retry する。click や送信は Retry しない。


14. 制限事項

初期実装では以下に対応しない。

複数ブラウザセッション / 複数ユーザー / HTTP Transport / REST API / DB / セッション永続化 / 自動ブラウザ復旧 / 複雑な Retry Policy / WebDriver Grid / 汎用 Selenium API / executeScript Tool / 多段 iframe(1 階層のみ)/ Element Cache / Metrics / 承認フロー / 認証・認可

Available Tools

10 tools
browser_closeClose browserB

Close the browser session. Safe to call repeatedly.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are available so the description carries the full burden. It only says 'close the browser session' and 'safe to call repeatedly', but does not disclose whether this terminates all browser state or if there are side effects on open windows, tabs, or downloads. The behavioral context is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences that are front-loaded and to the point. Every sentence adds value: the first states the action, the second clarifies safety/repeatability. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and no output schema, the description covers the basic purpose and safety. However, it lacks details on what happens after closing (e.g., can browser_start reopen cleanly) or any cleanup behavior, which might be useful context for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema coverage is 100%. The description adds value by stating it is safe to call repeatedly, which implies no parameters are needed and calls are idempotent. With no parameters to explain, this is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool closes the browser session with a specific verb and resource. It distinguishes enough from siblings like 'navigate' which moves within a session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description notes it is safe to call repeatedly, which implies idempotency, but does not explicitly tell when to call it (e.g., end of a browsing task) or when not to (e.g., still need to interact). No sibling differentiation is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_startStart Edge IE ModeA

Start Microsoft Edge in IE Mode through IEDriverServer. Only one browser session exists; calling this while a session is running returns the existing one. Also use this to recover after a DRIVER_LOST error.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the full burden of behavioral disclosure. It transparently reveals that only one browser session exists, that calling the tool again returns the existing session, and that it can be used for recovery. This is strong for a start tool, but it could additionally mention potential side effects like timeouts or prerequisites for the IEDriverServer.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long, front-loading the primary purpose in the first sentence and adding behavioral nuance in the second. Every sentence provides essential information without redundancy or fluff, achieving maximum conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, no output schema, straightforward start action), the description is complete. It covers the core function, the singleton behavior, error recovery, and is sufficient for an AI agent to understand when and how to invoke the tool alongside its sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and 100% schema coverage, so the baseline is 4. The description adds no parameter information, which is appropriate since there are none to document. No additional semantic value is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool starts Microsoft Edge in IE Mode via IEDriverServer, using the specific verb 'Start' and the resource 'Microsoft Edge in IE Mode'. It also distinguishes itself from sibling tools by noting that only one browser session exists and that calling it again returns the existing session, which is unique among the provided sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool: normally to start the browser, and also to recover after a DRIVER_LOST error. It implicitly advises against calling it multiple times for new sessions by stating that subsequent calls return the existing session. However, it does not explicitly list alternatives or state when not to use it, though no alternative starting tool exists among the siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clickClick elementA

Click an element after waiting for it to be visible and enabled. This operation is never retried automatically, because a repeated click may submit or register data twice.

ParametersJSON Schema
NameRequiredDescriptionDefault
frameNoOptional iframe/frame to switch into first. One level of nesting is supported.
selectorYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses two key behaviors: waiting for the element to be visible and enabled, and the lack of automatic retry with a rationale. However, it does not mention timeout behavior, scroll-into-view, or what happens if the element is not found.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose, and every sentence adds value. No fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's potential side effects (e.g., triggering navigation or form submission), the description is minimal. It does not mention return values, scroll behavior, or failure modes. It is adequate for a simple click but lacks completeness for an AI agent to fully anticipate outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%, and the description adds no information about the parameters. It does not explain the 'frame' or 'selector' parameters beyond what is already in the schema. The description should compensate for the missing schema descriptions but fails to do so.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Click an element after waiting for it to be visible and enabled,' using a specific verb and resource. It distinguishes the tool from siblings like 'type' and 'select' by specifying the action and the precondition (visibility and enabled state).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly warns that the operation is never retried automatically because a repeated click may submit or register data twice. This gives a clear usage caution about retries, though it does not explicitly compare to alternative tools or state when not to use click.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_pageInspect pageA

Return the current URL, title, visible page text and the operable elements (a, button, input, textarea, select, iframe). The full HTML is never returned. Pass frame to inspect the contents of an iframe listed by a previous inspect_page call.

ParametersJSON Schema
NameRequiredDescriptionDefault
frameNoOptional iframe/frame to switch into first. One level of nesting is supported.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of disclosing behavior. It explicitly states 'The full HTML is never returned' and that frame must reference an iframe from a previous call. The read-only nature is implied by 'Return' but not stated outright, though this is likely sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two succinct sentences that front-load the main purpose and then add the iframe caveat. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has one well-described parameter, no output schema, and no annotations. The description covers the output, a key constraint (no full HTML), and iframe usage, making it reasonably complete for an inspection tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by clarifying that the frame must come from a previous inspect_page call, which is not stated in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') and enumerates exactly what is returned (URL, title, visible text, operable elements). This clearly distinguishes it from sibling tools like 'click' or 'navigate'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool (to inspect page state) and gives specific guidance for iframe usage ('Pass frame to inspect the contents of an iframe listed by a previous inspect_page call'). It doesn't explicitly exclude alternatives, but the context is clear given the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshotScreenshotB

Capture the current browser window as a PNG image.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears the full burden. It states the action (capture) and format (PNG) but omits crucial details: whether it modifies state, if a browser window must be open, what exactly 'current browser window' captures (viewport vs full page), and if there are side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single 9-word sentence, efficient and front-loaded. However, it could include additional essential context (e.g., 'captures the visible viewport area') without losing conciseness, making it slightly under-specified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and no output schema, the description should clarify what the tool returns (e.g., base64 PNG data). It only says 'as a PNG image' but doesn't confirm the output type. The scope of 'current browser window' is ambiguous, and prerequisites are missing, leaving the agent uncertain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, and the schema coverage is 100% (trivially). The description adds minimal meaning by specifying 'current browser window' as the implicit input. A baseline of 4 is appropriate given no parameters to document.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('capture') and resource ('current browser window') with a clear output format ('PNG image'). It is distinct from sibling tools like 'navigate' or 'inspect_page' which serve different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use screenshot versus alternatives. Despite having sibling tools (e.g., inspect_page, wait_for), no exclusions or context is given. An agent must infer use case from tool purpose alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

selectSelect optionB

Choose an option of an HTML element by visible text, value or index.

ParametersJSON Schema
NameRequiredDescriptionDefault
byYesHow to identify the option.
frameNoOptional iframe/frame to switch into first. One level of nesting is supported.
valueYesOption text, value, or zero-based index.
selectorYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It discloses that selection is based on visible text, value or index, which is helpful. However, it doesn't mention side effects (e.g., whether the change triggers JavaScript events), error handling (e.g., what if option not found), or scope (e.g., operates within current page context).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that efficiently communicates the core action and identification methods. It is front-loaded with the key verb and resource. No waste, though it could optionally add a brief usage hint without breaching conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters with nested objects and no output schema, the description is somewhat complete but lacks coverage of return behavior (e.g., what happens on success/failure), frame handling nuances, and edge cases. For a selection action in a browser automation context, more behavioral detail would be helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, meaning most parameters are documented in the schema. The description adds that selection can be by 'visible text, value or index', which maps to the 'by' enum, and that the 'value' parameter can be text or zero-based index. This provides modest added meaning beyond the schema, but the 'frame' and 'selector' objects remain documented primarily in schema, not description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'choose' and resource 'HTML <select> element', specifying three identification methods (visible text, value, index). This distinguishes it from sibling tools like click or type, though it doesn't explicitly contrast with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use by saying 'choose an option of an HTML <select> element', which suggests this is for dropdown selections. However, it does not provide explicit when-not-to-use guidance, mention prerequisites (e.g., element must exist), or compare with alternatives like click on an option directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

switch_windowSwitch windowA

Switch to another browser window or popup. Use target:"newest" after an action that opens a window, or index to select a window by its zero-based position.

ParametersJSON Schema
NameRequiredDescriptionDefault
indexNoZero-based window index.
targetNoSwitch to the newest window.
timeoutMsNoHow long to poll for a new window. Defaults to IE_MCP_TIMEOUT_MS.

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It mentions polling behavior via timeoutMs parameter but does not state if switching is destructive, if it requires a window to exist, what happens if the window is closed, or any state changes. The description does not disclose potential side effects or preconditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with the core purpose in the first sentence. The second sentence adds specific usage hints. It could potentially omit 'or popup' as redundant with 'window', but overall concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 3 unrequired parameters, no output schema, no annotations, the description covers the basic purpose and usage hints. However, it lacks details on return values, error scenarios (e.g., window not found), or behavior when switching to a window that fails to load.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds context for the 'target' parameter (use after action that opens window) and 'index' (zero-based position), but the timeoutMs parameter meaning is already clear from schema. No additional semantic value beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool switches to another browser window or popup, specifying the verb 'switch' and the resource 'browser window or popup'. It distinguishes itself from sibling tools like browser_start, browser_close, and navigate by focusing on window selection rather than creation, closure, or navigation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool: after an action that opens a window, use target 'newest', or use index to select by position. It implicitly distinguishes from sibling tools by indicating this is for window focus rather than content navigation or page interaction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

typeType textA

Type text into an input or textarea. Set clear to false to append instead of replacing.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesText to send to the element.
clearNoClear the field first. Default true.
frameNoOptional iframe/frame to switch into first. One level of nesting is supported.
selectorYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that the tool can clear or append text via the 'clear' parameter, which is good. However, it does not mention potential side effects (e.g., triggering change events), error conditions (element not found), or behavior when the element is not a text input. This is adequate but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at one sentence plus one usage tip. Every word earns its place, clearly stating the action and a key parameter behavior. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description should ideally mention return values (e.g., success indicator, element state). It doesn't, leaving that unclear. With nested objects (selector, frame) and no explanation of selector strategies beyond the schema enums, it completes the basic usage but misses context on what happens after typing (e.g., waits for stability, triggers events).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is high at 75%, so the schema documents most parameters well. The description adds value by explaining the 'clear' boolean behavior (append vs replace) beyond the schema's default value note. It doesn't add to 'selector' or 'frame' parameters, which are already well-described in the schema, so this is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool types text into an input or textarea, using a specific verb and resource. It distinguishes from siblings like 'click' or 'select' by targeting text entry specifically, but doesn't differentiate from a potential 'send_keys' equivalent if one existed among siblings, so a slight deduction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a key usage guideline: set clear to false to append instead of replacing text. This gives basic advice on when to use a parameter. However, it lacks guidance on when to use this tool versus alternatives like clicking an element first or waiting, and doesn't mention prerequisites (e.g., element must be visible/interactable).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_forWait for conditionA

Wait until a condition holds. present/visible/enabled/text require a selector; text/url/title require text, which is matched as a substring. Use this instead of sleeping after an action.

ParametersJSON Schema
NameRequiredDescriptionDefault
textNoExpected substring for text/url/title conditions.
typeYesCondition to wait for.
frameNoOptional iframe/frame to switch into first. One level of nesting is supported.
selectorNo
timeoutMsNoTimeout in milliseconds. Defaults to IE_MCP_TIMEOUT_MS.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description explains key behavioral traits: that present/visible/enabled/text require a selector, text/url/title require text matched as substring, and that it waits for the condition. With no annotations provided, the description carries the full burden of transparency. It lacks details on timeout behavior or error handling, but covers core usage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with two sentences, front-loading the key purpose and condition types. Every sentence adds value, avoiding any redundancy. The structure is efficient for an AI agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (5 parameters, nested objects, no output schema), the description is adequate. It explains the core waiting concept and parameter dependencies. However, it lacks details on return values or what happens on timeout/failure, which the schema alone doesn't cover. The sibling 'inspect_page' might share similar conditions, but no differentiation is made.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, so the schema already documents most parameters well. The description adds value by clarifying the relationship between condition types and required parameters (e.g., 'present/visible/enabled/text require a selector; text/url/title require text'). This bridges gaps between parameters, though it does not detail the 'frame' parameter beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool waits until a condition holds, with specific verb+resource ('Wait for condition'). It lists the condition types and distinguishes itself from sleeping after an action, which differentiates it from sibling tools like 'click' or 'navigate'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool ('Use this instead of sleeping after an action'), providing clear guidance on avoiding poor alternatives. However, it does not specify when not to use it or which sibling would be more appropriate for different scenarios, such as synchronous checks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 10 tool updatesv0.1.0
    • First observedbrowser_close
    • First observedbrowser_start
    • First observedclick
    • First observedinspect_page
    • First observednavigate
    • First observedscreenshot
    • First observedselect
    • First observedswitch_window
    • First observedtype
    • First observedwait_for

TDQS

A3.8/5.0

Scored across 10 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: session management (start/close), navigation, inspection, interaction (click, type, select), window switching, waiting, and screenshot. No two tools overlap in functionality.

Naming Consistency3/5

The naming pattern is inconsistent: some tools use a 'browser_' prefix (browser_start, browser_close), while others are bare verbs (navigate, click, type) or compound snake_case (inspect_page, switch_window, wait_for). This mix of styles could cause confusion.

Tool Count5/5

With 10 tools, the set is well-scoped for browser automation. It covers session lifetime, navigation, element interaction, inspection, window handling, and waiting without being bloated or too thin.

Completeness3/5

The tools cover fundamental browser actions but miss common features like back/forward navigation, JavaScript execution, alert handling, or cookie management. The set is functional for basic scenarios but has notable gaps for comprehensive automation.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Exposes Selenium WebDriver as an MCP server, enabling AI agents and LLMs to control real browsers for automation tasks like navigation, element interaction, and screenshot capture.
    22
    21 PyPI
    3
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables LLMs to drive Edge in IE mode for automating legacy IE-only web applications, supporting tasks like clicking, filling forms, and data extraction via Selenium.
    28
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables AI agents and MCP clients to automate web browsers via Selenium WebDriver, supporting Chrome, Firefox, and Edge in headless or visible mode with tools for navigation, interaction, content extraction, screenshots, and scripting.
    21
    31 npm
    MIT