Skip to main content
Glama

ie-mode-mcp

用于从 AI 代理通过 MCP (Model Context Protocol) 操作 Microsoft Edge IE 模式下运行的旧版 Web 应用程序的 MCP Server。

AI Agent ──(MCP / stdio)──> ie-mode-mcp ──> BrowserManager ──> selenium-webdriver
                                                                     │
                                                          IEDriverServer.exe
                                                                     │
                                                     Microsoft Edge (IE Mode)
                                                                     │
                                                       Legacy Web Application
  • 仅由 Node.js 22 / TypeScript / selenium-webdriver 构成(无 HTTP Server、DB、DI、Logging Framework)

  • MCP Transport 仅支持 stdio

  • 浏览器会话仅 1 个,WebDriver 操作完全顺序执行

  • 不返回完整 HTML,inspect_page 返回为 LLM 总结的屏幕信息

  • 无审批流程。调用 Tool 时立即执行操作


目录

  1. 快速开始

  2. 前提条件

  3. Windows 端预先设置

  4. 安装与构建

  5. 环境变量

  6. 启动方法

  7. 注册到 AI 代理

  8. Tool 参考

  9. 使用示例

  10. 错误与处理

  11. 日志

  12. 故障排除

  13. 开发

  14. 限制事项


Related MCP server: ie-mcp

1. 快速开始

在 Windows 上执行以下操作。

git clone https://github.com/sumikof/iedriver-mcp.git
cd iedriver-mcp
npm install
npm run build

# IEDriverServer.exe のパスと、遷移を許可する Origin を指定して起動
$env:IE_MCP_DRIVER_PATH = "C:\tools\IEDriverServer.exe"
$env:IE_MCP_ALLOWED_ORIGINS = "http://legacy01.local"
node dist/index.js

如果 stderr 输出 {"level":"info","event":"started","transport":"stdio"} 则表示启动成功。 通常无需手动启动,而是通过 AI 代理端的 MCP 设置 自动启动。


2. 前提条件

项目

内容

OS

Windows 11 / Windows 10(已登录的交互式会话)

Node.js

22 以上

浏览器

Microsoft Edge(可使用 IE 模式)

Driver

IEDriverServer.exe(Selenium 4.x 系列。推荐 32bit 版)

  • IEDriverServer.exe 从 Selenium 下载页面 获取, 放置在任意文件夹(例如 C:\tools\)中。 由于 64bit 版存在已知限制,Selenium 官方推荐使用 32bit 版。

  • IEDriver 受 GUI、窗口焦点和原生事件的影响, 因此建议在专用的 Windows VM 或专用的 Windows 会话中使用。

  • 不假设在 Windows Service(Session 0)上运行浏览器的配置。

  • MCP Server、IEDriver 和 Edge 应在同一 Windows 环境中运行。


3. Windows 端预先设置

IEDriver 受环境设置影响很大。请先手动完成设置,然后再启动 MCP Server。

3.1 使 Edge 的 IE 模式可用

先通过 Edge 的手动操作确认目标站点可以在 IE 模式下打开。IE 模式通过以下 任一策略启用(位于 Software\Policies\Microsoft\Edge 下)。

策略(显示名称)

注册表值名称

Configure Internet Explorer integration

InternetExplorerIntegrationLevel

Configure the Enterprise Mode Site List

InternetExplorerIntegrationSiteList

Send all intranet sites to Internet Explorer

(在 Edge 77 及更高版本的组策略中设置)

具体配置取决于组织策略,详情请参考 Microsoft 的 IE 模式文档 并咨询本组织的管理员。确保 Windows / Edge 已应用最新更新。

3.2 IEDriver 要求的设置

项目

所需状态

本 Server 中的处理

浏览器缩放

100%

已设置 ignoreZoomSetting(true),因此非必需,但推荐 100%

保护模式(Protected Mode)

所有区域设置相同

未统一时启动会抛出异常。在 Internet 选项 → 安全 中统一

IEDriverServer 的位数

推荐 32bit

—

如果保护模式设置未统一,browser_start 将失败。IEDriver 的 introduceFlakinessByIgnoringProtectedModeSettings 会导致行为不稳定,因此未使用。


4. 安装与构建

npm install     # 依存パッケージの取得
npm run build   # TypeScript を dist/ へビルド

产物为 dist/index.js。构建后也可通过 npm start(= node dist/index.js)启动。


5. 环境变量

不使用配置文件(YAML / JSON),仅通过环境变量进行设置。

环境变量

说明

默认值

IE_MCP_EDGE_PATH

msedge.exe 的路径

未指定(IEDriver 自动检测)

IE_MCP_DRIVER_PATH

IEDriverServer.exe 的路径

未指定(从 PATH 搜索)

IE_MCP_ALLOWED_ORIGINS

允许 navigate 的 Origin 逗号分隔列表。* 表示无限制

*

IE_MCP_TIMEOUT_MS

元素搜索和等待的默认超时时间(ms)

10000

IE_MCP_EDGE_PATH=C:\Program Files (x86)\Microsoft\Edge\Application\msedge.exe
IE_MCP_DRIVER_PATH=C:\tools\IEDriverServer.exe
IE_MCP_ALLOWED_ORIGINS=http://legacy01.local,http://legacy02.local
IE_MCP_TIMEOUT_MS=10000
  • IE Driver 4.5.0 及更高版本会在未安装 IE 的环境(Windows 11 默认)中自动检测 Edge, 因此通常不需要 IE_MCP_EDGE_PATH。仅在自动检测失败时显式指定。

  • 如果优先考虑操作的可重复性,建议显式指定 IE_MCP_DRIVER_PATH。

  • IE_MCP_ALLOWED_ORIGINS 是用于防止误操作的简易限制,通过 Origin(scheme + host + port) 完全匹配进行判断。不进行路径级别的限制。


6. 启动方法

手动启动(用于确认操作)

PowerShell:

$env:IE_MCP_DRIVER_PATH = "C:\tools\IEDriverServer.exe"
$env:IE_MCP_ALLOWED_ORIGINS = "http://legacy01.local"
node dist/index.js

命令提示符:

set IE_MCP_DRIVER_PATH=C:\tools\IEDriverServer.exe
set IE_MCP_ALLOWED_ORIGINS=http://legacy01.local
node dist\index.js

通过 stdio 等待客户端连接。标准输入输出用于 MCP 协议, 因此在此状态下键盘输入不会有响应(正常)。所有日志输出到 stderr。 使用 Ctrl+C 退出(浏览器也会自动关闭)。

注意: 仅启动 MCP Server 不会启动浏览器。浏览器在代理调用 browser_start 时才会启动。

常规操作

AI 代理(MCP 客户端)将本 Server 作为子进程启动。无需手动启动。 进行下一章的设置。


7. 注册到 AI 代理

在 MCP 客户端的配置文件中添加以下内容。

{
  "mcpServers": {
    "ie-mode": {
      "command": "node",
      "args": ["C:\\ie-mode-mcp\\dist\\index.js"],
      "env": {
        "IE_MCP_DRIVER_PATH": "C:\\tools\\IEDriverServer.exe",
        "IE_MCP_EDGE_PATH": "C:\\Program Files (x86)\\Microsoft\\Edge\\Application\\msedge.exe",
        "IE_MCP_ALLOWED_ORIGINS": "http://legacy01.local,http://legacy02.local",
        "IE_MCP_TIMEOUT_MS": "10000"
      }
    }
  }
}
  • 路径需要在 JSON 中转义反斜杠(C:\\...)。

  • args 中指定构建后 dist/index.js 的绝对路径。

  • 对于 Claude Code,也可以使用 claude mcp add 进行注册。

claude mcp add ie-mode --env IE_MCP_DRIVER_PATH=C:\tools\IEDriverServer.exe --env IE_MCP_ALLOWED_ORIGINS=http://legacy01.local -- node C:\ie-mode-mcp\dist\index.js

注册后,如果客户端能看到包括 browser_start 在内的 10 个 Tool,则表示连接成功。


8. Tool 参考

公开的 Tool 共 10 个。不公开 WebDriver 的低级 API(如 findElement / executeScript)。

Tool

输入

概要

browser_start

无

启动 Edge IE Mode。如果已启动则重用现有会话

browser_close

无

关闭浏览器。多次调用也不会出错

navigate

url

检查 URL 允许列表后导航

inspect_page

frame?

返回 URL / title / 屏幕文本 / 可操作元素

click

selector, frame?

等待显示和启用后点击

type

selector, frame?, text, clear?

向 input / textarea 输入

select

selector, frame?, by, value

选择 <select> 的 option

wait_for

type, selector?, frame?, text?, timeoutMs?

等待条件满足

switch_window

target:"newest" / index, timeoutMs?

切换到弹出窗口或另一个窗口

screenshot

无

返回当前屏幕的 PNG(MCP image content)

通用: Selector

{ "by": "id | name | css | xpath | linkText", "value": "searchButton" }

由于旧版 Web 应用中 name 和 xpath 使用频率较高,因此支持它们。

通用: frame(iframe 仅一层)

所有元素操作 Tool 都接受可选的 frame。指定后会先返回 defaultContent, 然后切换到该 frame,在其中搜索元素。

{
  "frame": { "by": "name", "value": "mainFrame" },
  "selector": { "by": "id", "value": "searchButton" }
}

browser_start

{}
{ "status": "ready", "reused": false }

reused: true 表示直接使用了现有会话。如果现有会话已失效,则会自动重新启动。

navigate

{ "url": "http://legacy01.local/customer" }
{ "url": "http://legacy01.local/customer", "title": "顧客検索" }

inspect_page

代理理解屏幕的主要 Tool。不返回完整 HTML,仅返回 URL / title / 显示文本 / 可操作元素(a button input textarea select iframe)。 隐藏元素和 type="hidden" 的 input 会被排除。

{ "frame": { "by": "name", "value": "mainFrame" } }
{
  "url": "http://legacy01.local/customer",
  "title": "顧客検索",
  "text": "顧客検索 顧客名 支店 検索",
  "elements": [
    { "tag": "input", "id": "customerName", "name": "customerName", "type": "text" },
    { "tag": "select", "id": "branch", "name": "branch", "text": "東京支店", "optionCount": 12 },
    { "tag": "button", "id": "searchButton", "text": "検索" },
    { "tag": "iframe", "name": "mainFrame" }
  ],
  "truncated": false
}
  • truncated: true 表示元素数量达到上限(300 个)而被截断。

  • 如果元素列表中包含 iframe,要查看其内容,需要指定 frame 再次调用。

click

{ "selector": { "by": "id", "value": "searchButton" } }
{ "url": "http://legacy01.local/customer", "title": "顧客検索" }

等待元素显示和启用后点击。click 不会自动重试(为了防止在注册、更新、提交 已成功的情况下再次点击导致重复处理)。

type

{
  "selector": { "by": "id", "value": "customerName" },
  "text": "山田太郎",
  "clear": true
}

如果 clear(默认 true)为 true,则先执行 clear() 再输入;为 false 则追加输入。

select

{
  "selector": { "by": "id", "value": "branch" },
  "by": "text",
  "value": "東京支店"
}
{ "text": "東京支店", "value": "13", "index": 2 }

by 可以是 text / value / index(index 从 0 开始)。

wait_for

不使用固定 sleep,而是显式等待。

{
  "type": "visible",
  "selector": { "by": "id", "value": "resultTable" },
  "timeoutMs": 10000
}

type

所需输入

条件

present

selector

元素存在于 DOM 中

visible

selector

元素已显示

enabled

selector

元素已显示且可操作

text

selector, text

元素的文本包含 text

url

text

当前 URL 包含 text

title

text

title 包含 text

省略 timeoutMs 时使用 IE_MCP_TIMEOUT_MS。

switch_window

{ "target": "newest" }
{ "index": 1 }
{ "url": "http://legacy01.local/detail", "title": "顧客詳細", "index": 1, "windowCount": 2 }

newest 会短时间轮询直到出现新的 Window Handle。如果未检测到,则切换到现有的最后一个 Window。

screenshot

{}

返回 PNG 图像(MCP 的 image content)。用于确认仅凭 DOM 无法判断的布局和错误画面。


9. 使用示例

基本循环

browser_start → navigate → inspect_page → click / type / select → wait_for → inspect_page

通过 inspect_page 了解屏幕 → 操作 → 通过 wait_for 等待结果 → 再次 inspect_page,重复此过程。

示例: 搜索客户“山田太郎”并打开详情画面

#

Tool

参数

1

browser_start

{}

2

navigate

{ "url": "http://legacy01.local/customer" }

3

inspect_page

{}

4

type

{ "selector": { "by": "id", "value": "customerName" }, "text": "山田太郎" }

5

select

{ "selector": { "by": "id", "value": "branch" }, "by": "text", "value": "東京支店" }

6

click

{ "selector": { "by": "id", "value": "searchButton" } }

7

wait_for

{ "type": "visible", "selector": { "by": "id", "value": "resultTable" } }

8

inspect_page

{}

9

click

{ "selector": { "by": "linkText", "value": "山田太郎" } }

10

wait_for

{ "type": "title", "text": "顧客詳細" }

11

inspect_page

{}

示例: 操作 iframe 内部

{"tool": "inspect_page", "args": {}}
{"tool": "inspect_page", "args": { "frame": { "by": "name", "value": "mainFrame" } }}
{"tool": "click", "args": {
  "frame": { "by": "name", "value": "mainFrame" },
  "selector": { "by": "id", "value": "searchButton" }
}}

每次操作都需要传递 frame 指定(因为内部每次都会返回 defaultContent 再切换, 状态不会保持)。

示例: 操作弹出窗口并返回原窗口

{"tool": "click",         "args": { "selector": { "by": "id", "value": "openPopup" } }}
{"tool": "switch_window", "args": { "target": "newest" }}
{"tool": "inspect_page",  "args": {}}
{"tool": "switch_window", "args": { "index": 0 }}

10. 错误与处理

错误不返回 Selenium 的 Stack Trace,而是返回以下代码(isError: true)。

{
  "error": "ELEMENT_NOT_FOUND",
  "message": "Element was not found: id=searchButton",
  "selector": { "by": "id", "value": "searchButton" }
}

错误代码

含义

处理

BROWSER_NOT_STARTED

浏览器未启动

调用 browser_start

ELEMENT_NOT_FOUND

元素或 frame 未找到

使用 inspect_page 确认实际元素,重新检查 Selector

TIMEOUT

wait_for 的条件未满足

重新检查条件和 timeoutMs。画面可能与预期不同

WINDOW_NOT_FOUND

指定的 Window 不存在

重新检查 switch_window 的 index

NAVIGATION_FAILED

导航失败

检查 URL、网络、认证

DRIVER_LOST

IEDriver / Edge 异常终止

通过 browser_start 重新启动(参见下文)

URL_NOT_ALLOWED

不在允许列表中的 Origin

重新检查 IE_MCP_ALLOWED_ORIGINS

INVALID_ARGUMENT

参数无效

确认 Tool 的输入规范

INTERNAL_ERROR

其他(包括启动失败)

检查 message 和 stderr 的日志

从 DRIVER_LOST 恢复

如果浏览器或 Driver 崩溃,内部 WebDriver 将被销毁,后续操作将返回 BROWSER_NOT_STARTED。不会自动恢复或自动重试前一步操作(为了防止重复注册等 副作用)。代理端需要重新调用 browser_start,通过 inspect_page 确认屏幕状态后 再继续操作。由于前一步操作可能已经成功,因此不应直接重新执行注册、更新类操作。


11. 日志

stdout 用于 MCP 协议,因此所有日志均输出到 stderr,每行一个 JSON。

{"level":"info","event":"started","transport":"stdio"}
{"level":"info","tool":"navigate","url":"http://legacy01.local/customer","durationMs":842}
{"level":"info","tool":"type","selector":{"by":"id","value":"password"},"textLength":16,"durationMs":128}
{"level":"error","tool":"click","selector":{"by":"id","value":"x"},"error":"ELEMENT_NOT_FOUND","message":"Element was not found: id=x","durationMs":5012}

不记录输入字符串本身、Cookie、认证信息、完整 HTML(type 仅记录字符数)。 如果需要保存到文件,请重定向 stderr。

node dist/index.js 2>> C:\logs\ie-mode-mcp.log

12. 故障排除

症状

确认事项

browser_start 出现 INTERNAL_ERROR

IE_MCP_DRIVER_PATH 是否正确。能否单独启动 IEDriverServer.exe

出现保护模式相关异常

Internet 选项 → 安全性 中统一所有区域的保护模式设置

出现缩放相关异常

将 Edge / IE 的缩放恢复为 100%

Edge 启动但未进入 IE 模式

确认 IE 模式的策略(站点列表等)。先手动确认能否以 IE 模式显示

操作卡住·无法点击元素

窗口是否最小化或非激活。远程桌面断开期间会不稳定

inspect_page 的元素为空

是否在 frame 内画面(指定 frame 重新获取)。通过 screenshot 确认实际画面

Agent 侧看不到 Tool

是否使用绝对路径指定了 dist/index.js。是否已执行 npm run build

标准输出无任何输出

正常。日志输出到 stderr

screenshot 对原因调查有效。可以确认仅靠 DOM 信息无法判断的状态(模态框、认证对话框、 渲染异常)。


13. 开发

src/
├─ index.ts      MCP Server のエントリーポイント(stdio)
├─ config.ts     環境変数と stderr ログ
├─ tools.ts      MCP Tool の Schema と Handler
├─ browser.ts    BrowserManager(Selenium / IEDriver 操作の集約)
├─ selectors.ts  Selector → Selenium の By 変換
└─ errors.ts     Selenium Error → MCP Error Code 変換
npm run build   # tsc でビルド
npm start       # node dist/index.js
  • MCP Tool 不直接操作 Selenium,必须通过 BrowserManager 进行。

  • 所有 WebDriver 操作均通过 Promise Chain 串行化,即使 Tool 被并行调用, 也只会逐条发送给 IEDriver。

  • 仅重试无副作用的操作(元素搜索·窗口句柄检测)。不重试 click 或提交操作。


14. 限制事项

初始实现不支持以下内容。

多浏览器会话 / 多用户 / HTTP Transport / REST API / DB / 会话持久化 / 自动浏览器恢复 / 复杂重试策略 / WebDriver Grid / 通用 Selenium API / executeScript Tool / 多层 iframe(仅支持一层)/ Element Cache / Metrics / 审批流 / 认证与授权

Available Tools

10 tools
browser_closeClose browserB

Close the browser session. Safe to call repeatedly.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are available so the description carries the full burden. It only says 'close the browser session' and 'safe to call repeatedly', but does not disclose whether this terminates all browser state or if there are side effects on open windows, tabs, or downloads. The behavioral context is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences that are front-loaded and to the point. Every sentence adds value: the first states the action, the second clarifies safety/repeatability. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and no output schema, the description covers the basic purpose and safety. However, it lacks details on what happens after closing (e.g., can browser_start reopen cleanly) or any cleanup behavior, which might be useful context for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema coverage is 100%. The description adds value by stating it is safe to call repeatedly, which implies no parameters are needed and calls are idempotent. With no parameters to explain, this is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool closes the browser session with a specific verb and resource. It distinguishes enough from siblings like 'navigate' which moves within a session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description notes it is safe to call repeatedly, which implies idempotency, but does not explicitly tell when to call it (e.g., end of a browsing task) or when not to (e.g., still need to interact). No sibling differentiation is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_startStart Edge IE ModeA

Start Microsoft Edge in IE Mode through IEDriverServer. Only one browser session exists; calling this while a session is running returns the existing one. Also use this to recover after a DRIVER_LOST error.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the full burden of behavioral disclosure. It transparently reveals that only one browser session exists, that calling the tool again returns the existing session, and that it can be used for recovery. This is strong for a start tool, but it could additionally mention potential side effects like timeouts or prerequisites for the IEDriverServer.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long, front-loading the primary purpose in the first sentence and adding behavioral nuance in the second. Every sentence provides essential information without redundancy or fluff, achieving maximum conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, no output schema, straightforward start action), the description is complete. It covers the core function, the singleton behavior, error recovery, and is sufficient for an AI agent to understand when and how to invoke the tool alongside its sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and 100% schema coverage, so the baseline is 4. The description adds no parameter information, which is appropriate since there are none to document. No additional semantic value is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool starts Microsoft Edge in IE Mode via IEDriverServer, using the specific verb 'Start' and the resource 'Microsoft Edge in IE Mode'. It also distinguishes itself from sibling tools by noting that only one browser session exists and that calling it again returns the existing session, which is unique among the provided sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool: normally to start the browser, and also to recover after a DRIVER_LOST error. It implicitly advises against calling it multiple times for new sessions by stating that subsequent calls return the existing session. However, it does not explicitly list alternatives or state when not to use it, though no alternative starting tool exists among the siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clickClick elementA

Click an element after waiting for it to be visible and enabled. This operation is never retried automatically, because a repeated click may submit or register data twice.

ParametersJSON Schema
NameRequiredDescriptionDefault
frameNoOptional iframe/frame to switch into first. One level of nesting is supported.
selectorYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses two key behaviors: waiting for the element to be visible and enabled, and the lack of automatic retry with a rationale. However, it does not mention timeout behavior, scroll-into-view, or what happens if the element is not found.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose, and every sentence adds value. No fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's potential side effects (e.g., triggering navigation or form submission), the description is minimal. It does not mention return values, scroll behavior, or failure modes. It is adequate for a simple click but lacks completeness for an AI agent to fully anticipate outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%, and the description adds no information about the parameters. It does not explain the 'frame' or 'selector' parameters beyond what is already in the schema. The description should compensate for the missing schema descriptions but fails to do so.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Click an element after waiting for it to be visible and enabled,' using a specific verb and resource. It distinguishes the tool from siblings like 'type' and 'select' by specifying the action and the precondition (visibility and enabled state).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly warns that the operation is never retried automatically because a repeated click may submit or register data twice. This gives a clear usage caution about retries, though it does not explicitly compare to alternative tools or state when not to use click.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_pageInspect pageA

Return the current URL, title, visible page text and the operable elements (a, button, input, textarea, select, iframe). The full HTML is never returned. Pass frame to inspect the contents of an iframe listed by a previous inspect_page call.

ParametersJSON Schema
NameRequiredDescriptionDefault
frameNoOptional iframe/frame to switch into first. One level of nesting is supported.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of disclosing behavior. It explicitly states 'The full HTML is never returned' and that frame must reference an iframe from a previous call. The read-only nature is implied by 'Return' but not stated outright, though this is likely sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two succinct sentences that front-load the main purpose and then add the iframe caveat. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has one well-described parameter, no output schema, and no annotations. The description covers the output, a key constraint (no full HTML), and iframe usage, making it reasonably complete for an inspection tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by clarifying that the frame must come from a previous inspect_page call, which is not stated in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') and enumerates exactly what is returned (URL, title, visible text, operable elements). This clearly distinguishes it from sibling tools like 'click' or 'navigate'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool (to inspect page state) and gives specific guidance for iframe usage ('Pass frame to inspect the contents of an iframe listed by a previous inspect_page call'). It doesn't explicitly exclude alternatives, but the context is clear given the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshotScreenshotB

Capture the current browser window as a PNG image.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears the full burden. It states the action (capture) and format (PNG) but omits crucial details: whether it modifies state, if a browser window must be open, what exactly 'current browser window' captures (viewport vs full page), and if there are side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single 9-word sentence, efficient and front-loaded. However, it could include additional essential context (e.g., 'captures the visible viewport area') without losing conciseness, making it slightly under-specified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and no output schema, the description should clarify what the tool returns (e.g., base64 PNG data). It only says 'as a PNG image' but doesn't confirm the output type. The scope of 'current browser window' is ambiguous, and prerequisites are missing, leaving the agent uncertain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, and the schema coverage is 100% (trivially). The description adds minimal meaning by specifying 'current browser window' as the implicit input. A baseline of 4 is appropriate given no parameters to document.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('capture') and resource ('current browser window') with a clear output format ('PNG image'). It is distinct from sibling tools like 'navigate' or 'inspect_page' which serve different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use screenshot versus alternatives. Despite having sibling tools (e.g., inspect_page, wait_for), no exclusions or context is given. An agent must infer use case from tool purpose alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

selectSelect optionB

Choose an option of an HTML element by visible text, value or index.

ParametersJSON Schema
NameRequiredDescriptionDefault
byYesHow to identify the option.
frameNoOptional iframe/frame to switch into first. One level of nesting is supported.
valueYesOption text, value, or zero-based index.
selectorYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It discloses that selection is based on visible text, value or index, which is helpful. However, it doesn't mention side effects (e.g., whether the change triggers JavaScript events), error handling (e.g., what if option not found), or scope (e.g., operates within current page context).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that efficiently communicates the core action and identification methods. It is front-loaded with the key verb and resource. No waste, though it could optionally add a brief usage hint without breaching conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters with nested objects and no output schema, the description is somewhat complete but lacks coverage of return behavior (e.g., what happens on success/failure), frame handling nuances, and edge cases. For a selection action in a browser automation context, more behavioral detail would be helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, meaning most parameters are documented in the schema. The description adds that selection can be by 'visible text, value or index', which maps to the 'by' enum, and that the 'value' parameter can be text or zero-based index. This provides modest added meaning beyond the schema, but the 'frame' and 'selector' objects remain documented primarily in schema, not description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'choose' and resource 'HTML <select> element', specifying three identification methods (visible text, value, index). This distinguishes it from sibling tools like click or type, though it doesn't explicitly contrast with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use by saying 'choose an option of an HTML <select> element', which suggests this is for dropdown selections. However, it does not provide explicit when-not-to-use guidance, mention prerequisites (e.g., element must exist), or compare with alternatives like click on an option directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

switch_windowSwitch windowA

Switch to another browser window or popup. Use target:"newest" after an action that opens a window, or index to select a window by its zero-based position.

ParametersJSON Schema
NameRequiredDescriptionDefault
indexNoZero-based window index.
targetNoSwitch to the newest window.
timeoutMsNoHow long to poll for a new window. Defaults to IE_MCP_TIMEOUT_MS.

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It mentions polling behavior via timeoutMs parameter but does not state if switching is destructive, if it requires a window to exist, what happens if the window is closed, or any state changes. The description does not disclose potential side effects or preconditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with the core purpose in the first sentence. The second sentence adds specific usage hints. It could potentially omit 'or popup' as redundant with 'window', but overall concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 3 unrequired parameters, no output schema, no annotations, the description covers the basic purpose and usage hints. However, it lacks details on return values, error scenarios (e.g., window not found), or behavior when switching to a window that fails to load.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds context for the 'target' parameter (use after action that opens window) and 'index' (zero-based position), but the timeoutMs parameter meaning is already clear from schema. No additional semantic value beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool switches to another browser window or popup, specifying the verb 'switch' and the resource 'browser window or popup'. It distinguishes itself from sibling tools like browser_start, browser_close, and navigate by focusing on window selection rather than creation, closure, or navigation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool: after an action that opens a window, use target 'newest', or use index to select by position. It implicitly distinguishes from sibling tools by indicating this is for window focus rather than content navigation or page interaction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

typeType textA

Type text into an input or textarea. Set clear to false to append instead of replacing.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesText to send to the element.
clearNoClear the field first. Default true.
frameNoOptional iframe/frame to switch into first. One level of nesting is supported.
selectorYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that the tool can clear or append text via the 'clear' parameter, which is good. However, it does not mention potential side effects (e.g., triggering change events), error conditions (element not found), or behavior when the element is not a text input. This is adequate but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at one sentence plus one usage tip. Every word earns its place, clearly stating the action and a key parameter behavior. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description should ideally mention return values (e.g., success indicator, element state). It doesn't, leaving that unclear. With nested objects (selector, frame) and no explanation of selector strategies beyond the schema enums, it completes the basic usage but misses context on what happens after typing (e.g., waits for stability, triggers events).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is high at 75%, so the schema documents most parameters well. The description adds value by explaining the 'clear' boolean behavior (append vs replace) beyond the schema's default value note. It doesn't add to 'selector' or 'frame' parameters, which are already well-described in the schema, so this is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool types text into an input or textarea, using a specific verb and resource. It distinguishes from siblings like 'click' or 'select' by targeting text entry specifically, but doesn't differentiate from a potential 'send_keys' equivalent if one existed among siblings, so a slight deduction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a key usage guideline: set clear to false to append instead of replacing text. This gives basic advice on when to use a parameter. However, it lacks guidance on when to use this tool versus alternatives like clicking an element first or waiting, and doesn't mention prerequisites (e.g., element must be visible/interactable).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_forWait for conditionA

Wait until a condition holds. present/visible/enabled/text require a selector; text/url/title require text, which is matched as a substring. Use this instead of sleeping after an action.

ParametersJSON Schema
NameRequiredDescriptionDefault
textNoExpected substring for text/url/title conditions.
typeYesCondition to wait for.
frameNoOptional iframe/frame to switch into first. One level of nesting is supported.
selectorNo
timeoutMsNoTimeout in milliseconds. Defaults to IE_MCP_TIMEOUT_MS.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description explains key behavioral traits: that present/visible/enabled/text require a selector, text/url/title require text matched as substring, and that it waits for the condition. With no annotations provided, the description carries the full burden of transparency. It lacks details on timeout behavior or error handling, but covers core usage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with two sentences, front-loading the key purpose and condition types. Every sentence adds value, avoiding any redundancy. The structure is efficient for an AI agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (5 parameters, nested objects, no output schema), the description is adequate. It explains the core waiting concept and parameter dependencies. However, it lacks details on return values or what happens on timeout/failure, which the schema alone doesn't cover. The sibling 'inspect_page' might share similar conditions, but no differentiation is made.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, so the schema already documents most parameters well. The description adds value by clarifying the relationship between condition types and required parameters (e.g., 'present/visible/enabled/text require a selector; text/url/title require text'). This bridges gaps between parameters, though it does not detail the 'frame' parameter beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool waits until a condition holds, with specific verb+resource ('Wait for condition'). It lists the condition types and distinguishes itself from sleeping after an action, which differentiates it from sibling tools like 'click' or 'navigate'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool ('Use this instead of sleeping after an action'), providing clear guidance on avoiding poor alternatives. However, it does not specify when not to use it or which sibling would be more appropriate for different scenarios, such as synchronous checks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 10 tool updatesv0.1.0
    • First observedbrowser_close
    • First observedbrowser_start
    • First observedclick
    • First observedinspect_page
    • First observednavigate
    • First observedscreenshot
    • First observedselect
    • First observedswitch_window
    • First observedtype
    • First observedwait_for

TDQS

A3.8/5.0

Scored across 10 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: session management (start/close), navigation, inspection, interaction (click, type, select), window switching, waiting, and screenshot. No two tools overlap in functionality.

Naming Consistency3/5

The naming pattern is inconsistent: some tools use a 'browser_' prefix (browser_start, browser_close), while others are bare verbs (navigate, click, type) or compound snake_case (inspect_page, switch_window, wait_for). This mix of styles could cause confusion.

Tool Count5/5

With 10 tools, the set is well-scoped for browser automation. It covers session lifetime, navigation, element interaction, inspection, window handling, and waiting without being bloated or too thin.

Completeness3/5

The tools cover fundamental browser actions but miss common features like back/forward navigation, JavaScript execution, alert handling, or cookie management. The set is functional for basic scenarios but has notable gaps for comprehensive automation.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Exposes Selenium WebDriver as an MCP server, enabling AI agents and LLMs to control real browsers for automation tasks like navigation, element interaction, and screenshot capture.
    22
    21 PyPI
    3
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables LLMs to drive Edge in IE mode for automating legacy IE-only web applications, supporting tasks like clicking, filling forms, and data extraction via Selenium.
    28
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables AI agents and MCP clients to automate web browsers via Selenium WebDriver, supporting Chrome, Firefox, and Edge in headless or visible mode with tools for navigation, interaction, content extraction, screenshots, and scripting.
    21
    31 npm
    MIT