Skip to main content
Glama

ie-mode-mcp

Microsoft Edge의 IE 모드에서 작동하는 레거시 웹 애플리케이션을 AI 에이전트에서 MCP (Model Context Protocol)를 통해 조작하기 위한 MCP Server입니다.

AI Agent ──(MCP / stdio)──> ie-mode-mcp ──> BrowserManager ──> selenium-webdriver
                                                                     │
                                                          IEDriverServer.exe
                                                                     │
                                                     Microsoft Edge (IE Mode)
                                                                     │
                                                       Legacy Web Application
  • Node.js 22 / TypeScript / selenium-webdriver만으로 구성됨 (HTTP Server, DB, DI, Logging Framework 없음)

  • MCP Transport는 stdio만 지원

  • 브라우저 세션은 1개만, WebDriver 조작은 완전 순차 실행

  • HTML 전체를 반환하지 않고, inspect_page가 LLM을 위해 요약한 화면 정보를 반환

  • 승인 흐름 없음. Tool을 호출한 시점에 조작을 실행


목차

  1. 퀵스타트

  2. 전제 조건

  3. Windows 측 사전 설정

  4. 설치 및 빌드

  5. 환경 변수

  6. 시작 방법

  7. AI Agent 등록

  8. Tool 레퍼런스

  9. 사용 예시

  10. 에러 및 대처

  11. 로그

  12. 트러블슈팅

  13. 개발

  14. 제한 사항


Related MCP server: ie-mcp

1. 퀵스타트

Windows에서 다음을 실행합니다.

git clone https://github.com/sumikof/iedriver-mcp.git
cd iedriver-mcp
npm install
npm run build

# IEDriverServer.exe のパスと、遷移を許可する Origin を指定して起動
$env:IE_MCP_DRIVER_PATH = "C:\tools\IEDriverServer.exe"
$env:IE_MCP_ALLOWED_ORIGINS = "http://legacy01.local"
node dist/index.js

{"level":"info","event":"started","transport":"stdio"}가 stderr에 출력되면 시작 성공입니다. 일반적으로 수동으로 시작하지 않고, AI Agent 측의 MCP 설정에서 자동 시작시킵니다.


2. 전제 조건

항목

내용

OS

Windows 11 / Windows 10 (로그인된 대화형 세션)

Node.js

22 이상

브라우저

Microsoft Edge (IE 모드 사용 가능)

Driver

IEDriverServer.exe (Selenium 4.x 계열. 32비트 버전 권장)

  • IEDriverServer.exe는 Selenium 다운로드 페이지에서 받아, 원하는 폴더 (예: C:\tools\)에 배치합니다. 64비트 버전에는 알려진 제약이 있으므로, Selenium 공식에서는 32비트 버전 사용을 권장합니다.

  • IEDriver는 GUI, 창 포커스, 네이티브 이벤트의 영향을 받으므로, 전용 Windows VM 또는 전용 Windows 세션에서 사용하는 것을 권장합니다.

  • Windows Service (Session 0)에서 브라우저를 작동시키는 구성은 가정하지 않습니다.

  • MCP Server와 IEDriver / Edge는 동일한 Windows 환경에서 작동시킵니다.


3. Windows 측 사전 설정

IEDriver는 환경 설정의 영향을 크게 받습니다. 먼저 수동으로 설정을 완료한 후 MCP Server를 시작합니다.

3.1 Edge의 IE 모드 사용 가능하게 하기

대상 사이트가 IE 모드로 열리는지, 먼저 Edge의 수동 조작으로 확인해 둡니다. IE 모드는 다음 정책 중 하나로 활성화합니다 (Software\Policies\Microsoft\Edge 아래).

정책 (표시 이름)

레지스트리 값 이름

Configure Internet Explorer integration

InternetExplorerIntegrationLevel

Configure the Enterprise Mode Site List

InternetExplorerIntegrationSiteList

Send all intranet sites to Internet Explorer

(Edge 77 이후 그룹 정책에서 설정)

구체적인 구성은 조직의 정책에 따라 달라지므로, 자세한 내용은 Microsoft의 IE 모드 문서 와 소속 조직의 관리자에게 확인하십시오. Windows / Edge는 최신 업데이트를 적용해 둡니다.

3.2 IEDriver가 요구하는 설정

항목

필요한 상태

본 Server에서의 처리

브라우저 줌

100%

ignoreZoomSetting(true)가 설정되어 있어 필수는 아니지만, 100% 권장

보호 모드 (Protected Mode)

모든 영역에서 동일한 설정

통일되지 않은 경우 시작 시 예외 발생. 인터넷 옵션 → 보안에서 통일

IEDriverServer 비트 수

32비트 권장

—

보호 모드 설정이 통일되지 않으면 browser_start가 실패합니다. IEDriver의 introduceFlakinessByIgnoringProtectedModeSettings는 동작이 불안정해지므로 사용하지 않습니다.


4. 설치 및 빌드

npm install     # 依存パッケージの取得
npm run build   # TypeScript を dist/ へビルド

산출물은 dist/index.js입니다. 빌드 후 npm start (= node dist/index.js)로도 시작할 수 있습니다.


5. 환경 변수

설정 파일 (YAML / JSON)은 사용하지 않고, 환경 변수로만 설정합니다.

환경 변수

설명

기본값

IE_MCP_EDGE_PATH

msedge.exe 경로

미지정 (IEDriver가 자동 검색)

IE_MCP_DRIVER_PATH

IEDriverServer.exe 경로

미지정 (PATH에서 탐색)

IE_MCP_ALLOWED_ORIGINS

navigate를 허용할 Origin의 쉼표 구분. *로 무제한

*

IE_MCP_TIMEOUT_MS

요소 검색 및 대기의 기본 타임아웃 (ms)

10000

IE_MCP_EDGE_PATH=C:\Program Files (x86)\Microsoft\Edge\Application\msedge.exe
IE_MCP_DRIVER_PATH=C:\tools\IEDriverServer.exe
IE_MCP_ALLOWED_ORIGINS=http://legacy01.local,http://legacy02.local
IE_MCP_TIMEOUT_MS=10000
  • IE Driver 4.5.0 이후는 IE 미탑재 환경 (Windows 11 기본)에서 Edge를 자동 검색하므로, IE_MCP_EDGE_PATH는 일반적으로 불필요합니다. 자동 검색에 실패하는 경우에만 명시적으로 지정합니다.

  • 운영의 재현성을 중시하는 경우 IE_MCP_DRIVER_PATH를 명시적으로 지정하는 것을 권장합니다.

  • IE_MCP_ALLOWED_ORIGINS는 오조작 방지용 간이 제한이며, Origin (scheme + host + port) 의 완전 일치로 판단합니다. 경로 단위의 제한은 하지 않습니다.


6. 시작 방법

수동 시작 (동작 확인용)

PowerShell:

$env:IE_MCP_DRIVER_PATH = "C:\tools\IEDriverServer.exe"
$env:IE_MCP_ALLOWED_ORIGINS = "http://legacy01.local"
node dist/index.js

명령 프롬프트:

set IE_MCP_DRIVER_PATH=C:\tools\IEDriverServer.exe
set IE_MCP_ALLOWED_ORIGINS=http://legacy01.local
node dist\index.js

stdio에서 클라이언트의 연결을 기다립니다. 표준 입출력이 MCP 프로토콜에 사용되므로, 이 상태에서 키보드 입력해도 응답이 없습니다 (정상). 모든 로그는 stderr에 출력됩니다. 종료는 Ctrl+C (브라우저도 자동으로 닫힘).

주의: MCP Server 시작만으로는 브라우저가 시작되지 않습니다. 브라우저는 Agent가 browser_start 를 호출한 시점에 시작됩니다.

일반 운영

AI Agent (MCP 클라이언트)가 본 Server를 자식 프로세스로 시작합니다. 수동 시작은 불필요합니다. 다음 장의 설정을 수행합니다.


7. AI Agent 등록

MCP 클라이언트의 설정 파일에 다음을 추가합니다.

{
  "mcpServers": {
    "ie-mode": {
      "command": "node",
      "args": ["C:\\ie-mode-mcp\\dist\\index.js"],
      "env": {
        "IE_MCP_DRIVER_PATH": "C:\\tools\\IEDriverServer.exe",
        "IE_MCP_EDGE_PATH": "C:\\Program Files (x86)\\Microsoft\\Edge\\Application\\msedge.exe",
        "IE_MCP_ALLOWED_ORIGINS": "http://legacy01.local,http://legacy02.local",
        "IE_MCP_TIMEOUT_MS": "10000"
      }
    }
  }
}
  • 경로는 JSON 내에서 백슬래시를 이스케이프합니다 (C:\\...).

  • args에는 빌드 후 dist/index.js의 절대 경로를 지정합니다.

  • Claude Code의 경우 claude mcp add로도 등록할 수 있습니다.

claude mcp add ie-mode --env IE_MCP_DRIVER_PATH=C:\tools\IEDriverServer.exe --env IE_MCP_ALLOWED_ORIGINS=http://legacy01.local -- node C:\ie-mode-mcp\dist\index.js

등록 후, 클라이언트 측에서 browser_start를 포함한 10개의 Tool이 보이면 연결 성공입니다.


8. Tool 레퍼런스

공개하는 Tool은 10개입니다. WebDriver의 저수준 API (findElement / executeScript 등)는 공개하지 않습니다.

Tool

입력

개요

browser_start

없음

Edge IE Mode를 시작합니다. 이미 시작된 경우 기존 세션 재사용

browser_close

없음

브라우저를 종료합니다. 여러 번 호출해도 오류가 발생하지 않음

navigate

url

URL Allowlist를 확인한 후 이동

inspect_page

frame?

URL / title / 화면 텍스트 / 조작 가능 요소 반환

click

selector, frame?

표시 및 활성화를 기다린 후 클릭

type

selector, frame?, text, clear?

input / textarea에 입력

select

selector, frame?, by, value

<select>의 option 선택

wait_for

type, selector?, frame?, text?, timeoutMs?

조건이 충족될 때까지 대기

switch_window

target:"newest" / index, timeoutMs?

팝업 및 다른 Window로 전환

screenshot

없음

현재 화면을 PNG (MCP image content)로 반환

공통: Selector

{ "by": "id | name | css | xpath | linkText", "value": "searchButton" }

레거시 웹 애플리케이션에서는 name과 xpath 사용 빈도가 높으므로 대응합니다.

공통: frame (iframe은 1계층)

모든 요소 조작 Tool은 임의의 frame을 받습니다. 지정하면 defaultContent로 돌아간 후 frame으로 전환하여 그 안에서 요소를 검색합니다.

{
  "frame": { "by": "name", "value": "mainFrame" },
  "selector": { "by": "id", "value": "searchButton" }
}

browser_start

{}
{ "status": "ready", "reused": false }

reused: true는 기존 세션을 그대로 사용했음을 나타냅니다. 기존 세션이 죽어있는 경우 자동으로 다시 시작합니다.

navigate

{ "url": "http://legacy01.local/customer" }
{ "url": "http://legacy01.local/customer", "title": "顧客検索" }

inspect_page

Agent가 화면을 이해하기 위한 주요 Tool입니다. HTML 전체를 반환하지 않고, URL / title / 표시 텍스트 / 조작 가능 요소 (a button input textarea select iframe)만 반환합니다. 비표시 요소와 type="hidden"의 input은 제외됩니다.

{ "frame": { "by": "name", "value": "mainFrame" } }
{
  "url": "http://legacy01.local/customer",
  "title": "顧客検索",
  "text": "顧客検索 顧客名 支店 検索",
  "elements": [
    { "tag": "input", "id": "customerName", "name": "customerName", "type": "text" },
    { "tag": "select", "id": "branch", "name": "branch", "text": "東京支店", "optionCount": 12 },
    { "tag": "button", "id": "searchButton", "text": "検索" },
    { "tag": "iframe", "name": "mainFrame" }
  ],
  "truncated": false
}
  • truncated: true는 요소가 상한 (300건)에서 잘렸음을 나타냅니다.

  • 요소 목록에 iframe이 포함된 경우, 그 내용을 보려면 frame을 지정하여 다시 호출합니다.

click

{ "selector": { "by": "id", "value": "searchButton" } }
{ "url": "http://legacy01.local/customer", "title": "顧客検索" }

표시 및 활성화될 때까지 기다린 후 클릭합니다. click은 자동 Retry하지 않습니다 (등록, 갱신, 전송이 이미 성공한 상태에서 재클릭으로 인한 이중 처리를 방지하기 위해).

type

{
  "selector": { "by": "id", "value": "customerName" },
  "text": "山田太郎",
  "clear": true
}

clear (기본값 true)가 true이면 clear() 후 입력, false이면 추가 입력합니다.

select

{
  "selector": { "by": "id", "value": "branch" },
  "by": "text",
  "value": "東京支店"
}
{ "text": "東京支店", "value": "13", "index": 2 }

by는 text / value / index (index는 0부터 시작).

wait_for

고정 sleep을 사용하지 않고 명시적으로 대기합니다.

{
  "type": "visible",
  "selector": { "by": "id", "value": "resultTable" },
  "timeoutMs": 10000
}

type

필요한 입력

조건

present

selector

요소가 DOM에 존재

visible

selector

요소가 표시됨

enabled

selector

요소가 표시되고 조작 가능

text

selector, text

요소 텍스트가 text 포함

url

text

현재 URL이 text 포함

title

text

title이 text 포함

timeoutMs 생략 시 IE_MCP_TIMEOUT_MS를 사용합니다.

switch_window

{ "target": "newest" }
{ "index": 1 }
{ "url": "http://legacy01.local/detail", "title": "顧客詳細", "index": 1, "windowCount": 2 }

newest는 새 Window Handle이 나타날 때까지 짧은 시간 폴링합니다. 검색할 수 없는 경우 현존하는 마지막 Window로 전환합니다.

screenshot

{}

PNG 이미지 (MCP의 image content)를 반환합니다. DOM만으로는 판단할 수 없는 레이아웃 및 오류 화면 확인에 사용합니다.


9. 사용 예시

기본 루프

browser_start → navigate → inspect_page → click / type / select → wait_for → inspect_page

inspect_page로 화면 파악 → 조작 → wait_for로 결과 대기 → 다시 inspect_page를 반복합니다.

예시: 고객 "야마다 타로" 검색 후 상세 화면 열기

#

Tool

인수

1

browser_start

{}

2

navigate

{ "url": "http://legacy01.local/customer" }

3

inspect_page

{}

4

type

{ "selector": { "by": "id", "value": "customerName" }, "text": "야마다 타로" }

5

select

{ "selector": { "by": "id", "value": "branch" }, "by": "text", "value": "도쿄 지점" }

6

click

{ "selector": { "by": "id", "value": "searchButton" } }

7

wait_for

{ "type": "visible", "selector": { "by": "id", "value": "resultTable" } }

8

inspect_page

{}

9

click

{ "selector": { "by": "linkText", "value": "야마다 타로" } }

10

wait_for

{ "type": "title", "text": "고객 상세" }

11

inspect_page

{}

예시: iframe 내부 조작

{"tool": "inspect_page", "args": {}}
{"tool": "inspect_page", "args": { "frame": { "by": "name", "value": "mainFrame" } }}
{"tool": "click", "args": {
  "frame": { "by": "name", "value": "mainFrame" },
  "selector": { "by": "id", "value": "searchButton" }
}}

frame 지정은 조작마다 매번 전달합니다 (내부에서 매번 defaultContent로 돌아간 후 전환하므로, 상태는 유지되지 않습니다).

예시: 팝업 조작 후 원래 Window로 돌아가기

{"tool": "click",         "args": { "selector": { "by": "id", "value": "openPopup" } }}
{"tool": "switch_window", "args": { "target": "newest" }}
{"tool": "inspect_page",  "args": {}}
{"tool": "switch_window", "args": { "index": 0 }}

10. 에러 및 대처

에러는 Selenium의 Stack Trace가 아닌, 다음 코드로 반환됩니다 (isError: true).

{
  "error": "ELEMENT_NOT_FOUND",
  "message": "Element was not found: id=searchButton",
  "selector": { "by": "id", "value": "searchButton" }
}

에러 코드

의미

대처

BROWSER_NOT_STARTED

브라우저 미시작

browser_start 호출

ELEMENT_NOT_FOUND

요소 또는 frame을 찾을 수 없음

inspect_page로 실제 요소 확인 후 Selector 재검토

TIMEOUT

wait_for 조건이 충족되지 않음

조건 및 timeoutMs 재검토. 화면이 예상과 다를 가능성

WINDOW_NOT_FOUND

지정 Window가 존재하지 않음

switch_window의 index 재검토

NAVIGATION_FAILED

이동 실패

URL, 네트워크, 인증 확인

DRIVER_LOST

IEDriver / Edge 비정상 종료

browser_start로 재시작 (아래 참조)

URL_NOT_ALLOWED

Allowlist 외부 Origin

IE_MCP_ALLOWED_ORIGINS 재검토

INVALID_ARGUMENT

인수 오류

Tool 입력 사양 확인

INTERNAL_ERROR

기타 (시작 실패 포함)

message와 stderr 로그 확인

DRIVER_LOST로부터 복구

브라우저 또는 Driver가 다운된 경우, 내부 WebDriver는 폐기되고 이후 조작은 BROWSER_NOT_STARTED가 됩니다. 자동 복구 및 직전 조작 자동 재실행은 하지 않습니다 (이중 등록 등의 부작용을 방지하기 위해). Agent 측에서 browser_start를 다시 호출하고, 화면 상태를 inspect_page로 확인한 후 조작을 재개합니다. 직전 조작이 이미 성공했을 가능성이 있으므로, 등록 및 갱신 계열의 조작을 그대로 재실행해서는 안 됩니다.


11. 로그

stdout은 MCP 프로토콜이 사용하므로, 로그는 모두 stderr에 JSON 1행으로 출력합니다.

{"level":"info","event":"started","transport":"stdio"}
{"level":"info","tool":"navigate","url":"http://legacy01.local/customer","durationMs":842}
{"level":"info","tool":"type","selector":{"by":"id","value":"password"},"textLength":16,"durationMs":128}
{"level":"error","tool":"click","selector":{"by":"id","value":"x"},"error":"ELEMENT_NOT_FOUND","message":"Element was not found: id=x","durationMs":5012}

입력 문자열 자체, 쿠키, 인증 정보, HTML 전체는 기록하지 않습니다 (type은 문자 수만). 파일에 남기려면 stderr를 리디렉션합니다.

node dist/index.js 2>> C:\logs\ie-mode-mcp.log

12. 트러블슈팅

증상

확인할 사항

browser_start가 INTERNAL_ERROR가 되는 경우

IE_MCP_DRIVER_PATH가 올바른지 확인. IEDriverServer.exe를 단독으로 실행할 수 있는지 확인

보호 모드 관련 예외가 발생하는 경우

인터넷 옵션 → 보안에서 모든 영역의 보호 모드 설정을 통일

확대/축소 관련 예외가 발생하는 경우

Edge/IE의 확대/축소를 100%로 되돌리기

Edge는 실행되지만 IE 모드가 되지 않는 경우

IE 모드 정책(사이트 목록 등)을 확인. 수동으로 IE 모드 표시가 가능한지 먼저 확인

조작이 멈추거나 요소를 클릭할 수 없는 경우

창이 최소화 또는 비활성화되지 않았는지 확인. 원격 데스크톱 연결이 끊어져 있으면 불안정해짐

inspect_page의 요소가 비어 있는 경우

프레임 내의 화면인지 확인(frame을 지정하여 재취득). screenshot으로 실제 화면 확인

Agent 측에 Tool이 보이지 않는 경우

dist/index.js를 절대 경로로 지정했는지 확인. npm run build를 실행했는지 확인

표준 출력에 아무것도 나오지 않는 경우

정상. 로그는 stderr로 출력됨

screenshot은 원인 조사에 유효. DOM 정보만으로는 판단할 수 없는 상태(모달, 인증 대화상자, 렌더링 깨짐)를 확인할 수 있음.


13. 개발

src/
├─ index.ts      MCP Server のエントリーポイント(stdio)
├─ config.ts     環境変数と stderr ログ
├─ tools.ts      MCP Tool の Schema と Handler
├─ browser.ts    BrowserManager(Selenium / IEDriver 操作の集約)
├─ selectors.ts  Selector → Selenium の By 変換
└─ errors.ts     Selenium Error → MCP Error Code 変換
npm run build   # tsc でビルド
npm start       # node dist/index.js
  • MCP Tool은 Selenium을 직접 다루지 않고, 반드시 BrowserManager를 경유한다.

  • 모든 WebDriver 조작은 Promise Chain으로 직렬화되어 있으며, Tool이 병렬로 호출되어도 IEDriver에는 1건씩만 전송된다.

  • 부작용이 없는 조작(요소 검색, Window Handle 검출)만 Retry한다. click이나 전송은 Retry하지 않는다.


14. 제한 사항

초기 구현에서는 다음을 지원하지 않는다.

다중 브라우저 세션 / 다중 사용자 / HTTP Transport / REST API / DB / 세션 영속화 / 자동 브라우저 복구 / 복잡한 Retry Policy / WebDriver Grid / 범용 Selenium API / executeScript Tool / 다단계 iframe(1계층만) / Element Cache / Metrics / 승인 흐름 / 인증·인가

Available Tools

10 tools
browser_closeClose browserB

Close the browser session. Safe to call repeatedly.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are available so the description carries the full burden. It only says 'close the browser session' and 'safe to call repeatedly', but does not disclose whether this terminates all browser state or if there are side effects on open windows, tabs, or downloads. The behavioral context is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences that are front-loaded and to the point. Every sentence adds value: the first states the action, the second clarifies safety/repeatability. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and no output schema, the description covers the basic purpose and safety. However, it lacks details on what happens after closing (e.g., can browser_start reopen cleanly) or any cleanup behavior, which might be useful context for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema coverage is 100%. The description adds value by stating it is safe to call repeatedly, which implies no parameters are needed and calls are idempotent. With no parameters to explain, this is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool closes the browser session with a specific verb and resource. It distinguishes enough from siblings like 'navigate' which moves within a session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description notes it is safe to call repeatedly, which implies idempotency, but does not explicitly tell when to call it (e.g., end of a browsing task) or when not to (e.g., still need to interact). No sibling differentiation is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_startStart Edge IE ModeA

Start Microsoft Edge in IE Mode through IEDriverServer. Only one browser session exists; calling this while a session is running returns the existing one. Also use this to recover after a DRIVER_LOST error.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the full burden of behavioral disclosure. It transparently reveals that only one browser session exists, that calling the tool again returns the existing session, and that it can be used for recovery. This is strong for a start tool, but it could additionally mention potential side effects like timeouts or prerequisites for the IEDriverServer.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long, front-loading the primary purpose in the first sentence and adding behavioral nuance in the second. Every sentence provides essential information without redundancy or fluff, achieving maximum conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, no output schema, straightforward start action), the description is complete. It covers the core function, the singleton behavior, error recovery, and is sufficient for an AI agent to understand when and how to invoke the tool alongside its sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and 100% schema coverage, so the baseline is 4. The description adds no parameter information, which is appropriate since there are none to document. No additional semantic value is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool starts Microsoft Edge in IE Mode via IEDriverServer, using the specific verb 'Start' and the resource 'Microsoft Edge in IE Mode'. It also distinguishes itself from sibling tools by noting that only one browser session exists and that calling it again returns the existing session, which is unique among the provided sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool: normally to start the browser, and also to recover after a DRIVER_LOST error. It implicitly advises against calling it multiple times for new sessions by stating that subsequent calls return the existing session. However, it does not explicitly list alternatives or state when not to use it, though no alternative starting tool exists among the siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clickClick elementA

Click an element after waiting for it to be visible and enabled. This operation is never retried automatically, because a repeated click may submit or register data twice.

ParametersJSON Schema
NameRequiredDescriptionDefault
frameNoOptional iframe/frame to switch into first. One level of nesting is supported.
selectorYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses two key behaviors: waiting for the element to be visible and enabled, and the lack of automatic retry with a rationale. However, it does not mention timeout behavior, scroll-into-view, or what happens if the element is not found.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose, and every sentence adds value. No fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's potential side effects (e.g., triggering navigation or form submission), the description is minimal. It does not mention return values, scroll behavior, or failure modes. It is adequate for a simple click but lacks completeness for an AI agent to fully anticipate outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%, and the description adds no information about the parameters. It does not explain the 'frame' or 'selector' parameters beyond what is already in the schema. The description should compensate for the missing schema descriptions but fails to do so.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Click an element after waiting for it to be visible and enabled,' using a specific verb and resource. It distinguishes the tool from siblings like 'type' and 'select' by specifying the action and the precondition (visibility and enabled state).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly warns that the operation is never retried automatically because a repeated click may submit or register data twice. This gives a clear usage caution about retries, though it does not explicitly compare to alternative tools or state when not to use click.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_pageInspect pageA

Return the current URL, title, visible page text and the operable elements (a, button, input, textarea, select, iframe). The full HTML is never returned. Pass frame to inspect the contents of an iframe listed by a previous inspect_page call.

ParametersJSON Schema
NameRequiredDescriptionDefault
frameNoOptional iframe/frame to switch into first. One level of nesting is supported.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of disclosing behavior. It explicitly states 'The full HTML is never returned' and that frame must reference an iframe from a previous call. The read-only nature is implied by 'Return' but not stated outright, though this is likely sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two succinct sentences that front-load the main purpose and then add the iframe caveat. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has one well-described parameter, no output schema, and no annotations. The description covers the output, a key constraint (no full HTML), and iframe usage, making it reasonably complete for an inspection tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by clarifying that the frame must come from a previous inspect_page call, which is not stated in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') and enumerates exactly what is returned (URL, title, visible text, operable elements). This clearly distinguishes it from sibling tools like 'click' or 'navigate'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool (to inspect page state) and gives specific guidance for iframe usage ('Pass frame to inspect the contents of an iframe listed by a previous inspect_page call'). It doesn't explicitly exclude alternatives, but the context is clear given the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshotScreenshotB

Capture the current browser window as a PNG image.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears the full burden. It states the action (capture) and format (PNG) but omits crucial details: whether it modifies state, if a browser window must be open, what exactly 'current browser window' captures (viewport vs full page), and if there are side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single 9-word sentence, efficient and front-loaded. However, it could include additional essential context (e.g., 'captures the visible viewport area') without losing conciseness, making it slightly under-specified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and no output schema, the description should clarify what the tool returns (e.g., base64 PNG data). It only says 'as a PNG image' but doesn't confirm the output type. The scope of 'current browser window' is ambiguous, and prerequisites are missing, leaving the agent uncertain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, and the schema coverage is 100% (trivially). The description adds minimal meaning by specifying 'current browser window' as the implicit input. A baseline of 4 is appropriate given no parameters to document.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('capture') and resource ('current browser window') with a clear output format ('PNG image'). It is distinct from sibling tools like 'navigate' or 'inspect_page' which serve different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use screenshot versus alternatives. Despite having sibling tools (e.g., inspect_page, wait_for), no exclusions or context is given. An agent must infer use case from tool purpose alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

selectSelect optionB

Choose an option of an HTML element by visible text, value or index.

ParametersJSON Schema
NameRequiredDescriptionDefault
byYesHow to identify the option.
frameNoOptional iframe/frame to switch into first. One level of nesting is supported.
valueYesOption text, value, or zero-based index.
selectorYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It discloses that selection is based on visible text, value or index, which is helpful. However, it doesn't mention side effects (e.g., whether the change triggers JavaScript events), error handling (e.g., what if option not found), or scope (e.g., operates within current page context).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that efficiently communicates the core action and identification methods. It is front-loaded with the key verb and resource. No waste, though it could optionally add a brief usage hint without breaching conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters with nested objects and no output schema, the description is somewhat complete but lacks coverage of return behavior (e.g., what happens on success/failure), frame handling nuances, and edge cases. For a selection action in a browser automation context, more behavioral detail would be helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, meaning most parameters are documented in the schema. The description adds that selection can be by 'visible text, value or index', which maps to the 'by' enum, and that the 'value' parameter can be text or zero-based index. This provides modest added meaning beyond the schema, but the 'frame' and 'selector' objects remain documented primarily in schema, not description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'choose' and resource 'HTML <select> element', specifying three identification methods (visible text, value, index). This distinguishes it from sibling tools like click or type, though it doesn't explicitly contrast with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use by saying 'choose an option of an HTML <select> element', which suggests this is for dropdown selections. However, it does not provide explicit when-not-to-use guidance, mention prerequisites (e.g., element must exist), or compare with alternatives like click on an option directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

switch_windowSwitch windowA

Switch to another browser window or popup. Use target:"newest" after an action that opens a window, or index to select a window by its zero-based position.

ParametersJSON Schema
NameRequiredDescriptionDefault
indexNoZero-based window index.
targetNoSwitch to the newest window.
timeoutMsNoHow long to poll for a new window. Defaults to IE_MCP_TIMEOUT_MS.

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It mentions polling behavior via timeoutMs parameter but does not state if switching is destructive, if it requires a window to exist, what happens if the window is closed, or any state changes. The description does not disclose potential side effects or preconditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with the core purpose in the first sentence. The second sentence adds specific usage hints. It could potentially omit 'or popup' as redundant with 'window', but overall concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 3 unrequired parameters, no output schema, no annotations, the description covers the basic purpose and usage hints. However, it lacks details on return values, error scenarios (e.g., window not found), or behavior when switching to a window that fails to load.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds context for the 'target' parameter (use after action that opens window) and 'index' (zero-based position), but the timeoutMs parameter meaning is already clear from schema. No additional semantic value beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool switches to another browser window or popup, specifying the verb 'switch' and the resource 'browser window or popup'. It distinguishes itself from sibling tools like browser_start, browser_close, and navigate by focusing on window selection rather than creation, closure, or navigation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool: after an action that opens a window, use target 'newest', or use index to select by position. It implicitly distinguishes from sibling tools by indicating this is for window focus rather than content navigation or page interaction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

typeType textA

Type text into an input or textarea. Set clear to false to append instead of replacing.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesText to send to the element.
clearNoClear the field first. Default true.
frameNoOptional iframe/frame to switch into first. One level of nesting is supported.
selectorYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that the tool can clear or append text via the 'clear' parameter, which is good. However, it does not mention potential side effects (e.g., triggering change events), error conditions (element not found), or behavior when the element is not a text input. This is adequate but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at one sentence plus one usage tip. Every word earns its place, clearly stating the action and a key parameter behavior. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description should ideally mention return values (e.g., success indicator, element state). It doesn't, leaving that unclear. With nested objects (selector, frame) and no explanation of selector strategies beyond the schema enums, it completes the basic usage but misses context on what happens after typing (e.g., waits for stability, triggers events).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is high at 75%, so the schema documents most parameters well. The description adds value by explaining the 'clear' boolean behavior (append vs replace) beyond the schema's default value note. It doesn't add to 'selector' or 'frame' parameters, which are already well-described in the schema, so this is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool types text into an input or textarea, using a specific verb and resource. It distinguishes from siblings like 'click' or 'select' by targeting text entry specifically, but doesn't differentiate from a potential 'send_keys' equivalent if one existed among siblings, so a slight deduction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a key usage guideline: set clear to false to append instead of replacing text. This gives basic advice on when to use a parameter. However, it lacks guidance on when to use this tool versus alternatives like clicking an element first or waiting, and doesn't mention prerequisites (e.g., element must be visible/interactable).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_forWait for conditionA

Wait until a condition holds. present/visible/enabled/text require a selector; text/url/title require text, which is matched as a substring. Use this instead of sleeping after an action.

ParametersJSON Schema
NameRequiredDescriptionDefault
textNoExpected substring for text/url/title conditions.
typeYesCondition to wait for.
frameNoOptional iframe/frame to switch into first. One level of nesting is supported.
selectorNo
timeoutMsNoTimeout in milliseconds. Defaults to IE_MCP_TIMEOUT_MS.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description explains key behavioral traits: that present/visible/enabled/text require a selector, text/url/title require text matched as substring, and that it waits for the condition. With no annotations provided, the description carries the full burden of transparency. It lacks details on timeout behavior or error handling, but covers core usage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with two sentences, front-loading the key purpose and condition types. Every sentence adds value, avoiding any redundancy. The structure is efficient for an AI agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (5 parameters, nested objects, no output schema), the description is adequate. It explains the core waiting concept and parameter dependencies. However, it lacks details on return values or what happens on timeout/failure, which the schema alone doesn't cover. The sibling 'inspect_page' might share similar conditions, but no differentiation is made.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, so the schema already documents most parameters well. The description adds value by clarifying the relationship between condition types and required parameters (e.g., 'present/visible/enabled/text require a selector; text/url/title require text'). This bridges gaps between parameters, though it does not detail the 'frame' parameter beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool waits until a condition holds, with specific verb+resource ('Wait for condition'). It lists the condition types and distinguishes itself from sleeping after an action, which differentiates it from sibling tools like 'click' or 'navigate'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool ('Use this instead of sleeping after an action'), providing clear guidance on avoiding poor alternatives. However, it does not specify when not to use it or which sibling would be more appropriate for different scenarios, such as synchronous checks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 10 tool updatesv0.1.0
    • First observedbrowser_close
    • First observedbrowser_start
    • First observedclick
    • First observedinspect_page
    • First observednavigate
    • First observedscreenshot
    • First observedselect
    • First observedswitch_window
    • First observedtype
    • First observedwait_for

TDQS

A3.8/5.0

Scored across 10 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: session management (start/close), navigation, inspection, interaction (click, type, select), window switching, waiting, and screenshot. No two tools overlap in functionality.

Naming Consistency3/5

The naming pattern is inconsistent: some tools use a 'browser_' prefix (browser_start, browser_close), while others are bare verbs (navigate, click, type) or compound snake_case (inspect_page, switch_window, wait_for). This mix of styles could cause confusion.

Tool Count5/5

With 10 tools, the set is well-scoped for browser automation. It covers session lifetime, navigation, element interaction, inspection, window handling, and waiting without being bloated or too thin.

Completeness3/5

The tools cover fundamental browser actions but miss common features like back/forward navigation, JavaScript execution, alert handling, or cookie management. The set is functional for basic scenarios but has notable gaps for comprehensive automation.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Exposes Selenium WebDriver as an MCP server, enabling AI agents and LLMs to control real browsers for automation tasks like navigation, element interaction, and screenshot capture.
    22
    21 PyPI
    3
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables LLMs to drive Edge in IE mode for automating legacy IE-only web applications, supporting tasks like clicking, filling forms, and data extraction via Selenium.
    28
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables AI agents and MCP clients to automate web browsers via Selenium WebDriver, supporting Chrome, Firefox, and Edge in headless or visible mode with tools for navigation, interaction, content extraction, screenshots, and scripting.
    21
    31 npm
    MIT