ie-mode-mcp
ie-mode-mcp
Servidor MCP para operar aplicaciones web heredadas que funcionan en el modo IE de Microsoft Edge desde agentes de IA a través de MCP (Model Context Protocol).
AI Agent ──(MCP / stdio)──> ie-mode-mcp ──> BrowserManager ──> selenium-webdriver
│
IEDriverServer.exe
│
Microsoft Edge (IE Mode)
│
Legacy Web ApplicationCompuesto únicamente por Node.js 22 / TypeScript / selenium-webdriver (sin servidor HTTP, base de datos, DI, ni framework de logging)
El transporte MCP es únicamente stdio
Solo hay una sesión de navegador, las operaciones de WebDriver son completamente secuenciales
No devuelve el HTML completo;
inspect_pagedevuelve información de pantalla resumida para el LLMSin flujo de aprobación. La operación se ejecuta en el momento de llamar a la herramienta
Índice
Related MCP server: ie-mcp
1. Inicio rápido
Ejecuta lo siguiente en Windows.
git clone https://github.com/sumikof/iedriver-mcp.git
cd iedriver-mcp
npm install
npm run build
# IEDriverServer.exe のパスと、遷移を許可する Origin を指定して起動
$env:IE_MCP_DRIVER_PATH = "C:\tools\IEDriverServer.exe"
$env:IE_MCP_ALLOWED_ORIGINS = "http://legacy01.local"
node dist/index.jsSi aparece {"level":"info","event":"started","transport":"stdio"} en stderr, el inicio es correcto. Normalmente no se inicia manualmente, sino que se inicia automáticamente desde la configuración MCP del agente de IA.
2. Requisitos previos
Elemento | Contenido |
SO | Windows 11 / Windows 10 (sesión interactiva iniciada) |
Node.js | 22 o superior |
Navegador | Microsoft Edge (que pueda usar el modo IE) |
Driver | IEDriverServer.exe (Selenium 4.x. Se recomienda la versión de 32 bits) |
IEDriverServer.exe se obtiene de la página de descarga de Selenium y se coloca en una carpeta cualquiera (ejemplo:
C:\tools\). La versión de 64 bits tiene limitaciones conocidas, por lo que Selenium oficial recomienda usar la versión de 32 bits.IEDriver se ve afectado por la GUI, el foco de ventana y los eventos nativos, por lo que se recomienda usarlo en una VM de Windows dedicada o una sesión de Windows dedicada.
No se contempla una configuración donde el navegador funcione en un servicio de Windows (Sesión 0).
El servidor MCP, IEDriver y Edge deben ejecutarse en el mismo entorno Windows.
3. Configuración previa en Windows
IEDriver se ve muy afectado por la configuración del entorno. Realiza la configuración manualmente primero antes de iniciar el servidor MCP.
3.1 Habilitar el modo IE de Edge
Verifica manualmente en Edge que el sitio objetivo se pueda abrir en modo IE. El modo IE se habilita mediante una de las siguientes políticas (bajo Software\Policies\Microsoft\Edge).
Política (nombre mostrado) | Nombre del valor del registro |
Configure Internet Explorer integration |
|
Configure the Enterprise Mode Site List |
|
Send all intranet sites to Internet Explorer | (Configurado mediante directiva de grupo desde Edge 77) |
La configuración concreta depende de la política de la organización; para más detalles, consulta la documentación del modo IE de Microsoft y al administrador de tu organización. Asegúrate de tener las últimas actualizaciones de Windows y Edge.
3.2 Configuración requerida por IEDriver
Elemento | Estado requerido | Tratamiento en este servidor |
Zoom del navegador | 100% | No es obligatorio porque ya se ha configurado |
Modo protegido (Protected Mode) | Configuración igual en todas las zonas | Si no está unificado, se lanzará una excepción al iniciar. Unificarlo en Opciones de Internet → Seguridad |
Bits de IEDriverServer | Se recomienda 32 bits | — |
Si la configuración del modo protegido no está unificada, browser_start fallará. No se utiliza introduceFlakinessByIgnoringProtectedModeSettings de IEDriver porque hace que el comportamiento sea inestable.
4. Instalación y compilación
npm install # 依存パッケージの取得
npm run build # TypeScript を dist/ へビルドEl resultado es dist/index.js. Después de compilar, también se puede iniciar con npm start (= node dist/index.js).
5. Variables de entorno
No se utiliza archivo de configuración (YAML / JSON), solo se configura mediante variables de entorno.
Variable de entorno | Descripción | Valor por defecto |
| Ruta de msedge.exe | No especificado (IEDriver lo detecta automáticamente) |
| Ruta de IEDriverServer.exe | No especificado (busca en |
| Orígenes permitidos para |
|
| Tiempo de espera predeterminado para búsqueda de elementos y esperas (ms) |
|
IE_MCP_EDGE_PATH=C:\Program Files (x86)\Microsoft\Edge\Application\msedge.exe
IE_MCP_DRIVER_PATH=C:\tools\IEDriverServer.exe
IE_MCP_ALLOWED_ORIGINS=http://legacy01.local,http://legacy02.local
IE_MCP_TIMEOUT_MS=10000A partir de IE Driver 4.5.0, Edge se detecta automáticamente en entornos sin IE (por defecto en Windows 11), por lo que normalmente no se necesita
IE_MCP_EDGE_PATH. Solo especificarlo explícitamente si la detección automática falla.Si se prioriza la reproducibilidad de la operación, se recomienda especificar explícitamente
IE_MCP_DRIVER_PATH.IE_MCP_ALLOWED_ORIGINSes una restricción simple para evitar operaciones incorrectas; se evalúa mediante coincidencia exacta del origen (scheme + host + port). No se restringe por ruta.
6. Método de inicio
Inicio manual (para verificar funcionamiento)
PowerShell:
$env:IE_MCP_DRIVER_PATH = "C:\tools\IEDriverServer.exe"
$env:IE_MCP_ALLOWED_ORIGINS = "http://legacy01.local"
node dist/index.jsSímbolo del sistema:
set IE_MCP_DRIVER_PATH=C:\tools\IEDriverServer.exe
set IE_MCP_ALLOWED_ORIGINS=http://legacy01.local
node dist\index.jsEspera conexiones del cliente mediante stdio. La entrada/salida estándar se usa para el protocolo MCP, por lo que no hay respuesta si se introduce texto en este estado (es normal). Todos los registros se emiten por stderr. Para salir, pulsa Ctrl+C (el navegador también se cierra automáticamente).
Nota: El inicio del servidor MCP no inicia el navegador. El navegador se inicia cuando el agente llama a
browser_start.
Operación normal
El agente de IA (cliente MCP) inicia este servidor como un proceso hijo. No es necesario iniciarlo manualmente. Realiza la configuración del siguiente capítulo.
7. Registro en el agente de IA
Añade lo siguiente al archivo de configuración del cliente MCP.
{
"mcpServers": {
"ie-mode": {
"command": "node",
"args": ["C:\\ie-mode-mcp\\dist\\index.js"],
"env": {
"IE_MCP_DRIVER_PATH": "C:\\tools\\IEDriverServer.exe",
"IE_MCP_EDGE_PATH": "C:\\Program Files (x86)\\Microsoft\\Edge\\Application\\msedge.exe",
"IE_MCP_ALLOWED_ORIGINS": "http://legacy01.local,http://legacy02.local",
"IE_MCP_TIMEOUT_MS": "10000"
}
}
}
}Las rutas deben escaparse con barras invertidas dentro del JSON (
C:\...).En
argsse especifica la ruta absoluta dedist/index.jsdespués de la compilación.En Claude Code, también se puede registrar con
claude mcp add.
claude mcp add ie-mode --env IE_MCP_DRIVER_PATH=C:\tools\IEDriverServer.exe --env IE_MCP_ALLOWED_ORIGINS=http://legacy01.local -- node C:\ie-mode-mcp\dist\index.jsDespués del registro, si el cliente ve las 10 herramientas, incluyendo browser_start, la conexión es correcta.
8. Referencia de herramientas
Se publican 10 herramientas. No se exponen las API de bajo nivel de WebDriver (como findElement / executeScript).
Herramienta | Entrada | Resumen |
| Ninguna | Inicia Edge en modo IE. Si ya está iniciado, reutiliza la sesión existente |
| Ninguna | Cierra el navegador. No da error aunque se llame varias veces |
|
| Navega después de verificar la lista blanca de URLs |
|
| Devuelve URL / título / texto de pantalla / elementos operables |
|
| Espera a que esté visible y habilitado, luego hace clic |
|
| Introduce texto en input / textarea |
|
| Selecciona una opción de |
|
| Espera hasta que se cumpla una condición |
|
| Cambia a una ventana emergente u otra ventana |
| Ninguna | Devuelve la pantalla actual como PNG (contenido de imagen MCP) |
Común: Selector
{ "by": "id | name | css | xpath | linkText", "value": "searchButton" }En las aplicaciones web heredadas, name y xpath se usan con frecuencia, por lo que se admiten.
Común: frame (iframe de 1 nivel)
Todas las herramientas de manipulación de elementos aceptan un frame opcional. Si se especifica, se vuelve a defaultContent y luego se cambia al frame, y se busca el elemento dentro de él.
{
"frame": { "by": "name", "value": "mainFrame" },
"selector": { "by": "id", "value": "searchButton" }
}browser_start
{}{ "status": "ready", "reused": false }reused: true indica que se ha reutilizado la sesión existente. Si la sesión existente está muerta, se reinicia automáticamente.
navigate
{ "url": "http://legacy01.local/customer" }{ "url": "http://legacy01.local/customer", "title": "顧客検索" }inspect_page
Es la herramienta principal para que el agente entienda la pantalla. No devuelve el HTML completo, solo la URL / título / texto visible / elementos operables (a, button, input, textarea, select, iframe). Los elementos ocultos y los input con type="hidden" se excluyen.
{ "frame": { "by": "name", "value": "mainFrame" } }{
"url": "http://legacy01.local/customer",
"title": "顧客検索",
"text": "顧客検索 顧客名 支店 検索",
"elements": [
{ "tag": "input", "id": "customerName", "name": "customerName", "type": "text" },
{ "tag": "select", "id": "branch", "name": "branch", "text": "東京支店", "optionCount": 12 },
{ "tag": "button", "id": "searchButton", "text": "検索" },
{ "tag": "iframe", "name": "mainFrame" }
],
"truncated": false
}truncated: trueindica que el número de elementos ha alcanzado el límite (300) y se ha truncado.Si la lista de elementos incluye un
iframe, para ver su contenido se debe llamar de nuevo especificando elframe.
click
{ "selector": { "by": "id", "value": "searchButton" } }{ "url": "http://legacy01.local/customer", "title": "顧客検索" }Espera a que el elemento esté visible y habilitado, luego hace clic. click no reintenta automáticamente (para evitar procesamiento duplicado si el registro, actualización o envío ya se ha completado y se vuelve a hacer clic).
type
{
"selector": { "by": "id", "value": "customerName" },
"text": "山田太郎",
"clear": true
}Si clear (por defecto true) es true, se ejecuta clear() antes de introducir el texto; si es false, se añade al texto existente.
select
{
"selector": { "by": "id", "value": "branch" },
"by": "text",
"value": "東京支店"
}{ "text": "東京支店", "value": "13", "index": 2 }by puede ser text / value / index (index empieza en 0).
wait_for
No utiliza sleep fijo, sino que espera explícitamente.
{
"type": "visible",
"selector": { "by": "id", "value": "resultTable" },
"timeoutMs": 10000
}
| Entrada requerida | Condición |
|
| El elemento existe en el DOM |
|
| El elemento es visible |
|
| El elemento es visible y operable |
|
| El texto del elemento contiene |
|
| La URL actual contiene |
|
| El título contiene |
Si se omite timeoutMs, se usa IE_MCP_TIMEOUT_MS.
switch_window
{ "target": "newest" }{ "index": 1 }{ "url": "http://legacy01.local/detail", "title": "顧客詳細", "index": 1, "windowCount": 2 }newest sondea brevemente hasta que aparezca un nuevo identificador de ventana. Si no se detecta, cambia a la última ventana existente.
screenshot
{}Devuelve una imagen PNG (contenido de imagen de MCP). Se usa para confirmar el diseño o las pantallas de error que no se pueden determinar solo con el DOM.
9. Ejemplos de uso
Bucle básico
browser_start → navigate → inspect_page → click / type / select → wait_for → inspect_pageRepite: inspect_page para entender la pantalla → operar → wait_for para esperar el resultado → inspect_page de nuevo.
Ejemplo: Buscar al cliente «Yamada Tarō» y abrir la pantalla de detalle
# | Herramienta | Argumentos |
1 |
|
|
2 |
|
|
3 |
|
|
4 |
|
|
5 |
|
|
6 |
|
|
7 |
|
|
8 |
|
|
9 |
|
|
10 |
|
|
11 |
|
|
Ejemplo: Operar dentro de un iframe
{"tool": "inspect_page", "args": {}}
{"tool": "inspect_page", "args": { "frame": { "by": "name", "value": "mainFrame" } }}
{"tool": "click", "args": {
"frame": { "by": "name", "value": "mainFrame" },
"selector": { "by": "id", "value": "searchButton" }
}}La especificación del frame se debe pasar en cada operación (porque internamente se vuelve a defaultContent y luego se cambia al frame, el estado no se mantiene entre operaciones).
Ejemplo: Operar una ventana emergente y volver a la ventana original
{"tool": "click", "args": { "selector": { "by": "id", "value": "openPopup" } }}
{"tool": "switch_window", "args": { "target": "newest" }}
{"tool": "inspect_page", "args": {}}
{"tool": "switch_window", "args": { "index": 0 }}10. Errores y soluciones
Los errores no devuelven el stack trace de Selenium, sino que se devuelven con el siguiente código (isError: true).
{
"error": "ELEMENT_NOT_FOUND",
"message": "Element was not found: id=searchButton",
"selector": { "by": "id", "value": "searchButton" }
}Código de error | Significado | Solución |
| El navegador no está iniciado | Llama a |
| No se encuentra el elemento o frame | Verifica los elementos reales con |
| No se cumplió la condición de | Revisa la condición y |
| La ventana especificada no existe | Revisa el |
| Falló la navegación | Verifica la URL, la red y la autenticación |
| IEDriver / Edge terminó anormalmente | Reinicia con |
| Origen fuera de la lista blanca | Revisa |
| Argumento inválido | Verifica las especificaciones de entrada de la herramienta |
| Otros (incluye fallo de inicio) | Verifica el |
Recuperación de DRIVER_LOST
Si el navegador o el driver se bloquean, el WebDriver interno se destruye y las operaciones posteriores devolverán BROWSER_NOT_STARTED. No se realiza recuperación automática ni reejecución automática de la operación anterior (para evitar efectos secundarios como registros duplicados). El agente debe volver a llamar a browser_start, verificar el estado de la pantalla con inspect_page y reanudar las operaciones. No se debe reejecutar directamente la operación anterior, ya que podría haberse completado.
11. Registros
stdout lo utiliza el protocolo MCP, por lo que todos los registros se emiten a stderr en una línea JSON.
{"level":"info","event":"started","transport":"stdio"}
{"level":"info","tool":"navigate","url":"http://legacy01.local/customer","durationMs":842}
{"level":"info","tool":"type","selector":{"by":"id","value":"password"},"textLength":16,"durationMs":128}
{"level":"error","tool":"click","selector":{"by":"id","value":"x"},"error":"ELEMENT_NOT_FOUND","message":"Element was not found: id=x","durationMs":5012}No se registra la cadena de entrada en sí, las cookies, la información de autenticación ni el HTML completo (de type solo se registra el número de caracteres). Si se desea guardar en un archivo, redirigir stderr.
node dist/index.js 2>> C:\logs\ie-mode-mcp.log12. Solución de problemas
Síntoma | Qué comprobar |
| ¿ |
Aparecen excepciones relacionadas con el modo protegido | Unificar la configuración del modo protegido para todas las zonas en Opciones de Internet → Seguridad |
Aparecen excepciones relacionadas con el zoom | Restablecer el zoom de Edge/IE al 100% |
Edge se inicia pero no entra en modo IE | Verificar las políticas del modo IE (lista de sitios, etc.). Confirmar primero si se puede mostrar manualmente en modo IE |
La operación se congela o no se puede hacer clic en un elemento | ¿La ventana está minimizada o inactiva? Se vuelve inestable durante la desconexión de escritorio remoto |
El elemento | ¿No es una pantalla dentro de un frame? (Reobtener especificando |
La herramienta no se ve en el lado del Agent | ¿Se ha especificado |
No aparece nada en la salida estándar | Es normal. Los registros se muestran en stderr |
screenshot es útil para investigar causas. Permite verificar estados que no se pueden determinar solo con información del DOM (modales, diálogos de autenticación, errores de renderizado).
13. Desarrollo
src/
├─ index.ts MCP Server のエントリーポイント(stdio)
├─ config.ts 環境変数と stderr ログ
├─ tools.ts MCP Tool の Schema と Handler
├─ browser.ts BrowserManager(Selenium / IEDriver 操作の集約)
├─ selectors.ts Selector → Selenium の By 変換
└─ errors.ts Selenium Error → MCP Error Code 変換npm run build # tsc でビルド
npm start # node dist/index.jsLa herramienta MCP no toca directamente Selenium, siempre pasa por
BrowserManager.Todas las operaciones de WebDriver están serializadas mediante una Promise Chain; incluso si la herramienta se llama en paralelo, solo se envía una solicitud a la vez a IEDriver.
Solo se reintentan operaciones sin efectos secundarios (búsqueda de elementos, detección de manejadores de ventana). No se reintentan
clickni envíos.
14. Limitaciones
La implementación inicial no admite lo siguiente:
Múltiples sesiones de navegador / Múltiples usuarios / Transporte HTTP / API REST / DB / Persistencia de sesión / Recuperación automática del navegador / Política de reintentos compleja / WebDriver Grid / API genérica de Selenium / Herramienta executeScript / iframes anidados (solo un nivel) / Caché de elementos / Métricas / Flujo de aprobación / Autenticación y autorización
Available Tools
10 toolsbrowser_closeClose browserB
Close the browser session. Safe to call repeatedly.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are available so the description carries the full burden. It only says 'close the browser session' and 'safe to call repeatedly', but does not disclose whether this terminates all browser state or if there are side effects on open windows, tabs, or downloads. The behavioral context is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences that are front-loaded and to the point. Every sentence adds value: the first states the action, the second clarifies safety/repeatability. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and no output schema, the description covers the basic purpose and safety. However, it lacks details on what happens after closing (e.g., can browser_start reopen cleanly) or any cleanup behavior, which might be useful context for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and schema coverage is 100%. The description adds value by stating it is safe to call repeatedly, which implies no parameters are needed and calls are idempotent. With no parameters to explain, this is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool closes the browser session with a specific verb and resource. It distinguishes enough from siblings like 'navigate' which moves within a session.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description notes it is safe to call repeatedly, which implies idempotency, but does not explicitly tell when to call it (e.g., end of a browsing task) or when not to (e.g., still need to interact). No sibling differentiation is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_startStart Edge IE ModeA
Start Microsoft Edge in IE Mode through IEDriverServer. Only one browser session exists; calling this while a session is running returns the existing one. Also use this to recover after a DRIVER_LOST error.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the full burden of behavioral disclosure. It transparently reveals that only one browser session exists, that calling the tool again returns the existing session, and that it can be used for recovery. This is strong for a start tool, but it could additionally mention potential side effects like timeouts or prerequisites for the IEDriverServer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, front-loading the primary purpose in the first sentence and adding behavioral nuance in the second. Every sentence provides essential information without redundancy or fluff, achieving maximum conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, no output schema, straightforward start action), the description is complete. It covers the core function, the singleton behavior, error recovery, and is sufficient for an AI agent to understand when and how to invoke the tool alongside its sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and 100% schema coverage, so the baseline is 4. The description adds no parameter information, which is appropriate since there are none to document. No additional semantic value is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool starts Microsoft Edge in IE Mode via IEDriverServer, using the specific verb 'Start' and the resource 'Microsoft Edge in IE Mode'. It also distinguishes itself from sibling tools by noting that only one browser session exists and that calling it again returns the existing session, which is unique among the provided sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool: normally to start the browser, and also to recover after a DRIVER_LOST error. It implicitly advises against calling it multiple times for new sessions by stating that subsequent calls return the existing session. However, it does not explicitly list alternatives or state when not to use it, though no alternative starting tool exists among the siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clickClick elementA
Click an element after waiting for it to be visible and enabled. This operation is never retried automatically, because a repeated click may submit or register data twice.
| Name | Required | Description | Default |
|---|---|---|---|
| frame | No | Optional iframe/frame to switch into first. One level of nesting is supported. | |
| selector | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses two key behaviors: waiting for the element to be visible and enabled, and the lack of automatic retry with a rationale. However, it does not mention timeout behavior, scroll-into-view, or what happens if the element is not found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose, and every sentence adds value. No fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's potential side effects (e.g., triggering navigation or form submission), the description is minimal. It does not mention return values, scroll behavior, or failure modes. It is adequate for a simple click but lacks completeness for an AI agent to fully anticipate outcomes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%, and the description adds no information about the parameters. It does not explain the 'frame' or 'selector' parameters beyond what is already in the schema. The description should compensate for the missing schema descriptions but fails to do so.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Click an element after waiting for it to be visible and enabled,' using a specific verb and resource. It distinguishes the tool from siblings like 'type' and 'select' by specifying the action and the precondition (visibility and enabled state).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly warns that the operation is never retried automatically because a repeated click may submit or register data twice. This gives a clear usage caution about retries, though it does not explicitly compare to alternative tools or state when not to use click.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_pageInspect pageA
Return the current URL, title, visible page text and the operable elements (a, button, input, textarea, select, iframe). The full HTML is never returned. Pass frame to inspect the contents of an iframe listed by a previous inspect_page call.
| Name | Required | Description | Default |
|---|---|---|---|
| frame | No | Optional iframe/frame to switch into first. One level of nesting is supported. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of disclosing behavior. It explicitly states 'The full HTML is never returned' and that frame must reference an iframe from a previous call. The read-only nature is implied by 'Return' but not stated outright, though this is likely sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two succinct sentences that front-load the main purpose and then add the iframe caveat. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has one well-described parameter, no output schema, and no annotations. The description covers the output, a key constraint (no full HTML), and iframe usage, making it reasonably complete for an inspection tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by clarifying that the frame must come from a previous inspect_page call, which is not stated in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and enumerates exactly what is returned (URL, title, visible text, operable elements). This clearly distinguishes it from sibling tools like 'click' or 'navigate'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool (to inspect page state) and gives specific guidance for iframe usage ('Pass frame to inspect the contents of an iframe listed by a previous inspect_page call'). It doesn't explicitly exclude alternatives, but the context is clear given the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshotScreenshotB
Capture the current browser window as a PNG image.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears the full burden. It states the action (capture) and format (PNG) but omits crucial details: whether it modifies state, if a browser window must be open, what exactly 'current browser window' captures (viewport vs full page), and if there are side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single 9-word sentence, efficient and front-loaded. However, it could include additional essential context (e.g., 'captures the visible viewport area') without losing conciseness, making it slightly under-specified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and no output schema, the description should clarify what the tool returns (e.g., base64 PNG data). It only says 'as a PNG image' but doesn't confirm the output type. The scope of 'current browser window' is ambiguous, and prerequisites are missing, leaving the agent uncertain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, and the schema coverage is 100% (trivially). The description adds minimal meaning by specifying 'current browser window' as the implicit input. A baseline of 4 is appropriate given no parameters to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('capture') and resource ('current browser window') with a clear output format ('PNG image'). It is distinct from sibling tools like 'navigate' or 'inspect_page' which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use screenshot versus alternatives. Despite having sibling tools (e.g., inspect_page, wait_for), no exclusions or context is given. An agent must infer use case from tool purpose alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
selectSelect optionB
Choose an option of an HTML element by visible text, value or index.
| Name | Required | Description | Default |
|---|---|---|---|
| by | Yes | How to identify the option. | |
| frame | No | Optional iframe/frame to switch into first. One level of nesting is supported. | |
| value | Yes | Option text, value, or zero-based index. | |
| selector | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses that selection is based on visible text, value or index, which is helpful. However, it doesn't mention side effects (e.g., whether the change triggers JavaScript events), error handling (e.g., what if option not found), or scope (e.g., operates within current page context).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that efficiently communicates the core action and identification methods. It is front-loaded with the key verb and resource. No waste, though it could optionally add a brief usage hint without breaching conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 parameters with nested objects and no output schema, the description is somewhat complete but lacks coverage of return behavior (e.g., what happens on success/failure), frame handling nuances, and edge cases. For a selection action in a browser automation context, more behavioral detail would be helpful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75%, meaning most parameters are documented in the schema. The description adds that selection can be by 'visible text, value or index', which maps to the 'by' enum, and that the 'value' parameter can be text or zero-based index. This provides modest added meaning beyond the schema, but the 'frame' and 'selector' objects remain documented primarily in schema, not description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'choose' and resource 'HTML <select> element', specifying three identification methods (visible text, value, index). This distinguishes it from sibling tools like click or type, though it doesn't explicitly contrast with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use by saying 'choose an option of an HTML <select> element', which suggests this is for dropdown selections. However, it does not provide explicit when-not-to-use guidance, mention prerequisites (e.g., element must exist), or compare with alternatives like click on an option directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
switch_windowSwitch windowA
Switch to another browser window or popup. Use target:"newest" after an action that opens a window, or index to select a window by its zero-based position.
| Name | Required | Description | Default |
|---|---|---|---|
| index | No | Zero-based window index. | |
| target | No | Switch to the newest window. | |
| timeoutMs | No | How long to poll for a new window. Defaults to IE_MCP_TIMEOUT_MS. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions polling behavior via timeoutMs parameter but does not state if switching is destructive, if it requires a window to exist, what happens if the window is closed, or any state changes. The description does not disclose potential side effects or preconditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core purpose in the first sentence. The second sentence adds specific usage hints. It could potentially omit 'or popup' as redundant with 'window', but overall concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 3 unrequired parameters, no output schema, no annotations, the description covers the basic purpose and usage hints. However, it lacks details on return values, error scenarios (e.g., window not found), or behavior when switching to a window that fails to load.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds context for the 'target' parameter (use after action that opens window) and 'index' (zero-based position), but the timeoutMs parameter meaning is already clear from schema. No additional semantic value beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool switches to another browser window or popup, specifying the verb 'switch' and the resource 'browser window or popup'. It distinguishes itself from sibling tools like browser_start, browser_close, and navigate by focusing on window selection rather than creation, closure, or navigation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when to use this tool: after an action that opens a window, use target 'newest', or use index to select by position. It implicitly distinguishes from sibling tools by indicating this is for window focus rather than content navigation or page interaction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
typeType textA
Type text into an input or textarea. Set clear to false to append instead of replacing.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Text to send to the element. | |
| clear | No | Clear the field first. Default true. | |
| frame | No | Optional iframe/frame to switch into first. One level of nesting is supported. | |
| selector | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that the tool can clear or append text via the 'clear' parameter, which is good. However, it does not mention potential side effects (e.g., triggering change events), error conditions (element not found), or behavior when the element is not a text input. This is adequate but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise at one sentence plus one usage tip. Every word earns its place, clearly stating the action and a key parameter behavior. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description should ideally mention return values (e.g., success indicator, element state). It doesn't, leaving that unclear. With nested objects (selector, frame) and no explanation of selector strategies beyond the schema enums, it completes the basic usage but misses context on what happens after typing (e.g., waits for stability, triggers events).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high at 75%, so the schema documents most parameters well. The description adds value by explaining the 'clear' boolean behavior (append vs replace) beyond the schema's default value note. It doesn't add to 'selector' or 'frame' parameters, which are already well-described in the schema, so this is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool types text into an input or textarea, using a specific verb and resource. It distinguishes from siblings like 'click' or 'select' by targeting text entry specifically, but doesn't differentiate from a potential 'send_keys' equivalent if one existed among siblings, so a slight deduction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a key usage guideline: set clear to false to append instead of replacing text. This gives basic advice on when to use a parameter. However, it lacks guidance on when to use this tool versus alternatives like clicking an element first or waiting, and doesn't mention prerequisites (e.g., element must be visible/interactable).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_forWait for conditionA
Wait until a condition holds. present/visible/enabled/text require a selector; text/url/title require text, which is matched as a substring. Use this instead of sleeping after an action.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | Expected substring for text/url/title conditions. | |
| type | Yes | Condition to wait for. | |
| frame | No | Optional iframe/frame to switch into first. One level of nesting is supported. | |
| selector | No | ||
| timeoutMs | No | Timeout in milliseconds. Defaults to IE_MCP_TIMEOUT_MS. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explains key behavioral traits: that present/visible/enabled/text require a selector, text/url/title require text matched as substring, and that it waits for the condition. With no annotations provided, the description carries the full burden of transparency. It lacks details on timeout behavior or error handling, but covers core usage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise with two sentences, front-loading the key purpose and condition types. Every sentence adds value, avoiding any redundancy. The structure is efficient for an AI agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (5 parameters, nested objects, no output schema), the description is adequate. It explains the core waiting concept and parameter dependencies. However, it lacks details on return values or what happens on timeout/failure, which the schema alone doesn't cover. The sibling 'inspect_page' might share similar conditions, but no differentiation is made.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80%, so the schema already documents most parameters well. The description adds value by clarifying the relationship between condition types and required parameters (e.g., 'present/visible/enabled/text require a selector; text/url/title require text'). This bridges gaps between parameters, though it does not detail the 'frame' parameter beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool waits until a condition holds, with specific verb+resource ('Wait for condition'). It lists the condition types and distinguishes itself from sleeping after an action, which differentiates it from sibling tools like 'click' or 'navigate'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool ('Use this instead of sleeping after an action'), providing clear guidance on avoiding poor alternatives. However, it does not specify when not to use it or which sibling would be more appropriate for different scenarios, such as synchronous checks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.1.0- First observed
browser_close - First observed
browser_start - First observed
click - First observed
inspect_page - First observed
navigate - First observed
screenshot - First observed
select - First observed
switch_window - First observed
type - First observed
wait_for
TDQS
Scored across 10 tools
Each tool has a clearly distinct purpose: session management (start/close), navigation, inspection, interaction (click, type, select), window switching, waiting, and screenshot. No two tools overlap in functionality.
The naming pattern is inconsistent: some tools use a 'browser_' prefix (browser_start, browser_close), while others are bare verbs (navigate, click, type) or compound snake_case (inspect_page, switch_window, wait_for). This mix of styles could cause confusion.
With 10 tools, the set is well-scoped for browser automation. It covers session lifetime, navigation, element interaction, inspection, window handling, and waiting without being bloated or too thin.
The tools cover fundamental browser actions but miss common features like back/forward navigation, JavaScript execution, alert handling, or cookie management. The set is functional for basic scenarios but has notable gaps for comprehensive automation.
Maintenance
Related MCP Connectors
Let AI agents query data and act across all your business apps via MCP.
Governed app access for AI agents: 1,000+ apps & 12,000+ tools via Code Mode MCP.
Stealth web browser for agents: search, fetch, click, download and type in persistent MCP sessions.
Live browser debugging for AI assistants — DOM, console, network via MCP.
Related MCP Servers
- AlicenseAqualityCmaintenanceExposes Selenium WebDriver as an MCP server, enabling AI agents and LLMs to control real browsers for automation tasks like navigation, element interaction, and screenshot capture.2221 PyPI3MIT
- AlicenseAqualityCmaintenanceEnables LLMs to drive Edge in IE mode for automating legacy IE-only web applications, supporting tasks like clicking, filling forms, and data extraction via Selenium.28MIT
- AlicenseAqualityCmaintenanceMCP server providing browser automation for AI agents, enabling actions like clicking and typing, structured data extraction, content validation, multi-step task execution, and memory enrichment from web pages.8344 npmMIT
- AlicenseAqualityCmaintenanceEnables AI agents and MCP clients to automate web browsers via Selenium WebDriver, supporting Chrome, Firefox, and Edge in headless or visible mode with tools for navigation, interaction, content extraction, screenshots, and scripting.2131 npmMIT