selenium-mcp
A Selenium MCP server that lets an AI agent drive real Chrome/Firefox/Edge browsers for automation, testing, and scraping — 39 tools are exposed in this schema (the README advertises 41).
Browser lifecycle —
start_browser,stop_browser, plus parallel multi-session management (session_create,session_select,session_list,session_destroy), each with browser choice, headless mode, window size, and timeout tuning.Navigation —
navigate(and legacyopen_url),get_current_url,get_title, and page source retrieval.Element discovery —
find_element,wait_for_element,wait_until_visible, andcapture_pagefor page snapshots with stable element refs (e1, e2, …).Interaction —
click,retry_click,interact(hover/double/right-click),type,press_key,upload_file.Reading & assertions —
get_text,get_attribute,assert_text,assert_visible,assert_attributewith equals/contains/regex modes for built-in test assertions.Scripting & batching — arbitrary synchronous
execute_script, andbatch_executefor up to 10 chained steps (navigate, wait, click, type, script) in one call.Windows, frames, dialogs —
window(tabs, windows, switching, closing),frame(switch/parent/default),alert(accept, dismiss, read, send text).Cookies —
add_cookie,get_cookies,delete_cookie.Capture —
take_screenshotas PNG with optional base64 output or disk save.Persistent selector hints —
selector_hint_save/get/list/deleteto remember working locators per domain across runs.Extras — MCP resources
browser-status://currentandaccessibility://current, optional NDJSON tool-call tracing, and strict zod-validated, structured error responses tuned for LLM agents.
⚠️ Note: several README 0.3.0 tools are absent from this schema — select_option, scroll, history, wait_for_page, and the resizable-viewport window actions — while deprecated duplicates open_url and wait_until_visible are still present.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@selenium-mcpOpen Chrome, go to example.com, take a screenshot."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
selenium-mcp
Selenium MCP server for AI agents — 41 tools for real-browser automation: navigation, clicking, typing, assertions, screenshots, multi-session management, page snapshots with stable element refs, persistent selector hints, and batched multi-step execution.
Built with TypeScript, the official MCP SDK, and Selenium WebDriver — strict zod input validation, explicit waits, and structured responses designed for LLM agents.
One-Click Install
What's new in 0.3.0
New tools:
select_optionfor dropdowns,scroll(including infinite-scroll pages and scrollable panels),historyfor back, forward, and refresh, andwait_for_pageto wait for a URL or title after a redirect.Responsive testing:
windowcan now resize the viewport to an exact size, such as a 390x844 phone, or maximize it.Better tool choice by AI agents: every tool and parameter now explains what it does, what it returns, and when to use a similar tool instead.
Breaking:
open_urlandwait_until_visiblewere removed as duplicates. Usenavigate, andwait_for_elementwithvisible: true.Verified releases: every release is tested end to end against a real browser and published with npm provenance.
See the changelog for details.
Related MCP server: openmcp
Setup
claude mcp add selenium -- npx -y @gaforov/selenium-mcp@latestAdd to your client's MCP config (e.g. claude_desktop_config.json or .cursor/mcp.json):
{
"mcpServers": {
"selenium": {
"command": "npx",
"args": ["-y", "@gaforov/selenium-mcp@latest"]
}
}
}code --add-mcp '{"name":"selenium","command":"npx","args":["-y","@gaforov/selenium-mcp@latest"]}'goose session --with-extension "npx -y @gaforov/selenium-mcp@latest"Settings → Tools → AI Assistant → Model Context Protocol → Add, with command npx and arguments -y @gaforov/selenium-mcp@latest. Full walkthrough in docs/CLIENT_INTEGRATION.md.
git clone https://github.com/gaforov/selenium-mcp.git
cd selenium-mcp
npm install
npm run buildThen point your MCP client at node /absolute/path/to/selenium-mcp/dist/server.js.
Example Usage
Ask your AI agent:
Use selenium-mcp to open Chrome, go to https://example.com, read the page title, take a screenshot, and close the browser.
The agent chains start_browser → navigate → get_title → take_screenshot → stop_browser on its own — no scripting needed.
Requirements
Node.js 20+
Chrome, Firefox, or Edge installed (Selenium Manager provisions the matching driver automatically)
How it compares
Most Selenium MCP servers wrap WebDriver's basic commands. This one adds the layer that makes agents reliable:
Capability | selenium-mcp | Typical Selenium MCP servers |
Page snapshot with stable element refs ( | ✅ | rare |
Persistent per-domain selector memory ( | ✅ | ❌ |
Parallel multi-session browsing | ✅ | rare |
Batched multi-step execution in one call | ✅ | ❌ |
Built-in test assertions | ✅ | some |
Tool-call tracing (NDJSON audit log) | ✅ | ❌ |
Strict input validation + structured errors | ✅ | varies |
Every tool and parameter described for AI agents (enforced by tests) | ✅ | varies |
End-to-end tests against a real browser in CI | ✅ | some |
Why selenium-mcp
Snapshot-first workflows —
capture_pagereturns a page snapshot with stable element refs the agent can act on directly, no brittle selector guessingSelector hints — persist working locators per domain so repeat automations get faster and more reliable over time
Batched execution —
batch_executeruns constrained multi-step sequences in a single tool call, cutting round-tripsMulti-session — create, select, list, and destroy parallel browser sessions
Agent-friendly errors — every response is structured and validated with zod, so agents can recover instead of stalling
Optional tracing — NDJSON trace of every tool call for debugging and auditing
Tools (41)
Category | Tools |
Browser lifecycle |
|
Navigation |
|
Element discovery |
|
Interaction |
|
Reading |
|
Assertions |
|
Scripting |
|
Selector hints |
|
Windows & context |
|
Cookies |
|
Capture |
|
Full parameter documentation: docs/TOOL_REFERENCE.md
MCP Resources
browser-status://current— live browser/session statusaccessibility://current— accessibility snapshot of the current page
Optional Tracing
Enable lightweight NDJSON tracing of all tool calls:
SELENIUM_MCP_TRACE=true
SELENIUM_MCP_TRACE_PATH=./logs/selenium-mcp-trace.ndjsonIf SELENIUM_MCP_TRACE_PATH is omitted, the default is logs/selenium-mcp-trace.ndjson.
Documentation
Contributing
Contributions are welcome — bug reports, feature requests, and pull requests. See CONTRIBUTING.md to get started.
npm run typecheck
npm run build
npm testLicense
MIT. See LICENSE.
Available Tools
41 toolsadd_cookieA
Set a cookie in the browser, e.g. a session or feature-flag cookie to skip a login screen or switch on a test mode. Browsers only accept cookies for the site that is currently open, so navigate to a page on that domain first. The page does not see the cookie until its next request, so refresh (history) or navigate afterwards. Setting an existing name replaces that cookie.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Cookie name, e.g. 'session_id'. | |
| path | No | URL path the cookie applies to (default '/'). | |
| value | Yes | Cookie value. | |
| domain | No | Domain the cookie applies to, e.g. '.example.com' to include subdomains. Defaults to the current page's host. | |
| expiry | No | Expiry as a Unix timestamp in seconds. Omit for a session cookie that ends when the browser closes. | |
| secure | No | Only send the cookie over HTTPS. | |
| httpOnly | No | Hide the cookie from page JavaScript (document.cookie). | |
| sameSite | No | SameSite policy: Strict, Lax, or None (None also requires secure: true). Omit for the browser default. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and does so well: it discloses domain-scoping behavior, the deferred visibility of the cookie until the next request, and the replace-existing-name semantics. These are behaviors an agent could not infer otherwise.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Nine short clauses that each add distinct operational information, front-loaded with the core action and followed by prerequisites and post-conditions. No redundant restatement of the tool name or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter mutation tool with no annotations and no output schema, the description covers the essential behavioral context: prerequisite navigation, deferred cookie visibility, refresh requirement, and replacement semantics. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 8 parameters are fully documented in the schema with examples (e.g., '.example.com' for domain, Unix timestamp for expiry). The description elaborates on the domain/current-page constraint but adds no per-parameter syntax beyond what the schema already provides. Baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Set a cookie in the browser') with concrete examples of intent (session, feature-flag, skip login). Clearly distinguishes from the sibling get_cookies and delete_cookie mutators.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use examples and a critical prerequisite ('navigate to a page on that domain first'). Also states the post-condition action ('refresh or navigate afterwards') so the agent knows what to do next.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
alertA
Handle a native browser dialog opened by alert(), confirm(), or prompt(). While one is open, other page actions fail until it is handled. get_text reads the message; accept clicks OK and dismiss clicks Cancel, which close the dialog and let the page act on the answer (this cannot be undone); send_text types into a prompt() before you accept it. Every action returns the dialog's text. Fails if no dialog is open. Custom in-page modals are not native dialogs: use click on their buttons instead.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | For send_text: the text to type into the prompt() field. Ignored for other actions. | |
| action | Yes | get_text = read the message; accept = OK; dismiss = Cancel; send_text = type into a prompt() (then accept). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so richly: it discloses the blocking behavior (other page actions fail until handled), that accept/dismiss are irreversible ('this cannot be undone'), that every action returns the dialog's text, and the send_text→accept sequencing. This is exactly the behavioral context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense paragraph that is front-loaded with the purpose before moving to per-action semantics, failure mode, and the sibling distinction. Every sentence carries distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although there is no output schema, the description explicitly states that every action returns the dialog's text, and with no annotations the blocking/irreversibility semantics are fully covered. Nothing an agent needs in order to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by mapping each enum value to an outcome (accept=OK, dismiss=Cancel) and adding the sequencing constraint that send_text types into a prompt() 'before you accept it'. Marginal but real added value over the already-complete schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Handle) and a precisely scoped resource (native browser dialog opened by alert(), confirm(), or prompt()), and explicitly distinguishes it from custom in-page modals, which it routes to the click sibling. An agent can identify this tool and its boundary versus siblings without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use (a native dialog is open, which blocks other page actions), when-not (custom in-page modals are not native dialogs — use click on their buttons), and a failure condition ('Fails if no dialog is open'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
assert_attributeA
Test step: check that an element's attribute or property equals (default), contains, or matches a regular expression, e.g. that a button is disabled or an input's value is 'standard_user'. Waits for the element to exist; a missing attribute counts as an empty string. Fails as a tool error showing expected vs actual.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | equals = exact match (default); contains = substring; matches = JavaScript regular expression. | equals |
| name | Yes | Attribute or property to check, e.g. 'value', 'disabled', 'href', or 'class'. | |
| expected | Yes | Expected value or pattern, e.g. 'true' for a disabled button. | |
| selector | Yes | How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector. | |
| timeoutMs | No | How long to wait for the element to exist, in milliseconds (default 10000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the implicit wait for element existence, the edge-case rule that a missing attribute equals an empty string, and the failure semantics (tool error showing expected vs actual). Remaining gaps are minor, such as interaction with timeoutMs beyond the schema default.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the core assertion behavior, followed by edge-case and failure semantics. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully explains the success/failure outcome (tool error with expected vs actual) and the wait/missing-attribute behavior, which is enough to invoke it correctly. It could say slightly more about how the wait interacts with timeoutMs, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (mode, name, expected, selector, timeoutMs) is already documented in the schema. The description restates the mode semantics and examples but adds no syntax or format detail beyond what the schema provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (check/assert) applied to a specific resource (an element's attribute or property) with concrete examples ('a button is disabled', an input's value). It is clearly distinguishable from siblings like get_attribute (reads a value) and assert_text/assert_visible (different assertions).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Test step:' prefix and the enumeration of comparison modes give clear context for when this assertion tool applies, and the examples make the use case concrete. However, it never names alternatives or exclusions (e.g. use get_attribute to read rather than assert, or assert_visible for visibility), leaving that routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
assert_textA
Test step: check that an element's visible text equals, contains (default), or matches a regular expression. Waits for the element to be visible, then checks once (it does not wait for the text to change). Passes with the actual text, or fails as a tool error showing expected vs actual, so it works as an acceptance check. Use get_text to just read text without a pass/fail.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | equals = exact match; contains = substring (default); matches = JavaScript regular expression. | contains |
| trim | No | Trim whitespace from the actual text before comparing (default true). | |
| expected | Yes | Text or pattern the element's text should match, e.g. 'Epic sadface: Username is required'. | |
| selector | Yes | How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector. | |
| timeoutMs | No | How long to wait for the element to become visible, in milliseconds (default 10000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden and does so well: it discloses the visibility wait, the single-check behavior (no retry on text change), the pass payload (actual text), and the failure mode (tool error with expected vs actual). This is rich behavioral context beyond anything in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the core action, and each sentence adds distinct value (semantics, wait behavior, result contract, alternative). No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a nested-selector assertion tool with no output schema and no annotations, the definition covers the pass/fail contract, timing behavior, and sibling routing. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all five parameters including mode enum, trim default, and timeout. The description restates the mode options but adds no syntax or format detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('check that an element's visible text equals, contains, or matches a regex') and immediately differentiates itself from the sibling get_text. An agent can identify the tool's role without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Frames the tool as an acceptance check and names the alternative explicitly ('Use get_text to just read text without a pass/fail'), giving a clear when-to-use and when-not condition. The single-check semantics ('does not wait for the text to change') further disambiguate it from wait_for_element.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
assert_visibleA
Test step: check that an element becomes visible within timeoutMs, e.g. a success banner or a cart badge. Passes as soon as it is displayed; fails as a tool error if it is missing or stays hidden. Similar to wait_for_element with visible: true, but phrased as a pass/fail assertion.
| Name | Required | Description | Default |
|---|---|---|---|
| selector | Yes | How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector. | |
| timeoutMs | No | How long to wait for the element to become visible before failing, in milliseconds (default 10000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the behavioral burden and does so well: it states the pass condition ('passes as soon as it is displayed') and the failure mode ('fails as a tool error if it is missing or stays hidden'). It does not say whether the page/browser state is preserved on failure or whether the error message identifies the selector, leaving a small gap for a mutation-adjacent, blocking assertion tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero filler, front-loaded with the purpose, then the pass/fail semantics, then the sibling comparison. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description communicates the return-by-exception semantics directly, and the rich input schema covers parameters. What is missing is failure-side context (does the browser stay open, is a screenshot available after the tool error), which matters for an assertion tool inside a test flow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the selector shape, locator strategies and timeoutMs bounds/default. The description adds only that the timeout governs 'becomes visible' timing, which is a minor enrichment over the schema's own text. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+condition: check that an element becomes visible within timeoutMs. It names concrete examples (success banner, cart badge) and explicitly distinguishes itself from the sibling wait_for_element and from the pass/fail assertion family (assert_text, assert_attribute).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Frames the tool as a 'test step' and explicitly contrasts it with 'wait_for_element with visible: true, but phrased as a pass/fail assertion', which tells the agent when to pick this over the wait tool. It stops short of spelling out the inverse condition (use wait_for_element when you don't want the run to fail), so selection guidance is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_executeA
Run up to 10 steps in one call. Supported actions: navigate, wait_for_element, wait_for_page, click, type, select_option, and execute_script, each with the same fields as the standalone tool (selectors only, not capture_page refs). Use it for known linear flows such as a login or form fill, to save round trips. By default it stops at the first failing step; the result lists every executed step with its details or error.
| Name | Required | Description | Default |
|---|---|---|---|
| steps | Yes | Ordered steps, 1-10. Each has an action plus that action's fields, e.g. { action: 'type', selector: { by: 'id', value: 'user-name' }, text: 'standard_user' }. | |
| stopOnError | No | Stop at the first failing step (default true). Set false to run every step and collect all errors. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it discloses the 10-step cap, the default fail-fast semantics, that stopOnError can be flipped to collect all errors, and the result shape ('lists every executed step with its details or error'). It also flags a real constraint the schema cannot express — selectors only, no capture_page refs. It stops short of saying what happens to browser state after a mid-batch failure or whether a session must already exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the capability and its limits before moving to usage and failure behavior. Every sentence carries distinct information — actions supported, when to prefer it, and what happens on failure — with no filler or repetition of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex orchestration tool with zero annotation coverage and no output schema, the definition covers the essentials: step cap, allowed actions, failure semantics, and result contents. The remaining gaps are the unstated preconditions (an active browser session) and the side effects left behind by a partially executed batch, which matter for a mutation-heavy tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning beyond the schema: it states the accepted action vocabulary, that step fields mirror the standalone tools, and — critically — that capture_page references are not valid selectors here, a restriction absent from the JSON Schema. The 1-10 bounds and stopOnError default are restated rather than added, which keeps this from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and scope ('Run up to 10 steps in one call') and enumerates exactly which actions are batched (navigate, wait_for_element, wait_for_page, click, type, select_option, execute_script). By tying each action back to 'the same fields as the standalone tool', it cleanly distinguishes batch_execute from the ~38 single-step siblings like click, type, and select_option.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear positive context — 'known linear flows such as a login or form fill, to save round trips' — which implicitly rules out flows that need branching or intermediate inspection. However, it never explicitly states when NOT to batch (e.g. when a later step depends on reading an earlier step's result, use the standalone tools), so the exclusion is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capture_pageA
Snapshot the page's visible interactive elements and headings (links, buttons, inputs, selects, textareas, ARIA roles, h1-h4) with stable refs e1, e2, ... and a reusable selector for each. Pass a ref to click, type, get_text, select_option, interact, or scroll instead of guessing a selector. Refs belong to this snapshot, so capture again after navigation or major page changes. Prefer this over screenshots or get_page_source to understand what is on the page.
| Name | Required | Description | Default |
|---|---|---|---|
| maxElements | No | Maximum elements to return, 10-500 (default 200). Lower it on very large pages to keep the result small. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and it delivers the single most critical behavioral trait: refs are scoped to this snapshot and go stale after navigation or major page changes. It doesn't cover permissions, rate limits, or result-shape details like whether truncation occurs when maxElements is hit, so it falls short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Densely front-loaded: what is captured, the ref convention, how to consume refs, the staleness caveat, and the routing away from screenshot/get_page_source. Every sentence carries actionable information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-output-schema tool the description does explain the returned artifacts (stable refs and reusable selectors) plus the invalidation rule, which is enough for correct invocation. Minor gap: no note on truncation behavior on large pages or what happens if a page has no interactive elements.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is only one parameter (maxElements) and the schema already documents it fully at 100% coverage, including default 200 and the 10-500 range with tuning guidance. The description adds nothing about it, which is acceptable given the schema does the work—baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (snapshot) and resource (the page's visible interactive elements and headings), then enumerates exactly what is captured (links, buttons, inputs, selects, textareas, ARIA roles, h1-h4) and what is produced (e1, e2 refs plus selectors). It also differentiates itself from siblings like take_screenshot and get_page_source.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent to pass a ref to click, type, get_text, select_option, interact, or scroll instead of guessing selectors, mandates re-capturing after navigation or major page changes, and names screenshots/get_page_source as inferior alternatives for understanding page content.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clickA
Click an element: waits until it is visible and enabled, then clicks it (scrolling it into view). Target it by selector or by a ref from capture_page. If clicks fail intermittently because of overlays, animations, or re-rendering, use retry_click; for double-click, right-click, or hover, use interact. Fails with the reason if the element does not become clickable within timeoutMs.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Element ref from the latest capture_page result (e.g. 'e12'). Use instead of selector; refs go stale after navigation, so capture again if one is rejected. | |
| selector | No | How to find the element, e.g. { by: 'id', value: 'user-name' } or { by: 'css', value: '#login-button' }. Provide either selector or ref. | |
| timeoutMs | No | How long to wait for the element to become visible and enabled, in milliseconds (default 10000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does well: it discloses that the element is scrolled into view, that it waits for visibility and enablement, and that failure returns the reason. It does not describe what a successful click returns, which is a minor remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each front-loaded with actionable information: primary behavior first, then alternative routing, then failure mode. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and full input-schema coverage, the description covers the essential behavioral contract: waiting, scrolling, targeting, and failure reporting. It would be fully complete if it briefly stated the success return or side-effect expectation, but nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters already have detailed schema-level documentation (including ref staleness after navigation). The description adds only high-level targeting context and does not meaningfully extend parameter semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (click an element), the waiting behavior (visible and enabled), and the two targeting mechanisms (selector or ref). It also names sibling tools (retry_click, interact) for adjacent use cases, so an agent can distinguish it without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use this tool versus alternatives: use retry_click for intermittent failures from overlays/animations/re-rendering, and use interact for double-click, right-click, or hover. It also names the fallback behavior on timeout, giving complete routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_cookieA
Delete one cookie by name, or every cookie for the current page when name is omitted. Deleting all cookies usually logs the user out and resets consent banners and preferences; it cannot be undone. The page notices on its next request, so refresh (history) or navigate afterwards.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Name of the cookie to delete, e.g. 'session_id'. Omit to delete ALL cookies for the current page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden and does so well: it flags the operation as irreversible ('cannot be undone'), names concrete side effects (logging the user out, resetting consent banners and preferences), and explains the deferred timing of the effect with a required follow-up action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences ordered by importance: what it does, what gets destroyed, then how to make the change take effect. No padding, and the destructive warning is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter destructive tool with no annotations and no output schema, everything an agent needs is present: the effect, its irreversibility, the collateral consequences, and the post-call action required to observe the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter's description already documents the omit-to-delete-all semantics, so the schema does the heavy lifting. The description restates the omitted-name behavior but adds no new syntax or format detail beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (delete_cookie) and immediately distinguishes two modes: one cookie by name, or all cookies for the current page when name is omitted. This clearly separates it from siblings like add_cookie and get_cookies without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational context: omit the name to clear everything, and refresh (history) or navigate afterwards because the page only notices on its next request. It does not explicitly name or exclude alternatives like get_cookies or add_cookie, so it stops short of full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
execute_scriptA
Run synchronous JavaScript in the current page and return its result (use a return statement); arguments are available as arguments[0], arguments[1], .... Useful for reading several values in one call, or page state no other tool exposes (localStorage, computed styles, element counts). The script runs with the page's privileges and can change the page; for user actions such as clicking and typing, prefer click and type so real events fire.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | JSON values passed to the script as arguments[0], arguments[1], ... (default none). | |
| script | Yes | JavaScript function body, e.g. 'return document.querySelectorAll(".inventory_item").length;'. Return plain JSON values (strings, numbers, booleans, arrays, objects). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well: it discloses synchronous execution, that arguments are positional, that the script 'runs with the page's privileges and can change the page' (mutation risk), and the real-event caveat. It stops short of error behavior, sandboxing/isolation, or what happens on exceptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and return contract, then usage, then the safety/alternative caveat. No filler and no repetition of sibling concerns.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must explain return values — it does ('return its result', plain JSON values via the schema). For an arbitrary-code-execution tool with zero annotations, it covers scope, mutation risk, and alternatives well, though error/failure semantics remain unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reiterates the positional-argument convention and the return-statement requirement, which the schema already documents; it adds emphasis but little new syntax or format detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('Run synchronous JavaScript in the current page and return its result') and immediately pins down the contract with '(use a return statement)'. It is unmistakably distinct from siblings like get_text or get_attribute, since it is the only arbitrary-code escape hatch.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives positive selection criteria ('reading several values in one call, or page state no other tool exposes (localStorage, computed styles, element counts)') and an explicit exclusion with a named alternative ('for user actions such as clicking and typing, prefer click and type so real events fire'). Both when-to-use and when-not-to-use are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_elementA
Look up one element and describe it: tag, visible text, and whether it is displayed and enabled. Waits for it to exist. Use it to confirm a selector matches the intended element, or to inspect an element before acting. To discover elements without knowing a selector, use capture_page; to wait for something to appear, use wait_for_element.
| Name | Required | Description | Default |
|---|---|---|---|
| selector | Yes | How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector. | |
| timeoutMs | No | How long to wait for the element to exist, in milliseconds (default 10000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose two non-obvious behaviors: it blocks until the element exists, and it reports existence state plus displayed/enabled flags. It does not say what happens when the element never appears (error type, mention of timeoutMs as the bound), which is the one meaningful gap for a waiting lookup.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all load-bearing: what it returns first, then when to use it, then the two sibling alternatives. No restatement of the name or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description usefully enumerates the returned fields, which is the right compensation. For a selector-based lookup with a fully documented schema, the only unstated item is failure behavior on timeout, a minor completeness gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (nested selector with by/value, timeoutMs) are fully documented in the schema itself. The description only implicitly touches timeoutMs via 'waits for it to exist' and adds no locator-strategy guidance beyond what the schema already gives.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('look up one element') plus the exact return fields (tag, visible text, displayed/enabled), which no sibling provides. It also names the two siblings an agent might confuse it with, so the tool is distinguishable without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit use cases ('confirm a selector matches the intended element, or inspect an element before acting') and explicit alternatives with the condition that selects them: capture_page when no selector is known, wait_for_element when only waiting is needed. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
frameA
Move into or out of an iframe. Elements inside an iframe (embedded widgets, rich-text editors, payment fields) cannot be found by other tools until you switch into it. switch = enter a frame by selector, index, or nameOrId; parent = go up one level; default = return to the main page. The switch lasts until you change it again or a new page loads, so switch back with default when you are done inside the frame.
| Name | Required | Description | Default |
|---|---|---|---|
| index | No | For switch: zero-based position of the frame on the page. Used when no selector is given. | |
| action | Yes | switch = enter a frame (needs selector, index, or nameOrId); parent = up one level; default = back to the main page. | |
| nameOrId | No | For switch: the frame's name or id attribute. Used when neither selector nor index is given. | |
| selector | No | For switch: the <iframe> element, e.g. { by: 'css', value: 'iframe#editor' }. Most reliable way to pick a frame. | |
| timeoutMs | No | How long to wait for the frame element when switching by selector, in milliseconds (default 10000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose the key non-obvious trait: the switch is stateful and persists until changed or a new page load. That is exactly the kind of side effect an agent must know. It does not cover failure behavior (e.g., frame not found, timeout handling beyond the schema's timeoutMs) or return values, so it falls short of full transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the purpose and impact on other tools, then the action semantics, then the lifetime rule. No filler; every sentence contributes actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a stateful 5-parameter tool with a nested selector object, no annotations, and no output schema, the description covers purpose, per-action semantics, and the state lifetime plus how to revert. It omits error/timeout behavior and any indication of what a successful call returns, which keeps it just short of fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents action, index, nameOrId, selector, and timeoutMs with the precedence rules ('used when no selector is given'). The description restates the action meanings and the selector/index/nameOrId alternatives but adds no format, precedence, or fallback detail beyond what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Move into or out of an iframe') and immediately scopes it against siblings by noting elements inside a frame cannot be found by other tools until you switch in. An agent can distinguish this from window, click, or find_element without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger (elements inside embedded widgets, rich-text editors, payment fields are invisible to other tools) and a clear cleanup rule ('switch back with default when you are done inside the frame'). Each action value is defined inline, so the when-to-use and how-to-exit guidance is complete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_attributeA
Read an attribute or property of an element, e.g. an input's value, a link's href, or whether it is disabled, checked, or aria-expanded. Waits for the element to exist (it does not need to be visible). Returns the value, or null if it is not set; the live property wins when one exists, so 'value' gives the text currently in an input. Use get_text for visible text, and assert_attribute to check a value as a test step.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Attribute or property name, e.g. 'value', 'href', 'disabled', 'checked', 'class', or 'aria-expanded'. | |
| selector | Yes | How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector. | |
| timeoutMs | No | How long to wait for the element to exist, in milliseconds (default 10000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the existence-wait (independent of visibility), the null return when unset, and the property-beats-attribute precedence rule. It does not mention default/hard timeout behavior or failure modes when the element never appears, which the schema only partially covers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the core action and examples, then the wait/return semantics, then the sibling routing. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must explain returns, and it does ('the value, or null if it is not set'), alongside wait semantics and alternatives. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning beyond it by explaining the live-property-wins rule ('value' gives the text currently in an input), which clarifies how the 'name' parameter resolves at runtime.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and resource ('attribute or property of an element') and illustrates with concrete cases (value, href, disabled, checked, aria-expanded). It also explicitly names the sibling it is not (get_text) so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit alternatives with selection conditions: get_text for visible text, assert_attribute for value assertions as a test step. It also states the wait/visibility precondition ('Waits for the element to exist (it does not need to be visible)'), leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_cookiesA
Read the cookies the browser holds for the current page: all of them, or one by name. Includes httpOnly cookies that page JavaScript cannot see. Returns each cookie's name, value, domain, path, expiry, and flags. Useful to check that a login created a session cookie, or to copy one into another session with add_cookie.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Name of one cookie to read, e.g. 'session_id'. Omit to read all cookies for the current page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that read includes httpOnly cookies invisible to page JavaScript, and enumerates the returned fields (name, value, domain, path, expiry, flags). Read-only nature is implied but the safety profile is not explicitly declared.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences, front-loaded with the operation and scope, then capability disclosure, then use cases. No filler; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with no output schema, it covers scope, return values, and motivation adequately. Only the explicit read-only/safety signal is left to inference, which is minor given the 'read' framing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the sole 'name' parameter including the omit-to-read-all behavior and an example. The description restates this dual mode, adding little beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) and resource (cookies) with clear scope ('the current page') and dual mode ('all of them, or one by name'). It is immediately distinguishable from siblings add_cookie and delete_cookie.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete use cases ('check that a login created a session cookie, or to copy one into another session with add_cookie') and routes to the add_cookie alternative. Stops short of stating when-not to use it, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_current_urlA
Return the URL of the current tab, including any query string and fragment. Use it to confirm where a click or redirect landed; to wait until the URL changes, use wait_for_page.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavior burden. It tells the agent the return includes query string and fragment, which is meaningful. However it does not state whether this is a synchronous read, whether it reflects the active frame vs top-level document, or what happens when no page is loaded — gaps for a zero-arg read tool with no annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the return definition followed immediately by the usage/alternative. No filler words, every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-param read tool with no output schema, the description covers the return shape (URL plus query string and fragment) and the primary use case. It stops short of covering mode of operation (sync/async, frame scope, no-page behavior), which would be needed for a perfect score given the absence of annotations and output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline is 4. The description appropriately focuses on behavior rather than params; the 'current tab' scope is implicit context that helps interpret the call in a multi-tab session, but no param compensation is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Return') and resource ('the URL of the current tab'), and specifies the scope ('including any query string and fragment'). It is clearly distinguishable from its sibling get_title and from navigate/wait_for_page without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the use case ('confirm where a click or redirect landed') and names the alternative for the adjacent need ('to wait until the URL changes, use wait_for_page'). Both when-to-use and which-alternative are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_page_sourceA
Return the page's current HTML (the live DOM, including changes made by scripts), cut off at maxLength characters. Returns the HTML, its full length, and whether it was truncated. It is large and noisy: to find elements to act on, prefer capture_page; to read specific text or values, use get_text or get_attribute. Use this for raw markup such as meta tags, hidden fields, or inline data.
| Name | Required | Description | Default |
|---|---|---|---|
| maxLength | No | Maximum characters of HTML to return, 100-500000 (default 50000). Longer pages are truncated. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden, and it does well: it discloses that the result is the live DOM, that output is truncated at maxLength, the return shape (HTML, full length, truncated flag), and that the payload is 'large and noisy'. It does not quantify expected payload size or performance cost, which keeps it just under the top mark.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with what is returned before the caveats and alternatives. Every clause earns its place: none is redundant with the schema or with a sibling-name-only listing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description still tells the agent what comes back (HTML, full length, truncation flag) and how to route away when that is not what is wanted. Nothing needed to invoke this single-optional-parameter tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the maxLength description already gives range and default, so the schema does the heavy lifting. The description's 'cut off at maxLength characters' adds the truncation consequence but no format or syntax detail beyond the schema, making 3 the correct baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Return the page's current HTML'), and immediately qualifies it as the live DOM including script mutations, which is exactly the distinction an agent needs. It names the sibling alternatives (capture_page, get_text, get_attribute) so the tool is separable without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing: 'to find elements to act on, prefer capture_page; to read specific text or values, use get_text or get_attribute,' plus a positive case for raw markup (meta tags, hidden fields, inline data). Both when-to-use and when-to-prefer-alternatives are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_textA
Read the visible text of an element: what a user sees, not hidden text or an input's value. Waits for the element to be visible, then returns its text (trimmed by default). Target it by selector or by a ref from capture_page. To read an input's value or any attribute, use get_attribute; to check text as a pass/fail test step, use assert_text.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Element ref from the latest capture_page result (e.g. 'e12'). Use instead of selector; refs go stale after navigation, so capture again if one is rejected. | |
| trim | No | Remove leading and trailing whitespace (default true). | |
| selector | No | How to find the element, e.g. { by: 'id', value: 'user-name' } or { by: 'css', value: '#login-button' }. Provide either selector or ref. | |
| timeoutMs | No | How long to wait for the element to become visible, in milliseconds (default 10000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does so well for a read tool: it discloses that it waits for the element to become visible before returning, that it excludes hidden text, and that output is trimmed by default. It does not mention the timeout ceiling or pagination-style edge cases (e.g. what happens if the element never becomes visible), which the schema's timeoutMs only partially covers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences with zero waste: purpose, then visible-text qualifier, then targeting and sibling routing. Every clause earns its place and the most important constraint (visible, not hidden/value) comes first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by stating that it returns the element's text, trimmed by default. Combined with the visibility-wait behavior, cursor-targeting guidance, and sibling routing, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already fully documented in the schema. The description reinforces ref-vs-selector targeting and the trimmed-by-default behavior but adds no syntax or format detail beyond what the schema provides. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb+resource ('Read the visible text of an element') and immediately scopes it with the key qualifier 'what a user sees, not hidden text or an input's value.' It explicitly contrasts itself with get_attribute and assert_text, so an agent can distinguish it from siblings without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names both alternatives and the exact condition that selects each: get_attribute for input values or attributes, assert_text for pass/fail text assertions. It also tells the agent how to target the element (selector or a ref from capture_page), leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_titleA
Return the current page's title, the text shown in the browser tab. Use it to confirm which page is open; to wait until the title changes, use wait_for_page with titleContains.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, but this is a zero-parameter read-only getter, so the behavioral burden is inherently light. The description discloses what the value represents (the browser-tab title), which is the main behavioral fact an agent needs; it says nothing about failure modes when no page is open, keeping it just short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, and the return semantics are front-loaded before the alternative-tool hint. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must convey the return value, and it does ('the current page's title / text shown in the browser tab'). It does not cover edge cases such as a title-less or not-yet-loaded page, so it is complete for normal use but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to clarify, and it correctly omits any parameter discussion rather than inventing one.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Return the current page's title') and immediately disambiguates it from page content and URLs by clarifying it is 'the text shown in the browser tab'. This lets an agent distinguish it from siblings like get_text and get_current_url without inspecting any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete use case ('confirm which page is open') and explicitly names the alternative for the adjacent task: 'to wait until the title changes, use wait_for_page with titleContains.' That is a clear when-to-use plus a named sibling with the selecting condition, which is the strongest form of routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
historyA
Browser history navigation: go back to the previous page (like the Back button), go forward to the next page, or refresh/reload the current page (like F5). Waits for the page to load and returns the new URL and title. To open a specific URL, use navigate instead. Element refs from an earlier capture_page may be stale afterwards, so capture the page again before using refs. Refresh can reset unsaved form input.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | 'back' = previous page in this tab's history (Back button), 'forward' = next page (Forward button), 'refresh' = reload the current page (F5). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so description carries full burden. It discloses that it 'Waits for the page to load', returns the new URL and title, warns about stale element refs requiring recapture, and notes that refresh can reset unsaved form input. These are important behavioral details that go beyond what the schema provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each front-loaded with essential information: first sentence lists actions with analogies, second gives alternate tool, third warns about side effects. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description is nearly complete: it covers actions, return values (new URL and title), side effects (stale refs, form reset), and alternative tool. It could mention that it only works within the current tab's history or any limitations on forward navigation, but it's largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the enum values are fully described in the schema. The description lists the same three actions, adding no new syntax or format details, so baseline would be 3. However, it provides contextual analogies ('like the Back button') that may slightly aid understanding, warranting a 4 given the schema already does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific actions (back, forward, refresh) with analogies to browser buttons, making the purpose immediately clear. It doesn't differentiate itself from all siblings like get_current_url in the description, but the action enumeration and reference to navigate provide context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use this tool (back/forward/refresh navigation) and provides an alternative ('To open a specific URL, use navigate instead'). This covers the main routing decision, though it doesn't mention other siblings like get_current_url or press_key alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
interactA
Mouse actions beyond a plain click: double_click, right_click (opens a context menu), hover (reveals menus and tooltips), or click performed as a real mouse move-and-click. Waits for the element to be visible (hover) or visible and enabled (clicks). Target it by selector or by a ref from capture_page. For an ordinary click, prefer click.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Element ref from the latest capture_page result (e.g. 'e12'). Use instead of selector; refs go stale after navigation, so capture again if one is rejected. | |
| action | Yes | double_click | right_click | hover | click. 'click' moves the mouse onto the element first, which helps with elements that only react to real pointer movement. | |
| selector | No | How to find the element, e.g. { by: 'id', value: 'user-name' } or { by: 'css', value: '#login-button' }. Provide either selector or ref. | |
| timeoutMs | No | How long to wait for the element to become ready, in milliseconds (default 10000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does useful work: it discloses per-action readiness semantics (waits for visible on hover, visible-and-enabled on clicks) and the side effect that right_click opens a context menu. It does not cover failure behavior (what happens on timeout), whether the page can navigate as a result of a click, or what is returned, which are the remaining gaps for a state-mutating interaction tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action list, followed by readiness semantics and then the sibling routing rule. No filler, and the most decision-relevant information (what actions exist and which tool to use instead) comes first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter interaction tool with no annotations and no output schema, the description covers actions, targeting, and readiness adequately. It omits the post-action contract (return value, whether navigation may occur) and timeout behavior, which an agent would benefit from knowing before triggering a mutation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds real meaning beyond the schema by tying readiness/waiting behavior to the action value (hover vs clicks) and restating the selector-or-ref targeting choice. It stops short of adding timeout or locator-strategy nuance, so it is an incremental rather than a transformative gain.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific resource (mouse actions on an element) and enumerates the exact operations: double_click, right_click, hover, and click-as-real-mouse-move. It explicitly separates itself from the sibling 'click' tool with 'For an ordinary click, prefer click,' so an agent can disambiguate without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit routing rule to the alternative tool ('For an ordinary click, prefer click') and describes the conditions that select each action (right_click opens a context menu, hover reveals menus/tooltips, click helps elements that react to real pointer movement). This is as close to when/when-not guidance as a multi-action tool can get.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
press_keyA
Press one key on whichever element currently has focus, e.g. Enter to submit, Tab to move focus, Escape to close a dialog, or arrow keys in a list. Named keys: enter, tab, escape/esc, backspace, delete, space, arrowup, arrowdown, arrowleft, arrowright, home, end, pageup, pagedown; any other value is typed as literal text. To type into a specific field, use type, which focuses the field first.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Key name (case-insensitive), e.g. 'Enter', 'Tab', 'Escape', 'ArrowDown', or a single character. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that the key is delivered to the focused element rather than a addressed target, enumerates the recognized named keys, and reveals the important fallback that any other value is typed as literal text. It does not state what happens when nothing is focused or how failures surface, which is the remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all earning their place, with the core behavior front-loaded and the sibling routing placed last where it aids selection. No filler or restatement of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists and none is implied for a keystroke action, so return-value explanation is unnecessary. The accepted key vocabulary and literal-text fallback are covered; only edge behavior (no focused element, unsupported keys, focus loss) is left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real meaning beyond the schema: it enumerates the accepted named-key set and, crucially, defines the behavior of values outside that set (typed as literal text). That is semantics the schema's single 'key' string field does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (press) and resource (one key) plus the crucial scope restriction: it acts on whatever currently has focus, not a selector-addressed target. This cleanly distinguishes it from the sibling 'type', which focuses a field first.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage scenarios (Enter to submit, Tab to move focus, Escape to close a dialog, arrows in a list) and explicitly names the alternative tool plus the condition selecting it: 'To type into a specific field, use type, which focuses the field first.' No inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retry_clickA
Click with retries, for clicks that fail intermittently: the element is covered by a loading overlay or animation, or re-rendered (stale) between being found and clicked. Each attempt waits for the element to be visible and enabled, then clicks; failed attempts pause delayMs before the next. Returns the attempt that succeeded, or every attempt's error. Try click first; use this when click failed because of timing.
| Name | Required | Description | Default |
|---|---|---|---|
| delayMs | No | Pause between failed attempts, in milliseconds (default 250). | |
| attempts | No | Maximum click attempts, 1-10 (default 3). | |
| selector | Yes | How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector. | |
| timeoutMs | No | How long each attempt waits for the element to become visible and enabled, in milliseconds (default 10000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the retry loop, that each attempt waits for the element to be visible and enabled, the inter-attempt pause, and the return behavior on success vs. total failure. It omits any note on side effects of the click itself or whether the retry re-resolves the selector, but the core dynamics are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and failure modes, then mechanism, then return values. Dense and mostly waste-free; the middle clause enumerating failure modes is slightly verbose but earns its place by clarifying when to choose this tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description covers the return contract ("the attempt that succeeded, or every attempt's error"), which is what an agent needs. Combined with fully-covered parameters and explicit usage routing, it is essentially complete for a retry-click tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all four parameters are already documented, including delayMs and attempts. The description only loosely echoes these ("failed attempts pause delayMs"), adding little beyond what the schema states, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ("Click with retries") and immediately scopes it to "clicks that fail intermittently," naming the concrete failure modes (loading overlay, animation, stale re-render). This distinguishes it cleanly from the sibling `click` tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: "Try click first; use this when click failed because of timing." This states both the preferred alternative and the exact condition that selects this tool, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrollA
Scroll the page: scroll down or up by pixels, jump to the top or bottom, or scroll an element into view (to the middle of the screen). Also scrolls inside a scrollable container (chat panels, tables, sidebars with their own scrollbar) when you pass that container as selector/ref together with to or deltaX/deltaY. Returns the scroll position with atTop/atBottom flags. For lazy-loaded or infinite-scroll pages ('load more' on scroll), scroll to the bottom, wait_for_element for the new items, and repeat until atBottom stays true and nothing new appears. click and type already scroll their target into view, so use this to reveal content or trigger lazy loading.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | Jump to the top or bottom of the page, or of the container given by selector/ref. Scrolling to the bottom is how infinite-scroll pages load more items. | |
| ref | No | Element ref from capture_page (e.g. 'e12'); same meaning as selector. | |
| deltaX | No | Pixels to scroll horizontally: positive = right, negative = left. | |
| deltaY | No | Pixels to scroll vertically: positive = down, negative = up (e.g. 800 for about one screen). | |
| selector | No | Element to scroll to, brought to the middle of the screen. If you also pass to or deltaX/deltaY, this element is instead the scrollable container to scroll inside. Provide selector or ref, not both. | |
| timeoutMs | No | How long to wait for the selector/ref element to exist, in milliseconds (default 10000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the return value (scroll position with atTop/atBottom flags), the container-vs-element distinction, and the lazy-loading repetition pattern. It is thin on failure/timeout behavior beyond what the timeoutMs schema already states, keeping it just under a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense paragraph with the core verb front-loaded and no filler sentences; each clause adds information. It is on the long side and could be broken into a mode list, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with no output schema and no annotations, the description fills the critical gaps by stating the return shape and the lazy-load workflow. It is complete enough to invoke correctly, with only minor omissions around error handling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all six parameters including the dual role of selector/ref and the deltaY example. The description largely restates this, adding only the 'middle of the screen' target nuance and the container semantics already present in the schema. Baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Scroll the page') and enumerates the distinct scroll modes: pixel delta, top/bottom jump, element-into-view, and inside a scrollable container. It further distinguishes itself from sibling tools by noting that click and type already scroll their targets into view.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('reveal content or trigger lazy loading') and when-not-to-use ('click and type already scroll their target into view'). It also prescribes a concrete alternative workflow for infinite scroll, naming the sibling tool wait_for_element and the repeat-until-atBottom condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
select_optionA
Select an option in a native dropdown by visible text, value, or zero-based index, and return the resulting selection. Waits for the dropdown to be visible, then clicks the option like a user so change events fire; on a multi-select it adds to the current selection. If nothing matches, the error lists the available options so you can retry. Only for real elements; for custom dropdowns built from other elements, click the trigger and then the option.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Element ref from capture_page (e.g. 'e12') pointing at the <select>. Use instead of selector. | |
| text | No | Visible text of the option to select, matched exactly after trimming (e.g. 'Canada'). Provide exactly one of text, value, or index. | |
| index | No | Zero-based position of the option, counting every <option> including placeholders like 'Choose...'. Provide exactly one of text, value, or index. | |
| value | No | The option's value attribute (e.g. 'ca'). Provide exactly one of text, value, or index. | |
| selector | No | Locator for the <select> element, e.g. { by: 'id', value: 'country' }. Provide either selector or ref. | |
| timeoutMs | No | How long to wait for the dropdown to become visible, in milliseconds (default 10000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the wait-for-visible behavior, that it clicks like a user so change events fire, that multi-select adds to the existing selection, that mismatches error out with the available option list, and that it returns the resulting selection. It does not discuss timeout expiry behavior or read-only vs mutating implications in any depth, which is the only remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then behavior (wait, click semantics, multi-select, errors), then the exclusion. Four clauses, each carrying distinct information with no repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, so the description must cover behavior and returns — it states the return value (resulting selection) and the error format. Combined with a fully documented 6-parameter schema, an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (ref, text, value, index, selector, timeoutMs) is already fully documented including the exactly-one-of constraint and zero-based indexing. The description echoes the three selection modes ('visible text, value, or zero-based index') but adds no format or syntax detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (select) and resource (option in a native <select> dropdown) with a precise scope, and explicitly carves itself apart from the sibling 'click' path for custom dropdowns. An agent can distinguish it from click/interact without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit exclusion and the alternative: 'Only for real <select> elements; for custom dropdowns built from other elements, click the trigger and then the option.' It also states the pre-condition (waits for the dropdown to be visible) and the failure/retry path (error lists available options).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
selector_hint_deleteA
Permanently remove a saved selector hint from the hints file, e.g. when the site changed and the selector no longer works. This cannot be undone; save a corrected hint with selector_hint_save. Fails if no hint with that key exists for the domain.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Short name for the element, e.g. 'login_button' or 'search_box' (1-128 chars). | |
| domain | No | Site hostname, e.g. 'www.saucedemo.com'. Defaults to the hostname of the page currently open. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses irreversibility ('This cannot be undone'), the failure mode (key/domain mismatch), and the preferred follow-up path. It stops short of describing any file-location or permission prerequisites, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: action, reason/irreversibility, alternative, and error condition. The destructive nature is front-loaded in the first word.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter destructive tool with no annotations and no output schema, the description covers the essentials an agent needs: irreversibility, the failure precondition, and the replacement workflow. It omits only secondary details such as where the hints file lives or what auth is required.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both 'key' and 'domain' are already documented in the schema. The description only implies the domain-key pairing through the failure condition and adds no format or default details beyond the schema; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Permanently remove') and resource ('a saved selector hint from the hints file'), and its scope is immediately distinguishable from the sibling selector_hint_save/get/list operations without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger scenario ('when the site changed and the selector no longer works') and routes the agent to the correct alternative ('save a corrected hint with selector_hint_save'). It also documents the when-not condition ('Fails if no hint with that key exists for the domain').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
selector_hint_getA
Look up a saved selector by key for a site and return it, ready to pass as the selector of click, type, get_text, and similar tools. Fails if no hint with that key exists for the domain; use selector_hint_list to see what is saved.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Short name for the element, e.g. 'login_button' or 'search_box' (1-128 chars). | |
| domain | No | Site hostname, e.g. 'www.saucedemo.com'. Defaults to the hostname of the page currently open. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses the missing-key failure mode and that the returned value is directly consumable by click/type/get_text, but says nothing about persistence scope, session coupling, or whether the hint can be stale.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the lookup semantics and downstream use come first, then the failure/alternative guidance. No filler or repetition of structured data.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by stating what is returned (a selector ready for other tools) and what happens on miss. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, covering both 'key' and 'domain'. The description only restates that the lookup is per-site/domain; it adds no format or constraint detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource (look up a saved selector by key for a site) and states the output's purpose: it is ready to pass as the selector of click, type, get_text. This distinguishes it from selector_hint_list and the hint-mutation siblings without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives the failure condition ('fails if no hint with that key exists for the domain') and an explicit alternative ('use selector_hint_list to see what is saved'), which is exactly the when-to-use/alternate-route guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
selector_hint_listA
List saved selector hints (key, domain, and selector) for every site, or only for one domain. Check this at the start of a run on a familiar site to reuse known selectors. Does not need a browser to be running.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | No | Only list hints for this hostname, e.g. 'www.saucedemo.com'. Omit to list hints for every site. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose a meaningful operative fact: 'Does not need a browser to be running.' This is valuable given nearly every sibling tool requires an active browser session. It doesn't mention pagination or result-size behavior, but for a trivial read that gap is minor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and what is returned; the operational caveat is correctly placed last. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, but the description compensates by naming the returned fields (key, domain, selector). Combined with the browser-not-required note, an agent has what it needs to call and interpret this tool; only exhaustiveness/pagination details are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the parameter's own description already defines the domain filter and the omit-to-list-all behavior. 'For every site, or only for one domain' restates the schema rather than adding new meaning, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a precise verb+resource ('List saved selector hints') and enumerates the returned fields (key, domain, selector), plus the scope options (every site or one domain). This clearly separates it from the sibling selector_hint_get (single key), selector_hint_save, and selector_hint_delete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a concrete trigger: 'Check this at the start of a run on a familiar site to reuse known selectors.' That is clear when-to-use guidance, but it never names the alternative (selector_hint_get for a known key) or any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
selector_hint_saveA
Remember a selector that worked, under a short key for a site, so later runs can reuse it instead of rediscovering the element. Hints are saved to a JSON file on disk (.selenium-mcp/selector-hints.json in the server's working directory, or SELENIUM_MCP_SELECTOR_HINTS_PATH) and survive restarts. Save a hint after a selector has worked, e.g. after a successful click. Returns the saved hint.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Short name for the element, e.g. 'login_button' or 'search_box' (1-128 chars). | |
| domain | No | Site hostname, e.g. 'www.saucedemo.com'. Defaults to the hostname of the page currently open. | |
| selector | Yes | How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well: it discloses persistence location (.selenium-mcp/selector-hints.json or SELENIUM_MCP_SELECTOR_HINTS_PATH), durability across restarts, and the return value. These are non-obvious side effects an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads what the tool does, then storage location, then the usage trigger, then return value. Slightly dense but every sentence adds information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, persistence, keying, and return value for a 3-param nested-object tool with rich schema. No output schema exists, so explaining the return briefly is appropriate. Missing only explicit sibling/alternative routing, which is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema fully documents key, domain, and selector (including enum strategies and the default-to-current-hostname behavior). The description adds no parameter-level detail beyond what the schema provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (save/remember) and resource (selector hint keyed by site), and distinguishes itself from the sibling read tool selector_hint_get by clarifying the write direction. The keying and domain scoping is explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete when-to-use trigger ('Save a hint after a selector has worked, e.g. after a successful click'). Does not explicitly name selector_hint_get as the read counterpart or state when not to save, but the usage context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_createA
Open an additional, independent browser session (its own window, cookies, and login state) and make it the active one; all other tools act on the active session. Use it to test several users at once, such as a buyer and a seller, or to compare two states side by side. Switch between sessions with session_select. Takes the same options as start_browser. Returns the new session and overall status.
| Name | Required | Description | Default |
|---|---|---|---|
| browser | No | Which browser to launch: chrome (default), firefox, or edge. It must be installed on this machine. | chrome |
| headless | No | Run without a visible window (default false). | |
| sessionId | No | Optional readable id for the session, e.g. 'buyer' or 'admin' (1-128 chars). Defaults to a random UUID. | |
| windowSize | No | Initial window size, e.g. { width: 1440, height: 900 }. Change it later with the window tool's resize action. | |
| browserArgs | No | Extra command-line flags for the browser, e.g. ['--incognito']. Default none. | |
| scriptTimeoutMs | No | Maximum time an asynchronous script may run, in milliseconds (default 30000). | |
| pageLoadTimeoutMs | No | Maximum time a page load may take before navigate fails, in milliseconds (default 30000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and discloses the critical side effect: the new session becomes active and 'all other tools act on the active session'. It also states the return shape. It omits failure behavior, resource cost, and whether this spawns a new process vs. tab.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and its key trait, then usage scenarios, then the sibling routing, then the option/return note. Four tight sentences with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a session-creating tool with no output schema, the description covers purpose, active-session semantics, sibling routing, and return shape ('the new session and overall status'), while the schema fully documents parameters. Nothing essential to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all seven parameters (browser enum, headless, sessionId, windowSize, browserArgs, timeouts). The description only adds 'takes the same options as start_browser', which is useful context but not parameter detail. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (open a browser session) plus the distinguishing traits: 'additional, independent' with its own window, cookies, and login state, and that it becomes the active one. This clearly separates it from start_browser (first session) and session_select (switching).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage scenarios (test several users like buyer/seller, compare two states) and names the switching alternative (session_select). It does not explicitly contrast against start_browser as the 'when not to use', leaving that inference to the word 'additional'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_destroyA
Close one browser session by id and quit its browser; its cookies, login state, and open pages are lost. If it was the active session, another open session becomes active, or none if it was the last. To close just the active session, stop_browser does the same.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | Session id, as returned by session_create or session_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the destructive side effects (cookies, login state, and open pages are lost; the browser process quits) and the active-session reassignment behavior, including the last-session edge case. It omits return value/error semantics, but the core behavioral profile is covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, zero filler, front-loaded with the destructive effect and closing with alternative routing. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter mutation tool with no annotations and no output schema, the description covers purpose, destruction, and side effects adequately. Minor gaps remain around failure modes and whether an id validation error occurs, so it is strong but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter is fully documented in the schema (id returned by session_create or session_list). The description restates 'by id' but adds no format or syntax detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (close/destroy) and resource (one browser session) plus the concrete consequence 'quit its browser.' It is immediately distinguishable from stop_browser, which acts on the active session rather than a specified id.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to an alternative: 'To close just the active session, stop_browser does the same.' That is a clear when-to-use-this-vs-that statement, though it gives no guidance on error conditions (unknown id, already-closed session) or whether the target must first be selected.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_listA
List every open browser session (id, browser, headless, window size, start time) and which one is active. Use it to find session ids for session_select or session_destroy, or to check what is still running.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It compensates by disclosing the exact shape of the result (the fields and the active marker), effectively standing in for the missing output schema, and 'List' signals a non-mutating read. It stops short of stating whether the list can be empty or whether sessions from other clients are included.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The what (list with field breakdown) is front-loaded and the routing guidance follows immediately, so an agent gets value from the first clause onward.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only listing tool with no output schema and no annotations, the description covers the return shape and the intended follow-up actions. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing to disambiguate; baseline 4 applies. The description correctly adds no parameter guidance because none is possible.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List every open browser session') and enumerates exactly what each entry contains (id, browser, headless, window size, start time, active flag). This cleanly separates it from the session_create/session_select/session_destroy siblings, which mutate rather than enumerate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the downstream alternatives and the condition that selects them: use this to find ids for session_select or session_destroy, or to check what is still running. The agent does not have to infer when this tool is the right entry point.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_selectA
Make an existing browser session the active one, so every following tool call acts on that browser. The other sessions stay open exactly as they were. Use session_list to see the available ids.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | Session id, as returned by session_create or session_list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose key behavior: other sessions remain open untouched, and the selection governs every following tool call (ambient state). It omits what happens on an invalid or destroyed session id, which is a notable gap for a state-changing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and then the side-effect and the id source; nothing redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter selector with no output schema, the description covers purpose, cross-call effect, and id sourcing adequately; only error behavior on an invalid id is left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single sessionId parameter is already documented as coming from session_create or session_list, so the description adds only a pointer to session_list rather than new semantics. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Make an existing browser session the active one') plus the effect on subsequent calls, which cleanly separates it from session_list, session_create, and session_destroy.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operating context (multi-session switching) and points to session_list for obtaining ids, but never states when this should be preferred over session_create/start_browser or what to do when no session is currently open.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_browserA
Start a browser session (Chrome, Firefox, or Edge) that all other tools act on. Call this first. Only one session can be started this way: call stop_browser before starting another, or use session_create to run several browsers at once. The matching driver is downloaded automatically by Selenium Manager. Returns the session status.
| Name | Required | Description | Default |
|---|---|---|---|
| browser | No | Which browser to launch: chrome (default), firefox, or edge. It must be installed on this machine. | chrome |
| headless | No | Run without a visible window (default false). Use true for CI or background runs, false when the user wants to watch. | |
| windowSize | No | Initial window size, e.g. { width: 1440, height: 900 }. Change it later with the window tool's resize action. | |
| browserArgs | No | Extra command-line flags for the browser, e.g. ['--incognito'] or ['--lang=de']. Default none. | |
| scriptTimeoutMs | No | Maximum time an asynchronous script may run, in milliseconds (default 30000). | |
| pageLoadTimeoutMs | No | Maximum time a page load may take before navigate fails, in milliseconds (default 30000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full disclosure burden. It discloses two valuable non-obvious behaviors: the session is a singleton that all other tools act on, and the driver is auto-downloaded by Selenium Manager. It stops short of failure modes (e.g. browser not installed beyond the schema note) or auth/permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five tight sentences, front-loaded with the core action and the ordering instruction before the constraints and alternatives. Every sentence carries a distinct, useful fact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter, all-optional tool with no output schema, the description covers the key invocation contract (call first, singleton, driver auto-fetch) and notes the return value ('Returns the session status.'). Slightly light on what the returned status contains, but adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema with defaults, ranges, and examples. The description adds no syntax or format detail beyond what the schema provides, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Start) and resource (a browser session) with the concrete browser options, and explicitly distinguishes itself from stop_browser and session_create in the same breath. An agent can identify this tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit sequencing guidance ('Call this first'), the singleton constraint, and named alternatives with the conditions that select them ('call stop_browser before starting another, or use session_create to run several browsers at once'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop_browserA
Close the active browser session and quit its browser; its cookies, login state, and open pages are lost. With several sessions open, only the active one closes and another becomes active; use session_destroy to close a specific one. Safe to call when nothing is running. Call it when you are done, so no browser is left open.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries the behavioral burden: it warns that cookies, login state, and open pages are lost, explains multi-session behavior (only active closes, another becomes active), and states that calling it when nothing is running is safe.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and destructive consequence, then efficiently layers in multi-session behavior, the alternative tool, and safe no-op guidance. Every sentence earns its place with no wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool with no annotations, the description is complete: it covers the operation, side effects, multi-session routing, safety, and recommended usage timing. Nothing needed for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters and the schema is empty, so parameter semantics are not applicable. The baseline for a zero-parameter tool is 4, and the description appropriately adds no unnecessary parameter detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: close the active browser session and quit its browser. It also distinguishes the tool from its closest sibling by explicitly naming session_destroy for closing a specific session.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance ('Call it when you are done'), a safe no-op condition ('Safe to call when nothing is running'), and the alternative for a different need ('use session_destroy to close a specific one').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
take_screenshotA
Capture a PNG screenshot of the visible viewport. Returns it as base64 and/or saves it to savePath (folders are created as needed; an existing file is overwritten). Use it when visual layout or appearance matters, or to show the user a step. To check text or state, prefer capture_page, get_text, or the assert tools: they are exact and much smaller than an image.
| Name | Required | Description | Default |
|---|---|---|---|
| savePath | No | Absolute path to write the PNG to, e.g. '/home/me/shots/login.png' or 'C:/shots/login.png'. Omit to not save. | |
| includeBase64 | No | Include the PNG as base64 in the result (default true). Set false when saving to disk to keep the response small. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the return format (base64 and/or disk), that folders are created as needed, and that an existing file is overwritten – a clear destructive-on-disk side effect. It omits any note on permissions or capture timing, but the important behavioral traits are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all front-loaded: purpose first, then behavior/side effects, then the routing guidance. Every sentence earns its place and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a screenshot tool with no output schema: the description explains both the returned payload (base64) and the disk-write behavior, and gives clear routing relative to the many sibling read/inspect tools. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and both parameters are fully documented in the schema, including the savePath example paths and the includeBase64 default. The description reinforces the base64/savePath duality but adds no syntax or format detail beyond the schema, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Capture a PNG screenshot of the visible viewport'), including the scope ('visible viewport') that distinguishes it from a full-page capture. It also names the sibling capture_page and the text-based alternatives, so an agent can route between them without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('when visual layout or appearance matters, or to show the user a step') and explicit when-not with named alternatives ('To check text or state, prefer capture_page, get_text, or the assert tools'). The exclusion criterion ('exact and much smaller than an image') gives a concrete decision rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
typeA
Type text into an input, textarea, or other editable element. Waits for it to be visible, clears the current value first (clearFirst: false appends instead), then sends the text as keystrokes so the page's input events fire; submit: true presses Enter afterwards. Target it by selector or by a ref from capture_page. For dropdowns use select_option; for single keys like Tab or Escape use press_key.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Element ref from the latest capture_page result (e.g. 'e12'). Use instead of selector; refs go stale after navigation, so capture again if one is rejected. | |
| text | Yes | The text to type, e.g. 'standard_user'. | |
| submit | No | Press Enter after typing (default false), e.g. to submit a search box or login form. | |
| selector | No | How to find the element, e.g. { by: 'id', value: 'user-name' } or { by: 'css', value: '#login-button' }. Provide either selector or ref. | |
| timeoutMs | No | How long to wait for the field to become visible, in milliseconds (default 10000). | |
| clearFirst | No | Clear the field before typing (default true). Set false to append to the existing value. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the visibility wait, that the field is cleared first (clearFirst: false appends), that text is sent as keystrokes so the page's input events fire, and that submit: true presses Enter. It does not cover failure behavior (e.g., what happens if the element is not found within timeoutMs) or the return payload, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the core action and behavior, then the submit/clearing modifiers, then the sibling routing. No filler and nothing buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation-style tool with no annotations and no output schema, the description covers action, waiting, clearing, keystroke semantics, submit behavior, and alternative tools. It leaves failure/retry behavior and the response undefined, which is a modest gap rather than a blocker.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents ref, text, submit, selector, timeoutMs and clearFirst in detail; baseline is 3. The description restates clearFirst/submit semantics rather than adding new syntax or format information, so it adds little beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Type text into an input, textarea, or other editable element') and explicitly disambiguates from siblings by naming select_option for dropdowns and press_key for single keys. An agent can distinguish this from click/press_key/select_option without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing: use this for editable text fields, use select_option for dropdowns, use press_key for single keys like Tab/Escape. It also states the two targeting modes (selector or a ref from capture_page), so when-to-use and how-to-target are both covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_fileA
Attach a local file to a file input () without opening the operating system's file picker. Waits for the input to exist; it may be hidden, as styled upload buttons often hide the real input, so target the input itself. The file must exist on the machine running this server. After attaching, click the page's upload or submit button if it has one.
| Name | Required | Description | Default |
|---|---|---|---|
| filePath | Yes | Absolute path to the file on the machine running this server, e.g. '/home/me/report.pdf' or 'C:/Users/me/report.pdf'. | |
| selector | Yes | How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector. | |
| timeoutMs | No | How long to wait for the file input to exist, in milliseconds (default 10000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full burden. It does well: explains waiting for input existence, that hidden inputs are common, that the file must exist on the server's machine, and the post-upload click expectation. Doesn't cover whether the attach triggers events or what happens on failure, but covers the behavioral essentials an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, all earning their place, front-loaded with the action and the picker-avoidance point. No filler or repetition of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param nested-schema tool with no output schema and no annotations, the description covers the core workflow and edge cases (hidden input, server-side file, follow-up click). Missing a note on return value/errors and any permission or security context, but broadly complete for the browser-automation domain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with rich parameter docs (selector object, filePath examples, timeoutMs default/range). The description adds the server-side file location fact ('on the machine running this server'), reinforcing filePath semantics, but otherwise the schema fully documents parameters. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: attach a local file to a file input. Distinguishes from sibling 'type' or 'click' by clarifying it bypasses the OS file picker and targets the hidden input directly, which is the key differentiator for upload tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives context on when to use (attach without picker) and what to do afterward ('click the page's upload or submit button if it has one'), plus the constraint 'the file must exist on the machine running this server'. No explicit when-not guidance or named alternatives, but strong for the category.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_elementA
Wait for an element to exist in the page, or with visible: true to also be displayed, then return its tag and state (displayed, enabled). Use it before acting on content that loads later: spinners finishing, lazy lists, dialogs, single-page-app transitions. click, type, and get_text already wait for their own target, so use this to wait for something else first. For URL or title changes, use wait_for_page.
| Name | Required | Description | Default |
|---|---|---|---|
| visible | No | Also require the element to be displayed, not just present in the DOM (default false). | |
| selector | Yes | How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector. | |
| timeoutMs | No | How long to wait, in milliseconds (default 10000, max 60000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses the default existence-vs-visibility semantics, the returned state fields (displayed, enabled), and the timing behavior implied by 'wait'. It does not state what happens on timeout (error vs return) or block/polling behavior, which is a meaningful gap for a waiting primitive with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four front-loaded sentences, each earning its place: what it does and returns, when to use it, when not to, and where to go instead. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a 3-param waiting tool with full schema coverage and no output schema: purpose, semantics, and routing to siblings are all covered. The remaining gap is timeout/error behavior, which matters for a blocking primitive but is partially inferable from the timeoutMs parameter and default.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both selector and visible are already fully documented in the schema with examples and defaults. The description reinforces the visible toggle but adds no syntax or format detail beyond the schema. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (wait for element) with explicit scope: existence by default, visibility with visible:true, and return payload (tag and state). Clearly distinguishes itself from siblings like click/type/get_text which self-wait, and routes URL/title waits to wait_for_page.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use with concrete scenarios (spinners finishing, lazy lists, dialogs, SPA transitions), an explicit when-not-to-use note ('click, type, and get_text already wait for their own target'), and a named alternative for URL/title changes (wait_for_page). Nothing left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_pageA
Wait until the current page's URL or title meets a condition: URL contains text, URL matches a regular expression, and/or title contains text (all given conditions must hold). Use after an action that navigates or redirects (submitting a login form, clicking a link, a single-page-app route change) before checking the new page. Returns the final URL and title; on timeout, the error shows the URL and title the page actually had. To wait for an element to appear, use wait_for_element instead.
| Name | Required | Description | Default |
|---|---|---|---|
| timeoutMs | No | How long to wait, in milliseconds (default 10000, max 60000). | |
| urlMatches | No | Wait until the URL matches this JavaScript regular expression, e.g. '/orders/[0-9]+$'. | |
| urlContains | No | Wait until the URL contains this text, e.g. '/dashboard' or 'checkout-step-two'. | |
| titleContains | No | Wait until the page title contains this text (case-sensitive), e.g. 'Dashboard'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries full burden and does well: it discloses return behavior (final URL and title), timeout error content (the URL/title actually observed), and combined-condition semantics. Left implicit are the default/max timeout bounds (in schema) and blocking nature, so not quite a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: purpose/conditions first, usage context second, return/timeout and the sibling pointer last. No filler, front-loaded with the core capability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, yet the description explains what is returned and what a timeout error reveals. Combined with 100% schema coverage and no annotations needed for a read-only wait, an agent has everything required to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and each parameter documents its own semantics including examples and defaults, so the schema does the heavy lifting. The description adds the AND-combination rule across conditions, which is genuine value, but per-parameter meaning is otherwise redundant.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (wait) and resource (current page URL/title), enumerates the exact conditions (urlContains, urlMatches, titleContains) and their AND semantics. Clearly distinguished from sibling wait_for_element, which it explicitly names.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it: after a navigating/redirecting action (login submit, link click, SPA route change) and before checking the new page. Also routes the agent to wait_for_element for element waits, so both when-to-use and alternative are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
windowA
Manage browser tabs and windows, and resize the viewport. list = all window handles plus the current one; switch = go to a handle from list; switch_latest = go to the newest tab/window (e.g. after a link opened a new tab); new_tab / new_window = open a blank one and switch to it; close = close the current one (then switch to another handle). resize = set the viewport (page area) to width x height for responsive or mobile testing, e.g. 390x844 phone, 768x1024 tablet, 1920x1080 desktop (exact phone sizes use device emulation on Chrome/Edge); maximize = maximize the window. Returns window handles, or the resulting window and viewport sizes.
| Name | Required | Description | Default |
|---|---|---|---|
| width | No | Viewport width in CSS pixels for resize (e.g. 390 phone, 768 tablet, 1920 desktop). Required for resize. | |
| action | Yes | What to do: list | switch | switch_latest | new_tab | new_window | close | resize | maximize. | |
| handle | No | Window handle to switch to, as returned by action list. Required for switch; ignored otherwise. | |
| height | No | Viewport height in CSS pixels for resize (e.g. 844 phone, 1024 tablet, 1080 desktop). Required for resize. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it states side effects (new_tab/new_window open a blank one and switch to it; close closes the current one and then switches to another handle), the return values ('window handles, or the resulting window and viewport sizes'), and the Chrome/Edge-only caveat for exact phone sizes. It does not state permissions or rate limits, but nothing suggests those apply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the top-level scope and then a tight action-by-action breakdown. Every clause earns its place; there is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-action tool with no output schema, the description covers the main behavior, side effects, and return shape. It is slightly short of complete because it does not explain what happens if switch is called with a stale handle or whether resize affects only the current window.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by giving concrete action-to-parameter mappings (width/height for resize, handle for switch) and device-emulation guidance (exact phone sizes use device emulation on Chrome/Edge) that the schema does not contain.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('Manage browser tabs and windows, and resize the viewport') and then enumerates every action with a precise definition, so an agent can distinguish this from siblings like frame or alert that also affect context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Each action enum value is explained with a when-to-use condition (e.g. 'switch = go to a handle from list', 'switch_latest = go to the newest tab/window (e.g. after a link opened a new tab)'), which is strong. It does not name alternatives, but this tool has no direct sibling for tab/window management, so the gap is minor.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
39 tool updates
v0.3.0- Changed
add_cookie8 fields changed- added
Input schema / properties / domain / descriptionAdded value: +"Domain the cookie applies to, e.g. '.example.com' to include subdomains. Defaults to the current page's host." - added
Input schema / properties / expiry / descriptionAdded value: +"Expiry as a Unix timestamp in seconds. Omit for a session cookie that ends when the browser closes." - added
Input schema / properties / httpOnly / descriptionAdded value: +"Hide the cookie from page JavaScript (document.cookie)." - added
Input schema / properties / name / descriptionAdded value: +"Cookie name, e.g. 'session_id'." - added
Input schema / properties / path / descriptionAdded value: +"URL path the cookie applies to (default '/')." - added
Input schema / properties / sameSite / descriptionAdded value: +"SameSite policy: Strict, Lax, or None (None also requires secure: true). Omit for the browser default." - added
Input schema / properties / secure / descriptionAdded value: +"Only send the cookie over HTTPS." - added
Input schema / properties / value / descriptionAdded value: +"Cookie value."
- Changed
alert2 fields changed- added
Input schema / properties / action / descriptionAdded value: +"get_text = read the message; accept = OK; dismiss = Cancel; send_text = type into a prompt() (then accept)." - added
Input schema / properties / text / descriptionAdded value: +"For send_text: the text to type into the prompt() field. Ignored for other actions."
- Changed
assert_attribute7 fields changed- added
Input schema / properties / expected / descriptionAdded value: +"Expected value or pattern, e.g. 'true' for a disabled button." - added
Input schema / properties / mode / descriptionAdded value: +"equals = exact match (default); contains = substring; matches = JavaScript regular expression." - added
Input schema / properties / name / descriptionAdded value: +"Attribute or property to check, e.g. 'value', 'disabled', 'href', or 'class'." - added
Input schema / properties / selector / descriptionAdded value: +"How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector." - added
Input schema / properties / selector / properties / by / descriptionAdded value: +"Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText." - added
Input schema / properties / selector / properties / value / descriptionAdded value: +"The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath." - added
Input schema / properties / timeoutMs / descriptionAdded value: +"How long to wait for the element to exist, in milliseconds (default 10000)."
- Changed
assert_text7 fields changed- added
Input schema / properties / expected / descriptionAdded value: +"Text or pattern the element's text should match, e.g. 'Epic sadface: Username is required'." - added
Input schema / properties / mode / descriptionAdded value: +"equals = exact match; contains = substring (default); matches = JavaScript regular expression." - added
Input schema / properties / selector / descriptionAdded value: +"How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector." - added
Input schema / properties / selector / properties / by / descriptionAdded value: +"Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText." - added
Input schema / properties / selector / properties / value / descriptionAdded value: +"The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath." - added
Input schema / properties / timeoutMs / descriptionAdded value: +"How long to wait for the element to become visible, in milliseconds (default 10000)." - added
Input schema / properties / trim / descriptionAdded value: +"Trim whitespace from the actual text before comparing (default true)."
- Changed
assert_visible4 fields changed- added
Input schema / properties / selector / descriptionAdded value: +"How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector." - added
Input schema / properties / selector / properties / by / descriptionAdded value: +"Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText." - added
Input schema / properties / selector / properties / value / descriptionAdded value: +"The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath." - added
Input schema / properties / timeoutMs / descriptionAdded value: +"How long to wait for the element to become visible before failing, in milliseconds (default 10000)."
- Changed
batch_execute3 fields changed- added
Input schema / properties / steps / descriptionAdded value: +"Ordered steps, 1-10. Each has an action plus that action's fields, e.g. { action: 'type', selector: { by: 'id', value: 'user-name' }, text: 'standard_user' }." - changed
Input schema / properties / steps / items / oneOfPrevious value: -[ - { - "properties": { - "action": { - "const": "navigate", - "type": "string" - }, - "url": { - "format": "uri", - "type": "string" - } - }, - "required": [ - "action", - "url" - ], - "type": "object" - }, - { - "properties": { - "action": { - "const": "wait_for_element", - "type": "string" - }, - "selector": { - "properties": { - "by": { - "enum": [ - "css", - "xpath", - "id", - "name", - "class", - "className", - "tag", - "tagName", - "linkText", - "partialLinkText" - ], - "type": "string" - }, - "value": { - "minLength": 1, - "type": "string" - } - }, - "required": [ - "by", - "value" - ], - "type": "object" - }, - "timeoutMs": { - "default": 10000, - "maximum": 60000, - "minimum": 100, - "type": "integer" - }, - "visible": { - "default": false, - "type": "boolean" - } - }, - "required": [ - "action", - "selector" - ], - "type": "object" - }, - { - "properties": { - "action": { - "const": "click", - "type": "string" - }, - "selector": { - "properties": { - "by": { - "enum": [ - "css", - "xpath", - "id", - "name", - "class", - "className", - "tag", - "tagName", - "linkText", - "partialLinkText" - ], - "type": "string" - }, - "value": { - "minLength": 1, - "type": "string" - } - }, - "required": [ - "by", - "value" - ], - "type": "object" - }, - "timeoutMs": { - "default": 10000, - "maximum": 60000, - "minimum": 100, - "type": "integer" - } - }, - "required": [ - "action", - "selector" - ], - "type": "object" - }, - { - "properties": { - "action": { - "const": "type", - "type": "string" - }, - "clearFirst": { - "default": true, - "type": "boolean" - }, - "selector": { - "properties": { - "by": { - "enum": [ - "css", - "xpath", - "id", - "name", - "class", - "className", - "tag", - "tagName", - "linkText", - "partialLinkText" - ], - "type": "string" - }, - "value": { - "minLength": 1, - "type": "string" - } - }, - "required": [ - "by", - "value" - ], - "type": "object" - }, - "submit": { - "default": false, - "type": "boolean" - }, - "text": { - "type": "string" - }, - "timeoutMs": { - "default": 10000, - "maximum": 60000, - "minimum": 100, - "type": "integer" - } - }, - "required": [ - "action", - "selector", - "text" - ], - "type": "object" - }, - { - "properties": { - "action": { - "const": "execute_script", - "type": "string" - }, - "args": { - "default": [], - "items": { - "$ref": "#/definitions/__schema0" - }, - "type": "array" - }, - "script": { - "minLength": 1, - "type": "string" - } - }, - "required": [ - "action", - "script" - ], - "type": "object" - } -]New value: +[ + { + "properties": { + "action": { + "const": "navigate", + "type": "string" + }, + "url": { + "format": "uri", + "type": "string" + } + }, + "required": [ + "action", + "url" + ], + "type": "object" + }, + { + "properties": { + "action": { + "const": "wait_for_element", + "type": "string" + }, + "selector": { + "description": "How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector.", + "properties": { + "by": { + "description": "Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText.", + "enum": [ + "css", + "xpath", + "id", + "name", + "class", + "className", + "tag", + "tagName", + "linkText", + "partialLinkText" + ], + "type": "string" + }, + "value": { + "description": "The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath.", + "minLength": 1, + "type": "string" + } + }, + "required": [ + "by", + "value" + ], + "type": "object" + }, + "timeoutMs": { + "default": 10000, + "description": "How long to wait for the element, in milliseconds (default 10000, max 60000).", + "maximum": 60000, + "minimum": 100, + "type": "integer" + }, + "visible": { + "default": false, + "type": "boolean" + } + }, + "required": [ + "action", + "selector" + ], + "type": "object" + }, + { + "properties": { + "action": { + "const": "wait_for_page", + "type": "string" + }, + "timeoutMs": { + "default": 10000, + "description": "How long to wait for the element, in milliseconds (default 10000, max 60000).", + "maximum": 60000, + "minimum": 100, + "type": "integer" + }, + "titleContains": { + "minLength": 1, + "type": "string" + }, + "urlContains": { + "minLength": 1, + "type": "string" + }, + "urlMatches": { + "minLength": 1, + "type": "string" + } + }, + "required": [ + "action" + ], + "type": "object" + }, + { + "properties": { + "action": { + "const": "click", + "type": "string" + }, + "selector": { + "description": "How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector.", + "properties": { + "by": { + "description": "Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText.", + "enum": [ + "css", + "xpath", + "id", + "name", + "class", + "className", + "tag", + "tagName", + "linkText", + "partialLinkText" + ], + "type": "string" + }, + "value": { + "description": "The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath.", + "minLength": 1, + "type": "string" + } + }, + "required": [ + "by", + "value" + ], + "type": "object" + }, + "timeoutMs": { + "default": 10000, + "description": "How long to wait for the element, in milliseconds (default 10000, max 60000).", + "maximum": 60000, + "minimum": 100, + "type": "integer" + } + }, + "required": [ + "action", + "selector" + ], + "type": "object" + }, + { + "properties": { + "action": { + "const": "type", + "type": "string" + }, + "clearFirst": { + "default": true, + "type": "boolean" + }, + "selector": { + "description": "How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector.", + "properties": { + "by": { + "description": "Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText.", + "enum": [ + "css", + "xpath", + "id", + "name", + "class", + "className", + "tag", + "tagName", + "linkText", + "partialLinkText" + ], + "type": "string" + }, + "value": { + "description": "The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath.", + "minLength": 1, + "type": "string" + } + }, + "required": [ + "by", + "value" + ], + "type": "object" + }, + "submit": { + "default": false, + "type": "boolean" + }, + "text": { + "type": "string" + }, + "timeoutMs": { + "default": 10000, + "description": "How long to wait for the element, in milliseconds (default 10000, max 60000).", + "maximum": 60000, + "minimum": 100, + "type": "integer" + } + }, + "required": [ + "action", + "selector", + "text" + ], + "type": "object" + }, + { + "properties": { + "action": { + "const": "select_option", + "type": "string" + }, + "index": { + "maximum": 9007199254740991, + "minimum": 0, + "type": "integer" + }, + "selector": { + "description": "How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector.", + "properties": { + "by": { + "description": "Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText.", + "enum": [ + "css", + "xpath", + "id", + "name", + "class", + "className", + "tag", + "tagName", + "linkText", + "partialLinkText" + ], + "type": "string" + }, + "value": { + "description": "The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath.", + "minLength": 1, + "type": "string" + } + }, + "required": [ + "by", + "value" + ], + "type": "object" + }, + "text": { + "type": "string" + }, + "timeoutMs": { + "default": 10000, + "description": "How long to wait for the element, in milliseconds (default 10000, max 60000).", + "maximum": 60000, + "minimum": 100, + "type": "integer" + }, + "value": { + "type": "string" + } + }, + "required": [ + "action", + "selector" + ], + "type": "object" + }, + { + "properties": { + "action": { + "const": "execute_script", + "type": "string" + }, + "args": { + "default": [], + "items": { + "$ref": "#/definitions/__schema0" + }, + "type": "array" + }, + "script": { + "minLength": 1, + "type": "string" + } + }, + "required": [ + "action", + "script" + ], + "type": "object" + } +] - added
Input schema / properties / stopOnError / descriptionAdded value: +"Stop at the first failing step (default true). Set false to run every step and collect all errors."
- Changed
capture_page1 field changed- added
Input schema / properties / maxElements / descriptionAdded value: +"Maximum elements to return, 10-500 (default 200). Lower it on very large pages to keep the result small."
- Changed
click5 fields changed- added
Input schema / properties / ref / descriptionAdded value: +"Element ref from the latest capture_page result (e.g. 'e12'). Use instead of selector; refs go stale after navigation, so capture again if one is rejected." - added
Input schema / properties / selector / descriptionAdded value: +"How to find the element, e.g. { by: 'id', value: 'user-name' } or { by: 'css', value: '#login-button' }. Provide either selector or ref." - added
Input schema / properties / selector / properties / by / descriptionAdded value: +"Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText." - added
Input schema / properties / selector / properties / value / descriptionAdded value: +"The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath." - added
Input schema / properties / timeoutMs / descriptionAdded value: +"How long to wait for the element to become visible and enabled, in milliseconds (default 10000)."
- Changed
delete_cookie1 field changed- added
Input schema / properties / name / descriptionAdded value: +"Name of the cookie to delete, e.g. 'session_id'. Omit to delete ALL cookies for the current page."
- Changed
execute_script2 fields changed- added
Input schema / properties / args / descriptionAdded value: +"JSON values passed to the script as arguments[0], arguments[1], ... (default none)." - added
Input schema / properties / script / descriptionAdded value: +"JavaScript function body, e.g. 'return document.querySelectorAll(\".inventory_item\").length;'. Return plain JSON values (strings, numbers, booleans, arrays, objects)."
- Changed
find_element4 fields changed- added
Input schema / properties / selector / descriptionAdded value: +"How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector." - added
Input schema / properties / selector / properties / by / descriptionAdded value: +"Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText." - added
Input schema / properties / selector / properties / value / descriptionAdded value: +"The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath." - added
Input schema / properties / timeoutMs / descriptionAdded value: +"How long to wait for the element to exist, in milliseconds (default 10000)."
- Changed
frame7 fields changed- added
Input schema / properties / action / descriptionAdded value: +"switch = enter a frame (needs selector, index, or nameOrId); parent = up one level; default = back to the main page." - added
Input schema / properties / index / descriptionAdded value: +"For switch: zero-based position of the frame on the page. Used when no selector is given." - added
Input schema / properties / nameOrId / descriptionAdded value: +"For switch: the frame's name or id attribute. Used when neither selector nor index is given." - added
Input schema / properties / selector / descriptionAdded value: +"For switch: the <iframe> element, e.g. { by: 'css', value: 'iframe#editor' }. Most reliable way to pick a frame." - added
Input schema / properties / selector / properties / by / descriptionAdded value: +"Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText." - added
Input schema / properties / selector / properties / value / descriptionAdded value: +"The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath." - added
Input schema / properties / timeoutMs / descriptionAdded value: +"How long to wait for the frame element when switching by selector, in milliseconds (default 10000)."
- Changed
get_attribute5 fields changed- added
Input schema / properties / name / descriptionAdded value: +"Attribute or property name, e.g. 'value', 'href', 'disabled', 'checked', 'class', or 'aria-expanded'." - added
Input schema / properties / selector / descriptionAdded value: +"How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector." - added
Input schema / properties / selector / properties / by / descriptionAdded value: +"Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText." - added
Input schema / properties / selector / properties / value / descriptionAdded value: +"The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath." - added
Input schema / properties / timeoutMs / descriptionAdded value: +"How long to wait for the element to exist, in milliseconds (default 10000)."
- Changed
get_cookies1 field changed- added
Input schema / properties / name / descriptionAdded value: +"Name of one cookie to read, e.g. 'session_id'. Omit to read all cookies for the current page."
- Changed
get_page_source1 field changed- added
Input schema / properties / maxLength / descriptionAdded value: +"Maximum characters of HTML to return, 100-500000 (default 50000). Longer pages are truncated."
- Changed
get_text6 fields changed- added
Input schema / properties / ref / descriptionAdded value: +"Element ref from the latest capture_page result (e.g. 'e12'). Use instead of selector; refs go stale after navigation, so capture again if one is rejected." - added
Input schema / properties / selector / descriptionAdded value: +"How to find the element, e.g. { by: 'id', value: 'user-name' } or { by: 'css', value: '#login-button' }. Provide either selector or ref." - added
Input schema / properties / selector / properties / by / descriptionAdded value: +"Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText." - added
Input schema / properties / selector / properties / value / descriptionAdded value: +"The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath." - added
Input schema / properties / timeoutMs / descriptionAdded value: +"How long to wait for the element to become visible, in milliseconds (default 10000)." - added
Input schema / properties / trim / descriptionAdded value: +"Remove leading and trailing whitespace (default true)."
- Added
history - Changed
interact6 fields changed- added
Input schema / properties / action / descriptionAdded value: +"double_click | right_click | hover | click. 'click' moves the mouse onto the element first, which helps with elements that only react to real pointer movement." - added
Input schema / properties / ref / descriptionAdded value: +"Element ref from the latest capture_page result (e.g. 'e12'). Use instead of selector; refs go stale after navigation, so capture again if one is rejected." - added
Input schema / properties / selector / descriptionAdded value: +"How to find the element, e.g. { by: 'id', value: 'user-name' } or { by: 'css', value: '#login-button' }. Provide either selector or ref." - added
Input schema / properties / selector / properties / by / descriptionAdded value: +"Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText." - added
Input schema / properties / selector / properties / value / descriptionAdded value: +"The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath." - added
Input schema / properties / timeoutMs / descriptionAdded value: +"How long to wait for the element to become ready, in milliseconds (default 10000)."
- Changed
navigate1 field changed- added
Input schema / properties / url / descriptionAdded value: +"Absolute URL including the scheme, e.g. 'https://www.saucedemo.com'."
- Removed
open_url - Changed
press_key1 field changed- added
Input schema / properties / key / descriptionAdded value: +"Key name (case-insensitive), e.g. 'Enter', 'Tab', 'Escape', 'ArrowDown', or a single character."
- Changed
retry_click6 fields changed- added
Input schema / properties / attempts / descriptionAdded value: +"Maximum click attempts, 1-10 (default 3)." - added
Input schema / properties / delayMs / descriptionAdded value: +"Pause between failed attempts, in milliseconds (default 250)." - added
Input schema / properties / selector / descriptionAdded value: +"How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector." - added
Input schema / properties / selector / properties / by / descriptionAdded value: +"Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText." - added
Input schema / properties / selector / properties / value / descriptionAdded value: +"The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath." - added
Input schema / properties / timeoutMs / descriptionAdded value: +"How long each attempt waits for the element to become visible and enabled, in milliseconds (default 10000)."
- Added
scroll - Added
select_option - Changed
selector_hint_delete2 fields changed- added
Input schema / properties / domain / descriptionAdded value: +"Site hostname, e.g. 'www.saucedemo.com'. Defaults to the hostname of the page currently open." - added
Input schema / properties / key / descriptionAdded value: +"Short name for the element, e.g. 'login_button' or 'search_box' (1-128 chars)."
- Changed
selector_hint_get2 fields changed- added
Input schema / properties / domain / descriptionAdded value: +"Site hostname, e.g. 'www.saucedemo.com'. Defaults to the hostname of the page currently open." - added
Input schema / properties / key / descriptionAdded value: +"Short name for the element, e.g. 'login_button' or 'search_box' (1-128 chars)."
- Changed
selector_hint_list1 field changed- added
Input schema / properties / domain / descriptionAdded value: +"Only list hints for this hostname, e.g. 'www.saucedemo.com'. Omit to list hints for every site."
- Changed
selector_hint_save5 fields changed- added
Input schema / properties / domain / descriptionAdded value: +"Site hostname, e.g. 'www.saucedemo.com'. Defaults to the hostname of the page currently open." - added
Input schema / properties / key / descriptionAdded value: +"Short name for the element, e.g. 'login_button' or 'search_box' (1-128 chars)." - added
Input schema / properties / selector / descriptionAdded value: +"How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector." - added
Input schema / properties / selector / properties / by / descriptionAdded value: +"Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText." - added
Input schema / properties / selector / properties / value / descriptionAdded value: +"The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath."
- Changed
session_create9 fields changed- added
Input schema / properties / browser / descriptionAdded value: +"Which browser to launch: chrome (default), firefox, or edge. It must be installed on this machine." - added
Input schema / properties / browserArgs / descriptionAdded value: +"Extra command-line flags for the browser, e.g. ['--incognito']. Default none." - added
Input schema / properties / headless / descriptionAdded value: +"Run without a visible window (default false)." - added
Input schema / properties / pageLoadTimeoutMs / descriptionAdded value: +"Maximum time a page load may take before navigate fails, in milliseconds (default 30000)." - added
Input schema / properties / scriptTimeoutMs / descriptionAdded value: +"Maximum time an asynchronous script may run, in milliseconds (default 30000)." - added
Input schema / properties / sessionId / descriptionAdded value: +"Optional readable id for the session, e.g. 'buyer' or 'admin' (1-128 chars). Defaults to a random UUID." - added
Input schema / properties / windowSize / descriptionAdded value: +"Initial window size, e.g. { width: 1440, height: 900 }. Change it later with the window tool's resize action." - added
Input schema / properties / windowSize / properties / height / descriptionAdded value: +"Window height in pixels (240-4320)." - added
Input schema / properties / windowSize / properties / width / descriptionAdded value: +"Window width in pixels (320-7680)."
- Changed
session_destroy1 field changed- added
Input schema / properties / sessionId / descriptionAdded value: +"Session id, as returned by session_create or session_list."
- Changed
session_select1 field changed- added
Input schema / properties / sessionId / descriptionAdded value: +"Session id, as returned by session_create or session_list."
- Changed
start_browser8 fields changed- added
Input schema / properties / browser / descriptionAdded value: +"Which browser to launch: chrome (default), firefox, or edge. It must be installed on this machine." - added
Input schema / properties / browserArgs / descriptionAdded value: +"Extra command-line flags for the browser, e.g. ['--incognito'] or ['--lang=de']. Default none." - added
Input schema / properties / headless / descriptionAdded value: +"Run without a visible window (default false). Use true for CI or background runs, false when the user wants to watch." - added
Input schema / properties / pageLoadTimeoutMs / descriptionAdded value: +"Maximum time a page load may take before navigate fails, in milliseconds (default 30000)." - added
Input schema / properties / scriptTimeoutMs / descriptionAdded value: +"Maximum time an asynchronous script may run, in milliseconds (default 30000)." - added
Input schema / properties / windowSize / descriptionAdded value: +"Initial window size, e.g. { width: 1440, height: 900 }. Change it later with the window tool's resize action." - added
Input schema / properties / windowSize / properties / height / descriptionAdded value: +"Window height in pixels (240-4320)." - added
Input schema / properties / windowSize / properties / width / descriptionAdded value: +"Window width in pixels (320-7680)."
- Changed
take_screenshot2 fields changed- added
Input schema / properties / includeBase64 / descriptionAdded value: +"Include the PNG as base64 in the result (default true). Set false when saving to disk to keep the response small." - added
Input schema / properties / savePath / descriptionAdded value: +"Absolute path to write the PNG to, e.g. '/home/me/shots/login.png' or 'C:/shots/login.png'. Omit to not save."
- Changed
type8 fields changed- added
Input schema / properties / clearFirst / descriptionAdded value: +"Clear the field before typing (default true). Set false to append to the existing value." - added
Input schema / properties / ref / descriptionAdded value: +"Element ref from the latest capture_page result (e.g. 'e12'). Use instead of selector; refs go stale after navigation, so capture again if one is rejected." - added
Input schema / properties / selector / descriptionAdded value: +"How to find the element, e.g. { by: 'id', value: 'user-name' } or { by: 'css', value: '#login-button' }. Provide either selector or ref." - added
Input schema / properties / selector / properties / by / descriptionAdded value: +"Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText." - added
Input schema / properties / selector / properties / value / descriptionAdded value: +"The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath." - added
Input schema / properties / submit / descriptionAdded value: +"Press Enter after typing (default false), e.g. to submit a search box or login form." - added
Input schema / properties / text / descriptionAdded value: +"The text to type, e.g. 'standard_user'." - added
Input schema / properties / timeoutMs / descriptionAdded value: +"How long to wait for the field to become visible, in milliseconds (default 10000)."
- Changed
upload_file5 fields changed- added
Input schema / properties / filePath / descriptionAdded value: +"Absolute path to the file on the machine running this server, e.g. '/home/me/report.pdf' or 'C:/Users/me/report.pdf'." - added
Input schema / properties / selector / descriptionAdded value: +"How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector." - added
Input schema / properties / selector / properties / by / descriptionAdded value: +"Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText." - added
Input schema / properties / selector / properties / value / descriptionAdded value: +"The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath." - added
Input schema / properties / timeoutMs / descriptionAdded value: +"How long to wait for the file input to exist, in milliseconds (default 10000)."
- Changed
wait_for_element5 fields changed- added
Input schema / properties / selector / descriptionAdded value: +"How to find the element, e.g. { by: 'css', value: '#login-button' } or { by: 'id', value: 'user-name' }. Prefer id, name, or a short CSS selector." - added
Input schema / properties / selector / properties / by / descriptionAdded value: +"Locator strategy: css, xpath, id, name, class/className, tag/tagName, linkText, or partialLinkText." - added
Input schema / properties / selector / properties / value / descriptionAdded value: +"The locator for that strategy, e.g. '#login-button' for css, 'user-name' for id, or //button[@type=\"submit\"] for xpath." - added
Input schema / properties / timeoutMs / descriptionAdded value: +"How long to wait, in milliseconds (default 10000, max 60000)." - added
Input schema / properties / visible / descriptionAdded value: +"Also require the element to be displayed, not just present in the DOM (default false)."
- Added
wait_for_page - Removed
wait_until_visible - Changed
window5 fields changed- added
Input schema / properties / action / descriptionAdded value: +"What to do: list | switch | switch_latest | new_tab | new_window | close | resize | maximize." - changed
Input schema / properties / action / enumPrevious value: -[ - "list", - "switch", - "switch_latest", - "new_tab", - "new_window", - "close" -]New value: +[ + "list", + "switch", + "switch_latest", + "new_tab", + "new_window", + "close", + "resize", + "maximize" +] - added
Input schema / properties / handle / descriptionAdded value: +"Window handle to switch to, as returned by action list. Required for switch; ignored otherwise." - added
Input schema / properties / heightAdded value: +{ + "description": "Viewport height in CSS pixels for resize (e.g. 844 phone, 1024 tablet, 1080 desktop). Required for resize.", + "maximum": 4320, + "minimum": 240, + "type": "integer" +} - added
Input schema / properties / widthAdded value: +{ + "description": "Viewport width in CSS pixels for resize (e.g. 390 phone, 768 tablet, 1920 desktop). Required for resize.", + "maximum": 7680, + "minimum": 320, + "type": "integer" +}
14 tool updates
v0.2.0- Added
batch_execute - Added
capture_page - Changed
click2 fields changed- added
Input schema / properties / refAdded value: +{ + "minLength": 1, + "type": "string" +} - removed
Input schema / requiredRemoved value: -[ - "selector" -]
- Changed
get_text2 fields changed- added
Input schema / properties / refAdded value: +{ + "minLength": 1, + "type": "string" +} - removed
Input schema / requiredRemoved value: -[ - "selector" -]
- Changed
interact2 fields changed- added
Input schema / properties / refAdded value: +{ + "minLength": 1, + "type": "string" +} - changed
Input schema / requiredPrevious value: -[ - "action", - "selector" -]New value: +[ + "action" +]
- Added
selector_hint_delete - Added
selector_hint_get - Added
selector_hint_list - Added
selector_hint_save - Added
session_create - Added
session_destroy - Added
session_list - Added
session_select - Changed
type2 fields changed- added
Input schema / properties / refAdded value: +{ + "minLength": 1, + "type": "string" +} - changed
Input schema / requiredPrevious value: -[ - "selector", - "text" -]New value: +[ + "text" +]
29 tool updates
v0.1.2- First observed
add_cookie - First observed
alert - First observed
assert_attribute - First observed
assert_text - First observed
assert_visible - First observed
click - First observed
delete_cookie - First observed
execute_script - First observed
find_element - First observed
frame - First observed
get_attribute - First observed
get_cookies - First observed
get_current_url - First observed
get_page_source - First observed
get_text - First observed
get_title - First observed
interact - First observed
navigate - First observed
open_url - First observed
press_key - First observed
retry_click - First observed
start_browser - First observed
stop_browser - First observed
take_screenshot - First observed
type - First observed
upload_file - First observed
wait_for_element - First observed
wait_until_visible - First observed
window
TDQS
Scored across 41 tools
Almost every tool has a distinct purpose, and descriptions actively cross-reference alternatives (e.g. click vs interact vs retry_click, get_text vs assert_text vs get_attribute, wait_for_element vs assert_visible). There is mild overlap in the click/assert/wait families, but the descriptions consistently steer to the right choice, leaving little genuine ambiguity.
Names are uniformly snake_case, but the verb placement is mixed: most are verb_noun (get_text, add_cookie, take_screenshot) while others are noun_verb (session_list, selector_hint_save) and several are bare nouns (window, frame, alert, history). Still readable but not a single predictable pattern.
41 tools is well beyond the heavy range for a single server, even a broad browser-automation domain. Many assertion, wait, cookie, and session variants could be consolidated or parameterized rather than exposed as separate tools.
The surface covers the full browser-automation lifecycle: sessions, navigation, interaction, waits, assertions, cookies, windows/tabs, frames, native dialogs, file upload, JS execution, screenshots, and persisted selector hints. No obvious gaps or dead ends for the stated purpose.
Maintenance
Related MCP Connectors
Provides cloud browser automation capabilities using Stagehand and Browserbase, enabling LLMs to i…
Web search, browser automation, scraping, crawling and CAPTCHA solving for AI agents.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables browser automation using the Selenium WebDriver through MCP, supporting browser management, element location, and both basic and advanced user interactions.635 npm427MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to automate web tasks such as browsing, clicking, typing, and taking screenshots via the Model Context Protocol.1MIT
- FlicenseAqualityDmaintenanceEnables AI-powered browser automation controlled through natural language, integrating Playwright with the Model Context Protocol to perform web interactions like navigation, form filling, and screenshots.10-

Browseagent MCPofficial
AlicenseAqualityDmaintenanceEnables AI agents to control web browsers through the Model Context Protocol, supporting navigation, clicking, typing, and screenshots.124 npm1MIT