browsermcp-plus
browsermcp-plus lets MCP clients automate a connected real browser tab through a suite of browser-control tools.
Navigate to URLs, go back/forward, reload, and manage history.
Inspect pages with accessibility snapshots and text search for element references.
Click, hover, type, fill forms, select dropdown options, drag and drop, and press keys/shortcuts.
Upload local files, handle alert/confirm/prompt dialogs, run JavaScript, scroll, and take screenshots.
Read console logs and wait for time or for text to appear/disappear.
List, open, select, and close browser tabs.
Support multiple agents in parallel, each with its own browser tab.
The browsermcp-plus Chrome extension can be loaded in Arc, letting the server automate the user's real Arc browser profile — clicking, typing, uploading files, running JavaScript, navigating and reading pages in the connected tab.
The browsermcp-plus Chrome extension can be loaded in Brave, so the server can drive the user's real Brave browser profile — clicking, typing, uploading files, running JavaScript, navigating, scrolling and reading pages in the connected tab while staying logged in.
Provides browser_evaluate to run arbitrary JavaScript in the current page, optionally scoped to an element from the accessibility snapshot, and browser_get_console_logs to read the page console.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@browsermcp-plusgo to my GitHub notifications and summarize what needs attention"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
browsermcp-plus
Let AI apps drive your own browser — including file uploads, tabs and JavaScript.
browsermcp-plus is an MCP server and an open-source Chrome extension. Claude, Cursor, VS Code, Windsurf and any other MCP client can click, type, upload files and read pages in the tab you connect, using your real browser profile: you stay logged in, nothing runs in the cloud, and sites see a normal browser.
It started as a hardened fork of Browser MCP and stays compatible with its extension. Türkçe README
What's different from Browser MCP
Browser MCP | browsermcp-plus | |
File upload | ✗ | ✓ file inputs and "choose file" buttons, no OS dialog |
| freeze the extension | reported and handled |
Tokens per click on a large page | whole page (~14k on Wikipedia) | short report (~30) |
Forms | one call per field |
|
Run JavaScript, tabs, scrolling | ✗ | ✓ |
Extension source | closed | open source ( |
Server reachable from the network | yes, all interfaces | loopback only |
Any web page can connect to the server | yes | no, origin allowlist |
Several agents at once | each new client kills the previous one | every agent in its own tab, in parallel |
Builds from its own repository, tests, CI | ✗ | ✓ unit + real-browser end-to-end tests |
Related MCP server: Chrome MCP Server
Install
1. Extension — download browsermcp-plus-extension-*.zip from the
latest release
and unzip it, then in Chrome (or Edge, Brave, Arc…):
open
chrome://extensionsand enable Developer mode,click Load unpacked and choose the unzipped folder,
if you have the original Browser MCP extension, disable it.
2. Server — add it to your MCP client. Either run it straight from GitHub (needs Node.js 18+ and git):
{
"mcpServers": {
"browser": {
"command": "npx",
"args": ["-y", "github:Novpix/browsermcp-plus"]
}
}
}or download browsermcp-plus.cjs from the release — a single file with no
dependencies — and point the client at it:
{
"mcpServers": {
"browser": {
"command": "node",
"args": ["/path/to/browsermcp-plus.cjs"]
}
}
}Claude Code: claude mcp add browser -- npx -y github:Novpix/browsermcp-plus
3. Connect — open the tab you want to automate, click the extension icon
(Alt+J) and press Connect. The icon shows ON.
Tools
Tool | Description |
| Open a URL ( |
| History and reload |
| Accessibility snapshot with element refs (optionally one subtree) |
| Find elements by text without loading the whole snapshot |
| Act on an element from the snapshot |
| Type into an element (key by key for masked inputs), optionally pressing Enter |
| Fill many fields — text, checkboxes, radios, dropdowns, sliders, dates — in one call |
| Choose dropdown values |
| Drag one element onto another (incl. HTML5 drag and drop) |
| Keys and shortcuts ( |
| Upload local files through a file input or a "choose file" button |
| Accept or dismiss |
| Run JavaScript in the page, optionally on a snapshot element |
| Scroll by pixels or scroll an element into view |
| Work across tabs |
| Wait for time, or for text to appear/disappear |
| Read the page console |
| PNG of the visible viewport |
Actions (click, type, …) answer with a short report — URL, title, whether a new
page loaded, open dialogs, new tabs — instead of the whole page, which keeps
agents fast and their context small. Pass snapshot: true when you want the
page right away. Clicks on disabled or covered elements fail with an
explanation instead of silently hitting the overlay.
Tools carry MCP annotations (readOnlyHint, destructiveHint) so clients can
skip confirmation for read-only ones. With the original Browser MCP extension
everything except upload, forms, dialogs, evaluate, scroll and tabs works.
Options
--port <number> WebSocket port (default 9009)
--session-name <name> agent name shown in the browser (default: current folder)
--allow-origin <origin...> extra extension origins allowed to connect
--no-takeover don't share the browser with a running server
--kill-existing terminate a non-cooperating process on the port
--request-timeout <ms> timeout for a single browser action (default 30000)
--snapshot-max-chars <n> truncate large snapshots (default 80000, 0 = never)
--action-snapshots include the page snapshot in every action result
--verbose debug logging on stderrSeveral agents at once
Every Claude Code session (or other MCP client) that starts browsermcp-plus is an agent with its own tab, and agents work in parallel:
The tabs you connect are shared: an agent takes a free one, or gets a new tab next to them when all are busy. Tabs opened for an agent are grouped and labelled with its name (the project folder, or
--session-name).An agent cannot select or close another agent's tab; when an agent exits, its tab is freed for the next one.
The first server is the hub the extension connects to; the others join it. If the hub exits, another server takes its place automatically.
The extension popup shows every agent, its tab, what it is doing and its last error. "Stop all" disconnects everything.
Security model
The server listens on
127.0.0.1/::1only.Only the extension origins may connect, so web pages you visit cannot talk to the server or impersonate the extension.
The extension acts only on the tab you connect; Chrome shows its "is debugging this browser" bar while it does.
Page content returned by the tools is untrusted input for the model.
Development
npm install
npm run check # typecheck + unit tests + build
npm run test:e2e # Chromium + extension end to end (single and multi-agent)
npm run test:real # background tabs in a real, visible Chromium (needs a display)The wire protocol is documented in src/protocol.ts, the
extension in extension/README.md.
License
Apache-2.0. See LICENSE and NOTICE — based on Browser MCP, which was adapted from Playwright MCP. Not affiliated with the Browser MCP authors.
Available Tools
25 toolsbrowser_clickADestructive
Click an element. Fails with an explanation if the element is disabled or covered by another element (e.g. a cookie banner).
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | Exact target element reference from the page snapshot | |
| element | Yes | Human-readable element description used to obtain permission to interact with the element | |
| snapshot | No | Also return the page snapshot after the action (default: no). Use it when you need to see the result right away. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and openWorldHint=true, so the safety profile is covered. The description adds genuinely useful behavioral context beyond that: it discloses failure conditions (disabled element, element obscured by overlays like cookie banners), which helps an agent anticipate and diagnose failures.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core action followed by the failure behavior. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-output-schema tool, the definition covers the action and its key failure modes, and the schema fully documents parameters. It could add a note on when to use the snapshot flag or how failures surface, but it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (ref, element, snapshot) are already well documented in the schema. The description adds no parameter-level detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Click an element'), which is distinguishable from siblings like browser_hover and browser_type. However, it offers no explicit differentiation from those nearby interaction tools, which an agent must infer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, prerequisites, or alternatives relative to browser_hover, browser_type, or browser_fill_form. The failure note describes an outcome, not when this tool should be selected.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_dragCDestructive
Perform drag and drop between two elements
| Name | Required | Description | Default |
|---|---|---|---|
| endRef | Yes | Exact target element reference from the page snapshot | |
| snapshot | No | Also return the page snapshot after the action (default: no). Use it when you need to see the result right away. | |
| startRef | Yes | Exact source element reference from the page snapshot | |
| endElement | Yes | Human-readable target element description used to obtain the permission to interact with the element | |
| startElement | Yes | Human-readable source element description used to obtain the permission to interact with the element |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true and readOnlyHint=false, so the safety profile is partly covered. But the description adds nothing beyond the annotations: it does not warn that dragging can trigger arbitrary page behavior (file drops, state mutations), nor mention the snapshot-ref requirement or the optional post-action snapshot. With annotations present, some gap is tolerable, but zero added context warrants a low score.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, waste-free sentence that is appropriately sized for the tool. It is front-loaded but arguably too sparse to be considered exemplary given the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, open-world interaction with four required parameters and no output schema, the description should explain that element refs come from a page snapshot and that the action can trigger arbitrary page effects. None of that is present, leaving meaningful gaps despite good schema coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented, including the required start/end refs and the optional snapshot flag. The description only restates the 'two elements' idea and adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('drag and drop') and resource ('between two elements'), which clearly separates it from sibling tools like browser_click or browser_hover. However, it does not name or distinguish itself from any specific alternative, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as browser_click or browser_hover, nor any prerequisites (e.g. needing element refs from a snapshot). The agent must infer usage entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_evaluateADestructive
Run JavaScript in the page and return the JSON-serialisable result. Pass a function: () => document.title, or (element) => element.value together with a ref. Promises are awaited.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Element passed as the function's first argument | |
| element | No | Human-readable element description used to obtain permission to interact with the element | |
| function | Yes | JavaScript function source, e.g. `() => location.href` |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag readOnlyHint=false, destructiveHint=true and openWorldHint=true, and the description adds genuine behavioral detail beyond them: the result must be JSON-serialisable and Promises are awaited. However, it omits the most important behavioral caveat for a destructive tool with page-level privileges – that arbitrary JS can mutate or break the page – and says nothing about error behavior or the permission prompt implied by the 'element' parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: purpose first, then the two calling forms with inline examples, then the async caveat. No filler, and the most decision-relevant information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the return-value burden and does so adequately ('JSON-serialisable result', awaited Promises). Given a destructive, open-world tool operating on live pages, it is still missing failure modes and a warning about side effects on page state.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds contract detail the schema lacks: the function must be a function expression, it may take the ref element as its first argument (illustrated with `(element) => element.value`), and async results are awaited. It does not explain the 'element' permission-description parameter, so it is short of a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb+resource ('Run JavaScript in the page and return the JSON-serialisable result'), which unambiguously places it as the arbitrary-JS escape hatch among a sibling set of navigation/click/snapshot tools. It never names a sibling alternative, so the differentiation is implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the unique capability ('run JS in the page'), but there is no explicit when-to-use guidance, no statement of when the structured siblings (browser_snapshot, browser_click) should be preferred, and no preconditions such as needing a loaded page or permission. The invocation examples help a caller but do not advise on tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_file_uploadADestructive
Upload local files through a file input. ref may be the file input itself or the button that opens the file chooser; no dialog is shown.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | Exact target element reference from the page snapshot | |
| paths | Yes | Absolute paths of the files to upload (`~` is expanded) | |
| element | Yes | Human-readable element description used to obtain permission to interact with the element | |
| snapshot | No | Also return the page snapshot after the action (default: no). Use it when you need to see the result right away. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=false, destructive=true, and openWorld=true, so the safety profile is covered. The description adds one genuine behavioral fact, 'no dialog is shown,' explaining that the native chooser is bypassed. It does not elaborate on the destructive nature or permission requirements, so it is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action, then the targeting nuance. No filler; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the action, element targeting flexibility, and dialog behavior, with annotations handling the safety profile and the schema documenting parameters including the optional snapshot flag. Without an output schema the description could note the return behavior, but overall it is sufficient to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented. The description adds modest extra meaning for `ref` (file input OR triggering button), but says nothing about `paths` handling beyond what the schema states. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (upload) and resource (local files) plus the mechanism (through a file input), so an agent immediately knows what the tool does. No sibling tool performs uploads, so differentiation is implicit rather than stated, keeping it short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives actionable targeting guidance: `ref` may be the file input itself or the button that opens the chooser. This resolves a common ambiguity for the caller. It stops short of naming exclusions or contrasting with browser_type/browser_fill_form, but no real alternative exists.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_fill_formADestructive
Fill several form fields in one call: text fields, checkboxes, radio buttons, native dropdowns and sliders. Faster and more reliable than separate typing and clicking.
| Name | Required | Description | Default |
|---|---|---|---|
| fields | Yes | Fields to fill, in order | |
| snapshot | No | Also return the page snapshot after the action (default: no). Use it when you need to see the result right away. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=true, and openWorldHint=true, so the agent knows this mutates external page state. The description adds a general reliability claim but does not disclose what gets overwritten, whether existing values are replaced, or validation/failure behavior — useful context that would go beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the batch capability front-loaded and no wasted words. It is appropriately sized for a tool whose schema carries most of the detail, though it could be marginally sharper by naming the alternatives it beats.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a rich nested schema, full annotation coverage, and no output schema, the description covers the essentials an agent needs to call it correctly. The only real gap is that a destructiveHint=true mutation never explains what is replaced, but the core capability is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both the fields array and the snapshot flag are fully documented, including value format for each widget type. The description's widget enumeration (text, checkbox, radio, dropdown, slider) largely restates the type enum, adding no syntax or ordering detail beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Fill") and resource ("several form fields in one call") and enumerates the supported widget types. It implicitly contrasts with the separate typing/clicking siblings (browser_type, browser_click), though it does not name them. An agent can identify the tool's scope without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Faster and more reliable than separate typing and clicking" implies a preference over browser_type/browser_click, giving implied usage context. However, there is no explicit when-to-use/when-not (e.g., single-field cases) and no named alternatives, leaving the agent to infer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_findARead-only
Search the page snapshot for text (case-insensitive) and return only the matching elements with their refs and position in the page. Much smaller than a full snapshot.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Text to look for in element names, values, text and URLs |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so the safety profile is covered. The description adds genuinely useful behavior beyond that: matching is case-insensitive, the result is a filtered subset rather than the whole page, and results carry refs and positions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and the return contract; every clause carries information and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must describe returns, and it does (matching elements with refs and position). For a one-parameter read tool with annotations covering safety, nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema coverage on a single parameter, the baseline is 3, and the description earns an extra point by disclosing case-insensitive matching, which is behavioral semantics the schema does not state.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (page snapshot) plus the return shape: matching elements with refs and position. It implicitly distinguishes itself from browser_snapshot via the 'much smaller than a full snapshot' contrast, though it never names the sibling outright.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied through the size comparison to a full snapshot; the agent can infer this is the lighter-weight lookup alternative, but there is no explicit when-to-use/when-not-to-use statement or named alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_get_console_logsBRead-only
Get the console logs from the browser
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and openWorldHint=true, so safety is covered. The description adds no extra behavioral context, such as whether logs are cumulative, cleared, or from all tabs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single concise sentence that is front-loaded with the purpose. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no parameters and no output schema, the description is minimal but adequate. However, it could mention typical use cases or limitations (e.g., returns logs from the current page) to help an agent use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so baseline 4. No parameters to document, and the description appropriately focuses on the action.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb+resource (get console logs) and specifically names the browser as the source, distinguishing it from generic log retrieval. However, it doesn't differentiate from siblings like browser_snapshot, which might also capture console output.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. With siblings like browser_snapshot that may include console logs, the description should specify its unique role (e.g., retrieving only console logs for debugging).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_go_backB
Go back to the previous page
| Name | Required | Description | Default |
|---|---|---|---|
| snapshot | No | Also return the page snapshot after the action (default: yes). Use it when you need to see the result right away. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this mutates browser state (readOnlyHint=false) and interacts with an open world. The description adds nothing behavioral — notably it omits what happens when there is no history entry to go back to, or whether the call waits for the page to load before returning.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with zero filler. Nothing in it is redundant or padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter navigation tool with annotations covering the safety profile, the description is minimally viable but under-specified: it says nothing about failure when no prior page exists or about whether the resulting page state is returned.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single optional parameter with 100% schema description coverage, so the schema already explains the snapshot flag and its default. The description adds no parameter meaning, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and effect ('Go back to the previous page'), which is unambiguous and clearly distinct from browser_go_forward or browser_navigate. It does not explicitly name those siblings, but the action itself is self-identifying.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose implies when to use it (after navigating forward), but there is no explicit guidance on when to prefer it over browser_navigate, browser_reload, or browser_go_forward, and no mention of prerequisites such as existing history.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_go_forwardC
Go forward to the next page
| Name | Required | Description | Default |
|---|---|---|---|
| snapshot | No | Also return the page snapshot after the action (default: yes). Use it when you need to see the result right away. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=true, so the safety profile is partly covered. However, the description adds nothing beyond them: it does not say what happens when there is no forward history, whether it waits for the page to load, or that it operates on browser history rather than page content.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the action front-loaded and no wasted words. It is efficient, though borderline under-specified rather than genuinely concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter navigation tool whose parameter is fully documented in the schema and whose safety profile is covered by annotations, the description is minimally viable. It still omits the failure condition (no forward entry) and the history-vs-page distinction an agent would benefit from.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one optional parameter with 100% schema description coverage, so the schema already explains the snapshot flag and its default. The description adds no syntax or behavioral detail about the parameter, which matches the baseline of 3 when the schema does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: navigating the browser history forward one entry. It is clear enough to distinguish from browser_go_back, but it does not explicitly name or contrast with its siblings, and 'next page' is slightly ambiguous versus in-page pagination.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no condition stating that forward history must exist, and no mention of the browser_go_back/browser_navigate alternatives. The agent must infer all of this from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_handle_dialogADestructive
Accept or dismiss the JavaScript dialog (alert, confirm, prompt, beforeunload) that is blocking the page
| Name | Required | Description | Default |
|---|---|---|---|
| accept | Yes | true to press OK / accept, false to press Cancel / dismiss | |
| snapshot | No | Also return the page snapshot after the action (default: no). Use it when you need to see the result right away. | |
| promptText | No | Text to enter into a prompt() dialog before accepting |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, openWorldHint=true, and readOnlyHint=false, so the safety profile is covered. The description adds useful behavioral context by stating that the dialog is blocking the page and listing dialog types, though it does not detail side effects of accepting beyond what annotations imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with zero waste. It states the action and the key condition without unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given rich annotations and full schema coverage, the description is largely complete for selecting and invoking the tool. It could optionally mention the snapshot behavior or promptText only applying to prompt dialogs, but those are adequately covered by the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents accept, snapshot, and promptText in detail. The description adds no parameter-level meaning beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (accept/dismiss) and resource (JavaScript dialog), enumerates dialog types (alert, confirm, prompt, beforeunload), and identifies the blocking condition. It clearly distinguishes itself from siblings like browser_click or browser_press_key by focusing on modal dialogs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the usage condition clearly: use when a JavaScript dialog is blocking the page. It does not name alternatives or exclusions, but no sibling tool handles dialogs, so the context is sufficiently clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_hoverC
Hover over element on page
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | Exact target element reference from the page snapshot | |
| element | Yes | Human-readable element description used to obtain permission to interact with the element | |
| snapshot | No | Also return the page snapshot after the action (default: no). Use it when you need to see the result right away. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, and openWorldHint=true, so the safety profile is covered structurally. The description adds nothing beyond the action name: it doesn't explain that hovering may trigger tooltips, dropdown menus, or other hover-only UI that the agent must then act on, which is the key behavioral nuance for this tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence, front-loaded with the verb. It is efficient, though borderline under-specified rather than genuinely tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple interaction tool with full schema coverage and clear annotations, the description is minimally adequate. It omits the one thing that matters most here: what hovering accomplishes and whether a follow-up action or snapshot is typically needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so ref, element, and snapshot are all documented in the schema itself. The description contributes no additional parameter meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Hover over element on page'), so an agent immediately knows the action. However, it offers no differentiation from siblings like browser_click or browser_find, which also target elements by ref.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to hover versus clicking, finding, or evaluating. No prerequisites or typical follow-up (e.g., hover to reveal a menu before clicking) are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_press_keyADestructive
Press a key or shortcut, e.g. Enter, Escape, Tab, ArrowDown, PageDown (scrolls), Control+a or a single character
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Name of the key to press or a character to generate, such as `ArrowLeft` or `a` | |
| snapshot | No | Also return the page snapshot after the action (default: no). Use it when you need to see the result right away. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and openWorldHint=true, so the agent knows key presses can have side effects. The description adds a small behavioral detail ('PageDown scrolls') but omits that keys like Enter can submit forms or trigger navigation, which is the most relevant risk for a destructive-flagged tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the verb first and useful examples trailing; nothing is wasted and nothing needs to be skimmed past.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no output schema and annotations covering the safety profile, the description is nearly sufficient. It could be fuller by mentioning focus/prerequisite state and the optional snapshot behavior, but it is adequate for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented and the baseline is 3. The description supplements the schema with extra examples ('Control+a', single character) but adds no syntax or format rules beyond that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Press') and resource ('a key or shortcut') and lists concrete key examples including modifier combos. It clearly separates itself from browser_type/browser_click by being a discrete keystroke action, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied through examples (keyboard shortcuts, PageDown scrolling). There is no statement of when to prefer this over browser_type, browser_scroll, or browser_click, and no note about required focus or a prior snapshot.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_reloadA
Reload the current page by navigating to its current URL
| Name | Required | Description | Default |
|---|---|---|---|
| snapshot | No | Also return the page snapshot after the action (default: yes). Use it when you need to see the result right away. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds one useful behavioral detail — that reload is implemented as a navigation to the current URL, not a cached refresh — but says nothing about side effects like lost form state or pending requests.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short sentence with the action front-loaded and no filler. It is appropriately sized, though it is minimal enough that it could have carried a bit more routing guidance without bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter reload tool with full schema coverage and annotations present, the description gives an agent enough to invoke it correctly. No output schema exists, but a reload's result is unsurprising, so nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single optional 'snapshot' parameter is fully documented in the schema, including a default. The description contributes nothing about the parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Reload the current page') and clarifies the mechanism ('by navigating to its current URL'), which separates it from browser_navigate in practice. It does not explicitly name the sibling it is not, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage context is implied by 'current page' — you reload rather than navigate elsewhere — but there is no explicit when-to-use, when-not, or reference to alternatives like browser_navigate or browser_snapshot. Adequate but leaves routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_screenshotARead-only
Take a screenshot of the visible part of the current page
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description contributes the viewport-scoping detail ("visible part"), which is useful, but says nothing about the return format (image payload type), whether scrolling is needed to capture more, or whether it works during navigation/loading states.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler. The scoping qualifier arrives before the object it modifies, so the key constraint is read first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only capture tool with annotations covering safety, the description is nearly sufficient. The only real gap is that no output schema exists, so the description could have stated what is returned (an image of the viewport) rather than leaving the return shape implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. The description correctly implies no configuration is needed, and there is no parameter surface left undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ("Take a screenshot") and adds a meaningful scope qualifier ("visible part of the current page"), which rules out full-page capture. It does not, however, distinguish itself from the sibling browser_snapshot, which an agent could easily confuse with a visual capture tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the tool to reach for when it needs a visual rendering rather than page structure. There is no explicit when-to-use statement, no mention of browser_snapshot as the text/DOM alternative, and no note about prerequisites such as a page being loaded or navigation state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_scrollA
Scroll the page with the mouse wheel by a number of pixels, or scroll an element from the snapshot into view
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Element to scroll into view instead | |
| deltaX | No | Horizontal pixels; positive scrolls right | |
| deltaY | No | Vertical pixels; positive scrolls down | |
| element | No | Human-readable element description used to obtain permission to interact with the element | |
| snapshot | No | Also return the page snapshot after the action (default: no). Use it when you need to see the result right away. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds the useful behavioral fact that there are two distinct modes (viewport pixel scroll vs. bringing a snapshot element into view), but says nothing about the permission flow for the element parameter, whether ref and deltaX/deltaY are mutually exclusive, or what happens on an empty call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-formed sentence with zero filler, front-loading the primary action and immediately covering the secondary mode. Nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with no required parameters, no output schema, and existing annotations, the description covers the two modes adequately but leaves gaps: it does not explain the default behavior when called with no arguments, how ref and delta values interact, or that ref/element derive from a prior snapshot. Complete enough to call, but not fully specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema, including direction conventions and the snapshot default. The description adds no syntax or interaction detail beyond that, and notably omits that ref and deltaX/deltaY represent alternative modes. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (scroll) and resource (the page / an element from the snapshot) and clearly distinguishes its two operating modes: pixel-based mouse-wheel scrolling and scroll-into-view. It does not explicitly differentiate itself from siblings like browser_click or browser_press_key, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use/when-not guidance or mention of alternatives; usage is only implied (scroll when content is off-screen). The snapshot parameter's schema text carries a small usage hint ('Use it when you need to see the result right away'), but the top-level description offers no routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_select_optionCDestructive
Select an option in a dropdown
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | Exact target element reference from the page snapshot | |
| values | Yes | Array of values to select in the dropdown. This can be a single value or multiple values. | |
| element | Yes | Human-readable element description used to obtain permission to interact with the element | |
| snapshot | No | Also return the page snapshot after the action (default: no). Use it when you need to see the result right away. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=true, and openWorldHint=true, so the agent knows this mutates page state. The description adds nothing on top of that: it does not say the page state changes, whether prior selections are replaced or appended, or what the element/reference permission flow entails.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with zero waste and the action front-loaded. It is efficient, though the brevity is under-specification rather than deliberate conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter, three-required mutation tool with no output schema, the description omits key context: how values map to option values vs. visible labels, whether multiple values require a multi-select, and how the element string is used for permission. Annotations cover the safety profile but not the invocation details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (ref, values, element, snapshot) are already documented in the schema; baseline is 3. The description's only contribution is the word "dropdown", which loosely contextualizes "values" but adds no format or matching semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ("Select an option in a dropdown"), so an agent knows the action is a selection rather than a click or a free-text entry. It does not, however, distinguish this from close siblings like browser_click, browser_type, or browser_fill_form, which could also appear to interact with form controls.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus browser_click, browser_type, or browser_fill_form, and no prerequisites or exclusions. The only implied guidance is the word "dropdown", which the agent must infer means an HTML select-style control.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_snapshotARead-only
Capture accessibility snapshot of the current page. Use this for getting references to elements to interact with. Pass a ref to see only that part of the page.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | Only return the subtree of this element |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so the safety profile is covered. The description adds genuinely useful behavioral context: the tool's role as the producer of element refs that other tools consume, and the scoping effect of passing a ref.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero filler, with the primary purpose front-loaded and the parameter guidance last. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with no output schema, the description adequately covers purpose, usage, and the ref-scoping behavior. It could be slightly more complete by noting the ref format/prefix convention, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the ref property already documents "Only return the subtree of this element"; the description's "Pass a ref to see only that part of the page" largely restates it. Baseline 3 is appropriate when the schema carries the parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Capture accessibility snapshot of the current page") and uses vocabulary that distinguishes it from the screenshot sibling. It does not explicitly name browser_screenshot to rule it out, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use this for getting references to elements to interact with" gives concrete when-to-use context that links it to the click/type/hover siblings. There is no explicit when-not or named alternative, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_tab_closeADestructive
Close a tab (default: the connected one). Closing the connected tab connects the tab that becomes active.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No | Tab id from browser_tab_list |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose destructiveHint=true and openWorldHint=true, so the safety profile is covered. The description adds genuinely new behavioral context beyond the annotations: closing the connected tab re-connects to whichever tab becomes active, which matters for subsequent tool calls. It does not, however, mention irreversibility or what happens when no tab remains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the default and the side effect front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, non-read-only tool with a clear default and annotations covering destructiveness, the description is nearly complete. It stops short of covering edge cases such as closing the last remaining tab or the result of the operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single tabId parameter is documented as coming from browser_tab_list, so the schema carries the load. The description only clarifies the no-argument default, which adds marginal value over the schema baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Close) and resource (a tab), and immediately clarifies the default target ('the connected one'), which distinguishes it from browser_tab_new and browser_tab_select. An agent can identify the operation without consulting the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The default behavior implies when the tool is useful (closing the currently connected tab vs. a specific tabId), but there is no explicit when-to-use/when-not guidance or naming of alternatives such as browser_tab_select. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_tab_listARead-only
List open browser tabs. Your tab is marked connected; tabs used by other agents are marked and cannot be selected.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, openWorldHint). The description adds real behavioral context beyond them: which tab is marked connected and that other agents' tabs are unselectable, which is essential for interpreting the result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, lead statement first, constraint second; every clause carries information and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only listing with no output schema, the description covers purpose and the most important result semantics (selection eligibility). It could say slightly more about ordering or whether IDs are returned, but it is sufficient to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies and nothing is misleading.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List open browser tabs') and adds what the result contains (connected marker, other-agent markers). It is clearly distinguishable from browser_tab_select/new/close in intent, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The note that tabs used by other agents 'cannot be selected' implies you should call this before attempting browser_tab_select, but the when-to-use condition is left to inference and no alternative (e.g. snapshot) is mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_tab_newA
Open a new tab, optionally at a URL, and make it the connected tab
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to open (default about:blank) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover safety (readOnlyHint=false, openWorldHint=true, destructiveHint=false), and the description adds a genuinely useful behavioral fact beyond them: the new tab becomes 'the connected tab,' i.e. it changes the active/connected session context. It omits any note about existing tab state or popup/blocker behavior, hence not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that carries the action, the optional parameter, and the state-changing consequence with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tab-creation tool with full schema coverage and annotations carrying the safety profile, the description is essentially complete; only a pointer to the navigate/tab_select alternatives is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents the URL parameter and its about:blank default. The description merely restates that the URL is optional, adding no format or syntax detail beyond the schema. Baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Open a new tab') and adds a meaningful scope detail (optional URL, becomes the connected tab). It distinguishes the create-and-activate behavior from browser_tab_select, though it doesn't explicitly contrast with browser_navigate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus browser_navigate (navigate the current tab) or browser_tab_select/tab_list. The word 'optionally' hints at the parameter's optionality but gives no selector conditions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_tab_selectA
Work in another tab (ids from browser_tab_list). Tabs used by other agents cannot be selected.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | Yes | Tab id from browser_tab_list |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (not read-only, open-world, not destructive). The description adds a non-obvious behavioral constraint beyond that: tabs occupied by other agents are unavailable, which is concurrency context an agent cannot get from the schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero padding, with the id source and the availability restriction front-loaded. Nothing could be removed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tab-switching tool with annotations covering safety and no output schema, the description covers source of ids and the key restriction. It could note what state the agent sees after the switch, but the essentials are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single tabId is already documented as coming from browser_tab_list. The description repeats the same provenance without adding format or validation detail, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the resource (another tab) and the intent (work in it), with the id source named. The verb 'work in' is softer than the tool name's 'select', so it's clear but not maximally precise about what switching focus actually entails.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly points to browser_tab_list as the source of valid ids and states a real exclusion: tabs held by other agents cannot be selected. It does not distinguish this from browser_tab_new or say when a switch is preferable, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_typeADestructive
Type text into an editable element, replacing its content. To fill several fields at once, prefer browser_fill_form.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | Exact target element reference from the page snapshot | |
| text | Yes | Text to type into the element | |
| slowly | No | Type one key at a time, for inputs that format as you type (phone, card or date masks). Slower. | |
| submit | No | Whether to submit entered text (press Enter after) | |
| element | Yes | Human-readable element description used to obtain permission to interact with the element | |
| snapshot | No | Also return the page snapshot after the action (default: no). Use it when you need to see the result right away. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, and the description adds concrete context by stating it replaces existing element content, telling the agent what gets destroyed. It stops short of documenting auth, permissions, or limits, but that is a minor gap given annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, with the core action front-loaded and the sibling routing placed second. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-field typing tool with annotations covering safety and a fully documented schema, the definition is nearly complete; the only omission is any note about return/snapshot behavior, which the schema's 'snapshot' param partially covers.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the six parameters (ref, text, slowly, submit, element, snapshot) are already fully documented in the schema. The description adds no syntax or format detail beyond this, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (type) and resource (editable element) and clarifies the effect on existing content ('replacing its content'). It also names the sibling browser_fill_form, so an agent can distinguish the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to browser_fill_form when several fields must be filled at once, implying this tool is for single-field entry. It gives a clear positive condition and alternative but does not spell out exclusions or prerequisites beyond that.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_waitBRead-only
Wait for a specified time in seconds
| Name | Required | Description | Default |
|---|---|---|---|
| time | Yes | The time to wait in seconds (max 300) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds nothing beyond that, not saying whether the call blocks the browser session, whether it can be interrupted, or what happens at the boundary values.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with zero waste and the essential detail (seconds) front-loaded. Nothing could be trimmed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, fully documented, annotated tool this is minimally adequate. The notable missing piece is differentiation from browser_wait_for, which shares a nearly identical name and sits in the same 22-tool browser family.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter documents units, minimum 0, and maximum 300, so the schema carries the full burden. The description's "in seconds" merely restates the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Wait for a specified time in seconds" gives a clear verb and resource, and the word "time" distinguishes it conceptually from browser_wait_for's condition-based wait. However, it never names or contrasts with that sibling, so the agent must infer the difference from the tool name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no mention of when a fixed delay is preferable to browser_wait_for's condition polling, and no note about blocking behavior during the wait. The agent gets no routing help between the two wait tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_wait_forARead-only
Wait until text appears on (or disappears from) the page's accessibility snapshot, then return the snapshot
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | Text to wait for | |
| timeout | No | Maximum time to wait in seconds (default 10) | |
| textGone | No | Text to wait to disappear |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description usefully adds that the wait is evaluated against the accessibility snapshot (not the raw DOM) and that the snapshot is returned, but it omits timeout-failure behavior (error vs. silent return) and poll semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the trigger condition and ends with the return value. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly carries the return-value information ('then return the snapshot'). What remains thin is the timeout path — the agent does not learn what happens when the timeout elapses or how to combine text and textGone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so text, textGone and timeout are already fully documented with defaults and bounds; the baseline of 3 applies. The description only restates the appear/disappear duality and adds no syntax, precedence, or interaction detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: waiting on text presence/absence in the page's accessibility snapshot, and notes it returns the snapshot. It does not, however, distinguish itself from the sibling browser_wait (a fixed-duration wait), so an agent cannot fully disambiguate the two from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied — use it when you need to block until specific text appears or disappears — but there is no explicit when-to-use, when-not-to-use, or comparison against the sibling browser_wait that most obviously overlaps with this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
16 tool updates
v0.4.0- Changed
browser_click1 field changed- added
Input schema / properties / snapshotAdded value: +{ + "description": "Also return the page snapshot after the action (default: no). Use it when you need to see the result right away.", + "type": "boolean" +}
- Changed
browser_drag1 field changed- added
Input schema / properties / snapshotAdded value: +{ + "description": "Also return the page snapshot after the action (default: no). Use it when you need to see the result right away.", + "type": "boolean" +}
- Changed
browser_file_upload1 field changed- added
Input schema / properties / snapshotAdded value: +{ + "description": "Also return the page snapshot after the action (default: no). Use it when you need to see the result right away.", + "type": "boolean" +}
- Added
browser_fill_form - Added
browser_find - Changed
browser_go_back1 field changed- added
Input schema / properties / snapshotAdded value: +{ + "description": "Also return the page snapshot after the action (default: yes). Use it when you need to see the result right away.", + "type": "boolean" +}
- Changed
browser_go_forward1 field changed- added
Input schema / properties / snapshotAdded value: +{ + "description": "Also return the page snapshot after the action (default: yes). Use it when you need to see the result right away.", + "type": "boolean" +}
- Added
browser_handle_dialog - Changed
browser_hover1 field changed- added
Input schema / properties / snapshotAdded value: +{ + "description": "Also return the page snapshot after the action (default: no). Use it when you need to see the result right away.", + "type": "boolean" +}
- Changed
browser_navigate1 field changed- added
Input schema / properties / snapshotAdded value: +{ + "description": "Also return the page snapshot after the action (default: yes). Use it when you need to see the result right away.", + "type": "boolean" +}
- Changed
browser_press_key1 field changed- added
Input schema / properties / snapshotAdded value: +{ + "description": "Also return the page snapshot after the action (default: no). Use it when you need to see the result right away.", + "type": "boolean" +}
- Changed
browser_reload1 field changed- added
Input schema / properties / snapshotAdded value: +{ + "description": "Also return the page snapshot after the action (default: yes). Use it when you need to see the result right away.", + "type": "boolean" +}
- Changed
browser_scroll1 field changed- added
Input schema / properties / snapshotAdded value: +{ + "description": "Also return the page snapshot after the action (default: no). Use it when you need to see the result right away.", + "type": "boolean" +}
- Changed
browser_select_option1 field changed- added
Input schema / properties / snapshotAdded value: +{ + "description": "Also return the page snapshot after the action (default: no). Use it when you need to see the result right away.", + "type": "boolean" +}
- Changed
browser_snapshot1 field changed- added
Input schema / properties / refAdded value: +{ + "description": "Only return the subtree of this element", + "minLength": 1, + "type": "string" +}
- Changed
browser_type2 fields changed- added
Input schema / properties / slowlyAdded value: +{ + "description": "Type one key at a time, for inputs that format as you type (phone, card or date masks). Slower.", + "type": "boolean" +} - added
Input schema / properties / snapshotAdded value: +{ + "description": "Also return the page snapshot after the action (default: no). Use it when you need to see the result right away.", + "type": "boolean" +}
22 tool updates
v0.3.0- First observed
browser_click - First observed
browser_drag - First observed
browser_evaluate - First observed
browser_file_upload - First observed
browser_get_console_logs - First observed
browser_go_back - First observed
browser_go_forward - First observed
browser_hover - First observed
browser_navigate - First observed
browser_press_key - First observed
browser_reload - First observed
browser_screenshot - First observed
browser_scroll - First observed
browser_select_option - First observed
browser_snapshot - First observed
browser_tab_close - First observed
browser_tab_list - First observed
browser_tab_new - First observed
browser_tab_select - First observed
browser_type - First observed
browser_wait - First observed
browser_wait_for
TDQS
Scored across 25 tools
Most tools map to clearly distinct browser actions (navigation, interaction, tabs, waiting, inspection) with little overlap. Pairs like browser_type vs browser_fill_form and browser_wait vs browser_wait_for are differentiated by their descriptions, but the high tool count creates minor potential for confusion.
All 25 tools use a consistent browser_ prefix and snake_case, following a predictable verb or noun pattern (e.g., browser_navigate, browser_tab_close, browser_get_console_logs). No mixed conventions are present.
25 tools is on the heavy side for a browser automation server. While each tool has a purpose, several could be consolidated (e.g., wait/wait_for, type/fill_form), placing the count in the borderline-heavy range per the rubric.
The surface covers core browser lifecycle actions: navigation, tab management, interaction, waiting, screenshots, console logs, JS evaluation, and dialog handling. Minor gaps like direct URL/title access or network/cookie management exist but are workaroundable via browser_evaluate.
Maintenance
Related MCP Connectors
Browser MCP for logged-in tasks. Uses your Chrome — credentials stay local. Zero-token replay.
Hosted real Google Chrome MCP with per-user persistent state. Navigate, click, type, screenshot.
Live browser debugging for AI assistants — DOM, console, network via MCP.
Hosted MCP server connecting claude.ai, ChatGPT and other AI apps to your own computer
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAn MCP server that provides AI assistants with full control over a real browser session via a Chrome extension, supporting 36 tools for navigation, data extraction, and DOM manipulation. It bypasses bot detection by utilizing the user's active browser session, including cookies, authentication tokens, and installed extensions.14 npm3MIT
- AlicenseNot gradedqualityDmaintenanceAn extension-based MCP server that enables AI assistants to control your browser, leveraging existing sessions and login states for automation and content analysis. It provides over 20 tools for semantic tab search, interactive element manipulation, and network monitoring directly within your daily Chrome environment.MIT
- AlicenseNot gradedqualityNot gradedmaintenanceAn extension-based MCP server that enables AI assistants to control your existing Chrome browser, leveraging your active login states and settings for automation. It provides over 20 tools for tasks like semantic tab search, screen capture, network monitoring, and direct element interaction.-
- AlicenseBqualityCmaintenanceAn MCP server that provides AI models with full browser automation capabilities through Chrome. It enables navigation, interaction, screenshots, and complete DevTools access by bridging AI clients with a companion Chrome extension.9912 npm3Apache 2.0