browser-control-mcp-server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@browser-control-mcp-serverOpen https://news.ycombinator.com and take a screenshot"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
browser-control-mcp-server
MCP server that gives AI agents full browser control — navigate to any website, click, type, scroll, take screenshots, inspect DOM, and read console logs.
Works with any MCP-compatible client: Claude, Cursor, Windsurf, Cline, and more.
Features
Navigate to any URL — public sites, localhost, web apps
Screenshot pages and get visual feedback as PNG images
Click buttons, links, menus, dropdowns by CSS selector
Type into inputs, search bars, textareas, password fields
Scroll in any direction by pixel amount
Read DOM — get full HTML source for structural analysis
Console logs — read JS errors, warnings, and debug output
Custom viewports — test at desktop (1920x1080), tablet (768x1024), or mobile (375x812)
Session persistence — cookies, login state, history maintained across calls
Two modes: headless Puppeteer (background) or real Chrome browser (via CDP connect)
Anti-bot bypass — spoofed user agent, no webdriver flag
Related MCP server: Cloudflare Playwright MCP
Quick Start
1. Install
npm install -g browser-control-mcp-serverThat's it. The installer automatically configures browser-control in all detected AI clients on your system — no manual setup needed.
Supported clients (auto-configured on install): Claude Desktop · Claude Code · Cursor · Windsurf · VS Code (Copilot/Cline) · Zed · Continue · OpenCode · Cody
Works on macOS, Windows, and Linux.
After install you'll see something like:
[browser-control-mcp] Configuring AI clients...
✓ Claude Desktop configured
✓ Cursor configured
— Windsurf not found (skipped)
[browser-control-mcp] Done. Restart your AI client to activate browser-control.Just restart your AI client and the browser tools are ready.
Manual Configuration (if needed)
If auto-setup didn't catch your client, add this to your MCP config manually:
{
"mcpServers": {
"browser-control": {
"command": "npx",
"args": ["browser-control-mcp-server"]
}
}
}2. Use It
Tell your AI agent:
"Open https://example.com and take a screenshot"
"Fill out the contact form on my site and submit it"
"Check my website on mobile viewport and show me how it looks"
The agent will use the browser tools automatically.
Tools
Tool | Description |
| Choose headless or Chrome extension mode |
| Check connection status and active sessions |
| Open any URL with optional viewport size |
| Capture page as PNG image |
| Click element by CSS selector |
| Type text into input fields |
| Scroll page by pixel amount |
| Get current page URL |
| Get HTML (or plain text), optionally scoped to a selector |
| Read JS console output |
| Compact accessibility-tree view, cheaper than a full DOM dump |
| Extract visible text or an attribute, without raw HTML |
| Press keys and key combinations |
| Hover over an element |
| Select from a native |
| Wait for an element, text, or timeout |
| Accept or dismiss alert/confirm/prompt dialogs |
| Upload files to a file input |
| Drag and drop between elements |
| List, open, switch, or close tabs |
| Go back to the previous page |
| Run arbitrary JavaScript in the page |
| Get, set, delete, or clear cookies |
| Get, set, or clear localStorage/sessionStorage |
| Render the current page to a PDF file |
| List captured requests or block by resource type (headless only) |
| Configure a download directory and list downloaded files |
| Cheap page-size pre-check before calling |
| Save a named sequence of tool calls, replay it later |
| Set/clear HTTP Basic/Digest Auth credentials |
| Device presets, dark mode, timezone, geolocation, permissions, network/CPU throttling |
| List iframes; target one with |
| Fill multiple form fields (inputs, selects, checkboxes) in one call |
| Reload the page, with an optional cache-bypassing hard refresh |
| Go forward (counterpart to |
| Read uncaught JS exceptions, distinct from console logs |
| Read/write the system clipboard (auto-grants permission) |
| Full detail on one element: attributes, bounding box, styles |
| Search for elements by text/role, no CSS selector needed |
| List/delete named persistent Chrome profiles |
Browser Modes
Headless Mode (Default)
Opens an invisible background browser using Puppeteer. No setup required.
Default viewport: 1024x768
Custom viewport: pass
widthandheighttobrowser_navigateAnti-bot detection bypass included
Sessions identified by
sessionId— pass to all subsequent calls
Connect Mode (Optional)
Controls your real running Chrome browser via the Chrome DevTools Protocol (CDP). No extension required.
Setup:
Launch Chrome with the remote debugging port enabled:
# macOS
/Applications/Google\ Chrome.app/Contents/MacOS/Google\ Chrome --remote-debugging-port=9222
# Linux
google-chrome --remote-debugging-port=9222
# Windows
chrome.exe --remote-debugging-port=9222Then set connect mode in your AI agent:
browser_select_mode({ mode: "connect" })The server connects to your running Chrome instance and controls it directly — no extension install needed.
Examples
Check a website
browser_navigate({ url: "https://example.com" })
browser_screenshot()Test mobile layout
browser_navigate({ url: "https://example.com", width: 375, height: 812 })
browser_screenshot()Fill and submit a form
browser_navigate({ url: "https://example.com/contact" })
browser_type({ selector: "input[name=email]", text: "user@example.com" })
browser_type({ selector: "textarea[name=message]", text: "Hello!" })
browser_click({ selector: "button[type=submit]" })
browser_screenshot()Login to a site
browser_navigate({ url: "https://example.com/login" })
browser_type({ selector: "#username", text: "myuser" })
browser_type({ selector: "#password", text: "mypass" })
browser_click({ selector: "#login-btn" })
browser_screenshot()Debug JavaScript errors
browser_navigate({ url: "https://example.com" })
browser_console_logs()Tool Parameters
browser_navigate
Parameter | Type | Required | Default | Description |
| string | Yes | — | URL to navigate to |
| number | No | 1024 | Viewport width in pixels (headless only) |
| number | No | 768 | Viewport height in pixels (headless only) |
| string | No | — | Reuse existing headless session |
| string | No | — |
|
browser_click
Parameter | Type | Required | Description |
| string | Yes | CSS selector (e.g. |
| string | No | Headless session ID |
| string | No |
|
browser_type
Parameter | Type | Required | Description |
| string | Yes | CSS selector of input element |
| string | Yes | Text to type |
| string | No | Headless session ID |
| string | No |
|
browser_scroll
Parameter | Type | Required | Default | Description |
| number | No | 0 | Horizontal scroll (positive=right) |
| number | No | 0 | Vertical scroll (positive=down) |
| string | No | — | Headless session ID |
| string | No | — |
|
browser_screenshot, browser_get_url, browser_get_dom, browser_console_logs
Parameter | Type | Required | Description |
| string | No | Headless session ID |
| string | No |
|
browser_select_mode
Parameter | Type | Required | Description |
| string | No |
|
browser_status
No parameters.
Viewport Presets
Device | Width | Height |
Mobile (iPhone) | 375 | 812 |
Mobile (Android) | 360 | 800 |
Tablet (iPad) | 768 | 1024 |
Laptop | 1366 | 768 |
Desktop | 1920 | 1080 |
4K | 3840 | 2160 |
Requirements
Node.js >= 18.0.0
Chrome/Chromium (auto-downloaded by Puppeteer for headless mode)
Chrome browser + extension (for extension mode only)
Development
git clone https://github.com/yogesh-joshi-0333/browser-control-mcp-server.git
cd browser-control-mcp-server
npm install
npm run build
npm testLicense
Available Tools
40 toolsbrowser_authBrowser HTTP Basic AuthA
Set or clear HTTP Basic/Digest Auth credentials for the current page — the kind of login prompt the browser itself shows (a native dialog), not a login form on the page. Call this BEFORE browser_navigate to a protected URL (e.g. a password-protected staging site). Use clear:true to remove credentials afterward.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| clear | No | If true, clear any previously set credentials instead of setting new ones | |
| password | No | Basic Auth password (required unless clear:true) | |
| username | No | Basic Auth username (required unless clear:true) | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses the pre-navigate ordering constraint and the native-dialog nature of the credential, but is silent on persistence scope (does it apply only to the current page or all subsequent requests?), whether credentials survive navigation, and what happens on auth failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly written sentences. The core identity (native auth dialog) is front-loaded, followed by the timing instruction and the cleanup note; no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter, no-output-schema tool, the description covers identity, timing, and cleanup well, and the fully-described schema handles the parameter details. It falls just short of complete because credential persistence scope is never clarified, but an agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents mode, clear, username, password, and sessionId, including the 'required unless clear:true' constraint. The description reinforces clear:true but adds no syntax or format detail beyond the schema, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (set/clear) and resource (HTTP Basic/Digest Auth credentials), and crucially disambiguates from sibling tools by contrasting the native browser dialog against an on-page login form (browser_form_fill/browser_type). An agent can identify this tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs when to call it ('BEFORE browser_navigate to a protected URL'), gives a concrete scenario (password-protected staging site), and states the condition for the reverse operation (clear:true to remove credentials afterward). This is a full when/when-not/ordering guide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_clickBrowser ClickA
Click any element on the page by CSS selector — buttons, links, menus, dropdowns, checkboxes, tabs, etc. Examples: "#submit-btn", "button[type=submit]", ".nav-link", "a.login". Waits for the element to be visible, clicks it, then waits for the page to stabilize (handles AJAX, animations, re-renders). Use browser_get_dom first if you need to find the right selector. To click inside an iframe, pass frameIndex from browser_frames (this always uses the simple DOM-click path, since humanClick's mouse simulation only works on the main page). To reach into shadow DOM, prefix the selector with "pierce/" instead.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| selector | Yes | CSS selector of the element to click. Prefix with "pierce/" to reach into shadow DOM. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. | |
| frameIndex | No | Target a specific iframe by index (from browser_frames) instead of the main page. | |
| humanClick | No | If true (default), simulates full human-like mouse event chain: mouseover → mouseenter → mousemove → mousedown → focus → mouseup → click. Required for Select2/jQuery dropdowns, custom widgets, and JS-heavy UIs. Set false for simple links/buttons. Ignored (always simple click) when frameIndex is set. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the visibility wait, the click, and the post-click stabilization behavior (AJAX, animations, re-renders), plus the non-obvious constraint that iframe clicks always use the simple DOM-click path because mouse simulation only works on the main page. It stops short of describing failure behavior (e.g., what happens if the selector matches nothing or the element never becomes visible).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and examples, then layered constraints (iframes, shadow DOM). Dense but each sentence carries operational information; slightly long, and the parenthetical about humanClick could be tighter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-annotation, no-output-schema tool, the description covers the mechanism, the timing behavior, and the two main selection edge cases. It is nearly complete; only error/empty-match handling and return confirmation are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented in the schema and the baseline is 3. The description restates the 'pierce/' convention and the frameIndex→simple-click interaction, which largely duplicates what the schema already says rather than adding new format or constraint meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource ('Click any element on the page by CSS selector') and enumerates the target element types (buttons, links, menus, dropdowns, checkboxes, tabs). The inline selector examples make it immediately distinguishable from siblings like browser_hover, browser_select_option, or browser_type.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: 'Use browser_get_dom first if you need to find the right selector', names browser_frames as the source of frameIndex, and states the shadow-DOM escape hatch ('pierce/' prefix). It also implies when-not via the humanClick/frameIndex tradeoff.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_clipboardBrowser ClipboardA
Read or write the system clipboard from the page context. Automatically grants the clipboard-read/clipboard-write permissions for the current origin first — you do NOT need to call browser_emulate separately for this. Use to verify a "Copy to clipboard" button worked, or to paste text a page doesn't expose an input for.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| text | No | Text to write to the clipboard (required for "write") | |
| action | Yes | Action to perform | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does disclose one genuinely non-obvious side effect: permissions for the current origin are granted automatically. However, it says nothing about what read returns, what happens on an empty clipboard or denied permission, or how extension vs headless modes differ in outcome — and with no output schema, that return-shape information is a real gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: capability first, then the permission side effect / routing note, then concrete use cases. Nothing is repeated from the schema and the most actionable constraint comes early.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with full schema coverage and no annotations or output schema, the description covers capability, the unusual permission side effect, and concrete scenarios. It is slightly short of complete because it never hints at the read result or failure behavior, which an agent would want before relying on it for verification.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both enums are documented, so the schema already explains action, text-required-for-write, mode defaulting to extension, and the headless sessionId. The description adds only the phrase "from the page context" and does not clarify mode or sessionId tradeoffs, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a precise verb pair (read/write) and the exact resource (the system clipboard accessed from the page context), which matches the action enum and leaves no ambiguity about what happens. It also names a sibling (browser_emulate) so the agent can distinguish overlapping responsibility.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use cases are given (verifying a "Copy to clipboard" button, pasting text a page does not expose an input for) plus an explicit exclusion: you do NOT need to call browser_emulate separately. That is the strongest form of routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_console_logsGet Browser Console LogsA
Read JavaScript console output from the current page — includes console.log, console.warn, console.error, and console.info messages with timestamps. Use this to: debug JavaScript errors, check for failed API calls, find runtime exceptions, see application logs, or diagnose why something is not working on the page.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the message types and that entries are timestamped, which is genuinely useful, but says nothing about whether logs are cleared on read, whether output is truncated/paginated, or whether it covers only the current page versus all frames — meaningful gaps for a diagnostics tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the resource and scope, then the use cases. Every clause earns its place with no filler or restated metadata.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does some return-value work by naming message types and timestamps, and the use-case list frames the tool's role well. It is nearly complete, missing only retention/truncation and frame-scope details that would matter for a diagnostics tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both mode and sessionId are already documented in the schema, including the enum and the headless-session behavior. The description adds no parameter-level meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) and resource (JavaScript console output from the current page) and enumerates the message types captured (log, warn, error, info with timestamps). It does not, however, differentiate itself from the very close sibling browser_page_errors, leaving the agent to guess which to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Use this to:' list gives concrete scenarios (debugging JS errors, failed API calls, runtime exceptions, app logs, diagnosing page breakage), so the when-to-use context is clear. It stops short of naming alternatives like browser_page_errors or browser_network, which an agent comparing diagnostics tools would want.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_cookiesBrowser CookiesA
Get, set, delete, or clear cookies on the current page. Use this to inspect auth/session cookies, inject a saved session to skip a login flow, or clean up before testing a fresh session.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| name | No | Cookie name (required for "set" and "delete") | |
| path | No | Cookie path (optional for "set") | |
| value | No | Cookie value (required for "set") | |
| action | Yes | Action to perform | |
| domain | No | Cookie domain (optional for "set", defaults to the current page's domain) | |
| secure | No | Mark cookie Secure (optional for "set") | |
| expires | No | Unix timestamp in seconds when the cookie expires (optional for "set") | |
| httpOnly | No | Mark cookie HttpOnly (optional for "set") | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that the tool mutates state via set/delete/clear and frames the cleanup use case, but says nothing about the scope of clear (current page vs all domains), auth/permission needs, or reversibility of destructive actions. Adequate but materially incomplete for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, with the capability statement front-loaded and the use cases following. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter tool with no annotations and no output schema, the description covers intent well but omits the return shape of get, the meaning of mode/sessionId selection, and the blast radius of clear. It is workable but leaves gaps an agent must infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 10 parameters including enums and per-action requirements. The description adds no syntax, defaults, or format detail beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb set (get, set, delete, clear) plus resource (cookies on the current page), which is precise enough for an agent to act. It does not explicitly differentiate from the sibling browser_storage, which handles a related but distinct resource, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives three concrete usage scenarios (inspect auth/session cookies, inject a saved session to skip login, clean up before fresh testing), which is strong context for when to reach for this tool. It offers no exclusions or named alternatives (e.g., browser_storage or browser_auth), so it is clear but not complete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_downloadsBrowser DownloadsA
Configure where downloads are saved for this session, and list what has landed there. Call action "configure" once with a directory before triggering a download (e.g. clicking a download link), then action "list" afterward to see the files and their sizes.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| path | No | Absolute directory path to save downloads to (required for "configure") | |
| action | Yes | Action to perform | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose that configuration is session-scoped and that configure must precede the download trigger, which is useful behavioral context. However, it says nothing about whether the directory is validated, what happens if configure is called twice, or persistence beyond the session.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loading the capability and then the workflow. Every clause carries information. Slightly dense with the parenthetical example, but no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-action tool with no annotations and no output schema, the description covers the essential workflow and even hints at the list output (files and their sizes). Minor gaps remain on error handling and mode interaction, but an agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents path, action, mode and sessionId, and the enum values are self-describing. The description reinforces that path is the directory for 'configure' but adds no format or constraint detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states specific verbs and resources: it configures the download destination for the session and lists downloaded files. It is clearly differentiated from all siblings, none of which deal with downloads. It loses a point only because the dual configure/list purpose is slightly less crisp than a single-action tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit ordered guidance: call action 'configure' once with a directory before triggering a download (e.g. clicking a download link), then action 'list' afterward. The trigger condition and the sequence are both spelled out, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_drag_dropBrowser Drag & DropA
Drag an element and drop it onto another element. Use for kanban boards, sortable lists, reordering items, or any drag-and-drop interface. Simulates a full human drag: mousedown on source → mousemove to target → mouseup on target, with proper drag events dispatched.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. | |
| sessionId | No | Puppeteer session ID. Skips mode selection. | |
| sourceSelector | Yes | CSS selector of the element to drag. | |
| targetSelector | Yes | CSS selector of the element to drop onto. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden and does well by spelling out the event sequence (mousedown on source → mousemove to target → mouseup on target) and noting that drag events are dispatched. That tells an agent this is a real synthesized gesture rather than a synthetic property set. It does not say what happens on failure (e.g., source/target not found, element not draggable) or whether a session must already be attached.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the action front-loaded, followed by use cases and the mechanism. Nothing is redundant or padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description must stand alone; it covers what the tool does, when to use it, and how the gesture is simulated. Missing pieces are failure/return behavior and any session prerequisite, which are minor but not fully covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (including mode, sessionId, and both selectors) are already documented in the schema. The description adds useful conceptual framing of source vs target but no syntax, format, or default details beyond the schema; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Drag an element and drop it onto another element') and names concrete scenarios (kanban boards, sortable lists, reordering). An agent can immediately distinguish this from browser_click, browser_hover, or browser_select_option without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear positive usage context ('Use for kanban boards, sortable lists, reordering items, or any drag-and-drop interface'), which is more than most siblings offer. However, it gives no exclusions or alternatives — nothing says when to prefer browser_click or browser_hover over a drag sequence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_emulateBrowser EmulateA
Emulate device/environment conditions on the current page — device presets, dark mode, reduced motion, timezone, locale, geolocation, permissions, network throttling, and CPU throttling. Pass only the options you need; each is applied independently and the response lists what was actually changed. Geolocation and permissions are scoped to the page's current origin, so navigate first.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| device | No | A known device name, e.g. "iPhone 15 Pro", "Pixel 7", "iPad Mini". Sets viewport, user agent, and touch support together. | |
| locale | No | Locale/language, e.g. "fr-FR" — sent as the Accept-Language header (approximation; does not change navigator.language) | |
| timezone | No | IANA timezone id, e.g. "America/New_York", "Asia/Kolkata" | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. | |
| colorScheme | No | Emulate prefers-color-scheme | |
| cpuThrottle | No | CPU slowdown factor (e.g. 4 = 4x slower). Pass 1 to clear. | |
| geolocation | No | Coordinates to report for navigator.geolocation. Automatically grants the geolocation permission for the current origin. | |
| permissions | No | Permissions to grant for the current origin, e.g. ["notifications", "midi"]. For clipboard-read/clipboard-write specifically, use browser_clipboard instead — this generic path does not reliably grant clipboard permissions in headless Chrome. | |
| reducedMotion | No | Emulate prefers-reduced-motion | |
| networkThrottle | No | Named network condition preset, "offline" to disable the network, or "none" to clear throttling |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful traits: independent per-option application, a response that reports what actually changed, and origin-scoped geolocation/permissions. It omits persistence/reset semantics (whether emulation survives navigation or how to clear it) and any mode-selection interaction, which are the main remaining behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences that front-load the capability list, then the operating rule, then the one prerequisite. No filler and no repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter tool with no output schema, the description compensates by stating the response lists what changed and by flagging the origin-scoping prerequisite. Persistence across navigation and reset behavior are the only notable omissions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all 11 parameters are documented inline, so the schema already does the heavy lifting. The description's enumeration mirrors the schema rather than adding syntax, defaults, or interaction rules beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (emulate) plus resource (device/environment conditions on the current page) with an explicit enumeration of the condition categories covered. No sibling tool in the list does emulation, so an agent can route to it unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States that only needed options should be passed, that each is applied independently, and gives a real prerequisite ("navigate first" for origin-scoped geolocation/permissions). It stops short of naming alternatives or when-not-to-use cases in the description itself (the browser_clipboard carve-out lives in the schema).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_executeExecute JavaScriptA
Execute arbitrary JavaScript code on the current page and return the result. Use this for complex interactions that other tools cannot handle: jQuery/Select2 dropdowns (e.g. jQuery("#select").val("value").trigger("change")), reading JavaScript variables, calling page functions, manipulating complex widgets, dispatching custom events, or any DOM manipulation. The code runs in the page context with full access to window, document, jQuery, etc. Return a value from your code and it will be sent back as the result. Examples: Set Select2 dropdown: 'jQuery("#project").val("123").trigger("change")' | Read a value: 'document.querySelector("#total").textContent' | Fill multiple fields: 'jQuery("#name").val("John"); jQuery("#email").val("john@example.com"); "done"' | Trigger form submit: 'document.querySelector("form").submit()' | Check if element exists: '!!document.querySelector(".success-message")'
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | JavaScript code to execute in the page context. Has full access to window, document, jQuery ($), and all page globals. The return value of the last expression is sent back as the result. For async operations, return a Promise. | |
| mode | No | Force a specific mode. Defaults to extension. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningful work: it discloses the execution context (page context, full access to window/document/jQuery), how values are returned ('return a value from your code'), and async handling ('return a Promise'). It does not cover failure modes, timeouts, or the fact that arbitrary DOM mutation can leave page state altered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and the routing rule, then examples. The five pipe-separated examples are lengthy but each maps to a distinct real interaction pattern, so they earn their space. Minor redundancy: both the description and the schema restate jQuery/window access.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, so the description correctly explains return-value semantics; no annotations, so it also covers execution context. What remains thin is the mode/sessionId machinery for headless vs. extension operation, which is only in the schema and never tied back to a usage decision.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, and the description earns above it by supplying concrete code-body examples that demonstrate what belongs in the 'code' parameter (returning values, multi-statement bodies, expressions). mode and sessionId are left entirely to the schema, which is acceptable at full coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Execute arbitrary JavaScript code on the current page and return the result') and immediately scopes it against the rest of the toolset ('complex interactions that other tools cannot handle'). An agent can distinguish it from browser_click, browser_type, or browser_form_fill without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use conditions: Select2/jQuery dropdowns, reading JS variables, calling page functions, custom events, DOM manipulation. The phrase 'that other tools cannot handle' implies a fallback ordering against siblings, but no sibling is named explicitly and no when-not-to-use exclusion is stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_extractBrowser ExtractA
Extract structured text or attributes from the page without dumping raw HTML. With no selector, returns the page's visible text content (like a reader view). With a selector, returns the text (or a given attribute) of the matching element(s). Much smaller and more directly useful than browser_get_dom when you just need the content, not the markup.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| multiple | No | If true, return an array of results for every match instead of just the first (default false). | |
| selector | No | CSS selector to scope extraction to. Omit to extract the whole page's visible text. | |
| attribute | No | Attribute to extract instead of text, e.g. "href", "src", "value". Requires selector. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the burden. It discloses that output is smaller than raw HTML and how it behaves with/without a selector, but says nothing about permissions, whether it waits for the page to be stable, frame handling, or cost. Adequate but thin for an unannotated read/extract tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences front-loaded with the core value, then the two selector cases, then the sibling comparison. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete enough for a 5-param, no-required-params extraction tool with full schema coverage and no output schema: the two selector behaviors and the markup-vs-content tradeoff are covered. Could be richer on mode/session interaction and frame context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are fully documented in schema; baseline 3. The description only elaborates on selector and attribute behavior, adding little beyond what the schema already states for mode, multiple, and sessionId.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (extract) and resources (structured text/attributes from page), and explicitly names the sibling it replaces (browser_get_dom) with the distinguishing condition (need content, not markup). An agent can pick this over browser_get_dom without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditions for both modes: no selector returns visible reader-view text, selector returns element text or attribute. Explicitly contrasts with browser_get_dom. Lacks guidance on when to prefer other extraction siblings (browser_get_element, browser_find) or the mode/session interaction, so not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_file_uploadBrowser File UploadB
Upload one or more files to a file input element. Provide the CSS selector of the element and an array of absolute file paths to upload. Works with single and multiple file inputs. Examples: browser_file_upload({selector: "input[type=file]", paths: ["/home/user/document.pdf"]}) or multiple files: browser_file_upload({selector: "#photos", paths: ["/tmp/img1.png", "/tmp/img2.png"]})
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. | |
| paths | Yes | Array of absolute file paths to upload. | |
| selector | Yes | CSS selector of the <input type="file"> element. | |
| sessionId | No | Puppeteer session ID. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the mechanics but omits all behavioral traits: whether a change event is dispatched, whether existing selected files are replaced or appended, what happens if the selector matches nothing, and any session/mode implications. For a state-mutating browser tool this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core sentence is front-loaded and the examples illustrate the single-file and multi-file cases. The second example is somewhat redundant with the first, but nothing is wasted and the structure is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a moderately simple mutation tool this covers what it does and how to call it, but with no output schema and no annotations the agent still lacks information about return value or failure behavior. It is minimally complete rather than thorough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description goes slightly further by showing concrete call syntax with a realistic selector and path values, which concretizes the expected formats for selector and paths. It says nothing about the optional mode/sessionId params, keeping it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('Upload one or more files to a file input element') and clarifies scope with 'works with single and multiple file inputs'. It is immediately distinguishable from text-oriented siblings like browser_type and browser_form_fill by virtue of targeting an <input type="file">, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the purpose, and the examples clarify the calling pattern, but there is no explicit when-to-use/when-not-to-use guidance and no pointer to an alternative tool (e.g., browser_form_fill) for related cases. Adequate but leaves routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_findBrowser FindA
Search for element(s) by visible text and/or ARIA role, without needing a CSS selector or a full browser_snapshot dump. e.g. find({text: "sign in"}) or find({role: "button", text: "submit"}). Matches are tagged with a ref, usable as '[data-mcp-ref="f1"]' in browser_click/type/hover/select_option — same mechanism as browser_snapshot, but with an "f" prefix so the two never collide.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| role | No | Filter by ARIA role, e.g. "button", "link", "textbox", "heading". At least one of text/role is required. | |
| text | No | Substring (case-insensitive by default) to match against the element's accessible name/label/text. | |
| exact | No | If true, text must match exactly rather than as a substring (default false). | |
| multiple | No | If true, return all matches instead of stopping at the first (default false). | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. | |
| maxResults | No | Cap on results when multiple:true (default 10). Ignored when multiple is false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that matches are tagged with a ref using the '[data-mcp-ref="f1"]' attribute, that this is the same mechanism as browser_snapshot, and that the 'f' prefix prevents ref collisions between the two tools. That collision-avoidance detail is genuinely non-obvious. It doesn't state whether the operation is read-only or what happens in headless vs connect modes, leaving minor gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the purpose and immediately followed by usage examples and the ref mechanism. Every clause earns its place, though the ref/collision explanation is dense and could be slightly tighter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, no-annotation, no-output-schema tool, the description covers the essential behavior an agent needs: how to call it and what the return refs mean and where to use them. It omits discussion of multiple/maxResults result-cap behavior, but that is fully documented in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds concrete invocation examples – find({text: "sign in"}) and find({role: "button", text: "submit"}) – that demonstrate how text and role combine, which the schema alone doesn't show. It does not clarify mode/sessionId interaction (that sessionId skips mode selection is only in the schema).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search/find) and resource (elements by visible text and/or ARIA role), and explicitly contrasts itself with browser_snapshot by noting no CSS selector or full DOM dump is needed. An agent can distinguish it from sibling tools like browser_get_element and browser_snapshot without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to use it (locating elements by text/role when you lack a selector) and names the downstream tools that consume the emitted refs (browser_click/type/hover/select_option), plus the sibling it replaces (browser_snapshot). It stops short of stating explicit exclusions or prerequisites, so it's clear but not fully routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_form_fillBrowser Form FillA
Fill multiple form fields in one call: { "#email": "user@example.com", "#country": "IN" }. Automatically uses the right interaction per element — types into text inputs/textareas, uses native selection, and clicks checkboxes/radios (interpreting a truthy value as "check it"). Optionally click a submit button afterward.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| fields | Yes | Map of CSS selector -> value to fill. Non-empty. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. | |
| submitSelector | No | CSS selector of a submit button to click after all fields are filled. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful behavior: it auto-picks the right interaction per element (text input/textarea typing, native <select>, checkbox/radio clicks) and interprets truthy values as 'check it'. It stops short of covering failure behavior (unmatched selectors), waiting, or whether existing values are cleared.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with a concrete example, then the per-element behavior and submit option. No filler or redundant restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description is the only source of behavioral detail, yet it never says what the tool returns or how partial failures/selector misses are reported, which matters for an automated form fill with an optional submit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, and the description adds genuine semantics on top: the worked example clarifies the selector->value map shape and the truthy-value rule for checkboxes/radios is not captured in the schema. mode and submitSelector meaning still come entirely from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Fill multiple form fields') and clarifies the distinguishing trait of batching 'in one call', which separates it from single-action siblings like browser_type and browser_select_option. The concrete example map makes the intended payload unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'in one call' framing gives clear context for when to reach for this tool (filling several fields at once) and it explains the optional follow-up submit click. However, it never explicitly names alternatives such as browser_type/browser_select_option or states when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_framesBrowser FramesA
List all frames (iframes) on the current page, each with an index, url, and name. Pass that index as frameIndex to browser_click, browser_type, or browser_get_dom to target content inside that iframe — the main document is not part of this list navigation-wise, but IS included as index 0 if it has children. Note: shadow DOM (unlike iframes) needs no special tool — prefix any selector with "pierce/" (e.g. "pierce/.my-shadow-element") in browser_click/type/get_dom/hover/select_option to reach into shadow roots.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden and does substantial work: it clarifies the main-document edge case (included as index 0 if it has children) and distinguishes iframe handling from shadow DOM handling. It doesn't cover error cases (e.g., what happens if no frames exist) or whether frames from cross-origin iframes are reachable, which slightly limits it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then routing guidance, then a shadow-DOM caveat. The shadow DOM note, while helpful, is somewhat tangential for a tool whose job is listing frames, and adds length. Still tight and each sentence conveys actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description adequately describes what's returned and how to use it. The missing piece is error/empty-state behavior and cross-origin frame reachability, which would help an agent know when this tool won't be useful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so mode and sessionId are already fully documented in the schema, and neither is mentioned in the description. The description references frameIndex — a parameter of sibling tools, not this one — which is useful routing context but doesn't add semantics to this tool's own parameters. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List all frames (iframes) on the current page') and enumerates the returned fields (index, url, name). It distinguishes itself from DOM-focused siblings like browser_get_dom by scoping explicitly to frame enumeration.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent what to do next: 'Pass that index as frameIndex to browser_click, browser_type, or browser_get_dom.' Also handles the special case that the main document is index 0 only when it has children, and routes shadow DOM work to the pierce/ prefix instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_get_domGet Browser DOMA
Get the HTML source code of the current page (or a scoped subtree). Returns the rendered DOM including dynamically loaded content. Use this to: find CSS selectors for browser_click/browser_type, understand page structure, check element attributes and classes, inspect form fields, or analyze the page content as HTML or plain text. For large pages, prefer browser_snapshot (accessibility-tree view) or scope with selector — a full-page dump can be tens of thousands of characters.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| format | No | "html" (default) returns markup. "text" strips tags and returns visible text only — much smaller for reading content. | |
| selector | No | CSS selector to scope the result to a single subtree instead of the whole page. Prefix with "pierce/" to reach into shadow DOM. | |
| maxLength | No | Truncate the returned dom string to this many characters. Omit for no limit (default). | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. | |
| frameIndex | No | Read a specific iframe by index (from browser_frames) instead of the main page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full burden, and it discloses meaningful traits: the return is a *rendered* DOM including dynamically loaded content, and a full-page dump can be tens of thousands of characters (implying cost/size implications). It does not explicitly state the operation is read-only/non-mutating, but 'Get' and the described return value make that unambiguous.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose, then a compact list of use cases, then the large-page caveat with the alternative. Every sentence earns its place and the critical routing advice is not buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description explains what is returned (rendered HTML or tag-stripped text) and the size risk, and all six parameters are documented in the schema. An agent has enough to call it correctly and to know when to fall back to browser_snapshot.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (mode, format, selector, maxLength, sessionId, frameIndex) is already documented in the schema. The description reinforces selector scoping and the smaller text format, but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Get the HTML source code of the current page (or a scoped subtree)', and clarifies it returns the rendered DOM with dynamically loaded content. It also names browser_snapshot as the sibling to use instead for large pages, letting an agent distinguish it without comparing schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly enumerates when to use it ('find CSS selectors for browser_click/browser_type, understand page structure, check element attributes...') and when not to ('For large pages, prefer browser_snapshot ... or scope with selector'). Both the condition and the alternative are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_get_elementBrowser Get ElementA
Get full detail on ONE element by selector: tag name, all attributes, bounding box, visibility, text content, and a curated set of computed CSS properties. Complements browser_extract (text/attribute only) and browser_snapshot (a whole page of elements) — use this when you need to inspect a single specific element closely, e.g. checking why something looks wrong, or confirming an element is actually visible/enabled before interacting.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| styles | No | Which computed CSS properties to include (default: display, visibility, color, background-color, font-size, font-weight). Pass your own list for anything else, e.g. ["z-index", "opacity"]. | |
| selector | Yes | CSS selector of the element to inspect. Prefix "pierce/" to reach into shadow DOM. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return content thoroughly and its read-only nature is strongly implied by 'get' and the pre-interaction inspection use case, but it never explicitly states there are no side effects, nor does it mention mode/session requirements or any rate limits. Adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and return surface, then the sibling routing and use case. Two dense sentences with no filler; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description fully compensates by enumerating return values and giving routing plus use-case context. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented in the schema, including the styles default list and the pierce/ shadow-DOM prefix. The description references the selector and 'a curated set of computed CSS properties' but adds no syntax or format detail beyond what the schema supplies, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get full detail on ONE element by selector') and enumerates the exact return surface (tag name, attributes, bounding box, visibility, text, computed CSS). It explicitly distinguishes itself from browser_extract and browser_snapshot, so an agent can route without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names both alternative siblings and the conditions that select them ('text/attribute only', 'a whole page of elements'), then states when to use this tool ('need to inspect a single specific element closely') with concrete examples. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_get_urlGet Browser URLA
Get the current URL of the browser page. Useful to verify navigation succeeded, check the current location after redirects, or confirm which page the browser is on.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. 'Get' implies a read-only, side-effect-free call and 'current URL' names the returned value, but it does not explicitly state the absence of side effects, session requirements, or behavior when no page is loaded.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, the core action front-loaded before the usage examples. No wasted phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with no output schema, the description conveys both the action and the return value (the current URL). It is nearly complete, missing only edge-case behavior such as what is returned when no page or session exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both mode and sessionId are already documented in the schema. The description adds no meaning about these parameters, which is acceptable at this coverage level but is the baseline rather than added value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: get the current URL of the browser page. An agent can distinguish this from browser_navigate or browser_get_dom, though the description never names a sibling to sharpen the contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete use contexts — verifying navigation succeeded, checking location after redirects, confirming the current page. Clear when-to-use, but no when-not-to-use or explicit alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_handle_dialogHandle Browser DialogA
Handle JavaScript alert(), confirm(), and prompt() dialogs. Call this BEFORE triggering an action that shows a dialog (e.g. before clicking a delete button that shows a confirm). The handler will auto-accept or auto-dismiss the next dialog that appears. For prompt() dialogs, provide promptText to enter text before accepting. If a dialog is already showing, it will be handled immediately.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. | |
| action | Yes | Accept or dismiss the dialog | |
| sessionId | No | Puppeteer session ID. Skips mode selection. | |
| promptText | No | Text to enter if the dialog is a prompt() |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that the handler auto-accepts/auto-dismisses the NEXT dialog (one-shot semantics), that promptText is applied before accepting, and that an already-visible dialog is handled immediately. It omits edge behavior such as timeouts or what happens if no dialog appears, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, front-loaded with the purpose before the usage timing, with no redundant or filler content. Every sentence contributes actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating browser-interaction tool with no annotations and no output schema, the description covers purpose, timing, and the prompt-text special case well. It could go slightly further on failure/timeout behavior, but nothing critical for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents action, mode, sessionId, and promptText. The description adds only a light reinforcement of promptText's condition ('if the dialog is a prompt()'), so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource pairing ('Handle JavaScript alert(), confirm(), and prompt() dialogs'), covering the full scope of dialog types and clearly distinguishing it from sibling tools like browser_click or browser_type.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to call it ('BEFORE triggering an action that shows a dialog'), gives a concrete example (clicking a delete button that shows a confirm), and explains the alternative case (calling when a dialog is already showing). This is the kind of when/when-not guidance that lets an agent route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_hoverBrowser HoverA
Hover over an element to trigger dropdown menus, tooltips, hover effects, or any mouseover-activated content. The mouse moves to the element center and stays there. Use this before browser_click when a menu only appears on hover, or to reveal hidden UI elements.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| selector | Yes | CSS selector of the element to hover over. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the mechanics ('mouse moves to the element center and stays there'), which tells the agent the hover is positional and persistent rather than a synthetic event. It does not disclose failure behavior when the selector matches nothing, whether it waits for the element, or whether hover state persists across subsequent calls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, then the mechanics, then the routing guidance. No filler and nothing that restates the tool name or title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter, no-output-schema, no-annotation tool, the description covers purpose, usage routing, and core mechanics adequately. Remaining gaps are edge behaviors (missing element, timing, persistence of hover state) rather than anything required to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and every parameter (mode, selector, sessionId) already has its own description in the schema. The description adds no parameter-level detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (hover) on a specific resource (an element) and enumerates the concrete outcomes it produces: dropdown menus, tooltips, hover effects, mouseover-activated content. It also distinguishes itself from the closely related browser_click by naming that sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit usage condition and sequencing rule: 'Use this before browser_click when a menu only appears on hover, or to reveal hidden UI elements.' That covers when to use it and the alternative it pairs with, but offers no when-not guidance (e.g., for elements that respond to click directly).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_keyboardBrowser KeyboardA
Press keyboard keys like Enter, Tab, Escape, arrow keys, or key combinations with Ctrl/Shift/Alt modifiers. Use to: submit forms (Enter), navigate between fields (Tab), close modals (Escape), select all text (Ctrl+A), copy/paste, or navigate autocomplete menus (ArrowDown/ArrowUp). If selector is provided, focuses that element first.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Key to press. Examples: "Enter", "Tab", "Escape", "ArrowDown", "ArrowUp", "ArrowLeft", "ArrowRight", "Backspace", "Delete", "Space", "a", "1", etc. | |
| mode | No | Force a specific mode. Defaults to extension. | |
| selector | No | CSS selector to focus before pressing key. If omitted, key is sent to currently focused element. | |
| modifiers | No | Modifier keys to hold. Examples: ["Control", "a"] for select-all, ["Control", "c"] for copy. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that keys go to the currently focused element unless selector is provided, which is helpful. However, it doesn't mention whether this works across modes (headless/connect/extension), what happens if the element isn't focusable, or any timing/waiting behavior after the keypress.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with zero waste: purpose, use cases, and selector behavior. Front-loaded with the core action and key examples. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description adequately covers purpose, common use cases, and selector behavior. It could be more complete by mentioning mode implications or error handling, but for a keyboard press tool it covers the essential context an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters including key, mode, selector, modifiers, and sessionId. The description mentions modifiers (Ctrl/Shift/Alt) and selector behavior, but adds little beyond what the schema provides. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (press) and resource (keyboard keys), then enumerates concrete scenarios like submitting forms, navigating fields, closing modals, and autocomplete menus. Clearly distinguishable from siblings like browser_type (text input) and browser_click (mouse actions) through its keyboard-focused scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear when-to-use guidance with examples (Enter to submit, Tab to navigate, Escape to close, Ctrl+A to select). It also clarifies behavior when selector is provided vs. omitted. However, it does not explicitly contrast with sibling browser_type or browser_form_fill, which might overlap for text entry scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_macroBrowser MacroB
Save a named sequence of tool calls once, then replay it by name later against any session — collapses a repeated multi-step flow (e.g. "login") into a single tool call instead of N round trips.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Merged into every step's args when running | |
| name | No | Macro name (required for "save", "run", "delete") | |
| steps | No | Steps to save (required for "save") | |
| action | Yes | Action to perform | |
| sessionId | No | Merged into every step's args when running, so one macro can target any session |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It explains the concept but says nothing about persistence scope (are macros stored globally or per session?), whether replay is atomic or partial on failure, or what list/delete do. For a stateful tool with no annotations, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the save-then-replay lifecycle and includes a concrete example. Efficient, though it devotes words to the save/run path while omitting the list/delete actions the schema exposes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the core save/replay mental model but is thin for a stateful, multi-action tool with no annotations and no output schema. It omits what happens when a macro step fails, where macros are stored, and how the list/delete actions behave.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are documented in the schema itself. The description reinforces only the 'name' and replay mechanics, adding little beyond the structured fields, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb and resource: 'Save a named sequence of tool calls once, then replay it by name'. The macro concept is distinct from all browser_* siblings. However, the description frames the tool only around save/run, while the schema supports list and delete actions too, leaving the full purpose slightly undersold.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an implied use case ('collapses a repeated multi-step flow') and an example ('login'), which suggests when to reach for a macro. But there is no explicit when-not guidance, no mention of the list/delete actions, and no routing to alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_networkBrowser NetworkA
Inspect or control network traffic for a headless session: list captured requests (url, method, resourceType, status), clear the log, or block requests by resource type (e.g. "image", "font", "stylesheet", "media") to speed up pages and reduce noise. Blocking is only available in headless mode — connect mode (the user's real Chrome tab) never intercepts requests, to avoid adding latency or risk to their live browsing.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| limit | No | Max entries to return for "list" (default 100), most recent last. | |
| action | Yes | Action to perform | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. | |
| resourceTypes | No | Resource types for "block", e.g. ["image", "font", "stylesheet", "media"]. Omit for "unblock" to clear all blocking. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well: it discloses the headless-vs-connect behavioral asymmetry, explains the rationale (avoiding latency/risk to live browsing), and notes that block speeds up pages and reduces noise. It stops short of describing return payload shape or error behavior when blocking is attempted in connect mode.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences that are front-loaded with the three actions before the mode caveat. The second sentence is slightly long but every clause (headless-only, connect mode never intercepts, rationale) earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter tool with no output schema, the description plus the fully documented schema cover the essentials. The main remaining gap is that return-value shape for list (pagination/limit interaction) is left entirely implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds genuine meaning: it explains what block resourceTypes values look like and clarifies the connect-mode semantics tied to the mode parameter. The one gap is that it never mentions the limit or sessionId parameters, leaving the schema to do that work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (inspect/control) and resource (network traffic) and enumerates the three capabilities: list captured requests with fields, clear the log, and block by resource type. An agent can distinguish this from sibling tools like browser_console_logs or browser_stats without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly states that blocking is only available in headless mode and that connect mode never intercepts requests, which tells the agent when the block action will work versus fail. It does not explicitly route between list/clear/block, but the action enum and stated purposes make the selection clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_page_errorsGet Browser Page ErrorsA
Read uncaught JavaScript exceptions thrown by the page (window "error" events / unhandled errors in page scripts). Different from browser_console_logs: console_logs only captures explicit console.log/warn/error calls, this captures real crashes — a script that threw and stopped running, even if it never called console.error itself. Check this whenever a page seems "stuck" or a feature silently does nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden, and it does useful work by defining what is semantically captured (window 'error' events, uncaught exceptions that stopped script execution) rather than just restating the name. It stops short of covering capture/return mechanics, session prerequisites, or whether errors are drained per call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what the tool reads, then the sibling differentiation, then the trigger. Slightly chatty with quoted 'error' and em-dash asides, but every sentence carries information and nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A zero-required-parameter read tool with full schema coverage and no output schema: the description covers what the data represents and when to reach for it. It omits return shape (e.g., an array of error entries) and scoping (current page only), which would fully close the loop.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (mode, sessionId) are documented in the schema itself, so the description adds no parameter meaning. Baseline 3 applies since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (read uncaught JS exceptions from the page) and immediately names the sibling it must not be confused with, browser_console_logs. An agent can distinguish the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger condition ('whenever a page seems stuck or a feature silently does nothing') and names the alternative tool plus the precise distinction that selects it (console calls vs. real thrown errors). No when-not case is stated, but the alternative routing is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_pdfBrowser PDF ExportA
Render the current page to a PDF file on disk (not returned inline — PDFs are too large for tool-call context). Returns the saved path and file size. Headless mode only.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| path | Yes | Absolute file path to save the PDF to. | |
| format | No | Paper size (default "A4"). | |
| landscape | No | Landscape orientation (default false). | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that the PDF is written to disk rather than returned inline, that the response is a saved path plus file size, and the headless-only constraint. It omits overwrite behavior if the path exists and any permission/render-completion details, and 'Headless mode only' sits awkwardly against a schema that offers a 'connect' mode.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with zero filler; the most non-obvious fact (not returned inline) and the return shape are both stated up front.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so describing the return as path plus file size is necessary and present. The one gap is the unresolved relationship between 'Headless mode only' and the schema's connect/extension mode options, which an agent could find confusing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents path, format, landscape, mode, and sessionId; the description adds no syntax or format detail beyond what is there. Baseline 3 applies when structured data does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('Render the current page to a PDF file on disk') and immediately distinguishes itself from the neighboring browser_screenshot by naming the output artifact (PDF file on disk vs. an inline image). An agent can select it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides one hard usage constraint ('Headless mode only') but never says when to prefer this over browser_screenshot or how it interacts with browser_select_mode. Usage context is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_profilesBrowser ProfilesA
List or delete named persistent browser profiles. A profile is a Chrome user-data directory whose cookies, localStorage, and login state survive across separate MCP server runs — unlike a normal session, which is wiped once it ends. Create one implicitly by passing profile: "name" to browser_navigate on a brand-new session (letters, digits, dash, underscore only in the name).
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Profile name (required for "delete") | |
| action | Yes | Action to perform |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does real work: it defines a profile as a Chrome user-data directory whose cookies, localStorage, and login state survive across separate MCP server runs, contrasting it with a session that is wiped. It does not state whether 'delete' is irreversible, what happens to a profile in active use, or error behavior for a missing name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action, then the definition, then the creation path. Every sentence adds distinct information an agent needs and none is redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no output schema, the description covers purpose, resource semantics, and the creation path. It leaves the shape of the 'list' result (e.g., profile names) implicit, which is a minor gap given the simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds value beyond the schema: the profile-name character constraint (letters, digits, dash, underscore) appears nowhere in the schema, and it reinforces that name is required for 'delete'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific pair of verbs (list/delete) against a clearly named resource (named persistent browser profiles) and defines the resource itself. It implicitly distinguishes itself from browser_navigate by explaining that profile creation happens there, not here.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear routing guidance: manage profiles here, but create one by passing `profile: "name"` to browser_navigate on a new session. There's no explicit statement of when NOT to use it, but the create-vs-manage split is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_reloadBrowser ReloadA
Reload the current page. Equivalent to pressing the browser's refresh button. Use ignoreCache:true for a hard refresh that bypasses the browser cache (equivalent to Ctrl+Shift+R) — useful when testing a deployed change that the page might otherwise serve stale/cached.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. | |
| ignoreCache | No | If true, bypass the cache (hard refresh). Default false (normal reload). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It explains the effect of ignoreCache=true and its real-world equivalent (Ctrl+Shift+R), which is helpful. However, it does not disclose other behavioral traits such as whether the page state is preserved, what happens to sessions, or any side effects. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loads the core action before detailing the ignoreCache option. It wastes no words, though the parenthetical examples could be slightly redundant. Every sentence contributes to understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple reload tool with no output schema and full schema coverage, the description provides enough context to invoke it correctly, especially the ignoreCache semantics. It could mention sessionId or mode briefly, but their absence is not critical given the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters. The description adds meaning for ignoreCache by clarifying the hard refresh behavior and use case, but does not mention sessionId or mode, which are documented only in the schema. The baseline is 3 when schema coverage is high.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Reload the current page.' This is unambiguous and easily distinguishable from navigation siblings like browser_navigate or browser_navigate_back. It does not explicitly differentiate from all siblings, but the purpose is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the ignoreCache option (testing a deployed change), but it does not offer explicit when-to-use vs. when-not-to-use guidance or name alternative tools for page refreshing. The implied usage is sufficient for this simple operation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_screenshotBrowser ScreenshotA
Capture a screenshot of the current page as a PNG image. Returns the actual rendered visual — you can see and describe layout, design, text, images, errors, and any visual content. Use this after browser_navigate to see a page, after browser_click to verify what happened, or anytime you need to visually inspect the browser. Pass sessionId to screenshot a specific headless session.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the return medium (a real PNG of the rendered page), what becomes visible (layout, design, text, images, errors), and that sessionId targets a specific headless session. It omits any note on capture latency, permission/attach requirements, or failure modes when no page is loaded.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action and result, then usage triggers, then the one parameter that matters. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a screenshot tool with no output schema and zero required parameters, the description covers output nature and invocation context well. A brief note on default mode behavior (extension vs headless) or empty/blank-page results would close the remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both mode and sessionId are already documented in the schema; the description's note about passing sessionId to target a headless session largely restates that. Baseline 3 is appropriate since the schema does the heavy lifting and the description adds little new about 'mode'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Capture a screenshot of the current page as a PNG image') and characterizes the output as the rendered visual, which cleanly separates it from DOM-oriented siblings such as browser_get_dom or browser_snapshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names trigger points and siblings: after browser_navigate, after browser_click, or any time visual inspection is needed. It does not state when NOT to use it (e.g., prefer browser_snapshot/browser_get_dom for structural data), so it falls just short of a full when/when-not treatment.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_scrollScroll Browser PageA
Scroll the page by pixel amount. Use positive y to scroll down, negative y to scroll up. Examples: scroll down 500px → {y: 500}, scroll up → {y: -500}, scroll to bottom → {y: 99999}, scroll right → {x: 500}. Returns the final scroll position. Useful for viewing content below the fold, triggering lazy-loaded images, or reaching elements further down the page.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | Horizontal scroll amount in pixels. | |
| y | No | Vertical scroll amount in pixels. | |
| mode | No | Force a specific mode. Defaults to extension. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose that the tool 'Returns the final scroll position' and that it can trigger lazy loading, which is real behavioral value. However, it says nothing about timing (whether it waits for content to load), whether the scroll is instant or animated, or the mode/session implications of an extension vs headless target — notable gaps for a page-mutating interaction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The critical instruction (pixel amount plus sign convention) is front-loaded, and each sentence serves a purpose — direction semantics, concrete call examples, return value, and motivating use cases. The four inline examples are slightly repetitive but cheap and instructive rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, direction semantics, return value, and usage motivation for a 4-parameter tool with a fully documented schema and no output schema. The only shortfall is the absence of guidance on mode/sessionId (extension vs headless), which the schema does describe, so the definition is essentially self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description exceeds it by defining sign conventions ('positive y to scroll down, negative y to scroll up') and giving worked examples such as {y: 99999} for bottom and {x: 500} for right. It adds no meaning for the mode or sessionId parameters, which keeps this short of a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names a specific verb ('Scroll') and resource ('the page') plus the unit of control ('by pixel amount'), which no sibling tool duplicates among the 30+ browser_* tools. An agent can distinguish this from browser_snapshot, browser_navigate, or browser_get_element without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete when-to-use conditions: viewing content below the fold, triggering lazy-loaded images, reaching elements further down the page. It stops short of any when-not guidance or explicit alternatives (e.g., using browser_wait_for or browser_find instead to reach an element), so it is clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_select_modeSelect Browser ModeA
Check available browser modes and set the session default. CALL THIS FIRST before using any other browser tool. Two modes: "connect" (connects to user's Chrome via debug port) or "headless" (invisible background Puppeteer browser, always available). Call without params to see what is available. Call with mode="headless" or mode="connect" to set the default for all subsequent calls. If only headless is available, it is auto-selected.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Set the default browser mode for this session. Omit to just query available options. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden, and it discloses meaningful behavior: connect uses the user's Chrome via a debug port, headless is an invisible Puppeteer instance that is always available, the setting persists for the session, and headless auto-selects when it is the only option. Missing error/fallback behavior when connect fails, which is the main residual gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the critical instruction ('CALL THIS FIRST'), and every sentence carries actionable content. Five sentences is slightly more than needed; the auto-select sentence could be merged with the mode definitions without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description appropriately hints at the query return ('call without params to see what is available') and covers both invocation paths and the session-scoping behavior. It stops short of describing the returned shape or how a stale session default is reset, which for a stateful session tool would be useful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes beyond the enum by explaining the operational consequences of each value and clarifying that omitting mode performs a pure query rather than a set. That meaning is not inferable from the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (check available modes and set the session default) on a specific resource, and immediately distinguishes itself from all 38 sibling browser tools by being the mandatory first call. An agent can tell exactly what this does and where it sits in the workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Strong ordering guidance: 'CALL THIS FIRST before using any other browser tool,' plus explicit instructions for the query form (omit mode) and the configuration form (mode=headless/connect), and the auto-selection fallback. It lacks any when-not-to-use condition or failure-path guidance (e.g., what to do if connect is unavailable).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_select_optionBrowser Select OptionA
Select an option from a native HTML dropdown by value or visible label text. Also works with Select2/jQuery dropdowns — automatically detects and triggers the correct change events. Examples: browser_select_option({selector: "#country", value: "US"}) or browser_select_option({selector: "#country", label: "United States"}). For complex custom dropdowns that don't use , use browser_execute with jQuery instead.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| label | No | Visible text of the option to select. | |
| value | No | Option value attribute to select. | |
| selector | Yes | CSS selector of the <select> element. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full burden. It usefully discloses that it auto-detects Select2/jQuery and triggers the correct change events, which is real behavioral value. However it says nothing about failure modes (option not found), whether the select is synchronized/focused, or what is returned, leaving notable gaps for a mutation-style UI action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose, then a compatibility caveat, then compact runnable examples, then the alternative-tool escape hatch. Every sentence carries information; nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with only one required param and full schema coverage, the description covers the core routing and usage model well and no output schema needs explaining. Minor gap: it doesn't mention preconditions such as the page/dropdown being loaded or focus requirements.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the two inline examples concretely demonstrate the value-vs-label distinction and show the selector pairing, adding practical meaning beyond the field descriptions. It does not explain mode/sessionId interaction, but those are documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (select) and resource (option in a native HTML <select>) with the two supported addressing modes (value or visible label). It also distinguishes itself from sibling browser_click/type by clarifying the dropdown-specific scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the fallback for unsupported cases: complex custom dropdowns that don't use <select> should go to browser_execute with jQuery. Also asserts compatibility with Select2/jQuery dropdowns, so the agent knows when this tool applies and when it does not.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_snapshotBrowser Accessibility SnapshotA
Get a compact accessibility-tree view of the page: interactive and semantic elements (links, buttons, inputs, headings, landmarks) as a small indented text list, each tagged with a stable ref like [ref=e12]. Far cheaper than browser_get_dom for finding what to interact with — use this first when exploring an unfamiliar page. Every ref is also a valid CSS selector — pass [data-mcp-ref="e12"] directly as the selector argument to browser_click, browser_type, browser_hover, or browser_select_option. Pass diff:true to get only what changed since your last browser_snapshot call in this session — much cheaper than a full snapshot for checking the result of a click/type.
| Name | Required | Description | Default |
|---|---|---|---|
| diff | No | If true, return only elements added/removed since the last browser_snapshot call for this session, instead of the full tree. First call always returns the full snapshot. | |
| mode | No | Force a specific mode. Defaults to extension. | |
| maxNodes | No | Maximum number of elements to include (default 300). Lower it for a shorter list, raise it for exhaustive coverage of a large page. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does most of it well: it discloses the return format (small indented text list), the stable ref mechanism, and the session-scoped diff behavior including the first-call caveat. It stops short of covering failure modes, permissions, or how mode/sessionId interact in headless vs extension contexts, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads what the tool returns, then when to prefer it, then the ref-to-selector handoff, then the diff optimization. Each sentence delivers distinct, actionable information with no filler despite the density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations exist, so the description must describe returns and behavior itself, and it does: output shape, ref format, cross-tool selector usage, and diff semantics are all covered. An agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description goes beyond the schema by explaining the cross-tool consequence of refs ('every ref is also a valid CSS selector — pass `[data-mcp-ref="e12"]` directly as the selector argument to browser_click ...'), which meaningfully informs how parameters of this and sibling tools connect.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get a compact accessibility-tree view of the page') and enumerates the element kinds returned. It explicitly distinguishes itself from the sibling browser_get_dom by calling itself 'far cheaper ... for finding what to interact with', so an agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('use this first when exploring an unfamiliar page'), names the alternative it supersedes (browser_get_dom), and describes a distinct conditional mode ('Pass diff:true ... for checking the result of a click/type'). This is precisely the routing guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_statsBrowser Page StatsA
Cheap pre-check for page size and complexity: DOM node count, approximate HTML length, and approximate visible-text length. Call this BEFORE browser_get_dom on an unfamiliar page to decide whether a full dump is safe, or whether you should scope it (selector/maxLength) or use browser_snapshot/browser_extract instead.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the 'cheap' cost profile and that values are 'approximate', which manages agent expectations about precision. Missing operational details like auth requirements or rate limits, but for a read-only stats tool the key trait (cheapness/purpose) is communicated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the tool's value proposition, then immediately followed by actionable guidance. Every clause earns its place with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete enough for a 2-param, zero-required pre-check tool with no output schema. The description tells the agent what metrics come back and how to act on them, though it doesn't specify the return format (e.g., whether stats are numbers or objects), which is a minor gap for a tool with no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents mode and sessionId fully, including the enum and the headless-session relationship. The description adds no parameter information beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (cheap pre-check for page size/complexity) and enumerates exactly what it returns (DOM node count, HTML length, visible-text length). This clearly distinguishes it from browser_get_dom, browser_snapshot, and browser_extract by naming them as alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use: 'Call this BEFORE browser_get_dom on an unfamiliar page.' Also gives the decision tree: whether a full dump is safe, whether to scope with selector/maxLength, or whether to use browser_snapshot/browser_extract instead. This is textbook routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_statusBrowser StatusA
Check browser readiness: reports whether a Chrome debug port is available (for connect mode) and lists all active Puppeteer session IDs. No parameters needed. Use this to verify the server is running and see which browser sessions are available before interacting with pages.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| connectAvailable | Yes | Whether a Chrome debug port is available for connect mode |
| headlessSessions | Yes | List of active headless session IDs |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does disclose its read-only observational behavior ('Check...reports whether...lists all active session IDs') plus the connect-mode concept behind the debug port. It stops short of explicitly affirming no state is mutated or describing refresh/staleness of the report, but the signals are clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core purpose, then the concrete outputs, then the usage context. No filler or redundant restatement of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is not the description's responsibility. For a zero-parameter diagnostic tool, the description supplies purpose, outputs, and invocation timing, leaving nothing an agent needs before calling it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the 'No parameters needed' note is confirmatory rather than additive. Baseline 4 applies for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Check) and resource (browser readiness) and precisely enumerates what it reports: Chrome debug port availability for connect mode and all active Puppeteer session IDs. An agent can distinguish it from sibling inspection tools like browser_tabs or browser_profiles without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: 'use this to verify the server is running and see which browser sessions are available before interacting with pages', establishing a pre-flight usage pattern. It does not, however, name any sibling alternative or state when not to use it, which matters in a 38-tool family.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_storageBrowser StorageA
Get, set, or clear localStorage/sessionStorage on the current page. Use this alongside browser_cookies to save and restore full session state (auth tokens, app state) and skip repeated login flows across calls.
| Name | Required | Description | Default |
|---|---|---|---|
| key | No | Storage key (required for "set") | |
| area | No | Which storage area. Required for "set". Omit for "clear" to clear both. | |
| mode | No | Force a specific mode. Defaults to extension. | |
| value | No | Storage value (required for "set") | |
| action | Yes | Action to perform | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It never discloses that 'set'/'clear' are mutating (and that 'clear' wipes data), what permissions or loaded-page state are required, or what 'get' returns for a missing key. Only the abstract rationale (skip login flows) is added.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the core action/resource statement front-loaded before the integration guidance. Nothing needs trimming.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With six parameters, no annotations, and no output schema, the description should do more: it doesn't explain what 'get' returns, what 'clear' affects by default, or how mode/sessionId interact with execution. The identity and one integration path are covered, but an agent lacks enough to call it confidently in edge cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with enums on action/area/mode, so the schema already documents key, value, area, mode, action, and sessionId. The description adds no format or syntax detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States concrete verbs (get/set/clear) against a specific resource (localStorage/sessionStorage on the current page), and explicitly distinguishes itself from browser_cookies as the companion tool for full session state. An agent can identify what it does and where it sits relative to siblings without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear use case: persist/restore auth tokens and app state to skip repeated logins, and names the co-used sibling browser_cookies. It does not state when NOT to use it (e.g., per-call ephemeral state vs. durable storage) or offer an alternative for simple cookie reads.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_tabsBrowser TabsA
Manage browser tabs — list open tabs, create new tabs, switch between tabs, close tabs, or screenshot a specific tab WITHOUT switching to it. Use this for: OAuth flows that open popups, links that open in new tabs, multi-page workflows, comparing two open tabs side by side, or cleaning up tabs. Manages pages within the Puppeteer browser session.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to open in new tab (for "new" action). Defaults to "about:blank" | |
| mode | No | Force a specific mode. | |
| index | No | Tab index for "switch" or "close" actions (0-based) | |
| action | Yes | Action to perform | |
| sessionId | No | Puppeteer session ID. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose two useful traits: screenshots occur without switching focus, and the tool operates on pages within the Puppeteer browser session. However, it omits whether 'close' is destructive/irreversible, whether tab indices shift after a close, and any session/mode preconditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, with the action list and the differentiator front-loaded before the use-case block. The use-case enumeration is slightly list-heavy but each item earns its place by mapping to a real scenario.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter, action-dispatching tool with no output schema and no annotations, the description covers the action space and usage contexts well. It falls short only on return-value expectations (what 'list' yields) and destructive semantics of 'close', which an agent would benefit from knowing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter, giving a baseline of 3. The description adds the 'screenshot without switching' nuance and confirms session-scoped tab management, but does not clarify index stability or mode/sessionId interaction beyond what the schema says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource (manage browser tabs) and enumerates the full action set — list, create, switch, close, screenshot — which maps directly onto the 'action' enum. It also differentiates itself from the browser_screenshot sibling by highlighting screenshots taken WITHOUT switching tabs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Use this for:' clause gives five concrete triggering scenarios (OAuth popups, new-tab links, multi-page workflows, side-by-side comparison, cleanup), which is strong positive guidance. It never states when NOT to use the tool or names a specific alternative sibling for e.g. plain screenshots, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_typeBrowser TypeA
Type text into any input field, textarea, search bar, or contenteditable element by CSS selector. Use this to fill out forms, enter search queries, write messages, type login credentials, etc. Examples: browser_type({selector: "input[name=email]", text: "user@example.com"}), browser_type({selector: "#search", text: "search query"}). Combine with browser_click to submit forms after filling them. To type into an element inside an iframe, pass frameIndex from browser_frames. To reach into shadow DOM, prefix the selector with "pierce/" (e.g. "pierce/input#name") — no frameIndex needed for that.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. Defaults to extension. | |
| text | Yes | Text to type into the element. | |
| selector | Yes | CSS selector of the element to type into. Prefix with "pierce/" to reach into shadow DOM. | |
| sessionId | No | Puppeteer session ID for headless mode. Skips mode selection. | |
| frameIndex | No | Target a specific iframe by index (from browser_frames) instead of the main page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses targeting behavior for iframes and shadow DOM and implies it does not submit forms by itself, but it does not explain whether typing replaces or appends existing text, whether input/change events fire, or what errors or waiting behavior to expect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and then adds practical examples, sibling interaction, and targeting edge cases. The examples and iframe/shadow DOM notes are relevant and earn their place without excessive repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter browser action with no annotations and no output schema, the description covers purpose, common use cases, examples, iframe targeting, shadow DOM targeting, and form submission handoff. It still omits mutation details such as text replacement versus append, event firing, and return/error behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds useful examples for selector and text and clarifies that shadow DOM selectors do not need frameIndex. It does not add meaning for mode or sessionId beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: type text into input fields, textareas, search bars, or contenteditable elements. It distinguishes related behavior by explaining that it fills fields and should be combined with browser_click to submit forms, and it calls out iframe and shadow DOM targeting.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear usage contexts such as filling out forms, entering search queries, writing messages, and typing login credentials. It also explains how to combine with browser_click and how to target iframes or shadow DOM, but it does not explicitly state when to prefer alternatives like browser_form_fill or browser_keyboard.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_wait_forWait For ConditionA
Wait for a condition before proceeding — an element to appear, text to become visible, or text to disappear. Use this to handle dynamic content, AJAX loading, animations, or any async page updates. Examples: wait for a success message after form submit, wait for a loading spinner to disappear, wait for a modal to open. Timeout defaults to 10 seconds.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Force a specific mode. | |
| text | No | Wait for this text to appear anywhere on the page | |
| timeout | No | Maximum wait time in milliseconds | |
| visible | No | If true, wait for element to be visible (not just in DOM) | |
| selector | No | CSS selector to wait for in the DOM | |
| textGone | No | Wait for this text to disappear from the page | |
| sessionId | No | Puppeteer session ID. Skips mode selection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the blocking/wait nature and the 10-second default timeout, but does not say what happens when a condition is never met — whether it errors, returns false, or aborts the session. That timeout-failure behavior is exactly what an agent needs to handle the tool safely.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences plus a compact example list. The core purpose is front-loaded, and each example illustrates a distinct wait mode rather than padding. The timeout note closes with a useful default.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a seven-parameter tool with no output schema and no annotations, the description covers purpose, context, and examples well. It omits timeout-failure semantics and session/return expectations, which is a meaningful but not crippling gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all seven parameters (mode, text, timeout, visible, selector, textGone, sessionId) are already documented. The description implicitly maps its three condition types to selector/text/textGone, adding marginal conceptual value but no syntax, format, or precedence details beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('wait') and resource ('condition'), then enumerates the three concrete condition types: element appearing, text becoming visible, text disappearing. An agent can distinguish this from sibling action tools like browser_click or browser_navigate without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear triggering contexts (dynamic content, AJAX loading, animations, async page updates) and three worked examples. However, it offers no explicit 'when not to use' guidance or named alternatives, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
40 tool updates
v1.5.0- First observed
browser_auth - First observed
browser_click - First observed
browser_clipboard - First observed
browser_console_logs - First observed
browser_cookies - First observed
browser_downloads - First observed
browser_drag_drop - First observed
browser_emulate - First observed
browser_execute - First observed
browser_extract - First observed
browser_file_upload - First observed
browser_find - First observed
browser_form_fill - First observed
browser_frames - First observed
browser_get_dom - First observed
browser_get_element - First observed
browser_get_url - First observed
browser_handle_dialog - First observed
browser_hover - First observed
browser_keyboard - First observed
browser_macro - First observed
browser_navigate - First observed
browser_navigate_back - First observed
browser_navigate_forward - First observed
browser_network - First observed
browser_page_errors - First observed
browser_pdf - First observed
browser_profiles - First observed
browser_reload - First observed
browser_screenshot - First observed
browser_scroll - First observed
browser_select_mode - First observed
browser_select_option - First observed
browser_snapshot - First observed
browser_stats - First observed
browser_status - First observed
browser_storage - First observed
browser_tabs - First observed
browser_type - First observed
browser_wait_for
TDQS
Scored across 40 tools
Several inspection tools (browser_get_dom, browser_snapshot, browser_extract, browser_get_element, browser_find, browser_stats) all retrieve page content with subtle differences, making misselection likely. Interaction and navigation tools are mostly distinct, but the large surface adds ambiguity. Descriptions help, but the overlap is still notable.
All 40 tools use the browser_ prefix with snake_case verb/noun naming (e.g., browser_navigate, browser_click, browser_get_dom). There are no deviations in case or structure. The pattern is highly predictable throughout.
40 tools is well above the 15-tool guideline and feels excessive for a browser control server. While the domain is broad, many capabilities could be consolidated (e.g., the inspection tools). The large count forces agents to navigate a heavy menu for common actions.
The surface covers navigation, interaction, inspection, network, storage, cookies, emulation, dialogs, uploads/downloads, tabs, frames, and more, with no obvious major gaps. A minor missing piece is an explicit session-close tool, though closing all tabs is a workaround. Overall coverage is strong.
Maintenance
Related MCP Connectors
Undetectable cloud browser sessions for AI agents and scrapers. Navigate, extract, click, captcha.
AI-powered browser automation — navigate, click, fill forms, and extract data from any website.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to control a browser through a set of tools, allowing them to perform web automation tasks like navigation, typing, clicking, and taking screenshots.-
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to control a browser through a set of tools, allowing them to perform web automation tasks like navigation, typing, clicking, and taking screenshots.-
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to control a browser through a set of tools, allowing them to perform web automation tasks like navigation, typing, clicking, and taking screenshots.-
- FlicenseNot gradedqualityCmaintenanceEnables AI assistants to control a browser for web automation tasks such as navigation, typing, clicking, and taking screenshots.-