devtools-fleet-mcp
Manages a fleet of per-agent Chrome browsers with isolated profiles, saved logins, origin allowlists, and lifecycle management, while exposing Google's full DevTools toolset through the wrapped chrome-devtools-mcp.
Provides Lighthouse audit capabilities on web pages through the wrapped chrome-devtools-mcp toolset, enabling performance, accessibility, and SEO analysis.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@devtools-fleet-mcpStart a browser from my staging-admin login and check localhost:3000 for console errors"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Give every one of your AI agents its own Chrome, running Google's full DevTools toolset, already logged in, and fenced to the sites it should touch.
Why This Exists
Google's chrome-devtools-mcp is the best browser toolset an agent can have when the job is debugging a web app: console, network, performance traces, Lighthouse, heap snapshots. It stops working the moment you run more than one agent:
Agent two starts →
The browser is already running for …, because every server shares one Chrome profileSwitch on
--isolated→ every agent starts logged out and fights the wp-admin, staging or SSO login again, every timeThe MCP connection drops → the browser, its tabs and its login die with it
Ten agents are running → ten invisible headless Chromes. Nothing lists them, nothing cleans up after a crash
Other multi-agent browser tools solve pooling and logins, but on Playwright or through a browser extension. None of them carries Google's DevTools toolset.
Related MCP server: Chrome DevTools MCP
The Solution
devtools-fleet-mcp sits in front of chrome-devtools-mcp and manages the browsers. It doesn't reimplement a single browser tool: Google's tools are passed through unchanged, under the same names.
agent session A ─stdio─▶ devtools-fleet-mcp ─▶ chrome-devtools-mcp --browserUrl ─CDP─▶ Chrome A (own profile)
agent session B ─stdio─▶ devtools-fleet-mcp ─▶ chrome-devtools-mcp --browserUrl ─CDP─▶ Chrome B (own profile)
│
└──▶ ~/.devtools-fleet/ ◀── devtools-fleet CLI: ls, show, kill, gc, login, statesOne browser per agent, zero config. The first browser tool call launches that session's own Chrome. Up to 10 at once by default.
Browsers outlive the connection. An agent that reconnects inside the same session gets its browser back, tabs open. A small background reaper closes the browsers nobody comes back for.
Saved logins. You log in once, by hand, in a real window (2FA, SSO and captchas are fine). Any number of agents can then start from that login in parallel, each in its own browser.
Origin allowlists. A browser started from a saved login can only reach that login's origins. Typed URLs, link clicks, redirects and popups elsewhere are blocked. Strict mode locks the whole network in Chrome itself.
Recovery. If Chrome or chrome-devtools-mcp dies, the next tool call brings it back and tells the agent what happened.
Watch any agent live.
devtools-fleet show <id>opens Chrome's DevTools inspector on that agent's page without disturbing it.
Quick Start
Requires Node.js 22.12+ and Google Chrome (or point chromePath at another Chromium build).
Install as a Claude Code plugin (recommended)
The plugin installs the MCP server and a skill that teaches agents the fleet workflow: check state_list, start from a saved login, respect the allowlist, recover after a crash.
Inside Claude Code:
/plugin marketplace add pluginslab/devtools-fleet-mcp
/plugin install devtools-fleet@pluginslab-devtools-fleetOr from your terminal:
claude plugin marketplace add pluginslab/devtools-fleet-mcp
claude plugin install devtools-fleet@pluginslab-devtools-fleetRestart Claude Code and check that devtools-fleet shows up under /mcp.
Already using chrome-devtools-mcp? Remove it (claude mcp remove chrome-devtools, or disable its plugin). devtools-fleet exposes the same tools under the same names, and two servers with identical tool names confuse agents.
Updating: /plugin marketplace update pluginslab-devtools-fleet, then restart. The server runs through npx, so it picks up new npm releases on its own.
Install the CLI
The plugin gives agents the MCP tools. Saving logins, watching agents and cleaning up are done by you, from the CLI:
npm install -g devtools-fleet-mcp
devtools-fleet doctorNo global install? Every command also works as npx -p devtools-fleet-mcp devtools-fleet <command>.
Other MCP clients
Claude Code without the plugin:
claude mcp add devtools-fleet -- npx -y devtools-fleet-mcpAny other client (.mcp.json, Cursor, Claude Desktop, …):
{
"mcpServers": {
"devtools-fleet": {
"command": "npx",
"args": ["-y", "devtools-fleet-mcp"]
}
}
}First test
Open two Claude Code sessions and ask both:
"Open http://localhost:3000 and check the console for errors."
Each gets its own Chrome; neither hits the profile lock. Then, in a terminal:
devtools-fleet lsFirst saved login
devtools-fleet login staging-admin https://staging.example.com/wp-login.phpLog in in the window that opens, come back to the terminal, press Enter, confirm the allowlist. Then ask an agent:
"Using the staging-admin state, open the Plugins page and run a Lighthouse audit on it."
The agent calls state_list, then browser_start({ state: "staging-admin" }), and starts already logged in.
MCP Tools
devtools-fleet adds six tools. Everything else (navigate_page, take_snapshot, click, evaluate_script, list_network_requests, performance_start_trace, lighthouse_audit, …) is chrome-devtools-mcp's, with the same names and arguments, so its documentation and skills apply as-is.
browser_start
Starts this session's browser, or picks it up again after a reconnect. Optional: any browser tool starts one with defaults.
→ browser_start({ state: "staging-admin", url: "https://staging.example.com/wp-admin/" })
← Started: browser 3f9c2a, state "staging-admin", headless, stable
allowed origins: https://staging.example.com
tabs: https://staging.example.com/wp-admin/
uptime 2sParameter | Type | Description |
| string | Saved login to start from |
| boolean | Default from config ( |
| string | e.g. |
| string |
|
| string | Open this in the first tab |
| string[] | Without a state: restrict this browser voluntarily |
browser_status
This session's browser: id, state, allowlist, mode, open tabs, uptime, blocked requests.
→ browser_status()
← browser 3f9c2a, state "staging-admin", headless, stable
allowed origins: https://staging.example.com
tabs: https://staging.example.com/wp-admin/plugins.php
uptime 312s, 2 request(s) blocked so farbrowser_restart
Relaunches keeping cookies, storage and tabs. { headless: false } gives a visible window, for example so a person can solve a captcha.
→ browser_restart({ headless: false })
← Restarted with a visible window; reopened 2 tab(s). Page ids have changed: call list_pages.browser_stop
Closes the browser and deletes its temporary profile.
→ browser_stop()
← Closed browser 3f9c2a.state_save
Saves the current browser's login under a name. Parameters: name, allowedOrigins (defaults to the open tabs' origins), strict, overwrite. An agent can only narrow its own allowlist, never widen it, and never overwrites a state a person created.
→ state_save({ name: "local-shop" })
← Saved state "local-shop": 4 cookie(s), localStorage for 1 origin(s). Allowed: http://localhost:3000.state_list
Saved states with origins, cookie counts and expiry. Never values.
→ state_list()
← staging-admin: https://staging.example.com | 6 cookies (0 expired) | saved 2026-10-01T09:12:44.512Z by cli
local-shop: http://localhost:3000 | 4 cookies (0 expired) | saved 2026-10-01T10:03:10.087Z by agentCLI Reference
# See every fleet browser: id, status, session, state, mode, current tab
devtools-fleet ls
devtools-fleet ls --json
# Watch an agent's browser live in Chrome's DevTools inspector (doesn't disturb it)
devtools-fleet show 3f9c2a
devtools-fleet show 3f9c2a --tab 2 --print
# Close browsers
devtools-fleet kill 3f9c2a
devtools-fleet kill --all
# Close browsers whose session is gone; --detached also closes ones waiting for a reconnect
devtools-fleet gc
devtools-fleet gc --detached
# Log in by hand and save a state
devtools-fleet login staging-admin https://staging.example.com/wp-login.php
devtools-fleet login shop http://localhost:3000/login --allow https://cdn.example.com --strict
# Manage states (never prints cookie or storage values)
devtools-fleet states
devtools-fleet state show staging-admin
devtools-fleet state rm staging-admin
devtools-fleet state import shop ./storageState.json --allow http://localhost:3000
# Check Node, Chrome, permissions and config; print effective config
devtools-fleet doctor
devtools-fleet configBrowser statuses: active (an agent is connected), detached (the connection dropped; waiting for that session to reconnect), orphan (the session is gone; closed on the next cleanup), starting.
How It Works
Process Management
On the first browser tool call, devtools-fleet:
Checks the cap under a machine-wide lock (
maxBrowsers, default 10). A full fleet gets a clear error naming the running sessions anddevtools-fleet gcLaunches Chrome itself, on a port it picks, with its own temporary profile
Restores the state, if one was asked for: cookies over CDP,
localStoragethrough a throwaway tab whose request devtools-fleet answers itself, so the site is never contactedSpawns chrome-devtools-mcp from its own pinned
node_moduleswith--browserUrl, and re-exports its tool list unchangedRegisters the browser in
~/.devtools-fleet/browsers/<id>.json
When the connection drops, Chrome keeps running. When the same client session reconnects, it finds its browser in the registry and re-adopts it, tabs intact. A reaper process (one per machine, started on demand, gone when idle) closes orphans at once and detached browsers after orphanTimeoutMinutes.
How Sessions Are Recognised
MCP clients start servers through wrappers (npx, npm exec, node, shells) and restart that whole chain on reconnect, while the client session itself (the claude process, an editor's extension host) keeps running. devtools-fleet walks up the process tree past the wrappers to that session process. The same session reconnecting finds the same process and gets its browser back; a second session in the same directory finds a different one and never shares.
DEVTOOLS_FLEET_SESSION=<label> replaces the lookup with a fixed label, useful in CI or containers.
Saved Logins
devtools-fleet login opens a real Chrome window. You log in, come back to the terminal, press Enter. devtools-fleet proposes an allowlist of where you started and where you ended up. Origins you only passed through, such as an SSO provider, are listed but left out: saving their cookies would hand agents your whole SSO session. Add one with --allow if the app really needs it.
What gets saved: cookies and localStorage for the allowed origins, in Playwright's storageState format. Playwright loads these files directly, and state import takes Playwright or agent-browser files the other way.
Rules that keep states safe to hand to an agent:
States are only written outside git work trees, with mode
0600in a0700directory.No listing, CLI or MCP, ever shows cookie or storage values.
A state is always bound to an allowlist. There is no "use these cookies anywhere" mode.
The Allowlist
Allowlist entries are full origins:
Entry | Matches |
| exactly that scheme, host and port |
| ports matter: |
| any subdomain, not |
Two layers enforce it:
Argument check.
new_pageandnavigate_pageare checked before anything reaches the browser; the agent gets a plain error.Navigation interception. devtools-fleet's own CDP connection intercepts navigations in every tab, popup and frame and fails the ones outside the list. A tab that can't be put under interception is closed rather than left unguarded.
A browser with an allowlist also refuses file: URLs and extra browser contexts (new_page's isolatedContext), since pages there would sit outside the guard.
Strict mode (login --strict) is enforced by Chrome itself: the browser launches with a proxy that goes nowhere, and only the allowed origins bypass it. Every other request fails in Chrome's network stack: fetch(), XHR, beacons, workers, WebSockets, QUIC. WebRTC is pinned to the dead proxy and DNS prefetching is off. The lock holds even while no agent is connected. It's off by default because it also blocks CDNs you haven't listed. At that layer an entry means host and port, so http and https on the same port aren't told apart there; the navigation guard still tells them apart.
What the allowlist does not do:
It limits where the agent goes, not what it does there. On an allowed origin,
evaluate_scriptcan read anything the page can, including non-httpOnly cookies. Treat a saved login like handing someone your session.Without strict mode, only navigations are guarded, and only while an agent is connected. A detached browser can still be navigated away by a script already running in its pages.
Files and Paths
chrome-devtools-mcp only reads and writes files (screenshots, traces, uploads) inside the client's workspace roots and the temp directory. devtools-fleet passes your client's roots through, so take_screenshot({ filePath: "<project>/shot.png" }) works as usual. Nothing may point into ~/.devtools-fleet: not a filePath, not an upload_file, not a file: URL, symlinks included.
Data Storage
~/.devtools-fleet/
config.json # Optional configuration (see below)
browsers/
<id>.json # One registry entry per running browser
profiles/ # Temporary Chrome profiles, deleted on close
states/ # Saved logins (0700 dir, 0600 files)
<name>.json
locks/ # mkdir locks for the cap check and adoptionDEVTOOLS_FLEET_HOME moves the whole directory.
Configuration
~/.devtools-fleet/config.json, every key optional. Environment variables override the file.
Key | Env | Default | |
|
|
| |
|
| none | e.g. |
|
|
| |
|
| detected | |
|
|
| across the machine |
|
|
| how long a dropped connection's browser waits for a reconnect |
|
|
| |
|
| extra chrome-devtools-mcp flags | |
|
| extra Chrome flags, e.g. |
chrome-devtools-mcp collects usage statistics by default and may send performance trace URLs to Google's CrUX API. devtools-fleet keeps its defaults. Opt out with "upstreamArgs": ["--no-usage-statistics", "--no-performance-crux"].
Limits
Tested on macOS and Linux. Windows should work apart from reconnect re-adoption, but is untested.
Saved logins cover cookies and
localStorage. Not IndexedDB,sessionStorageor service worker caches.Sites with device-bound sessions (Device Bound Session Credentials) won't accept copied cookies.
Google accounts are not supported. Google blocks automated sign-in and binds sessions to devices.
Headless can't switch to a window in place.
browser_restartrelaunches, and the pages reload.The debugging port is local-only but not authenticated, the same as chrome-devtools-mcp and every CDP tool: other processes running as your user can drive these browsers.
Companion Tools
devtools-fleet is the look step for the PluginsLab WordPress MCPs:
MCP | Purpose |
Verified hooks/filters/APIs for writing plugin code | |
Verified block schemas for generating content | |
Ephemeral WordPress instances for testing | |
devtools-fleet-mcp (this) | One DevTools-equipped Chrome per agent for looking at the result |
Together: author → validate → test → look.
Requirements
Node.js 22.12+
Google Chrome (stable, beta, dev or canary), or any Chromium via
chromePathmacOS or Linux
Development
npm install
npm test # unit tests, no browser
npm run test:integration # launches real Chrome; ~1 minute
npm run lintLicense
MIT
Available Tools
36 toolsbrowser_restartA
Relaunch this session's browser keeping cookies, localStorage and open tabs. Pass headless: false to get a visible window (e.g. so the human can watch or solve a captcha), headless: true to hide it again. Pages reload.
| Name | Required | Description | Default |
|---|---|---|---|
| headless | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses what survives (cookies, localStorage, open tabs) and one side effect ('Pages reload'), but omits what is lost (in-memory JS state, scroll position), whether it errors when no browser is running, and how long the relaunch takes. Partial disclosure, not a rich profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and preservation guarantee, then the parameter semantics, then the reload side effect. Every sentence earns its place and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no output schema, the description covers the action, preserved state, parameter semantics, and reload behavior. The remaining gap is the default value/behavior of the optional headless parameter when omitted, which an agent would want to know.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single parameter name ('headless') alone is ambiguous, but the description explains both boolean values with concrete intent ('visible window … headless: true to hide it again'), which is exactly the meaning the schema lacks. It does not state the default when the optional parameter is omitted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Relaunch this session's browser') plus the scope that separates it from a fresh browser_start: cookies, localStorage and open tabs are preserved. It doesn't name the sibling it is contrasted with, but the session-scoped framing makes the distinction inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a concrete use condition for one parameter value (headless: false so a human can watch or solve a captcha), but never says when to reach for browser_restart rather than browser_start, browser_stop, or a plain reload, nor what prerequisites exist (e.g. an active session). Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_startA
Start this session's own Chrome (isolated from every other agent). Optional: the first browser tool call starts one with defaults anyway. Call it explicitly to load a saved login (state), to pick headless/viewport/channel, or to open a URL. A browser started from a state can only visit that state's allowed origins. If this session already had a browser before a reconnect, it is picked up again with its tabs.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Open this URL in the first tab | |
| state | No | Name of a saved login state (see state_list) | |
| channel | No | ||
| headless | No | Run without a window (default from config, normally true) | |
| viewport | No | Window size, e.g. "1280x720" | |
| allowedOrigins | No | Without a state: restrict this browser to these origins, e.g. ["http://localhost:3000"] |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses per-agent isolation, that a state-bound browser is restricted to that state's allowed origins, and that a browser surviving a reconnect is re-adopted with its tabs. It doesn't cover failure modes or what happens if a browser is already running with conflicting settings, leaving a small gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and the optionality caveat, then the explicit reasons to call it. Dense but every sentence carries information; the reconnect sentence is slightly tangential but useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter, all-optional tool with no output schema and no annotations, the description covers identity, isolation, origin restrictions, and reconnect behavior well enough to call it correctly. Return-shape details and behavior on a second call are the only notable omissions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 83%, but the description adds real meaning: it groups the tuning parameters (headless/viewport/channel), explains the state parameter's origin-restriction consequence, and cross-references state_list. It still doesn't clarify the distinction/additivity of allowedOrigins versus a state's own allowed origins.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Start this session's own Chrome') and immediately scopes it as isolated from every other agent, which cleanly separates it from siblings like browser_status, browser_restart and browser_stop. An agent can identify the tool's action and its unique session-scoped nature without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says the call is optional because the first browser tool call starts one with defaults, then names the exact reasons to invoke it eagerly: load a saved login state, choose headless/viewport/channel, or open a URL. This is genuine when-to-use/when-not-to-use guidance with conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_statusA
Show this session's browser: id, saved state, allowlist, headless, open tabs, uptime.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose the concrete state dimensions returned, which signals this is a non-mutating introspection tool, but it never states that it is read-only, that it requires an active session, or what happens if no browser is running. Adequate but incomplete for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that leads with the verb and subject and then lists the returned fields compactly. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the returned state fields, which is the main thing an agent needs. For a zero-parameter, zero-risk status tool this is close to complete; only the precondition of an active browser session is unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there are no parameter semantics to document; the schema coverage is 100% and there is nothing for the description to compensate for. Baseline 4 applies for a 0-param schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Show) plus the resource (this session's browser) and enumerates the exact fields returned (id, saved state, allowlist, headless, open tabs, uptime). That distinguishes it from sibling read tools like list_pages and state_list. It stops short of explicitly naming a sibling it is not, so a 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no routing to alternatives. The phrase 'this session's browser' implies an inspection context, but the agent must infer that it precedes actions like browser_start or close_page. No exclusions or alternative conditions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_stopA
Close this session's browser and delete its temporary profile.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose a genuinely useful destructive side effect — the temporary profile is deleted — but says nothing about what happens to open pages, whether the profile deletion is irreversible, whether it errors if no session is running, or whether a subsequent browser_start is required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the primary action and the secondary side effect. No filler, no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter teardown tool with no output schema, the description covers the action and its cleanup effect adequately. It could be stronger by noting the fate of open pages and whether a restart is needed afterward, but nothing essential to invoking it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema is empty and the baseline is 4. There is nothing further the description could add on parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (close the browser) plus the notable side effect (delete the temporary profile). The scope word "this session's" distinguishes it from the sibling close_page, which closes a single page rather than the whole browser. It doesn't explicitly name an alternative, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance. The agent must infer that this is a teardown call for ending a browser session, and there is no statement about prerequisites (e.g. a running browser) or when browser_restart would be preferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clickC
Clicks on the provided element
| Name | Required | Description | Default |
|---|---|---|---|
| uid | Yes | The uid of an element on the page from the page content snapshot | |
| pageId | Yes | Targets a specific page by ID. | |
| dblClick | No | Set to true for double clicks. Default is false. | |
| includeSnapshot | No | Whether to include a snapshot in the response. Default is false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint=false annotation already signals a mutating action. The description adds nothing beyond that: it does not disclose side effects such as navigation, form submission, focus changes, or potential waits triggered by the click, which matter greatly for a browser-automation click.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or redundancy. It is efficient, though its brevity partly reflects under-specification rather than disciplined concision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a browser click with a non-read-only annotation and no output schema, the description is too thin: it omits side-effect behavior and how includeSnapshot affects the response, leaving an agent unable to anticipate consequences of the call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so uid, pageId, dblClick, and includeSnapshot are all documented in the schema. The phrase 'the provided element' adds no syntax or format detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (clicks) and resource (the provided element), so an agent can distinguish it from siblings like hover, drag, or fill. It does not, however, explicitly contrast itself with those siblings, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus hover, drag, type_text, or press_key, and no preconditions or exclusions stated. The agent must infer usage entirely from the name and schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_pageA
Closes the page by its index. The last open page cannot be closed.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | The ID of the page to close. Call list_pages to list pages. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply readOnlyHint=false, so the description carries most of the burden and does add real behavioral value by disclosing the failure case that the last open page cannot be closed. It stops short of saying whether page state is destroyed or recoverable, which would be useful for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, purpose front-loaded, constraint second. Nothing extraneous and nothing buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema and only a readOnlyHint annotation, the description covers purpose and the key failure condition adequately. It could note the effect on browser state after closing, but that is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already explains pageId fully, including the hint to call list_pages, so the baseline is 3. The description's 'by its index' wording adds no detail and mildly conflicts with the schema's 'ID' framing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Closes the page') that clearly distinguishes it from siblings like new_page, select_page, and navigate_page. The qualifier 'by its index' is slightly at odds with the schema's pageId terminology, but the core purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool's use (closing an open page) and adds one important constraint — the last open page cannot be closed — but never states when to prefer this over alternatives or what prerequisites exist. Usage is inferable rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dragC
Drag an element onto another element
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | Targets a specific page by ID. | |
| to_uid | Yes | The uid of the element to drop into | |
| from_uid | Yes | The uid of the element to drag | |
| includeSnapshot | No | Whether to include a snapshot in the response. Default is false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint=false annotation already signals this is a mutating operation, but the description adds nothing beyond that: it does not explain what dragging actually does to the page (DOM reordering, moving an element into a container), whether it waits for completion, or any side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler, and the core action is front-loaded. It is efficient, though arguably terse enough that it under-specifies the operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description should at minimum convey the effect of the drag operation and whether it involves waiting or verification. The schema covers parameters fully, but the behavioral context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so pageId, from_uid, to_uid, and includeSnapshot are all documented in the schema. The description adds no syntax or format detail beyond the schema, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (drag) and both resources (element onto another element), which cleanly distinguishes it from sibling tools like click, hover, and fill. It stops short of naming the uid-based identity of the elements, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use drag versus alternatives such as click or hover, and no prerequisites mentioned. The agent must infer that this is for drag-and-drop interactions from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
emulateC
Emulates various features on the target page.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | Targets a specific page by ID. | |
| viewport | No | Emulate device viewports '<width>x<height>x<devicePixelRatio>[,mobile][,touch][,landscape]'. 'touch' and 'mobile' to emulate mobile devices. 'landscape' to emulate landscape mode. | |
| userAgent | No | User agent to emulate. Set to empty string to clear the user agent override. | |
| colorScheme | No | Emulate the dark or the light mode. Set to "auto" to reset to the default. | |
| geolocation | No | Geolocation (`<latitude>,<longitude>`) to emulate. Latitude between -90 and 90. Longitude between -180 and 180. Omit to clear the geolocation override. | |
| extraHttpHeaders | No | Extra HTTP headers as a JSON string object, e.g. {"X-Custom": "value", "Authorization": "Bearer token"}. Headers are included into every HTTP request originating from the page and persist across navigations until cleared. Pass an empty string to clear all extra headers. | |
| cpuThrottlingRate | No | Represents the CPU slowdown factor. Omit or set the rate to 1 to disable throttling | |
| networkConditions | No | Throttle network. Omit to disable throttling. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, so this mutates page state, but the description says nothing about persistence, scope, or reversibility. It does not repeat or contradict the annotation, but it also adds no behavioral context beyond what the schema itself hints at.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence, so it is front-loaded and free of waste, but its brevity comes from under-specification rather than discipline. Structure is acceptable; content is thin.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool that persistently overrides page behavior across navigations, the description leaves the agent without any sense of scope or side effects. With no output schema and only a null readOnlyHint annotation, the description should carry more of the load.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the parameter docs are unusually rich (clearing overrides, geolocation ranges, header format), so the baseline of 3 applies. The description contributes nothing to parameter meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description restates the tool name almost literally ('emulate' → 'Emulates various features'), which is close to a tautology. It never names which features (viewport, user agent, geolocation, network throttling) despite the schema enumerating them, so an agent cannot distinguish it from siblings like resize_page or navigate_page without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of prerequisites (does a page need to exist first?), and no routing to or away from any sibling tool. The agent must infer the entire usage context from the parameter list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluate_scriptB
Evaluate a JavaScript function inside the target page. Returns the response as JSON, so returned values have to be JSON-serializable.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | An optional list of arguments to pass to the function. | |
| pageId | Yes | Targets a specific page by ID. | |
| filePath | No | The absolute or relative path to a file to save the script output to. If omitted, the output is returned inline. | |
| function | Yes | A JavaScript function declaration to be executed by the tool in the target page. Example without arguments: `() => document.title` or `async () => await fetch("example.com")`. Example with arguments: `(el) => el.innerText` | |
| dialogAction | No | Handle dialogs while execution. "accept", "dismiss", or string for response of window.prompt. Defaults to accept. | |
| waitForStableDom | No | Whether to wait for the DOM to settle. Pass false if the script only reads data. Defaults to true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=false correctly signals this is not a safe read-only operation, and the description's note that arbitrary JS runs in the page is consistent with that. It adds the useful constraint that return values must be JSON-serializable, but says nothing about side effects on the page, error behavior, sandboxing/security boundaries, or that dialogs are auto-accepted by default.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core purpose followed by the return-value constraint. No filler, though the JSON note could be folded more tightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description reasonably explains that output is JSON and can be saved via filePath (per schema). However, for a tool that executes arbitrary code in a browser page, the description omits side-effect/security context and error handling that an agent would want before invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so pageId, function, args, filePath, dialogAction, and waitForStableDom are all documented in the schema itself. The description adds no parameter-level meaning beyond the JSON-serializable constraint, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: evaluate a JavaScript function inside the target page, and notes the result is returned as JSON. This distinguishes it from browsing siblings like click, fill, or take_snapshot. It does not, however, contrast itself against the other data-reading tools in the family.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance, and no alternatives are named (e.g. use click/fill for element interaction, take_snapshot for page content). The JSON-serializability note is a constraint on the return value, not selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fillC
Type text into an input, text area or select an option from a element.
| Name | Required | Description | Default |
|---|---|---|---|
| uid | Yes | The uid of an element on the page from the page content snapshot | |
| value | Yes | The value to fill in. "true" or "false" for checkboxes and toggles, "true" for radio buttons. | |
| pageId | Yes | Targets a specific page by ID. | |
| includeSnapshot | No | Whether to include a snapshot in the response. Default is false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, implying mutation, but the description says nothing about whether the existing value is cleared or appended, whether input/change events fire, or that it requires a valid uid from a live page. The one useful behavioral nuance (select options) is also covered by the schema's value description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action front-loaded and no filler. It is efficient, though the brevity comes at the cost of the usage and behavioral detail an agent would benefit from.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating DOM interaction tool with no output schema and minimal annotations, the description omits prerequisites (element discovery via snapshot), what happens to pre-existing values, and the effect on the page. Important context for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters including the checkbox/radio 'true'/'false' convention and includeSnapshot are documented in the schema. The description adds no parameter detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete action (type text / select an option) and names the target resources (input, text area, <select>). However, it does not distinguish itself from close siblings like fill_form and type_text, which an agent must choose between for similar tasks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this instead of type_text or fill_form, nor any note that the target element must first be located via a page snapshot. The agent is left to infer the selection criteria entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fill_formA
Fill out multiple form elements (inputs, selects, checkboxes, radios) at once. ALWAYS prefer this tool over multiple individual 'fill' or 'click' calls when interacting with forms. It is significantly faster, more reliable, and reduces turn count. Example: Fill username, password, and check "Remember Me" in one call.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | Targets a specific page by ID. | |
| elements | Yes | Elements from snapshot to fill out. | |
| includeSnapshot | No | Whether to include a snapshot in the response. Default is false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, consistent with a form-filling mutation. The description adds useful context (faster, more reliable, reduces turn count) and the batch semantics, but discloses nothing about failure behavior, atomicity, or what happens when a uid is invalid. Adds some value but not rich behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and the preference rule, followed by a justified example. Slightly promotional ('significantly faster, more reliable') but each sentence still earns its place by steering tool selection.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description needn't explain return values, and the annotations are minimal enough that the description covers the mutation nature and batching benefit. It is complete enough to invoke correctly, though it omits error/atomicity behavior for a multi-element write.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents pageId, elements, uid, value, and includeSnapshot, including the 'true'/'false' conventions for checkboxes and radios. The description adds the example semantics but no new parameter meaning beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Fill out) and resource (multiple form elements, enumerating inputs, selects, checkboxes, radios) in one call. It clearly distinguishes itself from the sibling tools 'fill' and 'click' by emphasizing the batch nature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states 'ALWAYS prefer this tool over multiple individual fill or click calls when interacting with forms,' giving a direct when-to-use rule and naming the alternatives. A concrete example (username, password, Remember Me) further grounds the guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_console_messageARead-only
Gets a console message by its ID. You can get all messages by calling list_console_messages.
| Name | Required | Description | Default |
|---|---|---|---|
| msgid | Yes | The msgid of a console message on the page from the listed console messages | |
| pageId | Yes | Targets a specific page by ID. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations supply readOnlyHint=true, so the safety profile is already covered. The description adds essentially nothing behavioral beyond that: no return shape, no behavior for an unknown/out-of-range msgid, and no note that the ID must originate from a prior list call. Minimal added context over the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, purpose front-loaded, with the cross-reference to the sibling tool in the second sentence. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-param read tool with annotations and full schema coverage, the description is just barely sufficient. It omits what a message object contains and how failures (invalid ID) behave, which a slightly richer definition could have supplied without bloat.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both pageId and msgid are documented in the schema itself. The description adds no format, range, or sourcing detail beyond what the schema already says, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (gets) and resource (a console message) plus the identifier used (by its ID). It also names the sibling list_console_messages as the way to obtain all messages, so the agent can infer this is the single-message lookup, though the distinction is implicit rather than spelled out.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Points the agent to list_console_messages when it needs all messages, which effectively scopes when this tool applies (when the ID is already known). There is no explicit exclusion or precondition statement (e.g. what happens with an invalid msgid), so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_css_stylesARead-only
Retrieve matched CSS rules, inline styles, inherited styles, and cascade information for an element identified by its UID. Use this tool to debug why specific CSS properties are applied, overridden, or conflicting. Results are paginated and return 10 rules per page by default; use pageIdx to page through the remaining rules. Requires a UID from take_snapshot.
| Name | Required | Description | Default |
|---|---|---|---|
| uid | Yes | The uid of the element on the page from the page content snapshot to inspect CSS styles for | |
| pageId | Yes | Targets a specific page by ID. | |
| pageIdx | No | Page number to return (0-based). Defaults to 0 (the first page). | |
| pageSize | No | Maximum number of CSS rules to return per page. Defaults to 10. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered, and the description adds genuinely useful behavior: results are paginated at 10 rules per page and pageIdx advances through them, plus a hard dependency on take_snapshot for the UID. It stops short of describing ordering or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, purpose front-loaded, then usage, then pagination logistics. Every sentence carries information, though the pagination sentence could be marginally tighter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the work of naming the return content (matched, inline, inherited rules and cascade info) and pagination behavior, which is what an agent needs to call and interpret it. Missing only edge-case/error behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents uid, pageId, pageIdx, and pageSize. The description's mention of pagination and default page size largely restates schema defaults, adding only the practical hint to advance with pageIdx, which is the baseline 3 expectation when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Retrieve) and a well-scoped resource (matched CSS rules, inline, inherited styles, cascade info for an element), and it is clearly distinguishable from sibling tools like take_snapshot or evaluate_script. An agent can tell what this returns without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the use case (debug why CSS properties are applied, overridden, or conflicting) and states the prerequisite (a UID from take_snapshot). It does not name an alternative tool or state when not to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_network_requestB
Gets a network request by an optional reqid, if omitted returns the currently selected request in the DevTools Network panel. Useful for inspecting request headers (including 'Cookie') and response headers (including 'Set-Cookie' and directives).
| Name | Required | Description | Default |
|---|---|---|---|
| reqid | No | The reqid of the network request. If omitted returns the currently selected request in the DevTools Network panel. | |
| pageId | Yes | Targets a specific page by ID. | |
| requestFilePath | No | The absolute or relative path to a .network-request file to save the request body to. If omitted, the body is returned inline. | |
| responseFilePath | No | The absolute or relative path to a .network-response file to save the response body to. If omitted, the body is returned inline. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Description implies a read-only operation ('Gets', 'returns') but annotation readOnlyHint=false indicates write side effects. The schema shows file-saving parameters that write request/response bodies to disk, which the description never discloses. This is a direct contradiction of the implied read-only nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core operation and default behavior, then a useful access note. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Omits the file-saving parameters (requestFilePath, responseFilePath) and required pageId. For a tool with four parameters and no output schema, the description is incomplete about how the tool can be used and what side effects it has.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, fully documenting all four parameters. The description only repeats the reqid default already in the schema and adds no meaning beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb 'Gets' and resource 'network request', with optional reqid behavior. Clearly distinguishes from sibling list_network_requests by fetching a single request.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear use case: inspecting request and response headers including cookies. Does not name alternatives or specify when not to use, but context is sufficient for a simple getter.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handle_dialogC
If a browser dialog was opened, use this command to handle it
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Whether to dismiss or accept the dialog | |
| pageId | Yes | Targets a specific page by ID. | |
| promptText | No | Optional prompt text to enter into the dialog. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations set readOnlyHint=false, marking this as a mutating operation, and the description does not contradict that. But it discloses nothing beyond the annotation: it never says dialogs block page execution, that accept vs dismiss produce different outcomes, or what entering promptText does to the page state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the condition comes before the instruction. It is efficient, though its brevity is partly under-specification rather than deliberate economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter tool with full schema coverage and no output schema, the definition is minimally adequate. It omits useful context such as how to discover an open dialog, that only one dialog can be handled at a time, and what happens when no dialog is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (action, pageId, promptText) are documented in the schema itself. The description adds no parameter meaning beyond that, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (browser dialog) and a condition for acting, but 'use this command to handle it' largely restates the tool name rather than specifying what handling means. It gives no detail on the accept/dismiss outcome, though no sibling handles dialogs so differentiation is not the issue.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a clear trigger condition ('If a browser dialog was opened'), which is genuine when-to-use guidance. However, it offers no when-not guidance, no mention of validation steps (e.g., checking via snapshot) and no alternatives or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hoverC
Hover over the provided element
| Name | Required | Description | Default |
|---|---|---|---|
| uid | Yes | The uid of an element on the page from the page content snapshot | |
| pageId | Yes | Targets a specific page by ID. | |
| includeSnapshot | No | Whether to include a snapshot in the response. Default is false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations declare readOnlyHint=false, signalling this is a state-affecting interaction, but the description adds nothing beyond that: it does not say whether hovering triggers tooltips/menus, whether it waits for a response, or what side effects occur. With only a single annotation, the description is expected to carry more of the behavioral burden and does not.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single short sentence with no wasted words and is front-loaded, but it is terse to the point of under-specification rather than economical, restating the tool name with almost no added information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an element-interaction tool with no output schema, the description omits whether the call waits/blocks, what hover effects to expect, and how includeSnapshot affects the response. Given annotations cover only readOnlyHint, the description leaves key behavioral questions unanswered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so uid, pageId, and includeSnapshot are already documented in the schema. The description adds no meaning beyond the schema, which is the expected baseline of 3 when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (hover) and its target (the element referenced by the uid param), so an agent knows the basic action. It does not differentiate this from siblings like click, drag, or fill, and 'the provided element' is mildly circular since the element is identified by the parameter.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use hover versus click, fill, drag, or other element-interaction siblings, and no prerequisites or exclusions. The agent must infer usage entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lighthouse_auditA
Get Lighthouse score and reports for accessibility, SEO, best practices, and agentic browsing. This excludes performance. For performance audits, run performance_start_trace
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | "navigation" reloads & audits. "snapshot" analyzes current state. | navigation |
| device | No | Device to emulate. | desktop |
| pageId | Yes | Targets a specific page by ID. | |
| outputDirPath | No | Directory for reports. If omitted, uses temporary files. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotation readOnlyHint=false is consistent with the fact that audits produce report artifacts, so there is no contradiction. The description does communicate what is excluded, but it says nothing about report generation side effects, run duration, or how outputDirPath affects behavior, so it adds only modest context beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, zero filler, with the capability scope front-loaded before the exclusion and the alternative-tool routing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read/analysis tool with no output schema and a fully documented parameter set, the description covers scope, exclusions, and the alternative tool adequately. Minor gaps remain around what the returned score/report looks like and run-time behavior, but nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so mode, device, pageId, and outputDirPath are already fully documented in the schema (including enum semantics). The description adds no parameter-level information, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get') and resource ('Lighthouse score and reports') and enumerates the covered categories (accessibility, SEO, best practices, agentic browsing). It also states what is explicitly out of scope (performance) and names the sibling tool that handles it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit exclusion and an explicit routing rule: 'This excludes performance. For performance audits, run performance_start_trace.' An agent can select between this tool and performance_start_trace without inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_console_messagesBRead-only
List all console messages for the target page since the last navigation.
| Name | Required | Description | Default |
|---|---|---|---|
| types | No | Filter messages to only return messages of the specified resource types. When omitted or empty, returns all messages. | |
| pageId | Yes | Targets a specific page by ID. | |
| pageIdx | No | Page number to return (0-based). When omitted, returns the first page. | |
| pageSize | No | Maximum number of messages to return. When omitted, returns all messages. | |
| serviceWorkerId | No | Filter messages to only return messages of the specified service worker. | |
| includeStackTraces | No | Set to true to include the stack trace for each message when available. Increases the response size. | |
| includePreservedMessages | No | Set to true to return the preserved messages over the last 3 navigations. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true, so the safety profile is covered. The description usefully adds the temporal boundary ('since the last navigation') that defines what the list contains, but says nothing about pagination or preserved-message behavior beyond what the schema already documents.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. It is efficient and readable, though it does not exploit the space to add routing or behavioral context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only listing tool with 100% schema coverage and a readOnlyHint annotation, the description covers the essential purpose and scope. No output schema exists, yet the temporal scope gives the agent enough to call it correctly; only the sibling relationship is omitted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all seven parameters (types, pageId, pageIdx, pageSize, serviceWorkerId, includeStackTraces, includePreservedMessages) are documented in the schema itself. The description adds no parameter detail, which is acceptable given full coverage — baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (console messages) with an explicit temporal scope ('since the last navigation'). This is clearly distinguishable from the singular get_console_message by scope, though that sibling relationship is not called out in the text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, no prerequisites, and does not name the alternative get_console_message for retrieving a single message. The temporal scope is stated but that is a filter, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_network_requestsBRead-only
Lists the most recent requests for the target page since the last navigation.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | Targets a specific page by ID. | |
| pageIdx | No | Page number to return (0-based). When omitted, returns the first page. | |
| pageSize | No | Maximum number of requests to return. When omitted, returns all requests. | |
| resourceTypes | No | Filter requests to only return requests of the specified resource types. When omitted or empty, returns all requests. | |
| includePreservedRequests | No | Set to true to return the preserved requests over the last 3 navigations. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already tells the agent this is a safe read. The description adds the temporal scope ('since the last navigation'), which is real context beyond the annotations, but it omits return-volume behavior and the existence of preserved-request handling, leaving notable gaps for a list tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. The scope constraint is stated immediately and nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter read-only list tool with a fully documented schema but no output schema, the description covers the essential scope. The minor tension between 'most recent requests' and the schema's 'returns all requests when pageSize is omitted' is left unresolved, and no return shape is hinted at, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (pageId, pageIdx, pageSize, resourceTypes, includePreservedRequests) are already fully documented in the schema. The description adds no extra parameter meaning, which is the expected baseline when the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Lists) and resource (network requests) scoped to the target page, and the singular sibling get_network_request is implicitly contrasted by the plural naming. It stops short of explicitly differentiating itself from that sibling, so it lands just below the top tier.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'since the last navigation' implies a temporal scope but no explicit when-to-use, when-not-to-use, or alternative (e.g., get_network_request, list_console_messages) is named. The agent must infer when this tool is preferred over its sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_pagesBRead-only
Get a list of pages open in the browser.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so safety is covered by structured data. The description adds only that the result is a 'list of pages open in the browser,' which hints at the return shape but says nothing about ordering, page identifiers, or whether background/closed pages are included.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with no filler and the resource front-loaded after the verb. It is efficient, though it is arguably so terse that it omits useful scoping detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and no output schema, the description must carry the burden of explaining the return value. It says only 'a list of pages,' leaving unclear what fields each page entry carries (id, url, title) — information an agent likely needs before calling select_page or close_page.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is no parameter semantics to clarify, and the description correctly implies a no-argument call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'Get a list of pages open in the browser.' An agent can tell it retrieves open pages rather than modifying them. However, it does no sibling differentiation, e.g. how it relates to select_page, new_page, or close_page.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to prefer this over siblings like browser_status or select_page, nor any stated prerequisites. Use is only loosely implied by the name and the resource it describes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
new_pageB
Open a new tab and load a URL. Use project URL if not specified otherwise.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to load in a new page. | |
| timeout | No | Maximum wait time in milliseconds. If set to 0, the default timeout will be used. | |
| background | No | Whether to open the page in the background without bringing it to the front. Default is false (foreground). | |
| isolatedContext | No | If specified, the page is created in an isolated browser context with the given name. Pages in the same browser context share cookies and storage. Pages in different browser contexts are fully isolated (useful for clean-slate testing of cookies and authentication). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The only annotation is readOnlyHint=false, so the description carries most of the disclosure burden and does confirm the default-to-project-URL behavior. It does not, however, mention that focus moves to the foreground by default, how the new page relates to existing tabs, or what the caller receives back (e.g. a page identifier needed for follow-up calls).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler and the core action front-loaded. The second sentence is slightly ambiguous, but there is no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter page-creation tool with no output schema, the description is minimal: it omits how the newly created page is identified or referenced in subsequent calls, and doesn't hint at the interaction with background/isolatedContext behavior documented only in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (timeout, background, isolatedContext all documented in-schema), so the schema already does the heavy lifting. The description adds only the default-URL rule for the required url parameter and nothing about the other three parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb (open) plus resource (a new tab/page) and the action taken (load a URL), so an agent can distinguish it from navigate_page or select_page at a glance. It is clear, though it never names the siblings it is distinct from.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage cue is 'Use project URL if not specified otherwise,' which is an implicit default rather than a when-to-use rule. Nothing says when to open a new page versus navigating an existing one (navigate_page) or selecting among open pages (select_page).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
performance_analyze_insightBRead-only
Provides more detailed information on a specific Performance Insight of an insight set that was highlighted in the results of a trace recording.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | Targets a specific page by ID. | |
| insightName | Yes | The name of the Insight you want more information on. For example: "DocumentLatency" or "LCPBreakdown" | |
| insightSetId | Yes | The id for the specific insight set. Only use the ids given in the "Available insight sets" list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint=true annotation already tells the agent this is a safe read operation, so the description's main added value is scoping it to insight sets from a trace recording. It does not disclose return format, whether results are cached, or any depth/limit on the 'more detailed information', which would be useful given there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler or repetition. It could be marginally tighter ('Provides details on' rather than 'Provides more detailed information on'), but it is efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter read-only tool with full schema coverage, the description covers the essential purpose. However, with no output schema, it leaves the shape and depth of the returned insight data entirely unspecified, which is a real gap for an agent deciding whether this call will give it what it needs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (pageId, insightName, insightSetId) are already documented in the schema, including an example insight name. The description adds no parameter-level syntax or format guidance beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: retrieving 'more detailed information on a specific Performance Insight of an insight set' produced by a trace recording. This clearly separates it from siblings like performance_start_trace and performance_stop_trace, though the phrasing 'Provides more detailed information' is a little generic about the retrieval mechanics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'highlighted in the results of a trace recording' implies this is a follow-up tool used after a trace has produced insight sets, but it never names the alternative tool(s) or states explicit when/when-not conditions. Usage context is implied rather than prescribed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
performance_start_traceA
Start a performance trace on the target webpage. Use to find frontend performance issues, Core Web Vitals (LCP, INP, CLS), and improve page load speed.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | Targets a specific page by ID. | |
| reload | No | Determines if, once tracing has started, the target page should be automatically reloaded. Navigate the page to the right URL using the navigate_page tool BEFORE starting the trace if reload or autoStop is set to true. | |
| autoStop | No | Determines if the trace recording should be automatically stopped. | |
| filePath | No | The absolute file path, or a file path relative to the current working directory, to save the raw trace data. For example, trace.json.gz (compressed) or trace.json (uncompressed). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The only annotation is readOnlyHint=false, so the description carries most of the behavioral burden. It never states that the trace must be captured then stopped/exported, that the page should be navigated first, or where the raw trace data goes (that context lives only in the schema parameter text, not the description). No behavioral traits beyond the annotations are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and then the motivation. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and only a single annotation, the description is adequate on purpose but thin on workflow: it does not mention the trace must be stopped (performance_stop_trace) and analyzed, nor the reload/navigation prerequisite. Sufficient to start a trace, but not complete enough for an agent to use the full trace lifecycle correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents pageId, reload, autoStop, and filePath fully, including the 'navigate first' prerequisite. The description adds no parameter meaning of its own, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Start a performance trace on the target webpage') and reinforces it with concrete use-case goals (Core Web Vitals: LCP, INP, CLS). It is clearly distinguishable from take_screenshot or evaluate_script, though it never names its counterpart performance_stop_trace, so the lifecycle pairing is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context ('Use to find frontend performance issues... and improve page load speed'), telling the agent when this tool is appropriate. However, it names no alternatives (performance_analyze_insight, lighthouse_audit, performance_stop_trace) and provides no when-not guidance or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
performance_stop_traceB
Stop the active performance trace recording on the target webpage.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | Targets a specific page by ID. | |
| filePath | No | The absolute file path, or a file path relative to the current working directory, to save the raw trace data. For example, trace.json.gz (compressed) or trace.json (uncompressed). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The only annotation is readOnlyHint=false, marking this as a state-changing operation, and the description adds nothing beyond the name's implication. It does not say whether the recorded data is discarded when saved, whether filePath is optional, or what occurs if no trace is active — all material for a mutation tool whose annotation coverage is minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence naming the action, the resource and the target. Nothing redundant, nothing padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation with no output schema, the description is minimally adequate but silent on the outcomes that matter: whether trace data survives the stop, how filePath relates to persistence, and the required precondition of an active trace.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so pageId and filePath are already documented in the schema, and the baseline is 3. The description adds no parameter meaning at all — notably it omits that filePath is the mechanism by which the trace data is persisted rather than lost.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Stop) and resource (active performance trace recording) plus the target scope (target webpage). It is distinguishable from performance_start_trace by the verb alone, but it does not explicitly name the relationship to that sibling, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the pairing with performance_start_trace — an agent can infer this must be called after a trace is started — but the description never states that a trace must be active or what happens if none is running. No explicit when/when-not or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
press_keyA
Press a key or key combination. Use this when other input methods like fill() cannot be used (e.g., keyboard shortcuts, navigation keys, or special key combinations).
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | A key or a combination (e.g., "Enter", "Control+A", "Control++", "Control+Shift+R"). Modifiers: Control, Shift, Alt, Meta | |
| pageId | Yes | Targets a specific page by ID. | |
| includeSnapshot | No | Whether to include a snapshot in the response. Default is false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=false already tells the agent this mutates page state, so the safety profile is covered structurally. The description adds the useful 'fallback when fill() cannot be used' framing but says nothing about focus requirements, whether the key fires immediately, or what a failed press looks like. Some added value, but thin beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the core action front-loaded before the usage condition. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 3-parameter action tool with no output schema and full schema coverage, the description plus structured fields give enough to invoke it correctly. The only missing nuance is what the response contains when includeSnapshot is true, which the schema partially covers.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so key syntax, pageId targeting, and includeSnapshot defaults are all documented in the schema itself. The description adds no format or constraint detail beyond that, which is the correct baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Press a key or key combination') and immediately distinguishes the tool from alternative input paths like fill(). An agent can identify this as the keyboard-input primitive without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit trigger conditions in parentheses (keyboard shortcuts, navigation keys, special combinations) and names fill() as the alternative to fall back from. It does not mention the closely related type_text or hover/click siblings, so the routing is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resize_pageB
Resizes the page's window so that the page has specified dimension
| Name | Required | Description | Default |
|---|---|---|---|
| width | Yes | Page width | |
| height | Yes | Page height | |
| pageId | Yes | Targets a specific page by ID. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint=false annotation already signals a mutation, but the description adds almost no behavioral context beyond restating the action. It does not say whether width/height are pixels, whether the browser window or viewport is affected, or what side effects occur for the page.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is a single, front-loaded sentence with no wasted words. It is slightly awkward and omits units, but structurally it is efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter mutation tool with full schema coverage and no output schema, the description is minimally viable. However, it does not clarify units or whether the page viewport or actual browser window is resized, leaving some ambiguity for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with clear per-parameter descriptions for width, height, and pageId. The description adds no extra meaning beyond what the schema already provides, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ("Resizes"), a clear resource ("page's window"), and the outcome ("specified dimension"). It is unambiguous and easily distinguishable from browser automation siblings like new_page, emulate, or take_screenshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as emulate or new_page, nor any prerequisites like requiring an active page. Usage context is only implied by the tool name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
select_pageBRead-only
Select a page as a context for future tool calls.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | The ID of the page to select. Call list_pages to get available pages. | |
| bringToFront | No | Whether to focus the page and bring it to the top. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true is already declared, so the safety profile is covered. The description usefully adds that selection affects future tool calls (context effect), but does not disclose what happens to the previously selected page or whether multiple contexts are allowed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. Slightly terse rather than wasteful — nothing redundant, but also little added beyond the essentials.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 2-parameter, low-complexity tool with full schema coverage and no output schema, the definition is minimally sufficient. The 'context' concept and persistence across calls could be spelled out a bit more for an agent to use it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%; pageId documents its role and even routes to list_pages, and bringToFront documents its effect. The description adds nothing beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (select) and resource (page) and adds the scope/purpose 'as a context for future tool calls', which distinguishes it from siblings like new_page, close_page, or navigate_page. It's clear what it does, though it doesn't explicitly name the distinguishing siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (set a context for later calls), and the schema's pageId note points to list_pages for discovery. However, it doesn't say when to prefer this over navigate_page or new_page, nor whether it's required before other tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
state_listA
List saved login states: name, allowed origins, cookie counts, when saved and when the first cookie expires. Never shows cookie values.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden, and it does disclose a meaningful behavioral trait: cookie values are never exposed, only metadata such as cookie counts and first-expiry time. It stops short of covering anything else behavioral, but for a parameterless read tool the disclosure is substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the action and then lists returned fields, closing with the most important caveat. No filler and nothing misordered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the enumeration of returned fields (name, origins, cookie counts, saved time, first expiry) and the no-cookie-values caveat do the work the schema would otherwise do. The only minor gap is that questions of ordering, volume, or scoping of the listing are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema sets the baseline at 4. The description adds no parameter detail because there is none to add, and correctly focuses on the returned shape instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (saved login states), and enumerates the fields it returns. It is distinguishable from the sibling state_save by the list-vs-save framing, though the sibling is not named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The nature of the operation (read-only enumeration with no parameters) makes the intended use implied rather than stated. There is no explicit when-to-use guidance, no mention of state_save as the counterpart action, and no prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
state_saveA
Save the current browser's login (cookies + localStorage) as a named state, so other sessions can start already logged in. Only cookies for the allowed origins are kept. allowedOrigins defaults to the http(s) origins open in tabs; it can never be wider than this browser's own allowlist. States a person created with devtools-fleet login cannot be overwritten.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | State name: letters, digits, . _ - | |
| strict | No | Block every request (not just navigations) outside the allowlist when the state is used | |
| overwrite | No | ||
| allowedOrigins | No | Origins this state may be used on, e.g. ["https://staging.example.com"] |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers meaningful behavior: allowedOrigins defaults to open http(s) tab origins, can never exceed the browser's own allowlist, and states created via `devtools-fleet login` are protected from overwrite. It doesn't spell out what cookies outside the allowlist undergo (dropped silently) or auth prerequisites, but the disclosure is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then behavioral caveats in descending importance. Four tight sentences with no filler, though the allowlist constraint is restated in two adjacent sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a stateful save operation with no annotations and no output schema, the description covers the critical behaviors an agent needs (origin scoping, overwrite protection, default derivation). Only minor gaps remain around failure modes and return behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75% and the schema already documents name, strict, and allowedOrigins. The description adds genuine meaning beyond the schema by explaining allowedOrigins' default derivation and its hard upper bound, plus the overwrite protection behavior that bears on the overwrite flag.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+payload: 'Save the current browser's login (cookies + localStorage) as a named state.' The purpose is unambiguous and distinguishable from the sibling state_list, which the agent can infer handles the read side.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The rationale 'so other sessions can start already logged in' implies when this is useful, and constraints on allowedOrigins scope are given, but there is no explicit when-to-use vs when-not guidance and no mention of the sibling state_list as an alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
take_heapsnapshotA
Capture a heap snapshot of the target page. Use to analyze the memory distribution of JavaScript objects and debug memory leaks.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | Targets a specific page by ID. | |
| filePath | Yes | A path to a .heapsnapshot file to save the heapsnapshot to. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint=false annotation already signals this is not a pure read operation, and the filePath parameter conveys that output is written to disk. However, the description never mentions that it writes a .heapsnapshot file, whether it overwrites an existing path, or the cost/size implications of capturing a heap snapshot, so it adds little beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero waste, front-loading the action before the use case. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with full schema coverage and no output schema, the description is nearly sufficient: it explains what is captured and why. The only gap is that it does not disclose the disk-write behavior that the missing-file semantics might matter for.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with pageId and filePath both documented in the schema, so the baseline is 3. The description adds no further meaning about either parameter, such as pageId validity or .heapsnapshot file semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Capture a heap snapshot of the target page,' which tells an agent exactly what the tool produces. It does not explicitly distinguish itself from the similarly named sibling take_snapshot, which could cause selection ambiguity, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a rationale ('analyze the memory distribution of JavaScript objects and debug memory leaks') that implies when the tool is useful, but offers no explicit when-not conditions or named alternatives such as take_snapshot or performance_start_trace. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
take_screenshotC
Take a screenshot of the page or element.
| Name | Required | Description | Default |
|---|---|---|---|
| uid | No | The uid of an element on the page from the page content snapshot. If omitted, takes a page screenshot. | |
| format | No | Type of format to save the screenshot as. Default is "png" | png |
| pageId | Yes | Targets a specific page by ID. | |
| quality | No | Compression quality for JPEG and WebP formats (0-100). Higher values mean better quality but larger file sizes. Ignored for PNG format. | |
| filePath | No | The absolute path, or a path relative to the current working directory, to save the screenshot to instead of attaching it to the response. | |
| fullPage | No | If set to true takes a screenshot of the full page instead of the currently visible viewport. Incompatible with uid. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations supply only readOnlyHint=false, and the description adds nothing further. It never says the screenshot is attached to the response by default, that filePath writes to disk (which is the only thing that could explain a non-read-only hint), or anything about format/quality side effects. For a tool with effectively no behavioral annotation coverage, this is a substantial gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single clean sentence with no wasted words, and the core action is front-loaded. However, its extreme brevity means it conveys almost nothing beyond the tool name, so it reads as under-specified rather than efficiently scoped.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With six parameters, no output schema, and only a one-line description, the definition is incomplete. It omits whether the image is returned inline or saved via filePath, what the response looks like, and when fullPage is appropriate, leaving the agent to reconstruct behavior from the schema alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters including uid, format, quality, filePath and fullPage are already documented in the schema. The description's 'page or element' phrasing merely restates what the uid description already says, adding no syntax or format detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('take a screenshot') and covers both page-level and element-level capture. It does not, however, distinguish itself from siblings like take_snapshot or take_heapsnapshot, which an agent could plausibly confuse with this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus take_snapshot or the other capture tools, nor any mention of prerequisites such as needing a page already open or selected. The page-vs-element choice is left entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
take_snapshotA
Take a text snapshot of the target page based on the a11y tree. The snapshot lists page elements along with a unique identifier (uid). Always use the latest snapshot. Prefer taking a snapshot over taking a screenshot. The snapshot indicates the element selected in the DevTools Elements panel (if any).
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | Targets a specific page by ID. | |
| verbose | No | Whether to include all possible information available in the full a11y tree. Default is false. | |
| filePath | No | The absolute path, or a path relative to the current working directory, to save the snapshot to instead of attaching it to the response. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are minimal (only readOnlyHint=false), so the description carries most of the burden, and it delivers: element list with uid, staleness semantics, and the fact that the selected DevTools element is marked. It never explains why the operation is flagged non-read-only (e.g., file emission via filePath), which is the one behavioral gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, all front-loaded and functional: purpose, output shape, staleness rule, sibling preference. Slightly more than strictly necessary, but no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by describing what a snapshot returns and how uids work. For a 3-parameter read-style tool this is nearly complete; only the side-effect/read-only nuance is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so pageId, verbose, and filePath are already fully documented in the schema. The description adds no format or usage detail beyond that, which is the correct baseline when the schema does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('take a text snapshot of the target page based on the a11y tree') and immediately distinguishes the output from take_screenshot by describing it as a list of page elements with unique identifiers. An agent can tell exactly what it gets without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes between siblings: 'Prefer taking a snapshot over taking a screenshot.' It also gives a lifecycle rule ('Always use the latest snapshot') that tells the agent when a prior snapshot is stale, which is a real when-to-use constraint.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
type_textA
Type text using keyboard into a previously focused input
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The text to type | |
| pageId | Yes | Targets a specific page by ID. | |
| submitKey | No | Optional key to press after typing. E.g., "Enter", "Tab", "Escape" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint=false establishes this as a mutating action, which the description does not contradict. The description adds one behavioral trait beyond the annotations: it simulates keyboard input into an element that must already have focus, implying prior interaction state. It does not disclose what happens on failure, timing, or whether the text is appended versus replacing existing content.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. Every clause — verb, mechanism, target precondition — carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 3-parameter, no-output-schema tool with a fully documented schema and an annotation covering the safety profile, the description covers the essentials. The only gap is the absence of explicit differentiation from fill, which the agent must infer from wording alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so text, pageId, and submitKey are all documented in the schema itself. The description adds no syntax, format, or edge-case detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Type text") and the mechanism ("using keyboard"), plus a scope constraint ("into a previously focused input") that distinguishes it from a value-setting tool like fill. It does not name the closest sibling (fill) to route the agent, so it stops short of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Previously focused input" implies you must focus an element first (e.g. via click), which is useful implied guidance. But there is no explicit statement of when to use this versus fill/fill_form/press_key, all of which live in the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_fileC
Upload a file through a provided element.
| Name | Required | Description | Default |
|---|---|---|---|
| uid | Yes | The uid of the file input element or an element that will open file chooser on the page from the page content snapshot | |
| pageId | Yes | Targets a specific page by ID. | |
| filePaths | Yes | One or more files paths to upload. File paths have to be local to the browser instance (not the MCP). | |
| includeSnapshot | No | Whether to include a snapshot in the response. Default is false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false, which the one-line description is consistent with, but the description adds no behavioral context: no indication that it may trigger a file chooser, mutate page state, or what happens on failure. For a mutation tool with a bare annotation, more disclosure is warranted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no redundancy, front-loading the action. Its brevity comes at the cost of specificity, but there is no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 4 well-documented parameters and no output schema, the schema carries most of the load, so the description need not explain return values. It is still thin on the operational context (element type required, interaction with the file chooser) that an agent invoking a browser mutation would benefit from.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter (uid, pageId, filePaths, includeSnapshot) is already documented in the schema, including that paths must be local to the browser instance. The description adds nothing beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It states a specific verb and resource (upload a file), but "through a provided element" is vague and never names the uid/file-input concept that the schema requires. An agent can guess the general intent but cannot confidently distinguish it from siblings like fill or fill_form without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, no alternatives. The description does not say the element must be a file input or a chooser-opening element, nor how it relates to click or fill_form, leaving selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_forCRead-only
Wait for the specified text to appear on the selected page.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Non-empty list of texts. Resolves when any value appears on the page. | |
| pageId | Yes | Targets a specific page by ID. | |
| timeout | No | Maximum wait time in milliseconds. If set to 0, the default timeout will be used. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply readOnlyHint=true, so the description carries the behavioral burden. It says nothing about polling behavior, what happens when the wait times out, whether it errors or returns a status, or that the default timeout applies when 0 is passed (that detail lives only in the schema).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is efficient, though its brevity is partly the source of the missing behavioral detail rather than a mark of completeness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a blocking/waiting tool with no output schema and no annotations beyond readOnlyHint, the description omits the critical question of what the tool returns or does on timeout, and whether the page must be selected first. These are real gaps for an agent deciding how to handle the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents pageId, the text list with any-match semantics, and the timeout default behavior. The description adds no parameter syntax or format detail beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Wait') and resource ('specified text ... on the selected page'), which distinguishes it from sibling reads like take_snapshot or get_console_message. It does not, however, explicitly name an alternative tool or clarify how it relates to select_page.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to prefer this over polling with evaluate_script, take_snapshot, or re-reading state. No preconditions (e.g., a page must already be selected) and no when-not-to-use conditions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
36 tool updates
v0.1.0- First observed
browser_restart - First observed
browser_start - First observed
browser_status - First observed
browser_stop - First observed
click - First observed
close_page - First observed
drag - First observed
emulate - First observed
evaluate_script - First observed
fill - First observed
fill_form - First observed
get_console_message - First observed
get_css_styles - First observed
get_network_request - First observed
handle_dialog - First observed
hover - First observed
lighthouse_audit - First observed
list_console_messages - First observed
list_network_requests - First observed
list_pages - First observed
navigate_page - First observed
new_page - First observed
performance_analyze_insight - First observed
performance_start_trace - First observed
performance_stop_trace - First observed
press_key - First observed
resize_page - First observed
select_page - First observed
state_list - First observed
state_save - First observed
take_heapsnapshot - First observed
take_screenshot - First observed
take_snapshot - First observed
type_text - First observed
upload_file - First observed
wait_for
TDQS
Scored across 36 tools
Most tools target distinct resources or actions, such as take_snapshot vs take_screenshot and list_console_messages vs get_console_message. However, the input family (click, fill, fill_form, type_text, press_key) has some overlapping boundaries, though descriptions clarify preferred usage.
The set is predominantly snake_case and verb_noun (close_page, list_pages, take_snapshot). A minority use noun-first or noun-verb ordering (browser_status, state_save, performance_start_trace), which is a minor deviation but still readable.
36 tools is high, but the domain (full browser automation and DevTools) legitimately requires separate tools for lifecycle, pages, input, inspection, performance, network, console, and state. It is slightly over the typical 3-15 range but each tool earns its place.
The surface covers browser lifecycle, navigation, page/tab management, input, a11y snapshots, screenshots, console, network, CSS, performance tracing, Lighthouse, memory heap, dialogs, file upload, emulation, and login states. No obvious lifecycle gaps for a DevTools automation server.
Maintenance
Related MCP Connectors
- TabfleetOAuthcom.tabfleet
Launch, inspect, control, and share isolated cloud browsers for your agents.
Undetectable cloud browser sessions for AI agents and scrapers. Navigate, extract, click, captcha.
Run multi-step tasks in a real Chrome browser: persistent environments, live view, human takeover.
Hosted browser for AI agents: screenshots, post-JS DOM, console, WCAG. No install, no API key.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI coding assistants to control and inspect a live Chrome browser through DevTools for automated testing, performance analysis, debugging, and web scraping. Provides reliable browser automation using Puppeteer with comprehensive DevTools access.1,798,972 npm3Apache 2.0
- AlicenseNot gradedqualityDmaintenanceEnables AI coding assistants to control and inspect a live Chrome browser for automation, debugging, performance analysis, network monitoring, and DOM interaction through Chrome DevTools Protocol.1,798,972 npmApache 2.0
- AlicenseBqualityAmaintenanceControls a real Chrome browser for AI agents, enabling authenticated automation with parallel lanes, token-efficient page reads, and robust recovery mechanisms.1221,036 npm237MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI coding agents to control and inspect a live Chrome browser, providing reliable automation, in-depth debugging, and performance analysis through Chrome DevTools.1,798,972 npmApache 2.0