Skip to main content
Glama

Give every one of your AI agents its own Chrome, running Google's full DevTools toolset, already logged in, and fenced to the sites it should touch.

Why This Exists

Google's chrome-devtools-mcp is the best browser toolset an agent can have when the job is debugging a web app: console, network, performance traces, Lighthouse, heap snapshots. It stops working the moment you run more than one agent:

  • Agent two starts → The browser is already running for …, because every server shares one Chrome profile

  • Switch on --isolated → every agent starts logged out and fights the wp-admin, staging or SSO login again, every time

  • The MCP connection drops → the browser, its tabs and its login die with it

  • Ten agents are running → ten invisible headless Chromes. Nothing lists them, nothing cleans up after a crash

Other multi-agent browser tools solve pooling and logins, but on Playwright or through a browser extension. None of them carries Google's DevTools toolset.

Related MCP server: Chrome DevTools MCP

The Solution

devtools-fleet-mcp sits in front of chrome-devtools-mcp and manages the browsers. It doesn't reimplement a single browser tool: Google's tools are passed through unchanged, under the same names.

agent session A ─stdio─▶ devtools-fleet-mcp ─▶ chrome-devtools-mcp --browserUrl ─CDP─▶ Chrome A (own profile)
agent session B ─stdio─▶ devtools-fleet-mcp ─▶ chrome-devtools-mcp --browserUrl ─CDP─▶ Chrome B (own profile)
                               │
                               └──▶ ~/.devtools-fleet/  ◀── devtools-fleet CLI: ls, show, kill, gc, login, states
  1. One browser per agent, zero config. The first browser tool call launches that session's own Chrome. Up to 10 at once by default.

  2. Browsers outlive the connection. An agent that reconnects inside the same session gets its browser back, tabs open. A small background reaper closes the browsers nobody comes back for.

  3. Saved logins. You log in once, by hand, in a real window (2FA, SSO and captchas are fine). Any number of agents can then start from that login in parallel, each in its own browser.

  4. Origin allowlists. A browser started from a saved login can only reach that login's origins. Typed URLs, link clicks, redirects and popups elsewhere are blocked. Strict mode locks the whole network in Chrome itself.

  5. Recovery. If Chrome or chrome-devtools-mcp dies, the next tool call brings it back and tells the agent what happened.

  6. Watch any agent live. devtools-fleet show <id> opens Chrome's DevTools inspector on that agent's page without disturbing it.

Quick Start

Requires Node.js 22.12+ and Google Chrome (or point chromePath at another Chromium build).

The plugin installs the MCP server and a skill that teaches agents the fleet workflow: check state_list, start from a saved login, respect the allowlist, recover after a crash.

Inside Claude Code:

/plugin marketplace add pluginslab/devtools-fleet-mcp
/plugin install devtools-fleet@pluginslab-devtools-fleet

Or from your terminal:

claude plugin marketplace add pluginslab/devtools-fleet-mcp
claude plugin install devtools-fleet@pluginslab-devtools-fleet

Restart Claude Code and check that devtools-fleet shows up under /mcp.

Already using chrome-devtools-mcp? Remove it (claude mcp remove chrome-devtools, or disable its plugin). devtools-fleet exposes the same tools under the same names, and two servers with identical tool names confuse agents.

Updating: /plugin marketplace update pluginslab-devtools-fleet, then restart. The server runs through npx, so it picks up new npm releases on its own.

Install the CLI

The plugin gives agents the MCP tools. Saving logins, watching agents and cleaning up are done by you, from the CLI:

npm install -g devtools-fleet-mcp
devtools-fleet doctor

No global install? Every command also works as npx -p devtools-fleet-mcp devtools-fleet <command>.

Other MCP clients

Claude Code without the plugin:

claude mcp add devtools-fleet -- npx -y devtools-fleet-mcp

Any other client (.mcp.json, Cursor, Claude Desktop, …):

{
  "mcpServers": {
    "devtools-fleet": {
      "command": "npx",
      "args": ["-y", "devtools-fleet-mcp"]
    }
  }
}

First test

Open two Claude Code sessions and ask both:

"Open http://localhost:3000 and check the console for errors."

Each gets its own Chrome; neither hits the profile lock. Then, in a terminal:

devtools-fleet ls

First saved login

devtools-fleet login staging-admin https://staging.example.com/wp-login.php

Log in in the window that opens, come back to the terminal, press Enter, confirm the allowlist. Then ask an agent:

"Using the staging-admin state, open the Plugins page and run a Lighthouse audit on it."

The agent calls state_list, then browser_start({ state: "staging-admin" }), and starts already logged in.

MCP Tools

devtools-fleet adds six tools. Everything else (navigate_page, take_snapshot, click, evaluate_script, list_network_requests, performance_start_trace, lighthouse_audit, …) is chrome-devtools-mcp's, with the same names and arguments, so its documentation and skills apply as-is.

browser_start

Starts this session's browser, or picks it up again after a reconnect. Optional: any browser tool starts one with defaults.

→ browser_start({ state: "staging-admin", url: "https://staging.example.com/wp-admin/" })
← Started: browser 3f9c2a, state "staging-admin", headless, stable
  allowed origins: https://staging.example.com
  tabs: https://staging.example.com/wp-admin/
  uptime 2s

Parameter

Type

Description

state

string

Saved login to start from

headless

boolean

Default from config (true)

viewport

string

e.g. "1280x720"

channel

string

stable, beta, dev or canary

url

string

Open this in the first tab

allowedOrigins

string[]

Without a state: restrict this browser voluntarily

browser_status

This session's browser: id, state, allowlist, mode, open tabs, uptime, blocked requests.

→ browser_status()
← browser 3f9c2a, state "staging-admin", headless, stable
  allowed origins: https://staging.example.com
  tabs: https://staging.example.com/wp-admin/plugins.php
  uptime 312s, 2 request(s) blocked so far

browser_restart

Relaunches keeping cookies, storage and tabs. { headless: false } gives a visible window, for example so a person can solve a captcha.

→ browser_restart({ headless: false })
← Restarted with a visible window; reopened 2 tab(s). Page ids have changed: call list_pages.

browser_stop

Closes the browser and deletes its temporary profile.

→ browser_stop()
← Closed browser 3f9c2a.

state_save

Saves the current browser's login under a name. Parameters: name, allowedOrigins (defaults to the open tabs' origins), strict, overwrite. An agent can only narrow its own allowlist, never widen it, and never overwrites a state a person created.

→ state_save({ name: "local-shop" })
← Saved state "local-shop": 4 cookie(s), localStorage for 1 origin(s). Allowed: http://localhost:3000.

state_list

Saved states with origins, cookie counts and expiry. Never values.

→ state_list()
← staging-admin: https://staging.example.com | 6 cookies (0 expired) | saved 2026-10-01T09:12:44.512Z by cli
  local-shop: http://localhost:3000 | 4 cookies (0 expired) | saved 2026-10-01T10:03:10.087Z by agent

CLI Reference

# See every fleet browser: id, status, session, state, mode, current tab
devtools-fleet ls
devtools-fleet ls --json

# Watch an agent's browser live in Chrome's DevTools inspector (doesn't disturb it)
devtools-fleet show 3f9c2a
devtools-fleet show 3f9c2a --tab 2 --print

# Close browsers
devtools-fleet kill 3f9c2a
devtools-fleet kill --all

# Close browsers whose session is gone; --detached also closes ones waiting for a reconnect
devtools-fleet gc
devtools-fleet gc --detached

# Log in by hand and save a state
devtools-fleet login staging-admin https://staging.example.com/wp-login.php
devtools-fleet login shop http://localhost:3000/login --allow https://cdn.example.com --strict

# Manage states (never prints cookie or storage values)
devtools-fleet states
devtools-fleet state show staging-admin
devtools-fleet state rm staging-admin
devtools-fleet state import shop ./storageState.json --allow http://localhost:3000

# Check Node, Chrome, permissions and config; print effective config
devtools-fleet doctor
devtools-fleet config

Browser statuses: active (an agent is connected), detached (the connection dropped; waiting for that session to reconnect), orphan (the session is gone; closed on the next cleanup), starting.

How It Works

Process Management

On the first browser tool call, devtools-fleet:

  1. Checks the cap under a machine-wide lock (maxBrowsers, default 10). A full fleet gets a clear error naming the running sessions and devtools-fleet gc

  2. Launches Chrome itself, on a port it picks, with its own temporary profile

  3. Restores the state, if one was asked for: cookies over CDP, localStorage through a throwaway tab whose request devtools-fleet answers itself, so the site is never contacted

  4. Spawns chrome-devtools-mcp from its own pinned node_modules with --browserUrl, and re-exports its tool list unchanged

  5. Registers the browser in ~/.devtools-fleet/browsers/<id>.json

When the connection drops, Chrome keeps running. When the same client session reconnects, it finds its browser in the registry and re-adopts it, tabs intact. A reaper process (one per machine, started on demand, gone when idle) closes orphans at once and detached browsers after orphanTimeoutMinutes.

How Sessions Are Recognised

MCP clients start servers through wrappers (npx, npm exec, node, shells) and restart that whole chain on reconnect, while the client session itself (the claude process, an editor's extension host) keeps running. devtools-fleet walks up the process tree past the wrappers to that session process. The same session reconnecting finds the same process and gets its browser back; a second session in the same directory finds a different one and never shares.

DEVTOOLS_FLEET_SESSION=<label> replaces the lookup with a fixed label, useful in CI or containers.

Saved Logins

devtools-fleet login opens a real Chrome window. You log in, come back to the terminal, press Enter. devtools-fleet proposes an allowlist of where you started and where you ended up. Origins you only passed through, such as an SSO provider, are listed but left out: saving their cookies would hand agents your whole SSO session. Add one with --allow if the app really needs it.

What gets saved: cookies and localStorage for the allowed origins, in Playwright's storageState format. Playwright loads these files directly, and state import takes Playwright or agent-browser files the other way.

Rules that keep states safe to hand to an agent:

  • States are only written outside git work trees, with mode 0600 in a 0700 directory.

  • No listing, CLI or MCP, ever shows cookie or storage values.

  • A state is always bound to an allowlist. There is no "use these cookies anywhere" mode.

The Allowlist

Allowlist entries are full origins:

Entry

Matches

https://app.example.com

exactly that scheme, host and port

http://127.0.0.1:3000

ports matter: :3001 is a different origin

https://*.example.com

any subdomain, not example.com itself

Two layers enforce it:

  1. Argument check. new_page and navigate_page are checked before anything reaches the browser; the agent gets a plain error.

  2. Navigation interception. devtools-fleet's own CDP connection intercepts navigations in every tab, popup and frame and fails the ones outside the list. A tab that can't be put under interception is closed rather than left unguarded.

A browser with an allowlist also refuses file: URLs and extra browser contexts (new_page's isolatedContext), since pages there would sit outside the guard.

Strict mode (login --strict) is enforced by Chrome itself: the browser launches with a proxy that goes nowhere, and only the allowed origins bypass it. Every other request fails in Chrome's network stack: fetch(), XHR, beacons, workers, WebSockets, QUIC. WebRTC is pinned to the dead proxy and DNS prefetching is off. The lock holds even while no agent is connected. It's off by default because it also blocks CDNs you haven't listed. At that layer an entry means host and port, so http and https on the same port aren't told apart there; the navigation guard still tells them apart.

What the allowlist does not do:

  • It limits where the agent goes, not what it does there. On an allowed origin, evaluate_script can read anything the page can, including non-httpOnly cookies. Treat a saved login like handing someone your session.

  • Without strict mode, only navigations are guarded, and only while an agent is connected. A detached browser can still be navigated away by a script already running in its pages.

Files and Paths

chrome-devtools-mcp only reads and writes files (screenshots, traces, uploads) inside the client's workspace roots and the temp directory. devtools-fleet passes your client's roots through, so take_screenshot({ filePath: "<project>/shot.png" }) works as usual. Nothing may point into ~/.devtools-fleet: not a filePath, not an upload_file, not a file: URL, symlinks included.

Data Storage

~/.devtools-fleet/
  config.json          # Optional configuration (see below)
  browsers/
    <id>.json          # One registry entry per running browser
  profiles/            # Temporary Chrome profiles, deleted on close
  states/              # Saved logins (0700 dir, 0600 files)
    <name>.json
  locks/               # mkdir locks for the cap check and adoption

DEVTOOLS_FLEET_HOME moves the whole directory.

Configuration

~/.devtools-fleet/config.json, every key optional. Environment variables override the file.

Key

Env

Default

headless

DEVTOOLS_FLEET_HEADLESS

true

viewport

DEVTOOLS_FLEET_VIEWPORT

none

e.g. "1280x720"

channel

DEVTOOLS_FLEET_CHANNEL

stable

chromePath

DEVTOOLS_FLEET_CHROME_PATH

detected

maxBrowsers

DEVTOOLS_FLEET_MAX_BROWSERS

10

across the machine

orphanTimeoutMinutes

DEVTOOLS_FLEET_ORPHAN_TIMEOUT_MINUTES

15

how long a dropped connection's browser waits for a reconnect

launchTimeoutSeconds

DEVTOOLS_FLEET_LAUNCH_TIMEOUT_SECONDS

30

upstreamArgs

[]

extra chrome-devtools-mcp flags

chromeArgs

[]

extra Chrome flags, e.g. ["--no-sandbox"] in containers

chrome-devtools-mcp collects usage statistics by default and may send performance trace URLs to Google's CrUX API. devtools-fleet keeps its defaults. Opt out with "upstreamArgs": ["--no-usage-statistics", "--no-performance-crux"].

Limits

  • Tested on macOS and Linux. Windows should work apart from reconnect re-adoption, but is untested.

  • Saved logins cover cookies and localStorage. Not IndexedDB, sessionStorage or service worker caches.

  • Sites with device-bound sessions (Device Bound Session Credentials) won't accept copied cookies.

  • Google accounts are not supported. Google blocks automated sign-in and binds sessions to devices.

  • Headless can't switch to a window in place. browser_restart relaunches, and the pages reload.

  • The debugging port is local-only but not authenticated, the same as chrome-devtools-mcp and every CDP tool: other processes running as your user can drive these browsers.

Companion Tools

devtools-fleet is the look step for the PluginsLab WordPress MCPs:

MCP

Purpose

wp-devdocs-mcp

Verified hooks/filters/APIs for writing plugin code

wp-blockmarkup-mcp

Verified block schemas for generating content

wp-playground-mcp

Ephemeral WordPress instances for testing

devtools-fleet-mcp (this)

One DevTools-equipped Chrome per agent for looking at the result

Together: author → validate → test → look.

Requirements

  • Node.js 22.12+

  • Google Chrome (stable, beta, dev or canary), or any Chromium via chromePath

  • macOS or Linux

Development

npm install
npm test                 # unit tests, no browser
npm run test:integration # launches real Chrome; ~1 minute
npm run lint

License

MIT

Available Tools

36 tools
browser_restartA

Relaunch this session's browser keeping cookies, localStorage and open tabs. Pass headless: false to get a visible window (e.g. so the human can watch or solve a captcha), headless: true to hide it again. Pages reload.

ParametersJSON Schema
NameRequiredDescriptionDefault
headlessNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses what survives (cookies, localStorage, open tabs) and one side effect ('Pages reload'), but omits what is lost (in-memory JS state, scroll position), whether it errors when no browser is running, and how long the relaunch takes. Partial disclosure, not a rich profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and preservation guarantee, then the parameter semantics, then the reload side effect. Every sentence earns its place and nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with no output schema, the description covers the action, preserved state, parameter semantics, and reload behavior. The remaining gap is the default value/behavior of the optional headless parameter when omitted, which an agent would want to know.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single parameter name ('headless') alone is ambiguous, but the description explains both boolean values with concrete intent ('visible window … headless: true to hide it again'), which is exactly the meaning the schema lacks. It does not state the default when the optional parameter is omitted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Relaunch this session's browser') plus the scope that separates it from a fresh browser_start: cookies, localStorage and open tabs are preserved. It doesn't name the sibling it is contrasted with, but the session-scoped framing makes the distinction inferable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a concrete use condition for one parameter value (headless: false so a human can watch or solve a captcha), but never says when to reach for browser_restart rather than browser_start, browser_stop, or a plain reload, nor what prerequisites exist (e.g. an active session). Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_startA

Start this session's own Chrome (isolated from every other agent). Optional: the first browser tool call starts one with defaults anyway. Call it explicitly to load a saved login (state), to pick headless/viewport/channel, or to open a URL. A browser started from a state can only visit that state's allowed origins. If this session already had a browser before a reconnect, it is picked up again with its tabs.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoOpen this URL in the first tab
stateNoName of a saved login state (see state_list)
channelNo
headlessNoRun without a window (default from config, normally true)
viewportNoWindow size, e.g. "1280x720"
allowedOriginsNoWithout a state: restrict this browser to these origins, e.g. ["http://localhost:3000"]

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses per-agent isolation, that a state-bound browser is restricted to that state's allowed origins, and that a browser surviving a reconnect is re-adopted with its tabs. It doesn't cover failure modes or what happens if a browser is already running with conflicting settings, leaving a small gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the optionality caveat, then the explicit reasons to call it. Dense but every sentence carries information; the reconnect sentence is slightly tangential but useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter, all-optional tool with no output schema and no annotations, the description covers identity, isolation, origin restrictions, and reconnect behavior well enough to call it correctly. Return-shape details and behavior on a second call are the only notable omissions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 83%, but the description adds real meaning: it groups the tuning parameters (headless/viewport/channel), explains the state parameter's origin-restriction consequence, and cross-references state_list. It still doesn't clarify the distinction/additivity of allowedOrigins versus a state's own allowed origins.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Start this session's own Chrome') and immediately scopes it as isolated from every other agent, which cleanly separates it from siblings like browser_status, browser_restart and browser_stop. An agent can identify the tool's action and its unique session-scoped nature without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says the call is optional because the first browser tool call starts one with defaults, then names the exact reasons to invoke it eagerly: load a saved login state, choose headless/viewport/channel, or open a URL. This is genuine when-to-use/when-not-to-use guidance with conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_statusA

Show this session's browser: id, saved state, allowlist, headless, open tabs, uptime.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose the concrete state dimensions returned, which signals this is a non-mutating introspection tool, but it never states that it is read-only, that it requires an active session, or what happens if no browser is running. Adequate but incomplete for an unannotated tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that leads with the verb and subject and then lists the returned fields compactly. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the returned state fields, which is the main thing an agent needs. For a zero-parameter, zero-risk status tool this is close to complete; only the precondition of an active browser session is unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameter semantics to document; the schema coverage is 100% and there is nothing for the description to compensate for. Baseline 4 applies for a 0-param schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Show) plus the resource (this session's browser) and enumerates the exact fields returned (id, saved state, allowlist, headless, open tabs, uptime). That distinguishes it from sibling read tools like list_pages and state_list. It stops short of explicitly naming a sibling it is not, so a 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no routing to alternatives. The phrase 'this session's browser' implies an inspection context, but the agent must infer that it precedes actions like browser_start or close_page. No exclusions or alternative conditions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_stopA

Close this session's browser and delete its temporary profile.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose a genuinely useful destructive side effect — the temporary profile is deleted — but says nothing about what happens to open pages, whether the profile deletion is irreversible, whether it errors if no session is running, or whether a subsequent browser_start is required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the primary action and the secondary side effect. No filler, no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter teardown tool with no output schema, the description covers the action and its cleanup effect adequately. It could be stronger by noting the fate of open pages and whether a restart is needed afterward, but nothing essential to invoking it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema is empty and the baseline is 4. There is nothing further the description could add on parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (close the browser) plus the notable side effect (delete the temporary profile). The scope word "this session's" distinguishes it from the sibling close_page, which closes a single page rather than the whole browser. It doesn't explicitly name an alternative, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use guidance. The agent must infer that this is a teardown call for ending a browser session, and there is no statement about prerequisites (e.g. a running browser) or when browser_restart would be preferable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clickC

Clicks on the provided element

ParametersJSON Schema
NameRequiredDescriptionDefault
uidYesThe uid of an element on the page from the page content snapshot
pageIdYesTargets a specific page by ID.
dblClickNoSet to true for double clicks. Default is false.
includeSnapshotNoWhether to include a snapshot in the response. Default is false.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint=false annotation already signals a mutating action. The description adds nothing beyond that: it does not disclose side effects such as navigation, form submission, focus changes, or potential waits triggered by the click, which matter greatly for a browser-automation click.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler or redundancy. It is efficient, though its brevity partly reflects under-specification rather than disciplined concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a browser click with a non-read-only annotation and no output schema, the description is too thin: it omits side-effect behavior and how includeSnapshot affects the response, leaving an agent unable to anticipate consequences of the call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so uid, pageId, dblClick, and includeSnapshot are all documented in the schema. The phrase 'the provided element' adds no syntax or format detail beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (clicks) and resource (the provided element), so an agent can distinguish it from siblings like hover, drag, or fill. It does not, however, explicitly contrast itself with those siblings, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus hover, drag, type_text, or press_key, and no preconditions or exclusions stated. The agent must infer usage entirely from the name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

close_pageA

Closes the page by its index. The last open page cannot be closed.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageIdYesThe ID of the page to close. Call list_pages to list pages.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply readOnlyHint=false, so the description carries most of the burden and does add real behavioral value by disclosing the failure case that the last open page cannot be closed. It stops short of saying whether page state is destroyed or recoverable, which would be useful for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, purpose front-loaded, constraint second. Nothing extraneous and nothing buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema and only a readOnlyHint annotation, the description covers purpose and the key failure condition adequately. It could note the effect on browser state after closing, but that is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already explains pageId fully, including the hint to call list_pages, so the baseline is 3. The description's 'by its index' wording adds no detail and mildly conflicts with the schema's 'ID' framing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Closes the page') that clearly distinguishes it from siblings like new_page, select_page, and navigate_page. The qualifier 'by its index' is slightly at odds with the schema's pageId terminology, but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool's use (closing an open page) and adds one important constraint — the last open page cannot be closed — but never states when to prefer this over alternatives or what prerequisites exist. Usage is inferable rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dragC

Drag an element onto another element

ParametersJSON Schema
NameRequiredDescriptionDefault
pageIdYesTargets a specific page by ID.
to_uidYesThe uid of the element to drop into
from_uidYesThe uid of the element to drag
includeSnapshotNoWhether to include a snapshot in the response. Default is false.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint=false annotation already signals this is a mutating operation, but the description adds nothing beyond that: it does not explain what dragging actually does to the page (DOM reordering, moving an element into a container), whether it waits for completion, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler, and the core action is front-loaded. It is efficient, though arguably terse enough that it under-specifies the operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description should at minimum convey the effect of the drag operation and whether it involves waiting or verification. The schema covers parameters fully, but the behavioral context is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so pageId, from_uid, to_uid, and includeSnapshot are all documented in the schema. The description adds no syntax or format detail beyond the schema, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (drag) and both resources (element onto another element), which cleanly distinguishes it from sibling tools like click, hover, and fill. It stops short of naming the uid-based identity of the elements, but the action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use drag versus alternatives such as click or hover, and no prerequisites mentioned. The agent must infer that this is for drag-and-drop interactions from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

emulateC

Emulates various features on the target page.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageIdYesTargets a specific page by ID.
viewportNoEmulate device viewports '<width>x<height>x<devicePixelRatio>[,mobile][,touch][,landscape]'. 'touch' and 'mobile' to emulate mobile devices. 'landscape' to emulate landscape mode.
userAgentNoUser agent to emulate. Set to empty string to clear the user agent override.
colorSchemeNoEmulate the dark or the light mode. Set to "auto" to reset to the default.
geolocationNoGeolocation (`<latitude>,<longitude>`) to emulate. Latitude between -90 and 90. Longitude between -180 and 180. Omit to clear the geolocation override.
extraHttpHeadersNoExtra HTTP headers as a JSON string object, e.g. {"X-Custom": "value", "Authorization": "Bearer token"}. Headers are included into every HTTP request originating from the page and persist across navigations until cleared. Pass an empty string to clear all extra headers.
cpuThrottlingRateNoRepresents the CPU slowdown factor. Omit or set the rate to 1 to disable throttling
networkConditionsNoThrottle network. Omit to disable throttling.

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, so this mutates page state, but the description says nothing about persistence, scope, or reversibility. It does not repeat or contradict the annotation, but it also adds no behavioral context beyond what the schema itself hints at.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence, so it is front-loaded and free of waste, but its brevity comes from under-specification rather than discipline. Structure is acceptable; content is thin.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter tool that persistently overrides page behavior across navigations, the description leaves the agent without any sense of scope or side effects. With no output schema and only a null readOnlyHint annotation, the description should carry more of the load.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the parameter docs are unusually rich (clearing overrides, geolocation ranges, header format), so the baseline of 3 applies. The description contributes nothing to parameter meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name almost literally ('emulate' → 'Emulates various features'), which is close to a tautology. It never names which features (viewport, user agent, geolocation, network throttling) despite the schema enumerating them, so an agent cannot distinguish it from siblings like resize_page or navigate_page without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites (does a page need to exist first?), and no routing to or away from any sibling tool. The agent must infer the entire usage context from the parameter list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evaluate_scriptB

Evaluate a JavaScript function inside the target page. Returns the response as JSON, so returned values have to be JSON-serializable.

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNoAn optional list of arguments to pass to the function.
pageIdYesTargets a specific page by ID.
filePathNoThe absolute or relative path to a file to save the script output to. If omitted, the output is returned inline.
functionYesA JavaScript function declaration to be executed by the tool in the target page. Example without arguments: `() => document.title` or `async () => await fetch("example.com")`. Example with arguments: `(el) => el.innerText`
dialogActionNoHandle dialogs while execution. "accept", "dismiss", or string for response of window.prompt. Defaults to accept.
waitForStableDomNoWhether to wait for the DOM to settle. Pass false if the script only reads data. Defaults to true.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=false correctly signals this is not a safe read-only operation, and the description's note that arbitrary JS runs in the page is consistent with that. It adds the useful constraint that return values must be JSON-serializable, but says nothing about side effects on the page, error behavior, sandboxing/security boundaries, or that dialogs are auto-accepted by default.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core purpose followed by the return-value constraint. No filler, though the JSON note could be folded more tightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description reasonably explains that output is JSON and can be saved via filePath (per schema). However, for a tool that executes arbitrary code in a browser page, the description omits side-effect/security context and error handling that an agent would want before invoking it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so pageId, function, args, filePath, dialogAction, and waitForStableDom are all documented in the schema itself. The description adds no parameter-level meaning beyond the JSON-serializable constraint, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: evaluate a JavaScript function inside the target page, and notes the result is returned as JSON. This distinguishes it from browsing siblings like click, fill, or take_snapshot. It does not, however, contrast itself against the other data-reading tools in the family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance, and no alternatives are named (e.g. use click/fill for element interaction, take_snapshot for page content). The JSON-serializability note is a constraint on the return value, not selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fillC

Type text into an input, text area or select an option from a element.

ParametersJSON Schema
NameRequiredDescriptionDefault
uidYesThe uid of an element on the page from the page content snapshot
valueYesThe value to fill in. "true" or "false" for checkboxes and toggles, "true" for radio buttons.
pageIdYesTargets a specific page by ID.
includeSnapshotNoWhether to include a snapshot in the response. Default is false.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, implying mutation, but the description says nothing about whether the existing value is cleared or appended, whether input/change events fire, or that it requires a valid uid from a live page. The one useful behavioral nuance (select options) is also covered by the schema's value description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the action front-loaded and no filler. It is efficient, though the brevity comes at the cost of the usage and behavioral detail an agent would benefit from.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating DOM interaction tool with no output schema and minimal annotations, the description omits prerequisites (element discovery via snapshot), what happens to pre-existing values, and the effect on the page. Important context for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters including the checkbox/radio 'true'/'false' convention and includeSnapshot are documented in the schema. The description adds no parameter detail beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete action (type text / select an option) and names the target resources (input, text area, <select>). However, it does not distinguish itself from close siblings like fill_form and type_text, which an agent must choose between for similar tasks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this instead of type_text or fill_form, nor any note that the target element must first be located via a page snapshot. The agent is left to infer the selection criteria entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fill_formA

Fill out multiple form elements (inputs, selects, checkboxes, radios) at once. ALWAYS prefer this tool over multiple individual 'fill' or 'click' calls when interacting with forms. It is significantly faster, more reliable, and reduces turn count. Example: Fill username, password, and check "Remember Me" in one call.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageIdYesTargets a specific page by ID.
elementsYesElements from snapshot to fill out.
includeSnapshotNoWhether to include a snapshot in the response. Default is false.

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, consistent with a form-filling mutation. The description adds useful context (faster, more reliable, reduces turn count) and the batch semantics, but discloses nothing about failure behavior, atomicity, or what happens when a uid is invalid. Adds some value but not rich behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the preference rule, followed by a justified example. Slightly promotional ('significantly faster, more reliable') but each sentence still earns its place by steering tool selection.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description needn't explain return values, and the annotations are minimal enough that the description covers the mutation nature and batching benefit. It is complete enough to invoke correctly, though it omits error/atomicity behavior for a multi-element write.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents pageId, elements, uid, value, and includeSnapshot, including the 'true'/'false' conventions for checkboxes and radios. The description adds the example semantics but no new parameter meaning beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Fill out) and resource (multiple form elements, enumerating inputs, selects, checkboxes, radios) in one call. It clearly distinguishes itself from the sibling tools 'fill' and 'click' by emphasizing the batch nature.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states 'ALWAYS prefer this tool over multiple individual fill or click calls when interacting with forms,' giving a direct when-to-use rule and naming the alternatives. A concrete example (username, password, Remember Me) further grounds the guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_console_messageA
Read-only

Gets a console message by its ID. You can get all messages by calling list_console_messages.

ParametersJSON Schema
NameRequiredDescriptionDefault
msgidYesThe msgid of a console message on the page from the listed console messages
pageIdYesTargets a specific page by ID.

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations supply readOnlyHint=true, so the safety profile is already covered. The description adds essentially nothing behavioral beyond that: no return shape, no behavior for an unknown/out-of-range msgid, and no note that the ID must originate from a prior list call. Minimal added context over the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, purpose front-loaded, with the cross-reference to the sibling tool in the second sentence. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-param read tool with annotations and full schema coverage, the description is just barely sufficient. It omits what a message object contains and how failures (invalid ID) behave, which a slightly richer definition could have supplied without bloat.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both pageId and msgid are documented in the schema itself. The description adds no format, range, or sourcing detail beyond what the schema already says, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (gets) and resource (a console message) plus the identifier used (by its ID). It also names the sibling list_console_messages as the way to obtain all messages, so the agent can infer this is the single-message lookup, though the distinction is implicit rather than spelled out.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Points the agent to list_console_messages when it needs all messages, which effectively scopes when this tool applies (when the ID is already known). There is no explicit exclusion or precondition statement (e.g. what happens with an invalid msgid), so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_css_stylesA
Read-only

Retrieve matched CSS rules, inline styles, inherited styles, and cascade information for an element identified by its UID. Use this tool to debug why specific CSS properties are applied, overridden, or conflicting. Results are paginated and return 10 rules per page by default; use pageIdx to page through the remaining rules. Requires a UID from take_snapshot.

ParametersJSON Schema
NameRequiredDescriptionDefault
uidYesThe uid of the element on the page from the page content snapshot to inspect CSS styles for
pageIdYesTargets a specific page by ID.
pageIdxNoPage number to return (0-based). Defaults to 0 (the first page).
pageSizeNoMaximum number of CSS rules to return per page. Defaults to 10.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered, and the description adds genuinely useful behavior: results are paginated at 10 rules per page and pageIdx advances through them, plus a hard dependency on take_snapshot for the UID. It stops short of describing ordering or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, purpose front-loaded, then usage, then pagination logistics. Every sentence carries information, though the pagination sentence could be marginally tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the work of naming the return content (matched, inline, inherited rules and cascade info) and pagination behavior, which is what an agent needs to call and interpret it. Missing only edge-case/error behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents uid, pageId, pageIdx, and pageSize. The description's mention of pagination and default page size largely restates schema defaults, adding only the practical hint to advance with pageIdx, which is the baseline 3 expectation when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Retrieve) and a well-scoped resource (matched CSS rules, inline, inherited styles, cascade info for an element), and it is clearly distinguishable from sibling tools like take_snapshot or evaluate_script. An agent can tell what this returns without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames the use case (debug why CSS properties are applied, overridden, or conflicting) and states the prerequisite (a UID from take_snapshot). It does not name an alternative tool or state when not to use it, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_network_requestB

Gets a network request by an optional reqid, if omitted returns the currently selected request in the DevTools Network panel. Useful for inspecting request headers (including 'Cookie') and response headers (including 'Set-Cookie' and directives).

ParametersJSON Schema
NameRequiredDescriptionDefault
reqidNoThe reqid of the network request. If omitted returns the currently selected request in the DevTools Network panel.
pageIdYesTargets a specific page by ID.
requestFilePathNoThe absolute or relative path to a .network-request file to save the request body to. If omitted, the body is returned inline.
responseFilePathNoThe absolute or relative path to a .network-response file to save the response body to. If omitted, the body is returned inline.

TDQS

B3.4/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description implies a read-only operation ('Gets', 'returns') but annotation readOnlyHint=false indicates write side effects. The schema shows file-saving parameters that write request/response bodies to disk, which the description never discloses. This is a direct contradiction of the implied read-only nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core operation and default behavior, then a useful access note. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Omits the file-saving parameters (requestFilePath, responseFilePath) and required pageId. For a tool with four parameters and no output schema, the description is incomplete about how the tool can be used and what side effects it has.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, fully documenting all four parameters. The description only repeats the reqid default already in the schema and adds no meaning beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb 'Gets' and resource 'network request', with optional reqid behavior. Clearly distinguishes from sibling list_network_requests by fetching a single request.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a clear use case: inspecting request and response headers including cookies. Does not name alternatives or specify when not to use, but context is sufficient for a simple getter.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

handle_dialogC

If a browser dialog was opened, use this command to handle it

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesWhether to dismiss or accept the dialog
pageIdYesTargets a specific page by ID.
promptTextNoOptional prompt text to enter into the dialog.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations set readOnlyHint=false, marking this as a mutating operation, and the description does not contradict that. But it discloses nothing beyond the annotation: it never says dialogs block page execution, that accept vs dismiss produce different outcomes, or what entering promptText does to the page state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the condition comes before the instruction. It is efficient, though its brevity is partly under-specification rather than deliberate economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter tool with full schema coverage and no output schema, the definition is minimally adequate. It omits useful context such as how to discover an open dialog, that only one dialog can be handled at a time, and what happens when no dialog is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (action, pageId, promptText) are documented in the schema itself. The description adds no parameter meaning beyond that, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (browser dialog) and a condition for acting, but 'use this command to handle it' largely restates the tool name rather than specifying what handling means. It gives no detail on the accept/dismiss outcome, though no sibling handles dialogs so differentiation is not the issue.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides a clear trigger condition ('If a browser dialog was opened'), which is genuine when-to-use guidance. However, it offers no when-not guidance, no mention of validation steps (e.g., checking via snapshot) and no alternatives or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hoverC

Hover over the provided element

ParametersJSON Schema
NameRequiredDescriptionDefault
uidYesThe uid of an element on the page from the page content snapshot
pageIdYesTargets a specific page by ID.
includeSnapshotNoWhether to include a snapshot in the response. Default is false.

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations declare readOnlyHint=false, signalling this is a state-affecting interaction, but the description adds nothing beyond that: it does not say whether hovering triggers tooltips/menus, whether it waits for a response, or what side effects occur. With only a single annotation, the description is expected to carry more of the behavioral burden and does not.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence with no wasted words and is front-loaded, but it is terse to the point of under-specification rather than economical, restating the tool name with almost no added information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an element-interaction tool with no output schema, the description omits whether the call waits/blocks, what hover effects to expect, and how includeSnapshot affects the response. Given annotations cover only readOnlyHint, the description leaves key behavioral questions unanswered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so uid, pageId, and includeSnapshot are already documented in the schema. The description adds no meaning beyond the schema, which is the expected baseline of 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (hover) and its target (the element referenced by the uid param), so an agent knows the basic action. It does not differentiate this from siblings like click, drag, or fill, and 'the provided element' is mildly circular since the element is identified by the parameter.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use hover versus click, fill, drag, or other element-interaction siblings, and no prerequisites or exclusions. The agent must infer usage entirely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lighthouse_auditA

Get Lighthouse score and reports for accessibility, SEO, best practices, and agentic browsing. This excludes performance. For performance audits, run performance_start_trace

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo"navigation" reloads & audits. "snapshot" analyzes current state.navigation
deviceNoDevice to emulate.desktop
pageIdYesTargets a specific page by ID.
outputDirPathNoDirectory for reports. If omitted, uses temporary files.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotation readOnlyHint=false is consistent with the fact that audits produce report artifacts, so there is no contradiction. The description does communicate what is excluded, but it says nothing about report generation side effects, run duration, or how outputDirPath affects behavior, so it adds only modest context beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, zero filler, with the capability scope front-loaded before the exclusion and the alternative-tool routing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read/analysis tool with no output schema and a fully documented parameter set, the description covers scope, exclusions, and the alternative tool adequately. Minor gaps remain around what the returned score/report looks like and run-time behavior, but nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so mode, device, pageId, and outputDirPath are already fully documented in the schema (including enum semantics). The description adds no parameter-level information, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get') and resource ('Lighthouse score and reports') and enumerates the covered categories (accessibility, SEO, best practices, agentic browsing). It also states what is explicitly out of scope (performance) and names the sibling tool that handles it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit exclusion and an explicit routing rule: 'This excludes performance. For performance audits, run performance_start_trace.' An agent can select between this tool and performance_start_trace without inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_console_messagesB
Read-only

List all console messages for the target page since the last navigation.

ParametersJSON Schema
NameRequiredDescriptionDefault
typesNoFilter messages to only return messages of the specified resource types. When omitted or empty, returns all messages.
pageIdYesTargets a specific page by ID.
pageIdxNoPage number to return (0-based). When omitted, returns the first page.
pageSizeNoMaximum number of messages to return. When omitted, returns all messages.
serviceWorkerIdNoFilter messages to only return messages of the specified service worker.
includeStackTracesNoSet to true to include the stack trace for each message when available. Increases the response size.
includePreservedMessagesNoSet to true to return the preserved messages over the last 3 navigations.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true, so the safety profile is covered. The description usefully adds the temporal boundary ('since the last navigation') that defines what the list contains, but says nothing about pagination or preserved-message behavior beyond what the schema already documents.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It is efficient and readable, though it does not exploit the space to add routing or behavioral context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only listing tool with 100% schema coverage and a readOnlyHint annotation, the description covers the essential purpose and scope. No output schema exists, yet the temporal scope gives the agent enough to call it correctly; only the sibling relationship is omitted.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all seven parameters (types, pageId, pageIdx, pageSize, serviceWorkerId, includeStackTraces, includePreservedMessages) are documented in the schema itself. The description adds no parameter detail, which is acceptable given full coverage — baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (console messages) with an explicit temporal scope ('since the last navigation'). This is clearly distinguishable from the singular get_console_message by scope, though that sibling relationship is not called out in the text.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no prerequisites, and does not name the alternative get_console_message for retrieving a single message. The temporal scope is stated but that is a filter, not usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_network_requestsB
Read-only

Lists the most recent requests for the target page since the last navigation.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageIdYesTargets a specific page by ID.
pageIdxNoPage number to return (0-based). When omitted, returns the first page.
pageSizeNoMaximum number of requests to return. When omitted, returns all requests.
resourceTypesNoFilter requests to only return requests of the specified resource types. When omitted or empty, returns all requests.
includePreservedRequestsNoSet to true to return the preserved requests over the last 3 navigations.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already tells the agent this is a safe read. The description adds the temporal scope ('since the last navigation'), which is real context beyond the annotations, but it omits return-volume behavior and the existence of preserved-request handling, leaving notable gaps for a list tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. The scope constraint is stated immediately and nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter read-only list tool with a fully documented schema but no output schema, the description covers the essential scope. The minor tension between 'most recent requests' and the schema's 'returns all requests when pageSize is omitted' is left unresolved, and no return shape is hinted at, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters (pageId, pageIdx, pageSize, resourceTypes, includePreservedRequests) are already fully documented in the schema. The description adds no extra parameter meaning, which is the expected baseline when the schema carries the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Lists) and resource (network requests) scoped to the target page, and the singular sibling get_network_request is implicitly contrasted by the plural naming. It stops short of explicitly differentiating itself from that sibling, so it lands just below the top tier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'since the last navigation' implies a temporal scope but no explicit when-to-use, when-not-to-use, or alternative (e.g., get_network_request, list_console_messages) is named. The agent must infer when this tool is preferred over its sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_pagesB
Read-only

Get a list of pages open in the browser.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so safety is covered by structured data. The description adds only that the result is a 'list of pages open in the browser,' which hints at the return shape but says nothing about ordering, page identifiers, or whether background/closed pages are included.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with no filler and the resource front-loaded after the verb. It is efficient, though it is arguably so terse that it omits useful scoping detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and no output schema, the description must carry the burden of explaining the return value. It says only 'a list of pages,' leaving unclear what fields each page entry carries (id, url, title) — information an agent likely needs before calling select_page or close_page.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. There is no parameter semantics to clarify, and the description correctly implies a no-argument call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: 'Get a list of pages open in the browser.' An agent can tell it retrieves open pages rather than modifying them. However, it does no sibling differentiation, e.g. how it relates to select_page, new_page, or close_page.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to prefer this over siblings like browser_status or select_page, nor any stated prerequisites. Use is only loosely implied by the name and the resource it describes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

new_pageB

Open a new tab and load a URL. Use project URL if not specified otherwise.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to load in a new page.
timeoutNoMaximum wait time in milliseconds. If set to 0, the default timeout will be used.
backgroundNoWhether to open the page in the background without bringing it to the front. Default is false (foreground).
isolatedContextNoIf specified, the page is created in an isolated browser context with the given name. Pages in the same browser context share cookies and storage. Pages in different browser contexts are fully isolated (useful for clean-slate testing of cookies and authentication).

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The only annotation is readOnlyHint=false, so the description carries most of the disclosure burden and does confirm the default-to-project-URL behavior. It does not, however, mention that focus moves to the foreground by default, how the new page relates to existing tabs, or what the caller receives back (e.g. a page identifier needed for follow-up calls).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler and the core action front-loaded. The second sentence is slightly ambiguous, but there is no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter page-creation tool with no output schema, the description is minimal: it omits how the newly created page is identified or referenced in subsequent calls, and doesn't hint at the interaction with background/isolatedContext behavior documented only in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (timeout, background, isolatedContext all documented in-schema), so the schema already does the heavy lifting. The description adds only the default-URL rule for the required url parameter and nothing about the other three parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb (open) plus resource (a new tab/page) and the action taken (load a URL), so an agent can distinguish it from navigate_page or select_page at a glance. It is clear, though it never names the siblings it is distinct from.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage cue is 'Use project URL if not specified otherwise,' which is an implicit default rather than a when-to-use rule. Nothing says when to open a new page versus navigating an existing one (navigate_page) or selecting among open pages (select_page).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

performance_analyze_insightB
Read-only

Provides more detailed information on a specific Performance Insight of an insight set that was highlighted in the results of a trace recording.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageIdYesTargets a specific page by ID.
insightNameYesThe name of the Insight you want more information on. For example: "DocumentLatency" or "LCPBreakdown"
insightSetIdYesThe id for the specific insight set. Only use the ids given in the "Available insight sets" list.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint=true annotation already tells the agent this is a safe read operation, so the description's main added value is scoping it to insight sets from a trace recording. It does not disclose return format, whether results are cached, or any depth/limit on the 'more detailed information', which would be useful given there is no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler or repetition. It could be marginally tighter ('Provides details on' rather than 'Provides more detailed information on'), but it is efficient and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter read-only tool with full schema coverage, the description covers the essential purpose. However, with no output schema, it leaves the shape and depth of the returned insight data entirely unspecified, which is a real gap for an agent deciding whether this call will give it what it needs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (pageId, insightName, insightSetId) are already documented in the schema, including an example insight name. The description adds no parameter-level syntax or format guidance beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: retrieving 'more detailed information on a specific Performance Insight of an insight set' produced by a trace recording. This clearly separates it from siblings like performance_start_trace and performance_stop_trace, though the phrasing 'Provides more detailed information' is a little generic about the retrieval mechanics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'highlighted in the results of a trace recording' implies this is a follow-up tool used after a trace has produced insight sets, but it never names the alternative tool(s) or states explicit when/when-not conditions. Usage context is implied rather than prescribed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

performance_start_traceA

Start a performance trace on the target webpage. Use to find frontend performance issues, Core Web Vitals (LCP, INP, CLS), and improve page load speed.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageIdYesTargets a specific page by ID.
reloadNoDetermines if, once tracing has started, the target page should be automatically reloaded. Navigate the page to the right URL using the navigate_page tool BEFORE starting the trace if reload or autoStop is set to true.
autoStopNoDetermines if the trace recording should be automatically stopped.
filePathNoThe absolute file path, or a file path relative to the current working directory, to save the raw trace data. For example, trace.json.gz (compressed) or trace.json (uncompressed).

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The only annotation is readOnlyHint=false, so the description carries most of the behavioral burden. It never states that the trace must be captured then stopped/exported, that the page should be navigated first, or where the raw trace data goes (that context lives only in the schema parameter text, not the description). No behavioral traits beyond the annotations are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the action and then the motivation. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and only a single annotation, the description is adequate on purpose but thin on workflow: it does not mention the trace must be stopped (performance_stop_trace) and analyzed, nor the reload/navigation prerequisite. Sufficient to start a trace, but not complete enough for an agent to use the full trace lifecycle correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents pageId, reload, autoStop, and filePath fully, including the 'navigate first' prerequisite. The description adds no parameter meaning of its own, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Start a performance trace on the target webpage') and reinforces it with concrete use-case goals (Core Web Vitals: LCP, INP, CLS). It is clearly distinguishable from take_screenshot or evaluate_script, though it never names its counterpart performance_stop_trace, so the lifecycle pairing is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context ('Use to find frontend performance issues... and improve page load speed'), telling the agent when this tool is appropriate. However, it names no alternatives (performance_analyze_insight, lighthouse_audit, performance_stop_trace) and provides no when-not guidance or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

performance_stop_traceB

Stop the active performance trace recording on the target webpage.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageIdYesTargets a specific page by ID.
filePathNoThe absolute file path, or a file path relative to the current working directory, to save the raw trace data. For example, trace.json.gz (compressed) or trace.json (uncompressed).

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The only annotation is readOnlyHint=false, marking this as a state-changing operation, and the description adds nothing beyond the name's implication. It does not say whether the recorded data is discarded when saved, whether filePath is optional, or what occurs if no trace is active — all material for a mutation tool whose annotation coverage is minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence naming the action, the resource and the target. Nothing redundant, nothing padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation with no output schema, the description is minimally adequate but silent on the outcomes that matter: whether trace data survives the stop, how filePath relates to persistence, and the required precondition of an active trace.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so pageId and filePath are already documented in the schema, and the baseline is 3. The description adds no parameter meaning at all — notably it omits that filePath is the mechanism by which the trace data is persisted rather than lost.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Stop) and resource (active performance trace recording) plus the target scope (target webpage). It is distinguishable from performance_start_trace by the verb alone, but it does not explicitly name the relationship to that sibling, so it falls just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the pairing with performance_start_trace — an agent can infer this must be called after a trace is started — but the description never states that a trace must be active or what happens if none is running. No explicit when/when-not or alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

press_keyA

Press a key or key combination. Use this when other input methods like fill() cannot be used (e.g., keyboard shortcuts, navigation keys, or special key combinations).

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesA key or a combination (e.g., "Enter", "Control+A", "Control++", "Control+Shift+R"). Modifiers: Control, Shift, Alt, Meta
pageIdYesTargets a specific page by ID.
includeSnapshotNoWhether to include a snapshot in the response. Default is false.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=false already tells the agent this mutates page state, so the safety profile is covered structurally. The description adds the useful 'fallback when fill() cannot be used' framing but says nothing about focus requirements, whether the key fires immediately, or what a failed press looks like. Some added value, but thin beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, with the core action front-loaded before the usage condition. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 3-parameter action tool with no output schema and full schema coverage, the description plus structured fields give enough to invoke it correctly. The only missing nuance is what the response contains when includeSnapshot is true, which the schema partially covers.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so key syntax, pageId targeting, and includeSnapshot defaults are all documented in the schema itself. The description adds no format or constraint detail beyond that, which is the correct baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Press a key or key combination') and immediately distinguishes the tool from alternative input paths like fill(). An agent can identify this as the keyboard-input primitive without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger conditions in parentheses (keyboard shortcuts, navigation keys, special combinations) and names fill() as the alternative to fall back from. It does not mention the closely related type_text or hover/click siblings, so the routing is clear but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resize_pageB

Resizes the page's window so that the page has specified dimension

ParametersJSON Schema
NameRequiredDescriptionDefault
widthYesPage width
heightYesPage height
pageIdYesTargets a specific page by ID.

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint=false annotation already signals a mutation, but the description adds almost no behavioral context beyond restating the action. It does not say whether width/height are pixels, whether the browser window or viewport is affected, or what side effects occur for the page.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is a single, front-loaded sentence with no wasted words. It is slightly awkward and omits units, but structurally it is efficient and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter mutation tool with full schema coverage and no output schema, the description is minimally viable. However, it does not clarify units or whether the page viewport or actual browser window is resized, leaving some ambiguity for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with clear per-parameter descriptions for width, height, and pageId. The description adds no extra meaning beyond what the schema already provides, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ("Resizes"), a clear resource ("page's window"), and the outcome ("specified dimension"). It is unambiguous and easily distinguishable from browser automation siblings like new_page, emulate, or take_screenshot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as emulate or new_page, nor any prerequisites like requiring an active page. Usage context is only implied by the tool name and description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

select_pageB
Read-only

Select a page as a context for future tool calls.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageIdYesThe ID of the page to select. Call list_pages to get available pages.
bringToFrontNoWhether to focus the page and bring it to the top.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true is already declared, so the safety profile is covered. The description usefully adds that selection affects future tool calls (context effect), but does not disclose what happens to the previously selected page or whether multiple contexts are allowed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. Slightly terse rather than wasteful — nothing redundant, but also little added beyond the essentials.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 2-parameter, low-complexity tool with full schema coverage and no output schema, the definition is minimally sufficient. The 'context' concept and persistence across calls could be spelled out a bit more for an agent to use it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%; pageId documents its role and even routes to list_pages, and bringToFront documents its effect. The description adds nothing beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (select) and resource (page) and adds the scope/purpose 'as a context for future tool calls', which distinguishes it from siblings like new_page, close_page, or navigate_page. It's clear what it does, though it doesn't explicitly name the distinguishing siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage (set a context for later calls), and the schema's pageId note points to list_pages for discovery. However, it doesn't say when to prefer this over navigate_page or new_page, nor whether it's required before other tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

state_listA

List saved login states: name, allowed origins, cookie counts, when saved and when the first cookie expires. Never shows cookie values.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden, and it does disclose a meaningful behavioral trait: cookie values are never exposed, only metadata such as cookie counts and first-expiry time. It stops short of covering anything else behavioral, but for a parameterless read tool the disclosure is substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence that front-loads the action and then lists returned fields, closing with the most important caveat. No filler and nothing misordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the enumeration of returned fields (name, origins, cookie counts, saved time, first expiry) and the no-cookie-values caveat do the work the schema would otherwise do. The only minor gap is that questions of ordering, volume, or scoping of the listing are unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema sets the baseline at 4. The description adds no parameter detail because there is none to add, and correctly focuses on the returned shape instead.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (saved login states), and enumerates the fields it returns. It is distinguishable from the sibling state_save by the list-vs-save framing, though the sibling is not named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The nature of the operation (read-only enumeration with no parameters) makes the intended use implied rather than stated. There is no explicit when-to-use guidance, no mention of state_save as the counterpart action, and no prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

state_saveA

Save the current browser's login (cookies + localStorage) as a named state, so other sessions can start already logged in. Only cookies for the allowed origins are kept. allowedOrigins defaults to the http(s) origins open in tabs; it can never be wider than this browser's own allowlist. States a person created with devtools-fleet login cannot be overwritten.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesState name: letters, digits, . _ -
strictNoBlock every request (not just navigations) outside the allowlist when the state is used
overwriteNo
allowedOriginsNoOrigins this state may be used on, e.g. ["https://staging.example.com"]

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers meaningful behavior: allowedOrigins defaults to open http(s) tab origins, can never exceed the browser's own allowlist, and states created via `devtools-fleet login` are protected from overwrite. It doesn't spell out what cookies outside the allowlist undergo (dropped silently) or auth prerequisites, but the disclosure is strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then behavioral caveats in descending importance. Four tight sentences with no filler, though the allowlist constraint is restated in two adjacent sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateful save operation with no annotations and no output schema, the description covers the critical behaviors an agent needs (origin scoping, overwrite protection, default derivation). Only minor gaps remain around failure modes and return behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and the schema already documents name, strict, and allowedOrigins. The description adds genuine meaning beyond the schema by explaining allowedOrigins' default derivation and its hard upper bound, plus the overwrite protection behavior that bears on the overwrite flag.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+payload: 'Save the current browser's login (cookies + localStorage) as a named state.' The purpose is unambiguous and distinguishable from the sibling state_list, which the agent can infer handles the read side.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The rationale 'so other sessions can start already logged in' implies when this is useful, and constraints on allowedOrigins scope are given, but there is no explicit when-to-use vs when-not guidance and no mention of the sibling state_list as an alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

take_heapsnapshotA

Capture a heap snapshot of the target page. Use to analyze the memory distribution of JavaScript objects and debug memory leaks.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageIdYesTargets a specific page by ID.
filePathYesA path to a .heapsnapshot file to save the heapsnapshot to.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint=false annotation already signals this is not a pure read operation, and the filePath parameter conveys that output is written to disk. However, the description never mentions that it writes a .heapsnapshot file, whether it overwrites an existing path, or the cost/size implications of capturing a heap snapshot, so it adds little beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero waste, front-loading the action before the use case. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with full schema coverage and no output schema, the description is nearly sufficient: it explains what is captured and why. The only gap is that it does not disclose the disk-write behavior that the missing-file semantics might matter for.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with pageId and filePath both documented in the schema, so the baseline is 3. The description adds no further meaning about either parameter, such as pageId validity or .heapsnapshot file semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Capture a heap snapshot of the target page,' which tells an agent exactly what the tool produces. It does not explicitly distinguish itself from the similarly named sibling take_snapshot, which could cause selection ambiguity, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a rationale ('analyze the memory distribution of JavaScript objects and debug memory leaks') that implies when the tool is useful, but offers no explicit when-not conditions or named alternatives such as take_snapshot or performance_start_trace. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

take_screenshotC

Take a screenshot of the page or element.

ParametersJSON Schema
NameRequiredDescriptionDefault
uidNoThe uid of an element on the page from the page content snapshot. If omitted, takes a page screenshot.
formatNoType of format to save the screenshot as. Default is "png"png
pageIdYesTargets a specific page by ID.
qualityNoCompression quality for JPEG and WebP formats (0-100). Higher values mean better quality but larger file sizes. Ignored for PNG format.
filePathNoThe absolute path, or a path relative to the current working directory, to save the screenshot to instead of attaching it to the response.
fullPageNoIf set to true takes a screenshot of the full page instead of the currently visible viewport. Incompatible with uid.

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations supply only readOnlyHint=false, and the description adds nothing further. It never says the screenshot is attached to the response by default, that filePath writes to disk (which is the only thing that could explain a non-read-only hint), or anything about format/quality side effects. For a tool with effectively no behavioral annotation coverage, this is a substantial gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single clean sentence with no wasted words, and the core action is front-loaded. However, its extreme brevity means it conveys almost nothing beyond the tool name, so it reads as under-specified rather than efficiently scoped.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With six parameters, no output schema, and only a one-line description, the definition is incomplete. It omits whether the image is returned inline or saved via filePath, what the response looks like, and when fullPage is appropriate, leaving the agent to reconstruct behavior from the schema alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters including uid, format, quality, filePath and fullPage are already documented in the schema. The description's 'page or element' phrasing merely restates what the uid description already says, adding no syntax or format detail beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('take a screenshot') and covers both page-level and element-level capture. It does not, however, distinguish itself from siblings like take_snapshot or take_heapsnapshot, which an agent could plausibly confuse with this tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus take_snapshot or the other capture tools, nor any mention of prerequisites such as needing a page already open or selected. The page-vs-element choice is left entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

take_snapshotA

Take a text snapshot of the target page based on the a11y tree. The snapshot lists page elements along with a unique identifier (uid). Always use the latest snapshot. Prefer taking a snapshot over taking a screenshot. The snapshot indicates the element selected in the DevTools Elements panel (if any).

ParametersJSON Schema
NameRequiredDescriptionDefault
pageIdYesTargets a specific page by ID.
verboseNoWhether to include all possible information available in the full a11y tree. Default is false.
filePathNoThe absolute path, or a path relative to the current working directory, to save the snapshot to instead of attaching it to the response.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (only readOnlyHint=false), so the description carries most of the burden, and it delivers: element list with uid, staleness semantics, and the fact that the selected DevTools element is marked. It never explains why the operation is flagged non-read-only (e.g., file emission via filePath), which is the one behavioral gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, all front-loaded and functional: purpose, output shape, staleness rule, sibling preference. Slightly more than strictly necessary, but no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by describing what a snapshot returns and how uids work. For a 3-parameter read-style tool this is nearly complete; only the side-effect/read-only nuance is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so pageId, verbose, and filePath are already fully documented in the schema. The description adds no format or usage detail beyond that, which is the correct baseline when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('take a text snapshot of the target page based on the a11y tree') and immediately distinguishes the output from take_screenshot by describing it as a list of page elements with unique identifiers. An agent can tell exactly what it gets without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes between siblings: 'Prefer taking a snapshot over taking a screenshot.' It also gives a lifecycle rule ('Always use the latest snapshot') that tells the agent when a prior snapshot is stale, which is a real when-to-use constraint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

type_textA

Type text using keyboard into a previously focused input

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesThe text to type
pageIdYesTargets a specific page by ID.
submitKeyNoOptional key to press after typing. E.g., "Enter", "Tab", "Escape"

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotation readOnlyHint=false establishes this as a mutating action, which the description does not contradict. The description adds one behavioral trait beyond the annotations: it simulates keyboard input into an element that must already have focus, implying prior interaction state. It does not disclose what happens on failure, timing, or whether the text is appended versus replacing existing content.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. Every clause — verb, mechanism, target precondition — carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 3-parameter, no-output-schema tool with a fully documented schema and an annotation covering the safety profile, the description covers the essentials. The only gap is the absence of explicit differentiation from fill, which the agent must infer from wording alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so text, pageId, and submitKey are all documented in the schema itself. The description adds no syntax, format, or edge-case detail beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Type text") and the mechanism ("using keyboard"), plus a scope constraint ("into a previously focused input") that distinguishes it from a value-setting tool like fill. It does not name the closest sibling (fill) to route the agent, so it stops short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Previously focused input" implies you must focus an element first (e.g. via click), which is useful implied guidance. But there is no explicit statement of when to use this versus fill/fill_form/press_key, all of which live in the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_fileC

Upload a file through a provided element.

ParametersJSON Schema
NameRequiredDescriptionDefault
uidYesThe uid of the file input element or an element that will open file chooser on the page from the page content snapshot
pageIdYesTargets a specific page by ID.
filePathsYesOne or more files paths to upload. File paths have to be local to the browser instance (not the MCP).
includeSnapshotNoWhether to include a snapshot in the response. Default is false.

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false, which the one-line description is consistent with, but the description adds no behavioral context: no indication that it may trigger a file chooser, mutate page state, or what happens on failure. For a mutation tool with a bare annotation, more disclosure is warranted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no redundancy, front-loading the action. Its brevity comes at the cost of specificity, but there is no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 4 well-documented parameters and no output schema, the schema carries most of the load, so the description need not explain return values. It is still thin on the operational context (element type required, interaction with the file chooser) that an agent invoking a browser mutation would benefit from.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each parameter (uid, pageId, filePaths, includeSnapshot) is already documented in the schema, including that paths must be local to the browser instance. The description adds nothing beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

It states a specific verb and resource (upload a file), but "through a provided element" is vague and never names the uid/file-input concept that the schema requires. An agent can guess the general intent but cannot confidently distinguish it from siblings like fill or fill_form without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, no alternatives. The description does not say the element must be a file input or a chooser-opening element, nor how it relates to click or fill_form, leaving selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_forC
Read-only

Wait for the specified text to appear on the selected page.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesNon-empty list of texts. Resolves when any value appears on the page.
pageIdYesTargets a specific page by ID.
timeoutNoMaximum wait time in milliseconds. If set to 0, the default timeout will be used.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply readOnlyHint=true, so the description carries the behavioral burden. It says nothing about polling behavior, what happens when the wait times out, whether it errors or returns a status, or that the default timeout applies when 0 is passed (that detail lives only in the schema).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It is efficient, though its brevity is partly the source of the missing behavioral detail rather than a mark of completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a blocking/waiting tool with no output schema and no annotations beyond readOnlyHint, the description omits the critical question of what the tool returns or does on timeout, and whether the page must be selected first. These are real gaps for an agent deciding how to handle the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents pageId, the text list with any-match semantics, and the timeout default behavior. The description adds no parameter syntax or format detail beyond what the schema provides, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Wait') and resource ('specified text ... on the selected page'), which distinguishes it from sibling reads like take_snapshot or get_console_message. It does not, however, explicitly name an alternative tool or clarify how it relates to select_page.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to prefer this over polling with evaluate_script, take_snapshot, or re-reading state. No preconditions (e.g., a page must already be selected) and no when-not-to-use conditions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 36 tool updatesv0.1.0
    • First observedbrowser_restart
    • First observedbrowser_start
    • First observedbrowser_status
    • First observedbrowser_stop
    • First observedclick
    • First observedclose_page
    • First observeddrag
    • First observedemulate
    • First observedevaluate_script
    • First observedfill
    • First observedfill_form
    • First observedget_console_message
    • First observedget_css_styles
    • First observedget_network_request
    • First observedhandle_dialog
    • First observedhover
    • First observedlighthouse_audit
    • First observedlist_console_messages
    • First observedlist_network_requests
    • First observedlist_pages
    • First observednavigate_page
    • First observednew_page
    • First observedperformance_analyze_insight
    • First observedperformance_start_trace
    • First observedperformance_stop_trace
    • First observedpress_key
    • First observedresize_page
    • First observedselect_page
    • First observedstate_list
    • First observedstate_save
    • First observedtake_heapsnapshot
    • First observedtake_screenshot
    • First observedtake_snapshot
    • First observedtype_text
    • First observedupload_file
    • First observedwait_for

TDQS

B3.4/5.0

Scored across 36 tools

Disambiguation4/5

Most tools target distinct resources or actions, such as take_snapshot vs take_screenshot and list_console_messages vs get_console_message. However, the input family (click, fill, fill_form, type_text, press_key) has some overlapping boundaries, though descriptions clarify preferred usage.

Naming Consistency4/5

The set is predominantly snake_case and verb_noun (close_page, list_pages, take_snapshot). A minority use noun-first or noun-verb ordering (browser_status, state_save, performance_start_trace), which is a minor deviation but still readable.

Tool Count4/5

36 tools is high, but the domain (full browser automation and DevTools) legitimately requires separate tools for lifecycle, pages, input, inspection, performance, network, console, and state. It is slightly over the typical 3-15 range but each tool earns its place.

Completeness5/5

The surface covers browser lifecycle, navigation, page/tab management, input, a11y snapshots, screenshots, console, network, CSS, performance tracing, Lighthouse, memory heap, dialogs, file upload, emulation, and login states. No obvious lifecycle gaps for a DevTools automation server.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers