Skip to main content
Glama

TabWeave MCP

TabWeave attaches to an existing Chrome/Chromium instance over CDP and exposes 54 primary tools plus 21 compatibility aliases through newline-delimited MCP JSON-RPC on standard input/output. Tool names, schemas, parameters and results retain the baseline contract; the project and its internal bindings are renamed.

Install and use

Node.js 18 or newer.

npm ci
CDP_PORT=9222 npm start
npm test

The Chrome instance must already be running with a localhost debugging port. TabWeave does not launch a user's browser. A client configuration can invoke:

{
  "mcpServers": {
    "tabweave": {
      "command": "node",
      "args": ["/absolute/path/to/TabWeave/launch.cjs"],
      "env": { "CDP_PORT": "9222" }
    }
  }
}

Related MCP server: chrome-remote-debugging-mcp

Source organization

  • launch.cjs: process lifecycle and command entry point.

  • stdio_channel.cjs: ordered line exchange, EOF handling and parse errors.

  • session_kernel.cjs: CDP session, documents, frames, console events and dispatch.

  • action_catalog.cjs: cached tool descriptions and schemas.

  • argument_gate.cjs: canonical file-path checks and schema validation.

  • actions/: session, composition, page, input, observation and compatibility tools.

  • checks/: renamed validation and protocol tests.

The action modules share live session values through accessors; they do not copy mutable state when configuration, tabs or frames change. Page-evaluated functions remain self-contained and third-party Playwright names remain intact.

Behavior and boundaries

Navigation, forms, keyboard/mouse input, human/smart timing, screenshots, cookies, storage, frame/tab management, assertions, retry/batch/step composition and bounded nesting remain available. Screenshot buffers are returned in full.

Only the baseline supported JSON Schema checks are enforced: required fields, declared top-level types and enums. Numeric strings remain accepted for numeric fields; unimplemented schema features remain unvalidated. Canonical path checks remain check-time restrictions, not a filesystem sandbox. CDP must stay local; use an isolated profile for sensitive work.

The local regression suite uses a browser stub. Additional rewrite verification uses isolated Chrome with synthetic local pages; it does not establish behavior on every external site. See VALIDATION.md, ORIGIN.md and LICENSE.

Available Tools

75 tools
assertC

Unified assertion (text/visible/hidden/url/count/element)

ParametersJSON Schema
NameRequiredDescriptionDefault
textNo
typeYes
stateNo
patternNo
timeoutNo
expectedNo
operatorNo
selectorNo

TDQS

C2/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it discloses nothing about failure behavior, waiting/retry semantics, return values, or side effects. It is just a list of types, so an agent cannot predict how the tool behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single parenthetical is concise but severely under-specified. It is not front-loaded with any useful context, and the brevity comes at the cost of clarity rather than being efficient communication.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter tool with no annotations and no output schema, the description is completely inadequate. It fails to cover parameter usage, behavioral expectations, or output, leaving major gaps an agent must guess at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 8 parameters. The description maps the 'type' enum values to assertion kinds, which is some help, but it does not explain state, pattern, timeout, expected, operator, or selector, leaving most parameters undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the verb 'assert' and enumerates the assertion types (text/visible/hidden/url/count/element), which gives a rough idea of what it does. However, it doesn't state the tool's purpose in a full phrase or distinguish it from siblings like exists, is_visible, wait_for, or get_text, leaving ambiguity about its role in the suite.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives. For example, it doesn't say whether to prefer assert over exists or wait_for, or whether it waits for a condition before failing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

batchD

Batch execution

ParametersJSON Schema
NameRequiredDescriptionDefault
actionsYes

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing: no execution ordering, no default for stopOnError, no atomicity, no error reporting behavior. For a tool that orchestrates arbitrary other tools this is a critical gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words are not concise, they are under-specified; every sentence that would earn its place is absent. Brevity here costs the agent the information it needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a nested-object, dynamic-tool-reference schema, no annotations, and no output schema, the description should explain execution semantics and return behavior. It explains nothing, leaving the tool effectively uncallable from documentation alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single parameter is a nested array whose items reference a "tool" name and an arbitrary "args" object. The description says nothing about how tool names map to sibling tools, how args are shaped, or what stopOnError defaults to, so an agent cannot construct a valid call without guessing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Batch execution" restates the tool name without identifying the resource or the kind of actions being batched. An agent cannot distinguish it from siblings like run_steps or the individual action tools (click, fill, navigate) that it presumably wraps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use batch versus run_steps, versus invoking individual action tools repeatedly, and no mention of ordering or sequencing semantics.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

blurD

Out of focus elements

ParametersJSON Schema
NameRequiredDescriptionDefault
selectorYes

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure, and it discloses nothing: not whether this mutates page state, requires a visible/focusable element, errors on missing elements, or what it returns. A one-fragment description leaves the behavior entirely opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Brevity here reflects under-specification rather than conciseness. The fragment is short but carries no actionable information and is not front-loaded around any identifiable purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a required selector parameter, no annotations, and no output schema, the description supplies none of the information an agent needs to invoke it correctly. It is wholly inadequate for even a trivial tool in this large browser-automation suite.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one required parameter ('selector') with 0% schema description coverage, and the description adds no meaning whatsoever about its format, accepted selector syntax, or expected values. With no schema documentation and no descriptive compensation, the parameter is completely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Out of focus elements' does not state an action at all – it is a fragment that could describe removing focus from an element or applying a visual blur. It never names the verb 'blur' as an operation nor distinguishes this from the sibling 'focus' tool, leaving the agent to guess.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of the complementary 'focus' sibling, and no prerequisites or context. The agent receives no signal about when this tool is the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

checkC

Unified check of element status (exists/visible/hidden/enabled/disabled/checked/focused/editable/in_viewport)

ParametersJSON Schema
NameRequiredDescriptionDefault
textNoCheck text exists
stateNo
timeoutNo
selectorNo
class_nameNoCheck class name

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it says nothing about what the tool returns (boolean vs. raise/throw on failure), how timeout interacts with the check, or whether it auto-waits/retries. For a verification tool whose entire contract is the shape of its result, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the state list in a parenthetical – front-loaded and free of filler. It is arguably under-specified rather than verbose, but structurally it is tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and only 40% param coverage, the description should explain the return contract and how the check behaves on failure. It omits all of this, leaving the agent unable to interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 40%, with five parameters and undocumented selector, state, and timeout entries. The description merely mirrors the state enum already present in the schema and adds nothing about selector vs text vs class_name precedence or timeout semantics, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (check) and resource (element status) and enumerates the nine states it can evaluate, which is more precise than a vague 'check element'. However, it never differentiates itself from siblings that appear to do overlapping things (exists, is_visible, assert), so an agent cannot tell when this unified tool should win.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance at all. With siblings exists, is_visible, wait_for, assert, and find all plausibly covering the same ground, the description should say when to prefer this unified check over those specific tools, and it does not.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

checkboxD

Check box operation

ParametersJSON Schema
NameRequiredDescriptionDefault
checkedNo
selectorYes

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and supplies none of it. It does not state whether the operation sets, toggles, or verifies checked state, whether it is idempotent, what happens if the selector matches nothing, or what the failure mode is.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief, but the brevity reflects under-specification rather than conciseness — there is no front-loaded verb, resource detail, or constraint. Nothing was trimmed because nothing was said.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation-capable interaction tool with no annotations, no output schema, and zero parameter documentation, the description is completely inadequate. An agent has no information needed to invoke this correctly or interpret its result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters, so the description must compensate and does not. It never explains what 'selector' accepts or what the 'checked' boolean controls (set vs unset vs toggle), leaving the semantics of both parameters undefined.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Check box operation' is essentially a tautology of the tool name 'checkbox' — it restates the resource without naming a specific verb or describing the effect. It does not distinguish this tool from sibling tools such as 'check', 'set_checked', or 'toggle_checkbox', leaving the agent unable to tell them apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of alternatives. This is particularly costly here because several siblings ('check', 'set_checked', 'toggle_checkbox') appear to overlap, and the description gives no basis for choosing among them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cleanupC

Clear cache

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and falls short. It doesn't state which cache is cleared, whether the operation is destructive or reversible, or what state changes result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words with no filler, so nothing is wasted. But the extreme terseness reflects under-specification rather than disciplined conciseness, leaving the agent with little to work with.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter mutation tool with no annotations and no output schema, the description is too thin. It omits what is cleared, the effect on sessions or state, and any indication of return behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to clarify at the parameter level. The baseline of 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb+resource ('Clear cache'), which is clearer than a pure tautology. However, it fails to distinguish itself from close siblings like clear_cookies and clear_storage, leaving ambiguity about which cache is meant.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, no prerequisites, and no mention of alternatives such as clear_cookies or clear_storage. The agent must infer usage entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clear_cookiesC

Clear Cookies(alias)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. It says nothing about what gets cleared (all cookies, domain-specific, session vs persistent), whether the operation is destructive or reversible, or any side effects, making it behaviorally opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a mere fragment with no structure. The unexplained '(alias)' adds noise without earning its place, and while short, it is under-specified rather than concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation-style tool that clears browser state, the description is incomplete. It omits what cookies are affected, when to use it, and any behavioral context, which is especially important given there are no annotations or output schema to compensate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and an empty schema with 100% description coverage, so per the rubric baseline is 4. There are no parameter semantics to clarify, and the description adds nothing needed here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Clear Cookies(alias)' essentially restates the tool name, making it a tautology rather than a distinct statement of purpose. It does not differentiate this tool from siblings like clear_storage, cookies, or set_cookie, so an agent cannot tell when this specific tool is appropriate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as clear_storage or set_cookie. The description provides no context, exclusions, or prerequisites, leaving usage entirely to inference from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clear_storageD

Clear Storage (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNo

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It says nothing about what precisely gets destroyed (local vs session scope is only inferable from the enum), whether the action is reversible, or whether confirmation/permissions are needed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words plus a parenthetical is under-specification, not conciseness. There is no front-loaded purpose to evaluate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive-sounding mutation tool with no annotations, no output schema, and an undocumented parameter, the description is completely inadequate. It supplies none of the context an agent needs to call it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single "type" parameter, and the description adds nothing. The enum values local/session are the only semantic signal, and they come entirely from the schema, not the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Clear Storage (alias)" merely restates the tool name and appends a meaningless parenthetical. It does not state a distinct verb+resource or distinguish this tool from siblings like storage, get_storage, set_storage, or clear_cookies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative is given. An agent cannot tell from the description why it would call clear_storage rather than clear_cookies or the general storage tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clickC

Unified click (supports selector/text/coordinates, single/double/triple/right/long, fast/human/smart mode)

ParametersJSON Schema
NameRequiredDescriptionDefault
xNocoordinate X
yNocoordinate Y
modeNoClick mode
textNoFind the click target by text
typeNoClick type
indexNoNth element
timeoutNo
durationNoLong press duration (ms)
selectorNoCSS selector
if_existsNoClick only when present
wait_afterNoWait after clicking
hover_firstNoHover this selector first
scroll_firstNoScroll to the element first

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It mentions modes and click types that parallel the schema, but it does not disclose side effects, failure behavior, waiting semantics, or any other behavioral trait beyond what the structured fields already indicate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded parenthetical with no filler, which is highly concise. It is slightly terse for a 13-parameter tool, but every phrase contributes directly to describing the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 13 parameters, no annotations, and no output schema, this description is incomplete. It enumerates capabilities but omits how the many parameters interact, when to prefer one targeting or mode option over another, and what behavior to expect during or after a click.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 92%, so the schema already documents almost all parameters. The description groups and summarizes some parameter families (selector/text/coordinates, click types, modes), but adds no syntax, constraints, or format details beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific action (click) and enumerates the supported targeting methods, click types, and modes. It is clear what the tool does, but it does not explicitly differentiate itself from sibling tools like mouse, mouse_down, or hover.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lists capabilities but gives no guidance on when to use this tool versus alternatives, nor how to choose between selector, text, or coordinate targeting. There is no when-to-use or when-not-to-use context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

close_tabD

Close tab

ParametersJSON Schema
NameRequiredDescriptionDefault
indexNo

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure but states nothing. It does not indicate whether closing a tab is destructive, how unsaved state is handled, or whether the optional index defaults to the active tab.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but its brevity reflects under-specification rather than effective conciseness. It fails to include necessary context for a tool with an optional parameter, so the two-word phrase does not earn its place as adequate documentation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a browser automation tool with one optional parameter, no annotations, and no output schema, the description is completely incomplete. It does not explain which tab is closed, what happens if index is omitted, or any side effects, making it inadequate for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not mention the 'index' parameter at all. It provides no meaning, format, or constraint information for the single parameter, leaving it completely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Close tab' is a verbatim restatement of the tool name, making it tautological. While the underlying action is clear, the description adds no distinguishing detail beyond the name itself, so it falls into the tautology category.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like switch_tab or new_tab, nor any mention of prerequisites or context. The description provides no usage instructions at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

console_logsD

Console log

ParametersJSON Schema
NameRequiredDescriptionDefault
clearNo
limitNo

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing: not whether logs are retrieved or flushed, whether retrieval is destructive, how many entries are returned, or whether a page context is required. Nothing beyond the two-word fragment is available.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words is not conciseness but under-specification; there is no front-loaded verb or scope statement to anchor the agent. Nothing useful is communicated in the space used.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A two-parameter tool with no annotations, no output schema, and no parameter documentation leaves the agent unable to call it correctly or anticipate its return value. The definition is wholly inadequate for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for both parameters. The description does not explain that 'clear' likely empties the log buffer or that 'limit' caps returned entries, so the agent must guess at semantics and destructive side effects from bare property names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Console log" merely restates the tool name and never states a verb or what is actually done with the console log (retrieve? clear? stream?). It gives no basis for distinguishing it from siblings like get_text, snapshot, or eval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to call this tool, what preconditions exist, or which sibling it replaces. With dozens of observation-oriented siblings (get, get_text, snapshot, health_check), the agent has no routing information at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cookiesC

Cookie operation (get/set/delete/clear)

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
pathNo
valueNo
actionNo
domainNo
secureNo
expiresNo
httpOnlyNo

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it discloses almost nothing. It does implicitly signal that set/delete/clear are mutations, but says nothing about whether changes persist, what the return value looks like, or how the browser context is affected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single parenthetical line is economical with no wasted words, but the brevity here reflects under-specification rather than discipline. It is appropriately sized only in the sense that it says nothing that needs trimming.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter tool with zero schema coverage, no annotations, no output schema, and three competing siblings, this description is far too thin. An agent cannot determine parameter semantics, behavior, or sibling routing from what is written.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 8 parameters at 0% schema description coverage, the description must compensate and does not. The only parameter detail it offers is the action enum values, which are already fully documented in the schema, so it adds no meaning for name, path, value, domain, secure, expires, or httpOnly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (cookies) and enumerates the operations (get/set/delete/clear), so an agent can infer the general function. However, 'Cookie operation' is a generic framing, and it does no work to distinguish this multi-action tool from the overlapping siblings get_cookies, set_cookie, and clear_cookies that already exist in the toolset.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this composite tool versus the single-purpose siblings (get_cookies, set_cookie, clear_cookies). The agent is left to guess whether this is a unified alternative or a legacy duplicate, and there are no stated prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

countD

element count

ParametersJSON Schema
NameRequiredDescriptionDefault
selectorYes

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations and no output schema, the description carries full behavioral burden, yet 'element count' discloses nothing about the return value (an integer), how counts are computed, or any edge cases such as zero matches. No mutation, auth, or rate-limit context is given.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but this reflects under-specification rather than efficient conciseness. Two words do not earn their place because they fail to communicate anything beyond the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one required parameter, no parameter documentation, no output schema, and no annotations, the description is completely inadequate. It omits the selector format and the return value, both essential for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single required parameter 'selector' has 0% schema description coverage, and the description provides no meaning for it. It does not clarify whether the selector is a CSS selector, XPath, text, or some custom syntax, leaving the agent to guess.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'element count' is a fragment that essentially restates the tool name rather than stating a clear verb+resource relationship. It hints that the tool counts elements, but provides no full sentence or scope, and does not differentiate from siblings like get_text, exists, or get.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as get, exists, is_visible, or get_text. No context, prerequisites, or exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dialogD

Dialog processing

ParametersJSON Schema
NameRequiredDescriptionDefault
textNo
actionNo

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing. It does not say whether the call blocks until a dialog appears, whether it fails if no dialog is present, what accept vs dismiss does to the page, or what the text parameter feeds (likely a prompt input).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words is under-specification rather than conciseness; there is no front-loaded explanation because there is no content at all. Nothing is padded, but nothing is earned either.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A browser-automation dialog tool with no annotations, no output schema, 0% parameter coverage, and a two-word description is not complete enough for an agent to call it correctly. The critical context of what a dialog is and when one exists is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description adds no parameter meaning. The action enum (accept/dismiss) is somewhat self-explanatory, but the purpose of the text parameter is never explained, leaving half the parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Dialog processing" essentially restates the tool name without naming a verb or the resource's role in the browser session. It does not distinguish this from siblings like click or press_key, and an agent cannot tell whether it dismisses JS alert/confirm/prompt dialogs or handles UI modals.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisite (e.g. a dialog must be open), and no mention of alternatives among the many sibling tools. The agent is left to guess the trigger condition entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dragD

Drag and drop

ParametersJSON Schema
NameRequiredDescriptionDefault
to_xNo
to_yNo
from_xNo
from_yNo
offset_xNo
offset_yNo
to_selectorNo
from_selectorNo

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing: whether the drag is instant or animated, whether it fires intermediate mouse events, whether offsets are applied relative to selectors or coordinates, or what happens on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is brief, but the brevity reflects under-specification rather than conciseness — a shortened label is not a useful front-loaded summary for an 8-parameter gesture tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter, zero-annotation, no-output-schema tool, the definition is completely inadequate; an agent has no basis for choosing between coordinate and selector inputs or for understanding the gesture semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Eight parameters with 0% schema description coverage and the description adds no meaning at all. The interaction between from_x/from_y, to_x/to_y, offset_x/offset_y, and the *_selector alternatives is entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Drag and drop' merely restates the tool name and gives no verb-plus-resource specificity beyond the obvious. It does not distinguish this tool from siblings like mouse_down/mouse_up, move_mouse, or hover, which could plausibly be used for similar pointer manipulation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the many pointer-related alternatives (mouse, mouse_down, mouse_up, move_mouse, hover). Nothing indicates prerequisites, that a source and target are needed, or that it is a composite gesture.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

enter_frameC

Enter iframe

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutNo
selectorYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it discloses nothing about behavior: that this mutates the active frame context for later commands, what happens if the selector matches nothing, or how the 'timeout' parameter affects failure. Only the bare action is stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words is brevity born of under-specification rather than economy; there is nothing front-loaded or structured because there is essentially no content. It fails the 'every sentence earns its place' test by having no sentence at all.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A two-parameter, state-mutating tool with zero annotation coverage, no output schema, and no parameter documentation leaves the agent with nothing actionable beyond the name. The definition is not complete enough to invoke this tool reliably.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so everything about the two parameters must come from the description, and it says nothing. It is unclear whether 'selector' is a CSS selector, frame name, index, or URL, and 'timeout' is entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Enter iframe'), which maps to the sibling pair exit_frame/exit_all_frames well enough to distinguish it from other tools. However, it never clarifies how the frame is identified or whether this affects the browsing context for subsequent commands, so the purpose is only minimally articulated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to enter a frame versus using list_frames, exit_frame, or exit_all_frames, nor any prerequisite about being on a page first. The agent must infer the entire usage context from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evalD

Execute JavaScript

ParametersJSON Schema
NameRequiredDescriptionDefault
scriptYes
timeoutNo

TDQS

D1.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It says only 'Execute JavaScript' and omits execution context, permissions, side effects, timeout behavior, security implications, and return behavior. For a code-execution tool, this is a severe transparency gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short at two words, which is concise but not appropriately informative. This is under-specification rather than useful brevity, and it does not front-load any operational or safety context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no annotations, no output schema, and 0% parameter description coverage. The description supplies none of the missing context an agent would need for correct invocation, such as execution environment, timeout semantics, or safety constraints. It is inadequate for a code-execution tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and there are two parameters (script and timeout). The description does not explain either parameter, does not clarify that 'script' is required, and does not describe what 'timeout' controls. It fails to compensate for the undocumented schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource, 'Execute JavaScript', so an agent can infer it runs JavaScript. However, it gives no execution context (browser page, Node, sandbox) and does not distinguish itself from the large set of browser automation siblings. This makes the purpose minimally adequate but still vague about scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description contains no guidance on when to use this tool versus alternatives, no prerequisites, and no exclusions. It does not indicate whether it is for browser-context scripting, automation steps, or general JavaScript evaluation. This is a clear absence of usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

existsD

Check for existence (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutNo
selectorYes

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and delivers nothing: it does not say whether the check is immediate or waiting, whether it throws or returns a boolean, or what happens on timeout. The 'timeout' parameter hints at a wait, but the description never confirms or explains that behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but this is under-specification rather than conciseness — the one sentence spends its words on a tautology plus an unexplained '(alias)' marker instead of front-loading the actual check semantics.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% parameter documentation, the description must carry everything and carries nothing. An agent cannot reliably invoke this tool correctly from what is given.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across both parameters, and the description adds no meaning for 'selector' or 'timeout'. Neither the selector syntax nor the timeout unit/default is stated anywhere, leaving the agent to guess.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the name with a parenthetical '(alias)', which is closer to tautology than explanation. It never says what is being checked for existence (a DOM element? a file? a page?) or what the selector targets, so an agent cannot distinguish it from siblings like is_visible, count, or wait_for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given, and no alternative is named. The crowded sibling set (is_visible, count, wait_for, find) makes routing ambiguous, and the '(alias)' note implies a duplicate of another tool without saying which one or why to prefer either.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exit_all_framesC

Exit all iframes

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.6/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It reveals nothing about what 'exiting' entails—e.g., whether it resets the frame stack, affects page state, or requires specific prior conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at two words, front-loaded with the action. It contains no redundant or padding text, though its brevity borders on under-specification rather than optimal conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of annotations, output schema, and parameters, the description should provide minimal context to guide invocation. It fails to explain when to use this tool versus related siblings like exit_frame or what 'exit all iframes' semantically does, leaving an agent without enough information to confidently select it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, and the schema is empty. Per the rubric, a tool with 0 parameters receives a baseline of 4 for parameter semantics, as there are no parameters to document.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Exit all iframes' states a specific verb and resource, making the basic action clear. However, it essentially restates the tool name with only slight clarification ('iframes' vs 'frames') and adds no differentiation from the sibling exit_frame beyond the word 'all'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as exit_frame, enter_frame, or list_frames. It states what the tool does but not the context or conditions for invoking it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

exit_frameB

Exit the current iframe

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure, and it says nothing beyond the action. It does not state what happens on a nested frame (return to immediate parent vs. top level), what occurs if no frame is currently entered, or whether this is a state-mutating context switch.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single four-word sentence with zero waste and the action front-loaded. Nothing could be trimmed and nothing extraneous is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with no output schema or annotations, the definition is minimally viable, but it omits the frame-stack behavior (parent vs. top) that determines correct invocation in nested-frame scenarios. A few words on fallback behavior would make it complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, which is the baseline-4 case; there is nothing to document and the schema is empty at full coverage. No additional parameter meaning is needed or missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb ('Exit') and resource ('the current iframe'), so the agent knows exactly what action occurs. The word 'current' implicitly distinguishes it from the sibling exit_all_frames, though the distinction is not spelled out.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent can infer this is called after enter_frame or when navigating out of a frame. There is no explicit when-to-use or when-not-to-use guidance, and the siblings enter_frame, exit_all_frames and list_frames are not referenced to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fillD

Quick filling (direct replacement)

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
timeoutNo
selectorYes

TDQS

D1.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. The phrase 'direct replacement' hints at a set-value mechanism rather than simulated typing, which is a small behavioral clue. However, it omits side effects, whether input/change events fire, permission or auth needs, timeout behavior, and reversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short and front-loaded, but its brevity comes from missing essential information rather than efficient expression. It is under-specified for a tool with required parameters and no annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% parameter description coverage, the description needed to compensate by explaining behavior and parameters. It provides almost nothing, leaving the tool effectively undocumented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has three parameters (selector, text, timeout) with 0% description coverage. The description mentions none of them and adds no meaning beyond the bare schema, leaving all parameter semantics undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Quick filling (direct replacement)' largely restates the tool name 'fill' and does not specify what resource is being filled (e.g., an input field). It hints at a mechanism ('direct replacement') but does not clearly distinguish from sibling tools such as 'type' or 'fill_form'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like 'type' or 'fill_form'. No prerequisites, exclusions, or contextual routing are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fill_formC

Fill out forms in batches

ParametersJSON Schema
NameRequiredDescriptionDefault
fieldsYes
submitNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. 'Fill out forms in batches' implies a mutation (data entry), but it omits critical details: whether it submits forms, what happens to existing data, error handling, or any side effects. This is insufficient for a write tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. However, it is too terse for the complexity of the tool, and its brevity fails to convey necessary operational context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a nested object parameter, no annotations, and no output schema, the description is inadequate. It does not explain how to structure the 'fields' object, what 'submit' does, or what the tool returns, leaving significant gaps for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not mention either parameter. The 'fields' object structure and the 'submit' boolean are completely undocumented, leaving an agent with no guidance on expected input format or behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Fill out) and resource (forms) with a scope qualifier (in batches). It distinguishes the tool from its sibling 'fill' (likely single-field) by implying a form-level, batch operation. However, it does not explicitly differentiate from other form-related siblings or clarify what 'batches' means.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is provided. The description does not explain when to choose this tool over alternatives like 'fill', 'type', or 'batch', nor does it state any prerequisites or conditions. It only implies a batch context without elaboration.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

findD

Find elements

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNo
textNo
limitNo
valueNo
selectorNo
attributeNo

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. It says nothing about whether this is a read-only query, what happens if no elements match, whether results are paginated, or what the return shape is.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The two-word description is not concisely helpful; it is under-specified. Brevity here reflects a lack of content rather than efficient structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter tool with no annotations and no output schema, the description is completely inadequate. It provides no basis for an agent to invoke the tool correctly against any of its siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All six parameters (tag, text, limit, value, selector, attribute) have 0% schema description coverage, and the description adds no meaning for any of them. An agent cannot tell how these filters combine or which are required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Find elements' essentially restates the tool name with only a generic resource added. It does not specify what kind of elements, in what context, or how it differs from siblings like get, count, exists, or snapshot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as get, get_text, exists, is_visible, or wait_for. The agent is left to infer the appropriate context entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

focusD

Focus on elements

ParametersJSON Schema
NameRequiredDescriptionDefault
selectorYes

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral disclosure. It says nothing about whether this changes document.activeElement, scrolls into view, waits for the element, fails if not found, or returns anything.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but it is under-specified rather than concise. It does not front-load actionable information about behavior or selector syntax.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a browser automation action with no annotations, no output schema, and one undocumented required parameter, the description is inadequate. It omits selector format, side effects, error behavior, and distinctions from sibling focus-adjacent tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one required parameter, selector, with 0% schema description coverage. The description does not define whether selector is a CSS selector, XPath, text matcher, or how multiple matches are handled.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Focus on elements" uses the verb focus and a generic object, but does not specify DOM focus, selector matching, or what focusing does. It is barely distinguishable from siblings such as blur, click, and hover beyond the word "focus."

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no when-not-to-use guidance, no prerequisites, and no alternative tools. An agent cannot infer when this should be selected over blur, click, or hover.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

getC

Get element information (text/html/attribute/value/position/dimensions/styles/count/classes/tag/dataset)

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNo
outerNo
fallbackNo
selectorYes
attributeNo
propertiesNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It says nothing about what happens when the selector matches nothing, what 'fallback' or 'outer' do, or whether it errors, returns null, or falls back — all material for a read tool with a 'fallback' parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no wasted prose. The parenthetical enum recap is somewhat redundant against the schema but keeps the tool's scope scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter tool with no annotations and no output schema, the description is far too thin. It should at minimum explain the selector-matching/fallback behavior and clarify the relationship to the dedicated get_text/get_html/get_attribute siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 6 parameters. The parenthetical list merely re-states the 'type' enum already in the schema, leaving selector, outer, fallback, attribute, and properties completely undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource ('Get element information') and enumerates the kinds of data retrievable. However, it draws no boundary against strong siblings like get_text, get_html, and get_attribute, whose purposes it visibly overlaps with, so an agent cannot tell when to prefer this generic tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative routing is given. With overlapping siblings such as get_text, get_html, get_attribute, and count, the absence of any disambiguation guidance is a real gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_attributeD

Get attributes (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault
selectorYes
attributeYes

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only says "Get attributes (alias)" and says nothing about read-only behavior, return format, error handling, or whether any state is affected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short, but it is under-specified rather than concise. A single vague fragment does not front-load useful information or earn its place against the tool's two required parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two undocumented parameters, no annotations, and no output schema, the description is completely inadequate. It omits parameter meaning, return values, and any usage context needed to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain either required parameter. It mentions attributes generally but provides no meaning for selector or attribute, including expected syntax or accepted values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description says "Get attributes (alias)", which only partially restates the tool name and leaves the actual resource ambiguous. It does not explain that it retrieves an element attribute by selector, nor does it distinguish itself from sibling tools such as get_text, get_html, or get_url.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool instead of alternatives. With many sibling tools that also retrieve page data, the absence of any routing or context leaves the agent to infer usage entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_configD

Get configuration

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

D1.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure but fails to do so. It doesn't indicate whether this is a read-only operation, if it requires authentication, what the return format might be, or any potential side effects, making it inadequate for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

While concise with only two words, the description is under-specified rather than efficiently structured. It lacks essential details about what configuration is retrieved, making it too brief to be helpful, which is a flaw in content rather than a virtue of brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations, no output schema, and a vague description, this is completely inadequate. The tool's purpose and behavior are unclear, and the description fails to compensate for the missing structured data, leaving significant gaps in understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has 0 parameters with 100% schema description coverage, so the schema fully documents the absence of inputs. The description doesn't need to add parameter details, and it appropriately avoids unnecessary complexity, meeting the baseline for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get configuration' is a tautology that merely restates the tool name 'get_config' without adding specificity. It doesn't clarify what configuration is being retrieved, from where, or for what purpose, making it vague and minimally informative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like 'get_content' or 'get_health'. The description lacks any context about prerequisites, timing, or comparisons with sibling tools, leaving the agent with no usage direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_cookiesD

Get Cookies (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

D1.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure, and it delivers nothing: no indication that this is a non-mutating read, no description of the returned cookie structure, no scope (current page vs. all domains).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but this is under-specification rather than conciseness — two words that add no information beyond the tool name. Length is not earned by value here.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter getter with no output schema, an agent still needs to know what the cookie data looks like and which cookies are returned; the description supplies none of that. The definition is functionally incomplete even for a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so by the rubric baseline is 4; there is nothing for the description to clarify about inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get Cookies" merely restates the tool name get_cookies with no added specificity; the parenthetical "(alias)" is noise rather than a distinguishing detail. It does convey a verb+resource, but offers zero differentiation from the sibling tools cookies, set_cookie, and clear_cookies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No usage guidance whatsoever — nothing about when to call this versus the sibling cookies tool, set_cookie, or clear_cookies. The agent is left to infer intent purely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_htmlD

Get HTML (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault
outerNo
selectorYes

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it discloses nothing: not what HTML is returned (outerHTML of the matched node?), not whether it is a live read or a cached snapshot, and not the interaction with frames or page state. "(alias)" adds ambiguity rather than behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four words is short, but this is under-specification rather than conciseness — nothing is front-loaded because nothing is stated. The parenthetical alias note consumes space without conveying information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter retrieval tool with no annotations and no output schema, the description supplies none of the missing context: no return shape, no selector syntax, no semantics for outer, and no routing away from near-identical siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description explains neither parameter. The required selector and the boolean outer (which presumably toggles outerHTML vs innerHTML) are left entirely undocumented, so the agent must guess the meaning of the only optional flag.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get HTML" names a verb and a resource, but "(alias)" is cryptic — it does not say what it is an alias of, nor how it differs from siblings like get, get_text, get_attribute, or get_page. An agent cannot distinguish this tool's actual purpose from the surrounding retrieval tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus get_text, get_attribute, get, or get_page, all of which sit in the same sibling cluster. The "(alias)" note hints at duplication but never names the primary tool or the condition that selects it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pageC

Get page information (url/title/source/text/viewport)

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It only lists what can be retrieved, never stating whether the call is read-only, whether it waits for page load, what happens when the type parameter is omitted, or any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short parenthetical clause with the verb front-loaded and zero filler. Every word earns its place in the description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and no annotations, the description is too sparse. It leaves critical behavior unstated, especially what is returned when the optional 'type' parameter is omitted and what each mode actually produces.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description merely repeats the five enum values already present in the schema. It adds no meaning for ambiguous values such as 'source' or 'viewport' and does not explain optionality or the default behavior when 'type' is omitted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb + resource ('Get page information') with the retrievable data types enumerated, so the tool's core purpose is clear. It does not distinguish itself from sibling getters such as get_url, get_title, or get_text, which prevents a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no alternatives, and no conditions are given. The enum implies selectable fields but does not route an agent between this tool and more specific siblings like get_url or get_title.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_storageD

Get Storage (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNo
typeNo

TDQS

D1.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure and delivers nothing. It does not say whether the read is scoped to a key, what happens when the key is absent, or what form the returned value takes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is short, but this is under-specification rather than conciseness — brevity here reflects a missing description, not an efficient one.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A two-parameter read tool with no annotations, no output schema, and 0% parameter coverage should explain key scoping, the storage-type enum, and the return shape. None of that is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Two parameters exist (key, and a type enum of local/session) with 0% schema description coverage, and the description supplies no meaning for either. The enum values in particular are left completely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Get Storage (alias)" merely restates the tool name and adds a parenthetical that conveys nothing. It does not state what storage is read from, what the retrieval does, or how it differs from siblings like storage, get_config, or get_cookies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool versus the closely related siblings storage, set_storage, clear_storage, or get_cookies. No prerequisites, no exclusions, no context at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_textD

Get text (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault
fallbackNo
selectorYes

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It gives no information about read-only behavior, selector requirements, fallback behavior, return format, or error handling, leaving the agent with essentially no behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but it is under-specified rather than concise. Every part of the fragment is present but none of it earns its place by adding useful information beyond the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with undocumented parameters, no annotations, no output schema, and many similar siblings, the description is completely inadequate. An agent cannot reliably infer how to invoke it or interpret its results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has two parameters (selector required, fallback optional) with 0% description coverage. The description does not explain what selector targets, what fallback does, or when fallback is used, so it fails to compensate for the missing schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get text (alias)' restates the tool name and adds an unclear parenthetical. It does not specify what text is retrieved, from what element or page context, or how it differs from siblings like get, get_html, or get_title.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no alternatives mentioned, and no conditions for selecting this tool over the many sibling retrieval tools. The single word 'alias' does not provide actionable usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_titleC

Get title (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Get' implies a read operation, but the description does not state return format, error behavior, or any other behavioral trait.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and front-loaded, which is appropriate for a simple getter. However, '(alias)' is unexplained and does not clearly earn its place, leaving the structure ambiguous rather than maximally useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with no output schema, the description should at least clarify what title is returned and what the alias refers to. Instead it is too sparse to be considered complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema description coverage is 100%. Per the rubric, zero parameters gives a baseline of 4, and no additional parameter semantics are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get title (alias)' largely restates the tool name get_title, making it close to tautological. The parenthetical '(alias)' adds ambiguity rather than clarifying scope, and there is no differentiation from siblings like get_text, get_url, or get_page.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as get, get_text, or get_page. The description provides no context, exclusions, or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_urlC

Get URL (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It offers no details beyond the action itself: no read-only confirmation, no return format, no side-effect profile, and no indication of whether it returns the current page URL or fetches a URL.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and front-loads the verb and resource. However, '(alias)' is unexplained clutter that does not earn its place and does not help an agent select or invoke the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is low complexity with no parameters and no output schema, but the description still omits what URL is returned, how it differs from get_page, and any behavioral context. It is not complete enough for an agent to invoke it confidently without checking siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and 100% schema description coverage, so the baseline is 4. The description has no parameter information to add, and none is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb 'Get' and resource 'URL', which distinguishes it from siblings such as get_title, get_html, and get_text. However, the parenthetical '(alias)' is unexplained and does not clarify which tool it aliases or what scope it covers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No indication of when to use this tool versus get, get_page, get_title, or other read-oriented siblings. The description provides no context, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

go_backC

Back

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure, and it discloses nothing: no mention of history semantics, what happens at the start of history, whether it waits for navigation to settle, or what state it leaves the page in.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One word is technically short, but this is under-specification rather than conciseness — there is no front-loaded statement of purpose or any structure for an agent to latch onto.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a trivial history-navigation tool the bar is low, but with no annotations, no output schema, and a one-word description, an agent has no confirmation of scope, side effects, or expected post-condition. Only the tool name carries any meaning.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the schema is empty, so there is no parameter semantics to explain. Per the baseline for zero-param tools, a 4 is appropriate; the description adds nothing but nothing is required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The entire description is the single word "Back", which essentially restates the tool name go_back. It never identifies the resource or domain (browser navigation history) or distinguishes itself from siblings like go_forward or reload beyond the loose implication of the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative guidance at all. The closest thing to guidance is the implicit pairing with go_forward in the sibling list, which the description itself does nothing to establish.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

go_forwardD

Forward

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

D1.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations and no explanatory text, the description carries the full burden and discloses nothing. It doesn't say whether this moves browser history forward, whether it fails silently when no forward history exists, or what state the page ends in.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single word is not concise, it is under-specified. There is no front-loaded statement of what the tool does or returns, so brevity here costs the agent the information it needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A navigation tool with no annotations, no output schema, and no parameters depends entirely on the description for meaning, and the description provides none. An agent cannot tell whether this is a browser action, when it is valid, or what happens on failure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to clarify beyond the schema. Baseline 4 applies for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Forward" merely restates the tool name in noun form. It implies browser-history traversal when read against the sibling `go_back`, but it never states a verb or resource, so an agent must infer the purpose from the name alone rather than from the description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this over `go_back`, `navigate`, or `reload`. No conditions, prerequisites, or alternatives are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

health_checkD

Health Check

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

D1.6/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing — not what is checked, what the result means, whether it is read-only, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but this is under-specification rather than conciseness — a two-word phrase that conveys no actionable information and does not front-load anything useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and no description of the check's scope or return value, an agent cannot know what this tool actually does or what a result signifies. Completely inadequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies; there is no parameter meaning for the description to add or omit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Health Check" merely restates the tool name and is a tautology. It names no verb, resource, or scope, and does not distinguish itself from siblings like status, request_stats, or reconnect.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance whatsoever on when to use this tool versus the many related siblings (status, request_stats, reconnect, cleanup). The agent is left to infer usage entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

highlightC

Highlight elements (debugging)

ParametersJSON Schema
NameRequiredDescriptionDefault
colorNo
durationNo
selectorYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It hints that the tool is for debugging, but does not disclose whether highlighting is temporary or persistent, whether it modifies the page, required permissions, or what happens after the highlight ends.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short and front-loaded, but it is under-specified rather than concise. Every word earns its place only because there are so few words; the phrase is too sparse to guide correct invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a browser automation tool with three undocumented parameters, no annotations, no output schema, and many sibling tools. The description is far too thin to be complete; it omits param semantics, behavioral effects, and usage conditions entirely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has three parameters with 0% description coverage, and the description does not mention any of them. It does not explain the required selector, the meaning or format of color, or the units and default for duration. The description adds no semantic value beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb and resource ('Highlight elements') and adds a parenthetical context ('debugging'), so the basic action is clear. However, it does not distinguish this tool from sibling visual or inspection tools, and the fragmentary phrasing leaves the exact effect ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical '(debugging)' implies a debugging context, but there is no explicit guidance on when to use this tool versus alternatives like screenshot, snapshot, or eval. No prerequisites, exclusions, or alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hotkeyD

Key combination

ParametersJSON Schema
NameRequiredDescriptionDefault
keysYes

TDQS

D1.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure but offers none. It does not state whether this simulates a key chord, triggers an application hotkey, requires focus, or has side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but it is under-specified rather than concise. Two words do not earn their place because they fail to convey any actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one required parameter, no annotations, no output schema, and no schema descriptions, the description is completely inadequate. It omits purpose, usage, parameter format, and behavioral context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single required parameter 'keys' has 0% schema description coverage, so the description must explain its format (e.g., 'Control+C' or an array). It provides no parameter semantics at all, leaving the agent unable to format the input correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Key combination' is a noun phrase that restates the tool name rather than stating a specific action. It does not distinguish this tool from siblings like press_key or type, and provides no verb to indicate what operation is performed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as press_key or type. No context, prerequisites, or exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hoverD

Hover

ParametersJSON Schema
NameRequiredDescriptionDefault
selectorYes

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, and it discloses nothing: no mention of side effects, whether a selector must exist, hover duration, or what the call returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but this is under-specification rather than genuine conciseness. A single word cannot be considered well-structured or front-loaded content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a browser-automation action tool with no annotations, no output schema, and an undocumented parameter, the definition is completely inadequate — an agent cannot reliably invoke or distinguish it from the many sibling interaction tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage for its single required 'selector' parameter, so the description must compensate — yet it says nothing about what selector accepts (CSS? XPath? text?) or its format. The single word adds no semantic value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

Tautological: description restates name/title.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance whatsoever on when to use this tool versus alternatives like click, focus, or mouse. No preconditions, no exclusions, no routing hints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

is_visibleD

check visible (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault
selectorYes

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It does not state what 'visible' means, what the return value is, whether it waits, what happens if the element is missing, or any other behavioral trait.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short, but it is under-specified rather than usefully concise. It is front-loaded only in the sense that there is almost nothing to read, and it does not earn its place with actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a browser automation tool with one undocumented parameter, no annotations, and no output schema, the description is completely inadequate. It provides no context needed for correct invocation or interpretation of results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single required selector parameter has 0% schema description coverage, and the description does not explain its format, expected value, or meaning. With no output schema and no annotation support, this leaves parameter semantics undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'check visible (alias)' essentially restates the tool name is_visible and does not specify what resource is checked or how it differs from siblings like exists, get, or assert. It is a tautology rather than a clear verb+resource statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, when not to use it, or which sibling alternatives exist. The agent is left to infer everything from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_framesB

List all iframes

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden. 'List all iframes' conveys a read-only listing operation and scope ('all'), but it does not disclose return format, ordering, or handling of cross-origin frames. For a zero-parameter listing tool, this is minimally adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded phrase with no wasted words. It is appropriately sized for the tool's simplicity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple zero-parameter schema and no output schema, the description is sufficient for invocation but incomplete regarding what the listing returns and how those results map to follow-up tools like enter_frame. It meets the minimum viable bar without being rich.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there are no parameter semantics to explain. The description does not need to add parameter meaning, and the baseline for zero-parameter tools applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'List all iframes'. It clearly states what the tool does, though it does not differentiate itself explicitly from sibling frame tools such as enter_frame or exit_frame.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives like enter_frame, exit_frame, or list_tabs. The intended context is only implied by the tool name and purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_tabsB

List all tabs

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not state that this is read-only, whether any side effects occur, the return format, tab ordering, or whether the active tab is included—only the bare listing action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words, front-loaded with the action and scope. No wasted language, and it is appropriately sized for a zero-parameter list operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity, the description is minimally adequate: no parameters, no annotations, and no output schema. However, because there is no output schema, an agent receives no information about what 'list all tabs' returns (e.g., identifiers, titles, URLs) or their order, leaving a meaningful gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description correctly adds no parameter information, and the empty schema fully covers the input surface.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('tabs'), and adds the scope 'all'. It is clear enough for an agent to know the operation, but does not differentiate the tool from siblings like new_tab, switch_tab, or close_tab.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides no when-to-use or when-not-to-use guidance, no alternatives, and no prerequisites. The agent must infer that this is for retrieving the current tab list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mouseD

Mouse operation (move/down/up/click)

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
stepsNo
actionYes
buttonNo

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing: not whether x/y are absolute coordinates, whether missing coordinates reuse the current cursor position, what the default button is, or whether steps animates movement. All of this must be inferred from the schema alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but brevity here reflects under-specification rather than economy. The single fragment is front-loaded but conveys almost no actionable content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter input tool with no annotations, no output schema, and two enums, the definition is completely inadequate. An agent cannot call this correctly without guessing at coordinate semantics, defaults, and the distinction from sibling mouse tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Five parameters with 0% schema description coverage, and the description explains none of them. The action enum restated in prose duplicates the schema verbatim, while x, y, steps, and button remain completely undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description essentially restates the tool name: 'mouse' becomes 'Mouse operation'. Listing the action values (move/down/up/click) adds a little specificity, but it only mirrors the schema enum and gives no indication of what makes this tool distinct from overlapping siblings such as click, move_mouse, mouse_down, and mouse_up.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no exclusions, and no mention of the closely-related siblings. An agent has no basis for choosing this tool over click or move_mouse for the same operations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mouse_downD

Mouse press (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault
buttonNo

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing: no mention of whether this dispatches a low-level event, whether a release is required, or whether state persists. "Mouse press (alias)" adds no behavioral context at all.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words plus a qualifier is technically short, but the parenthetical "(alias)" is noise rather than signal, and the brevity stems from under-specification rather than efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation-style input tool with zero annotations, no output schema, and an undocumented parameter, the definition is completely inadequate. An agent has no basis for calling it correctly or knowing its relationship to mouse_up.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description never mentions the button parameter or its left/right/middle options. The enum is self-explanatory, which is the only reason this is not a 1, but the description compensates for none of the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Mouse press" does convey a verb and resource, but the parenthetical "(alias)" is unexplained and leaves the agent unsure whether this is the primary operation or a synonym of another tool. It does not differentiate itself from close siblings such as mouse_up, mouse, or click.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus mouse_up, mouse, hover, or click, and no indication of the prerequisite that a matching mouse_up must follow. The agent must infer all usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mouse_upC

mouse release (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault
buttonNo

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing: not the default button, whether a prior mouse_down is required, whether it blocks, or what it returns. "(alias)" hints that it maps to another operation but doesn't say which or how behavior differs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words plus a parenthetical is under-specification rather than conciseness. Nothing is front-loaded beyond the bare action name, and no sentence earns explanatory value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a UI-mutation tool with no annotations, no output schema, and an undocumented parameter, the description is far too thin. An agent lacks the state, preconditions, and default behavior needed to call it confidently alongside mouse_down and click.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter (button) with 0% schema description coverage and only an enum backing it; the description adds no meaning about default value, valid usage, or how button interacts with a preceding mouse_down. It leaves the agent to infer that omitting button defaults to left.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"mouse release" names a specific action, so an agent can infer it releases a mouse button, but the parenthetical "(alias)" is opaque — it doesn't say what it aliases or how it differs from mouse_down, click, or mouse. The purpose is guessable but not explicitly stated as a verb+resource with scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No indication of when to use this instead of the adjacent siblings mouse_down, mouse, click, or hover. The only hint is the vague "(alias)", which gives no selection criteria for an agent choosing among mouse tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

move_mouseD

Move mouse (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
stepsNo

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and delivers nothing: no indication of whether movement is simulated instantly or animated, whether the move is relative or absolute, whether it triggers hover events, or what it returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but the brevity reflects under-specification rather than conciseness; every sentence fails to earn its place because the text conveys no actionable information. The parenthetical '(alias)' is unexplained noise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter interaction tool with no annotations, no output schema, and zero schema documentation, the definition is wholly inadequate — an agent has no basis for choosing it over mouse, hover, or drag, or for supplying the steps argument correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across three parameters. The description adds no meaning to x, y, or the optional steps parameter — notably, steps (presumably interpolation count for animated movement) is completely unexplained in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description merely restates the tool name with the addition of '(alias)', which conveys no information about scope, target surface, or behavior. It does not distinguish move_mouse from siblings such as mouse, hover, drag, or mouse_down/mouse_up, leaving the agent to guess how they differ.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance whatsoever. The description does not say whether this is for viewport coordinates, page coordinates, or a relative move, nor does it mention any alternative tool for related actions like hover or drag.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

new_tabC

Open a new tab

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It doesn't say whether the new tab becomes the active/focused tab, what happens if url is omitted (blank tab?), or whether the call is reversible/closable. 'Open a new tab' restates the name without adding behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence with zero waste and no front-loading problem, but the brevity comes at the cost of missing specification rather than from disciplined compression. Adequate as a label, not as documentation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and an optional undocumented parameter, the description is the only source of behavioral information and supplies almost none. An agent cannot determine tab focus behavior or the blank-tab case before invoking.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single parameter (url) is not referenced in the description. Critically, url is optional, and the description does not explain the no-url case. With only one self-evident parameter, the gap is small but real, and the description does not compensate for it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Open a new tab'), which an agent can distinguish from the state-reading siblings like list_tabs or get_page. However it does nothing to separate itself from switch_tab or navigate, which are the closest siblings and are the most likely to be confused with it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance at all. In a sibling set containing navigate, switch_tab, close_tab and list_tabs, the agent gets no signal on when opening a new tab is preferred over navigating the current one or switching to an existing tab.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

press_keyD

Button

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
modifiersNo

TDQS

D1.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure, and it discloses nothing: no indication of what element receives the keypress, whether focus is required, whether modifiers are held simultaneously, or what errors can occur. "Button" conveys no behavioral information at all.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is extremely short, but that brevity is under-specification rather than conciseness; one word with zero information is not a well-structured, front-loaded definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter input tool with no annotations, no output schema, and 0% schema coverage, the description supplies nothing an agent needs to invoke it correctly. It is completely inadequate for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters (key, modifiers), and the description adds no meaning: it does not explain what key values are valid (e.g. name vs. code vs. single character) or how modifiers are applied. This is exactly the case where the description must compensate and fails to.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is the single word "Button", which neither states a verb nor the resource being acted on; the tool name press_key implies keyboard input, so the description is actively unhelpful rather than merely vague. It gives an agent no basis for distinguishing this tool from siblings like click, hotkey, or type.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance whatsoever about when to use press_key versus click, hotkey, type, or focus. An agent has no way to infer the intended invocation context from "Button".

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reconnectC

Reconnect browser

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and discloses nothing: it does not say whether reconnecting drops the current session, loses page state, requires the browser to be in a failed state, or has any side effects. Two words cannot cover a session-level operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words are technically concise but reflect under-specification rather than economy, the same failure mode as a bare "Process" description. Nothing is front-loaded because nothing meaningful is said.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters, no output schema, and no annotations, the description is the only source of behavioral information, and it omits when to call it, what state change occurs, and what the caller should expect afterward. It is insufficient for a lifecycle tool sitting beside health_check and reload.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the schema is empty, so there are no parameter semantics for the description to clarify. The baseline of 4 applies since the schema already fully specifies the absence of inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Reconnect browser" essentially restates the tool name with its resource and adds no differentiating detail. An agent cannot tell how this differs from siblings like reload, navigate, or health_check, which is the key ambiguity for a browser-session tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of the condition that warrants reconnecting (e.g., a dropped session or stale connection), and no reference to alternatives such as reload or health_check. Usage is only faintly implied by the verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reloadC

Refresh the page

ParametersJSON Schema
NameRequiredDescriptionDefault
timeoutNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and falls far short. It never mentions that a reload discards unsaved form/JS state, whether it blocks until load completes, or how the optional timeout interacts with the operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is extremely short and front-loaded, with no wasted words, but this brevity comes at the cost of under-specification rather than being earned conciseness. Three words is arguably too little for a tool with a documented parameter and real side effects.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a caveated parameter, no annotations, and no output schema, the description leaves critical gaps: side effects on page state, blocking behavior, and the meaning of timeout. It is not complete enough for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description says nothing about the 'timeout' parameter, so an agent cannot know its unit, default, or whether it caps page-load waiting. It fails to compensate for the undocumented parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase gives a specific verb ('Refresh') and resource ('the page'), so an agent can tell it is a page-reload operation. It offers no differentiation from near-siblings like navigate, go_back, or stop_loading, which all affect page navigation state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to reload versus using navigate/reconnect, nor any prerequisite or side-effect warning. The agent must infer everything about selection from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

request_statsC

Request statistics

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing: not whether this is a read-only inspection or a counter-resetting operation, not what requires permission, and not what the response contains. Two words cannot cover a diagnostic tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words are not concise, they are under-specified; nothing is front-loaded because nothing is stated. A single sentence naming the metric domain and read-only nature would cost almost nothing and remove the ambiguity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and a bare tautological description, an agent has no basis for deciding when to call this tool or how to interpret the result. It is completely inadequate even for a zero-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the schema is trivially complete at 100% coverage. Baseline 4 applies, with no param-related gap to compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description merely restates the tool name with no verb or scope: it never says what statistics (counters, timings, network, memory?) are collected, over what period, or from which browser/session. Against siblings like health_check, status, and console_logs it is impossible to tell which diagnostic tool this is.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no condition that distinguishes it from health_check or status, and no statement about prerequisites. The agent is left to guess the intended call context entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retryD

Automatic retry

ParametersJSON Schema
NameRequiredDescriptionDefault
argsNo
toolYes
delay_msNo
max_retriesNo
success_checkNo

TDQS

D1.5/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure and delivers none. It does not explain retry policy, whether calls block until success or exhaust retries, what happens on final failure, or what side effects are repeated on each attempt.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words is not conciseness but under-specification; there is nothing front-loaded because there is nothing substantive behind it. The brevity hides a complex, five-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a nested object parameter, a required tool target, a success-check expression, and no annotations or output schema, the definition is completely inadequate. An agent cannot reliably invoke this tool without external knowledge of its contract.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Five parameters (including a nested args object and a success_check expression) are documented at 0% schema coverage, and the description adds no meaning for any of them. The semantics of args, tool, delay_ms, max_retries, and success_check are entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Automatic retry" is essentially a restatement of the tool name rather than a description of what the tool does. It does not say what is being retried (a tool call? a page action? a step?), what "automatic" implies for the caller, or how it differs from siblings like run_steps, batch, or reconnect.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no conditions for choosing this over alternatives such as run_steps or batch, and no note about what kinds of operations are retryable. The agent is left to guess entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_stepsA

Intelligent multi-step execution - execute multiple operation steps in one call, with built-in automatic waiting and error retry. Suitable for multi-step processes such as login and form filling, 10 times faster than step-by-step calls

ParametersJSON Schema
NameRequiredDescriptionDefault
stepsYesArray of operation steps, executed in order
auto_waitNoAutomatically wait for the element to be ready (default true)
step_timeoutNoDefault timeout ms for each step (default 10000)
retry_on_failNoAutomatically retry on failure (default true)
stop_on_errorNoStop on error (default true)
max_step_retriesNoRetries after the first attempt, 0 to 10 (default 2)
return_intermediateNoReturns the intermediate step results (default false)

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the full burden. It discloses built-in automatic waiting and error retry, which is useful behavioral context, but says nothing about default failure behavior (stop_on_error), what is returned, or partial-failure semantics beyond what the schema already shows.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with what it does and then when to use it. The '10 times faster' claim is mildly promotional but still carries selection value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex orchestration tool with a nested steps array, no output schema, and no annotations, the description covers the core value but omits return-format and error-handling behavior that an agent would want before relying on it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all seven parameters including nested step fields. The description adds no parameter-level detail (e.g. step ordering, retry interaction), so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource ('execute multiple operation steps in one call') and names itself as multi-step execution. It distinguishes itself from step-by-step calls, though it does not address the sibling 'batch' tool, which could plausibly overlap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names concrete use cases (login, form filling) and the alternative (step-by-step calls) with a comparative claim ('10 times faster'). Missing an explicit exclusion, e.g. when a single 'batch'/'retry' sibling should be preferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshotD

Screenshot

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo
fullPageNo
selectorNo

TDQS

D1.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full behavioral burden, yet it discloses nothing: not the output format, whether a file is written, resolution, or any side effects. A one-word description leaves behavior entirely opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single word is technically brief but is under-specification rather than conciseness; there is no front-loaded intent or structure to speak of.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter tool with no annotations and no output schema, the description is completely inadequate. An agent cannot determine how to invoke it correctly or what it returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are three parameters (path, fullPage, selector). The description adds no meaning for any of them; it does not explain what path designates, that fullPage captures the entire scrollable page, or that selector scopes the capture. No compensation for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

Tautological: description restates name/title.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as snapshot, get_page, or get_html. No context, prerequisites, or exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scrollC

unified scrolling (supports top/bottom/element/text/coordinates/relative/direction/loading more)

ParametersJSON Schema
NameRequiredDescriptionDefault
byNo
toNotop/bottom/center or selector
textNoScroll to text
amountNo
withinNoScroll within the element
positionNo
directionNo
load_moreNo
max_scrollsNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it discloses almost nothing behavioral: whether scrolling waits for completion, how amounts are interpreted, what load_more does (trigger infinite scroll? loop?), or how max_scrolls bounds it. The mode list describes surface capability, not runtime traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense parenthetical with zero filler, and the core action is front-loaded in 'unified scrolling'. It is terse to the point of being a capability dump, but every token carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, 0-required, no-annotation, no-output-schema tool with nested objects, this description is far too thin. Key behaviors and parameter interactions (by vs position, amount, max_scrolls) are left for the agent to guess.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% across 9 parameters, so the description must compensate, and it does partially: coordinates maps to by/position, relative to offsets, text, direction, and loading more map to their params. However the crucial distinction between 'by' and 'position', the meaning of 'amount', and max_scrolls semantics remain undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb 'scrolling' plus 'unified' clearly identify the action, and the enumerated modes (top/bottom/element/text/coordinates/relative/direction/loading more) scope its capability set. It is distinguishable from siblings like mouse/hover/navigate, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lists what modes exist but gives no guidance on when to pick scroll over alternatives (e.g., mouse, hover, navigate, wait_for) or in what order to combine options. The 'loading more' hint implies an intent but does not state conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

selectD

drop-down selection

ParametersJSON Schema
NameRequiredDescriptionDefault
textNo
indexNo
valueNo
selectorYes

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, yet it discloses nothing about behavior — not whether selection replaces or adds, what happens on failure, or how the four parameters interact. 'drop-down selection' is a label, not a behavioral statement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but this is under-specification rather than conciseness; the single fragment cannot earn its place because it conveys essentially no information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter tool with no annotations, no output schema, and a same-domain sibling (select_option) that appears to do the same thing, this description is entirely inadequate to support correct invocation or disambiguation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Four parameters exist with 0% schema description coverage, and the description explains none of them. The mutual exclusivity of text/index/value versus the required selector is completely undocumented, leaving the agent to guess how to form a valid call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'drop-down selection' only gestures at the action and resource without a clear verb, and it does nothing to distinguish this tool from the sibling select_option, which appears to overlap entirely. An agent cannot tell which of the two to invoke from this text.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus select_option, checkbox, set_checked, or toggle_checkbox. No prerequisites or context are offered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

select_optionD

Select options (aliases)

ParametersJSON Schema
NameRequiredDescriptionDefault
valueYes
selectorYes

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral disclosure burden. It does not explain whether selecting an option mutates page state, whether it requires focus or visibility, what happens on failure, or what the return value is.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short and front-loaded, but it is under-specified rather than concise. The single sentence does not earn its place because it adds almost no information beyond the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% schema description coverage, the description is completely inadequate for a two-required-parameter interaction tool. It leaves the agent without the information needed to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two required parameters. The description does not clarify what 'selector' identifies or what 'value' should contain, so an agent gets no semantic help beyond the bare parameter names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Select options (aliases)' essentially restates the tool name and adds only the cryptic word 'aliases'. It does not distinguish this tool from siblings like select, set_checked, or toggle_checkbox, so an agent cannot confidently tell what specific operation this performs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, when not to use it, or which sibling tool it replaces. The description provides no context for choosing select_option over select, checkbox, or fill.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_checkedD

Set check (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault
checkedNo
selectorYes

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing. It does not say whether the operation is idempotent, whether the target must exist, what happens if 'checked' is omitted, or whether it waits for the element.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words is not conciseness but under-specification. There is nothing to front-load because no substantive information is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation tool with no annotations and no output schema, the description is completely inadequate; an agent has essentially no basis for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description adds no parameter meaning whatsoever. 'selector' (a CSS/selector string?) and 'checked' (the desired boolean state?) are both undocumented in schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Set check (alias)' gestures at setting a checkbox-like state but never states the resource explicitly or what element it operates on. With siblings like check, checkbox, and toggle_checkbox, the agent cannot tell from the description how this differs from them, and '(alias)' only adds ambiguity about what it aliases.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no conditions, no alternatives named. The agent gets no help choosing between set_checked, check, checkbox, and toggle_checkbox.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_configD

Setup configuration

ParametersJSON Schema
NameRequiredDescriptionDefault
fast_timeoutNo
long_timeoutNo
default_timeoutNo

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, yet it says nothing about whether this is a mutation, what gets overwritten, permission requirements, or persistence. 3 undocumented timeout parameters are left completely unexplained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words are technically short, but this is under-specification rather than conciseness. There is no front-loaded information an agent could act on, and no structure to speak of.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation-style config tool with no annotations, no output schema, and three undocumented parameters, the description is wholly inadequate. An agent has no basis for choosing it over get_config or for supplying correct values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description adds no meaning for fast_timeout, long_timeout, or default_timeout — no units (seconds/ms), no defaults, no interaction between them. The description provides zero compensating detail for three opaque numeric parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Setup configuration" essentially restates the tool name set_config without adding a specific verb+resource distinction. It doesn't differentiate from the sibling get_config, so an agent cannot tell what setting this tool does that get_config doesn't, or what part of the config it touches.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus get_config or set_debug, no prerequisites, and no mention of which configuration values are affected. Nothing indicates when a call is appropriate or when to avoid it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_debugC

Set debug mode

ParametersJSON Schema
NameRequiredDescriptionDefault
enabledYes

TDQS

C2.2/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden, and it discloses nothing: not whether debug mode persists, what it changes (logging verbosity, console capture), whether it is safe or reversible, or who can call it. A mutation-style setter with zero disclosed behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely short and front-loaded with no wasted words, but the brevity reflects under-specification rather than efficiency. It is appropriately sized only because it says almost nothing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is structurally simple (one boolean), but with no annotations and no output schema, the description should at least explain the observable effect of toggling debug mode and its persistence. That is entirely absent, leaving the agent unable to predict the call's consequences.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description says nothing about the single 'enabled' boolean beyond what the name implies. Booleans are largely self-documenting, which prevents a 1, but the description adds no value over the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a verb ('Set') and resource ('debug mode'), which is more than a tautology, but it doesn't distinguish itself from siblings like set_config or console_logs, nor does it say what debug mode governs. The purpose is inferable but thin.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No indication of when to enable debug mode, what it affects, or how it relates to set_config or console_logs. The agent gets no context for choosing this tool over the config setter.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_storageD

Set Storage (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
typeNo
valueNo

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are absent, so the description carries the full burden and discloses nothing: not that this mutates browser storage, not whether it overwrites or appends, not whether a page/context must exist first, and not what happens to an existing key.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but this is under-specification rather than conciseness — a bare name-and-alias string with no front-loaded purpose statement, so brevity buys nothing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter mutation tool with zero annotation coverage, no output schema, and an unexplained enum, the description omits everything an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all three parameters. The description does not explain that key is the storage key, that value is the payload, or that the local/session enum selects the storage backend — the only genuinely ambiguous choice in this schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Set Storage (alias)" only restates the tool name; it names no resource scope, no target (localStorage vs sessionStorage), and does not distinguish itself from siblings such as get_storage, clear_storage, or storage. The "(alias)" tag adds no actionable meaning.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no conditions, and no mention of the four closely related siblings (storage, get_storage, clear_storage, set_cookie) that an agent must choose between. The agent is left to infer everything from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

snapshotB

Get the page ARIA snapshot as YAML

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It only mentions the output format and says nothing about side effects, required page state, permissions, or whether it is a safe read operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no wasted words. It delivers the core information immediately and is appropriately sized for a simple getter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations mean the description should explain the return value and usage context more fully. It only says 'ARIA snapshot as YAML' without defining what that contains or when an agent should choose it over siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and an empty schema, so there are no parameter semantics to document. The baseline score of 4 applies for a no-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('page ARIA snapshot') with output format ('as YAML'). This clearly distinguishes it from sibling getters like get_html, get_text, and screenshot by specifying the ARIA accessibility tree as the target.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as get_html, get_text, or screenshot. It only states what it does, leaving selection criteria entirely to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

statusC

Get browser status

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden, and it discloses almost nothing: it does not say what fields the status contains, whether it errors when the browser is down, or how it differs from health_check. 'Get' weakly implies a non-mutating read, which is the only behavioral signal present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three words, front-loaded with the verb, and nothing redundant. The terseness is arguably excessive, but as a conciseness measure it is near-optimal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description should explain what 'status' returns or what the agent can do with it, especially alongside a sibling called health_check. Instead it leaves the return values and the readiness semantics entirely unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema-coverage baseline of 4 applies. There is nothing for the description to disambiguate beyond the trivial empty input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb+resource pairing ('Get browser status') is clear enough to identify a read of browser state, but 'status' is vague and the definition makes no attempt to distinguish itself from siblings such as health_check, get_url, or get_title. An agent cannot tell exactly what 'status' means without invoking it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no conditions, and no mention of alternatives like health_check. The description simply states the action, leaving the agent to infer that this is the go-to readiness probe.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stop_loadingC

Stop loading

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, and it discloses nothing: not whether an in-progress navigation is cancelled or merely paused, not what happens if no load is active, and not whether it throws or returns silently. "Stop loading" adds no context beyond the name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words are technically concise but this is under-specification rather than efficiency, mirroring the LOW calibration case for a single-word description. Nothing is front-loaded because nothing is said.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter control tool with no annotations and no output schema, the description should at minimum state what is stopped, what happens to the pending navigation, and any follow-up expectations. None of that is present, leaving the agent to guess.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and the schema is empty with 100% coverage, so there is nothing for the description to document. The baseline of 4 applies for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description "Stop loading" merely restates the tool name and adds no further specificity. It does convey a verb+resource pair by implication (stop an in-progress page load), but it does not differentiate this from siblings like wait, wait_for, reload, or navigate, all of which relate to load state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus waiting tools (wait, wait_for, wait_for_text) or navigation controls (navigate, reload, go_back). An agent must infer that this aborts an in-flight load, with no stated preconditions or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

storageC

Storage operation (get/set/remove/clear)

ParametersJSON Schema
NameRequiredDescriptionDefault
keyNo
typeNo
valueNo
actionNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does not disclose whether set/remove/clear mutate persistently, what clear actually wipes (all keys vs. one), whether the key/value are required for which action, or any auth/scope context. The mutation semantics are left entirely implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single terse fragment with no wasted words and the action list is front-loaded. But it is under-specified rather than genuinely concise, with no structure to indicate the per-action parameter requirements.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter, multi-mode tool with no annotations, no output schema, and zero schema description coverage, this definition is far too thin. An agent cannot tell which parameters each action requires or what a get returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All four parameters have 0% schema description coverage, so the description must compensate. It only echoes the four action values, leaving key, value, and the local/session type enum unexplained, and giving no hint that action dictates which other params are required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the resource (storage) and enumerates the four actions, so the general purpose is inferable. However, it is a generic multiplexer whose siblings get_storage, set_storage and clear_storage each do a subset of this more precisely, and the description does nothing to distinguish itself from them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this combined operation versus the dedicated get_storage/set_storage/clear_storage siblings, nor any note about required parameters per action. Nothing beyond the bare action list is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

switch_tabD

Switch tabs

ParametersJSON Schema
NameRequiredDescriptionDefault
indexYes

TDQS

D1.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It states nothing about what switching tabs changes, whether the index is zero-based, what happens if the tab does not exist, or how focus/state is affected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is only two words, which is terse but not sufficiently informative. Brevity here reflects under-specification rather than efficient communication.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tab-switching operation, the description still omits essential context: what index means, how it relates to tab order, and how it differs from adjacent tab-management tools. With no annotations and no output schema, this is incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one required parameter, index, with 0% description coverage. The description does not mention the parameter at all, leaving its meaning entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a verb and resource ("Switch tabs"), so the general action is inferable. However, it does not distinguish the tool from siblings like new_tab, close_tab, or list_tabs, nor does it specify that it activates a tab at a given index.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus new_tab, close_tab, list_tabs, or navigation tools. The agent must infer its role purely from the name and siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

toggle_checkboxD

Toggle checkbox (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault
selectorYes

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it discloses nothing: not the mutation semantics of toggling, not whether the element must be visible/in-frame, not what happens for disabled checkboxes, and not the return value. It is a bare label.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and front-loaded, but the brevity reflects under-specification rather than efficiency. "(alias)" is noise that consumes the only informative slot.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A mutating interaction tool with no annotations, no output schema, undocumented parameters, and a crowded set of competing siblings gets essentially no orientation from this description. It is not complete enough for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single required selector parameter, so the description must compensate and adds nothing — no syntax (CSS/XPath/text), no required/uniqueness semantics. An agent cannot know what to pass.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Toggle checkbox" essentially restates the tool name, and the only added word, "(alias)", does not clarify what resource or effect is involved. It never distinguishes itself from siblings like checkbox, set_checked, or check, which an agent must choose between.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus checkbox or set_checked, nor any prerequisite or context. The word "alias" hints at duplication with a sibling but actively muddies rather than directs selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

typeC

unified input (supports selector/label/placeholder/index, fast/human/slow mode)

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo
textYes
clearNo
delayNo
indexNoby index
labelNoFind by label
timeoutNo
selectorNo
if_existsNo
placeholderNoSearch by placeholder
press_enterNo

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions fast/human/slow modes, which hints at behavioral variation, but does not explain what those modes do, how typing affects the page, or what permissions/requirements exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short parenthetical phrase, so it is concise and front-loaded. However, its extreme brevity leaves it under-specified and fragmentary rather than efficiently informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter input tool with no annotations and no output schema, the description is far too sparse. It hints at locator options and typing modes but omits core behavior, parameter meanings, and usage context an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 27%, so the description must compensate for undocumented parameters. It names selector, label, placeholder, index, and mode, but adds no meaning beyond the schema and omits the required 'text' parameter plus clear, delay, timeout, if_exists, and press_enter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description says 'unified input' and lists supported locator strategies, which implies typing into an element, but it never states the specific action clearly. It does not differentiate this tool from sibling tools like 'fill' or 'fill_form'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives such as 'fill' or 'press_key'. It only lists capabilities, leaving context and exclusions entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_fileD

File upload

ParametersJSON Schema
NameRequiredDescriptionDefault
filesYes
selectorYes

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and delivers nothing. It does not say whether the selector targets an <input type=file>, whether the upload is triggered immediately or deferred to a form submit, whether multiple files are supported, or what happens on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two words is short but this is under-specification rather than conciseness; there is no substantive content to be front-loaded. Brevity here costs the agent all actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A two-parameter tool with 100% required params, zero schema descriptions, no annotations, and no output schema gets essentially no documentation. The definition is wholly inadequate for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description adds nothing. It never explains what 'selector' identifies or what form 'files' takes (string vs array, local path vs URL vs base64), despite the schema's oneOf union making this ambiguity acute.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"File upload" merely restates the tool name with no verb+resource specificity beyond the obvious. It does not distinguish this tool from siblings like 'fill', 'set_storage', or 'eval', which could also plausibly feed data into a page.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance whatsoever on when to use this tool, what prerequisites exist (e.g. a file input must already be located/focused), or how it relates to alternatives. The agent is left to guess entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

waitC

Unified waiting (supports time/element/disappear/text/URL/loading/network idle/DOM stable/JS conditions)

ParametersJSON Schema
NameRequiredDescriptionDefault
msNowait milliseconds
textNo
typeNo
stateNo
patternNoURL match
timeoutNo
selectorNo
expressionNoJS expression

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It lists condition types but does not state whether the tool blocks execution, what the default timeout is, what happens on timeout (error vs silent return), or what the return value is. These are critical for a wait operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no wasted words, and the core idea 'Unified waiting' is front-loaded. The parenthetical list is efficient, though it could be structured more clearly as a bullet list if longer. Appropriately concise for the amount of information conveyed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 8 parameters, low schema coverage (38%), no annotations, and no output schema, the description is far too sparse. It omits default timeout behavior, blocking semantics, error handling, and how to construct a wait condition. It is not complete enough for an agent to use the tool reliably.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 38% (8 parameters, 3 with descriptions). The description loosely maps some condition types but does not explain how to use selector, state, timeout, or how type interacts with other parameters. It adds minimal meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it is a 'unified waiting' tool and lists supported condition types, giving some clarity on what it does. However, it does not distinguish this tool from siblings like wait_for and wait_for_text, leaving ambiguity about when to choose each. The purpose is vague beyond the name itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The parenthetical merely lists supported conditions but offers no context, prerequisites, or exclusions. An agent has no basis for selecting this over wait_for or wait_for_text.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_forD

Wait for element (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault
stateNo
timeoutNo
selectorYes

TDQS

D1.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It says nothing about blocking behavior, timeout handling, failure modes, or what states are supported. This is a severe gap for a waiting tool with no other structured behavioral hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short and front-loaded, but it is under-specified rather than appropriately concise. For a tool with three parameters and no other documentation, this brevity leaves critical call information missing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (three parameters, no annotations, no output schema), the description is completely inadequate. It omits parameter meanings, waiting conditions, and any behavioral details needed for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not mention any of the three parameters (selector, state, timeout). The required 'selector' and optional 'state'/'timeout' are left entirely undocumented, so the description fails to compensate for the schema's lack of descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a verb ('wait') and a resource ('element'), and the '(alias)' parenthetical indicates it duplicates the 'wait' sibling. However, it does not clarify what condition the element is waited for (visible, attached, etc.), leaving the purpose only partially distinguished from related tools like 'wait_for_text'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The '(alias)' hint implies the tool is interchangeable with 'wait', but there is no explicit guidance on when to use this tool versus alternatives such as 'wait' or 'wait_for_text'. No context or conditions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_for_textD

Wait for text (alias)

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
timeoutNo

TDQS

D1.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description must carry the behavioral burden. It implies a blocking wait but does not state polling behavior, what happens on timeout, whether it returns the text, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely short to the point of under-specification; it is concise but fails to include necessary information, making the brevity a liability rather than an asset.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, no annotations, and undocumented parameters, the description is inadequate for correctly invoking this tool. An agent cannot determine required behavior from it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 0% description coverage for two parameters (text, timeout). The description mentions neither parameter nor their semantics (e.g., timeout units, text matching mode), leaving them completely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description 'Wait for text (alias)' essentially restates the tool name and adds only that it is an alias, without specifying what text is waited for (appearance, change) or how it differs from siblings like wait_for or wait. This provides minimal differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus wait, wait_for, assert, or get_text. Usage context is entirely absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 75 tool updatesv4.3.0
    • First observedassert
    • First observedbatch
    • First observedblur
    • First observedcheck
    • First observedcheckbox
    • First observedcleanup
    • First observedclear_cookies
    • First observedclear_storage
    • First observedclick
    • First observedclose_tab
    • First observedconsole_logs
    • First observedcookies
    • First observedcount
    • First observeddialog
    • First observeddrag
    • First observedenter_frame
    • First observedeval
    • First observedexists
    • First observedexit_all_frames
    • First observedexit_frame
    • First observedfill
    • First observedfill_form
    • First observedfind
    • First observedfocus
    • First observedget
    • First observedget_attribute
    • First observedget_config
    • First observedget_cookies
    • First observedget_html
    • First observedget_page
    • First observedget_storage
    • First observedget_text
    • First observedget_title
    • First observedget_url
    • First observedgo_back
    • First observedgo_forward
    • First observedhealth_check
    • First observedhighlight
    • First observedhotkey
    • First observedhover
    • First observedis_visible
    • First observedlist_frames
    • First observedlist_tabs
    • First observedmouse
    • First observedmouse_down
    • First observedmouse_up
    • First observedmove_mouse
    • First observednavigate
    • First observednew_tab
    • First observedpress_key
    • First observedreconnect
    • First observedreload
    • First observedrequest_stats
    • First observedretry
    • First observedrun_steps
    • First observedscreenshot
    • First observedscroll
    • First observedselect
    • First observedselect_option
    • First observedset_checked
    • First observedset_config
    • First observedset_cookie
    • First observedset_debug
    • First observedset_storage
    • First observedsnapshot
    • First observedstatus
    • First observedstop_loading
    • First observedstorage
    • First observedswitch_tab
    • First observedtoggle_checkbox
    • First observedtype
    • First observedupload_file
    • First observedwait
    • First observedwait_for
    • First observedwait_for_text

TDQS

D1.8/5.0

Scored across 75 tools

Disambiguation1/5

Many tools are explicit aliases or unified duplicates (cookies vs get_cookies/set_cookie/clear_cookies; storage vs get_storage/set_storage/clear_storage; mouse vs move_mouse/mouse_down/mouse_up; select vs select_option; checkbox vs set_checked/toggle_checkbox; get vs get_text/get_html/get_attribute/get_url/get_title). These overlapping purposes make misselection highly likely.

Naming Consistency3/5

Mostly snake_case, but the set mixes one-word noun/verb tools (click, wait, cookies), operation-style names, and alias variants with different conventions. Readable overall, but not a predictable verb_noun pattern throughout.

Tool Count1/5

75 tools is an extreme mismatch for the apparent scope. Many are redundant aliases, inflating the surface far beyond what is needed for browser automation.

Completeness4/5

The surface broadly covers navigation, tabs, frames, interactions, waits, assertions, cookies/storage, dialogs, upload, JavaScript eval, screenshots, and page snapshots. Minor gaps remain around network interception, download handling, and permission controls.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Connects local stdio MCP servers to an existing Chrome 144+ session, preserving the user's signed-in sessions, cookies, tabs, and extension environment without launching a second browser. It provides tab control, semantic snapshots, screenshots, pointer, keyboard, form selection, scrolling, navigation, and waiting tools.
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables MCP clients to drive individually named Chrome profiles, providing tools for tabs, navigation, page interaction, screenshots, JavaScript evaluation, console logs, and network inspection over stdio without a TCP port.
    -