Skip to main content
Glama
surpasses
by surpasses

firefox-mcp

Claude in Chrome, but for Firefox. An MCP server that attaches to your running, signed-in Firefox over WebDriver BiDi and gives Claude Code the same kind of tools: list and open tabs, navigate, screenshot, read the page as refs, click, type with real key events, fill forms, evaluate JS, read the console and network log.

No Playwright, no Puppeteer, no bundled browser. Node 22 (global WebSocket), the MCP SDK and zod. Everything else is the BiDi protocol Firefox ships with.

Setup

  1. Install once:

    git clone https://github.com/surpasses/firefox-mcp ~/firefox-mcp
    cd ~/firefox-mcp && npm install
    claude mcp add --scope user firefox -- node ~/firefox-mcp/server.mjs
  2. Firefox only reads the remote-debugging flag at startup, so relaunch it with the helper. Your normal profile, tabs and logins are kept (session restore):

    ~/firefox-mcp/bin/firefox-debug --restart

    Without --restart it just tells you to quit Firefox first. Firefox shows a small robot icon in the URL bar while the remote agent is on. It listens on 127.0.0.1 only.

  3. In Claude Code, the tools appear as mcp__firefox__*. Start with tabs_context.

Port defaults to 9222; override with FIREFOX_MCP_PORT for both the launcher and the server.

Related MCP server: Firefox Browser Bridge

Tools

Tool

What it does

tabs_context

List tabs: id, url, title, active marker. Ids are stable while the server runs.

tabs_create / tabs_close / tabs_activate

Open (optionally with a URL), close, focus a tab.

navigate

URL, back or forward. Waits for load (45 s cap).

screenshot

Viewport PNG. region crops, savePath also writes a file.

read_page

Elements with refs (ref_12, or f1:ref_3 inside frame 1), role, name, value, position. filter=all adds headings and text.

find

Regex over name/label/placeholder/id/href, or a CSS selector.

get_page_text

Visible text, all frames.

click / hover / drag

By ref (scrolled into view first) or by viewport coordinates. Modifiers, double and triple click.

type

Real key events into the focused element or a ref. clear selects-all first, pressEnter submits.

press_key

Enter, Tab, cmd+a, shift+Tab, Backspace Backspace Enter, with repeat.

form_input

Set a value directly (fires input/change): text, select, checkbox, radio, contenteditable.

scroll / scroll_to

Wheel scroll at a point or ref; scroll a ref into the centre.

evaluate

Run a JS expression, optionally inside a frame (frame: "f1").

read_console / read_network

Buffered since connect, regex-filterable, clear to reset.

resize_viewport

Set the viewport size (phone widths etc.).

wait

Sleep up to 10 s.

Frames and Fission

Firefox runs cross-origin iframes out of process ("Fission"). Input actions sent to the top frame's context reach the <iframe> element but are never routed into the child process, and focusing a child field by script then typing on the top frame does not work either.

This server therefore treats frames as first class: read_page and find walk every frame and prefix refs with the frame key; every click, hover, scroll and keystroke is dispatched on the context of the frame that owns the target (by ref, by hit-testing the coordinates against the frame rectangles, or by following the focused <iframe> chain for typing). Shopify's hosted checkout card fields, Stripe Elements and similar embedded forms work through the normal click and type tools.

Test

npm test

Boots a throwaway headless Firefox on port 9333 with its own profile (your real Firefox is untouched), serves a two-origin fixture, and checks tabs, navigation, snapshot, refs, form input, real typing, cross-origin iframe input, screenshots, page text, console and network capture, history and tab close.

Not there yet

  • GIF recording of a sequence of actions.

  • Natural-language find (the Chrome one calls a model); find here is regex or CSS.

  • Downloads and file uploads.

  • Multiple simultaneous BiDi clients: Firefox allows one session, so run one server at a time.

Available Tools

22 tools
clickA

Click an element by ref (scrolled into view first) or at viewport coordinates. count=2 double-clicks, count=3 triple-clicks (select paragraph). modifiers like 'shift' or 'cmd+shift'.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
refNoElement ref from read_page or find, e.g. ref_12 or f1:ref_3 (f1 = inside frame 1)
countNo
tabIdYesTab id from tabs_context
buttonNo
modifiersNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does disclose scroll-into-view, multi-click semantics, and modifier syntax. It does not describe return values or possible navigation side effects, but the disclosed behaviors are the most decision-relevant ones.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences front-load the core operation and then add only useful behavioral details. No filler or restatement of schema constraints.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Good coverage for the core click behavior, but the schema only requires tabId while a call also needs a target (ref or x/y), and no guidance is given for combining or prioritizing these. With no output schema, return/result behavior is also unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 29%, so the description must add meaning. It maps ref to element references, x/y to viewport coordinates, explains count semantics, and gives modifier examples. Button semantics and the precedence of ref vs coordinates are left to inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear operation: click an element by ref or at viewport coordinates, and differentiates from sibling tools by naming the action and target modes. It also covers click-count variants without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for the two input modes (ref vs coordinates) and special click variants (count=2, count=3, modifiers). It does not explicitly compare against hover/drag alternatives, but no misleading or exclusionary guidance is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dragC

Drag with the left mouse button from one point to another.

ParametersJSON Schema
NameRequiredDescriptionDefault
toXYes
toYYes
fromXYes
fromYYes
tabIdYesTab id from tabs_context

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full behavioral burden. It only notes the left mouse button, but does not disclose whether it presses and releases, requires a draggable element, or any side effects. Insufficient for an action that manipulates UI state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is concise, but it is under-specified rather than efficiently informative. It does not use the space to add value, so it earns a neutral score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 5 required parameters, no output schema, and no annotations, this description is incomplete. It fails to explain coordinate semantics, error behavior, or drag mechanics, leaving an agent unable to call it correctly in many scenarios.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% (tabId has a description). The description adds no meaning to fromX, fromY, toX, toY beyond their names, and does not clarify units, coordinate system, or relationship to viewport/page. With low coverage, the description should compensate but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear action ('drag') and resource ('mouse button'), and mentions the movement from one point to another. However, it does not distinguish it from sibling tools like click or hover, and 'one point to another' is vague about what is dragged.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No mention of when to use this tool versus alternatives. It does not state prerequisites (e.g., an element to drag) or contrast with click/scroll, leaving the agent to infer the appropriate context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evaluateA

Evaluate a JavaScript expression in the page (top frame by default) and return its value. Use frame (e.g. f1) to target an iframe. Never trigger alert/confirm/prompt.

ParametersJSON Schema
NameRequiredDescriptionDefault
frameNo
tabIdYesTab id from tabs_context
expressionYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It discloses top-frame default and warns against alert/confirm/prompt, which is useful. However, it does not mention that evaluating JS can mutate page state, trigger navigation, or that return values may need to be serializable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise, front-loaded sentences. The main purpose is stated first, followed by frame targeting and a critical behavioral warning. Every sentence earns its place with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a powerful evaluate tool with no annotations and no output schema, this is incomplete. It omits important guidance on expression syntax, promise/async handling, result serialization, and side-effect risk. An agent could invoke it incorrectly or misinterpret the returned value.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, so the description must compensate. It clarifies `frame` with an example format (e.g. f1) and explains that `expression` is evaluated and returned. But it doesn't give concrete expression examples or return-value expectations, and `tabId` is only described in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the verb ('Evaluate') and resource ('JavaScript expression in the page'), plus the return value behavior. It also distinguishes frame targeting from the top-frame default, differentiating it from sibling navigation/read tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives some context on how to target an iframe with `frame`, but it does not explicitly state when to choose evaluate over siblings like read_page or get_page_text, nor when not to use it. The warning about dialogs is a constraint, not a usage guideline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

findA

Find elements by text, label, placeholder, id, name or href (case-insensitive regex), or by CSS selector. Returns up to 20 refs per frame.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesRegex or plain text matched against name/label/placeholder/id/href
tabIdYesTab id from tabs_context
selectorNoCSS selector; when set, query is ignored

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full load. It discloses matching behavior, results are case-insensitive regex, applies to several attributes, honors selector over query, and returns up to 20 refs per frame, which is meaningful behavioral detail beyond the schema. It could go further, but it is already substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The content is tight, front-loaded, and contains no filler. The first sentence conveys the core search modes and attributes; the second sentence mentions resource limit. Everything included is useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For simple search tool with no output schema, the description covers the key behaviors and even the return cap. It still does not explain what 'refs' mean or how they are consumed by the follow-up tools, but that is largely contextual expectation rather than a lethal gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema descriptions cover 100% of parameters, so the baseline is 3. The description adds small semantic context like matching being case-insensitive and how selector overrides the query, but it does not explain tabId beyond the schema, so it doesn't push above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool finds elements by text, label, placeholder, id, name, href, or CSS selector, and gives a result cap of 20 refs per frame. This makes its purpose distinct from sibling interaction tools like hover/click and page-level functions like get_page_text.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use guidance is provided. It does not say that this tool should be used to obtain refs before action tools, nor does it mention any alternative tools to prefer in specific situations. The usage context is only implied by the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

form_inputA

Set a form control's value directly (fires input/change): text/textarea, select (option text or value), checkbox/radio (true/false), contenteditable. Prefer type for fields that validate on keystrokes (card numbers).

ParametersJSON Schema
NameRequiredDescriptionDefault
refYesElement ref from read_page or find, e.g. ref_12 or f1:ref_3 (f1 = inside frame 1)
tabIdYesTab id from tabs_context
valueYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It does so well by stating that the tool fires input/change events and directly sets values, and by clarifying value handling for select, checkbox/radio, and contenteditable. It does not mention broader side effects like validation or focus behavior, but it covers the key behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences deliver the core behavior, supported control types, and the key alternative routing. There is no filler or repetition of schema fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with no output schema, the description covers what the tool does, how values are interpreted, and when to use an alternative. An agent has enough information to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67%, with `ref` and `tabId` already documented in the schema. The description adds meaningful semantics for `value`, explaining how it maps to different control types: option text/value for select, true/false for checkbox/radio, and content for contenteditable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'Set a form control's value directly.' It enumerates the exact control types it supports, and the closing note differentiates it from `type`, a sibling tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to prefer `type` for fields that validate on keystrokes, such as card numbers. This gives the agent a clear rule for when not to use `form_input` and names the alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_page_textA

Visible text of the page (innerText), including frames. Up to ~80k chars per frame.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdYesTab id from tabs_context

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the behavioral burden. It discloses that only innerText is returned, that frames are included, and that output may be truncated around 80k chars per frame—useful behavioral context beyond the tool's name. It does not explicitly state it is read-only, but the operation is clearly non-mutating.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence conveys the core behavior, scope, and a key limitation with no filler. Every word adds value and the most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one well-documented parameter and no output schema, the description adequately explains what the caller will receive and its limitations. It could add a note about error behavior or when output might be empty, but the current coverage is sufficient for a simple read operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already fully documents the only parameter, tabId, including its source (tabs_context). The description adds no parameter-specific meaning, so the baseline score of 3 is appropriate given full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific operation—get the visible text of the page (innerText)—and adds clarifying scope details: it includes frames and has a per-frame character cap. This clearly distinguishes it from sibling tools like screenshot or find.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool should be used when an agent needs the visible text content of a page, but it gives no explicit guidance on when to prefer this over alternatives such as read_page or find. No exclusionary or routing information is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hoverA

Move the mouse over an element (by ref) or a point, to reveal menus and tooltips.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
refNoElement ref from read_page or find, e.g. ref_12 or f1:ref_3 (f1 = inside frame 1)
tabIdYesTab id from tabs_context

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden; it clearly discloses that the tool moves the mouse over an element or point and does not click. It does not enumerate side effects, but for a low-risk hover action the behavior is adequately transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler; every clause adds meaning (action, target, purpose).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and the description covers the core action, but it leaves gaps: x/y coordinate semantics, precedence if both ref and point are provided, and the effect of providing neither. These matter for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema documents ref and tabId but leaves x and y undescribed; the description compensates partially by saying a point can be used, which maps to x/y. It does not specify coordinate space or whether x and y must be supplied together.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Move') and names the exact resource ('an element (by ref) or a point'), with the intended effect 'to reveal menus and tooltips.' This clearly distinguishes hover from sibling actions like click or scroll.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states when to use it: whenever an element or coordinate needs a mouse-over to reveal menus/tooltips. It does not explicitly mention alternatives or exclusions, but the context is clear enough relative to click/scroll siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

press_keyA

Press keys: 'Enter', 'Tab', 'Escape', 'ArrowDown', 'cmd+a', 'shift+Tab', or several space-separated ('Backspace Backspace Enter'). repeat repeats the whole sequence.

ParametersJSON Schema
NameRequiredDescriptionDefault
keysYes
tabIdYesTab id from tabs_context
repeatNo

TDQS

A3.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It does explain the 'repeat' parameter's behavior but omits side effects (e.g., whether the keypress requires focus, if there are rate limits, or what happens on invalid keys). For a tool with no annotations, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, well-structured sentence that front-loads the purpose and provides concrete examples. Every word earns its place; there is no fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is sufficient for invoking the tool with valid key sequences and repeat, but lacks any error handling, success/failure signaling, or preconditions (e.g., active tab). Given no output schema and no annotations, an agent may be uncertain about edge cases. It is adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33% (only tabId has a description). The description compensates by thoroughly explaining the 'keys' format with examples and defining 'repeat' as repeating the whole sequence. It adds meaning beyond the schema for two of three parameters, though tabId is left to the schema's one-liner.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource ('Press keys') and immediately enumerates valid key names and formats, including multi-key sequences. It clearly distinguishes this low-level keyboard tool from siblings like 'type' (for text) and 'click' (for mouse). The purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies its usage for keyboard shortcuts and special keys, and the sibling list suggests alternatives for other actions. However, it does not explicitly state when not to use it or name a specific alternative for text entry. The context is clear but exclusions are missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_consoleA

Console messages and uncaught errors buffered since the server connected. pattern is a regex filter; clear=true empties the buffer after reading.

ParametersJSON Schema
NameRequiredDescriptionDefault
clearNo
limitNo
tabIdNoTab id from tabs_context; omit for all tabs
patternNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses a key side-effect (clear=true empties the buffer) and the buffered nature of the data source, which informs the user about data freshness and mutation. However, it does not state whether reading is non-destructive by default, though this is implied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences: the first front-loads the core purpose, and the second adds clarifications for two parameters. There is no filler, and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Without an output schema or annotations, the agent still needs to know the return format and the default behavior of 'clear'. The description does not describe the output shape or explain 'limit', so there are gaps in the information needed to invoke the tool confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%, so the description must compensate. It explains 'pattern' as a regex filter and 'clear' as emptying the buffer, but does not clarify 'limit' (e.g., number of messages). This partially compensates but leaves a parameter undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific resource ('console messages and uncaught errors') and a scope ('buffered since the server connected'). This clearly distinguishes it from sibling tools like read_page and read_network, which handle page content and network requests respectively.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives. The purpose implies using it when console messages are needed, but no exclusions or alternative tool mentions are provided, leaving usage to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_networkB

Network requests buffered since the server connected: method, status, mime, url. pattern filters url/status/mime by regex.

ParametersJSON Schema
NameRequiredDescriptionDefault
clearNo
limitNo
tabIdNoTab id from tabs_context; omit for all tabs
patternNo

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose the buffer window and filterable fields, but it omits the side effect of the `clear` parameter, whether invoking the tool can mutate the buffer, default/limit behavior, or what shape the response takes. This is a significant gap for a potentially state-changing read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, front-loaded with the resource and returned fields, followed by a focused explanation of the filter parameter. No wasted words or redundant restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and four parameters, this description is under-specified. It does not clarify return structure, buffer clearing semantics, limit defaults, or how tab filtering behaves beyond the schema snippet. An agent could call it, but not with full confidence about side effects and output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%, so the description must compensate. It explains `pattern` (regex over url/status/mime) but leaves `clear` and `limit` undocumented; `tabId` has a schema description but still lacks behavioral context. The description adds value for one parameter but does not cover the others.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource as buffered network requests and enumerates the fields returned (method, status, mime, url), with an additional note on the pattern filter. It is distinguishable from siblings like read_console and read_page by data type, though it does not explicitly name those alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: inspect network traffic buffered since connection, optionally filtered by regex. However, there is no explicit guidance on when to use this over read_console, read_page, or other siblings, and no exclusions or conditions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_pageA

List page elements with refs for click/type/form_input. filter=interactive (default) lists controls and links; filter=all adds headings and text. Crosses into iframes (refs prefixed fN:).

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdYesTab id from tabs_context
filterNo
maxItemsNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses key behavioral traits: default filter value, iframe crossing with fN: ref prefixes, and the difference between filter modes. 'List' implies read-only, though side effects are not explicitly ruled out.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences, front-loaded with the main purpose, followed by filter behavior and iframe handling. No unnecessary words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool, the description covers purpose, filtering, and iframe behavior. It misses maxItems semantics and return format details, but the core usage is clear. Without an output schema, a bit more on return structure could help, but it is sufficient for most calls.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (tabId), so the description must compensate. It explains the filter parameter in detail (default and modes) but says nothing about maxItems, leaving that parameter undocumented. Partial compensation for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and resource ('page elements with refs for click/type/form_input'), clearly distinguishing it from sibling tools like get_page_text or find. The filter explanation further clarifies the tool's scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies when to use it: before interactions like click/type/form_input, to obtain refs. It also explains filter modes and default behavior. However, it does not explicitly name alternative tools or state when not to use it, so it provides clear context without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resize_viewportA

Set the tab's viewport size in CSS pixels (e.g. 390x844 for phone).

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdYesTab id from tabs_context
widthYes
heightYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the behavioral burden. It clearly states the effect and uses a precise verb, but does not mention side effects like re-layout, page reload, persistence, or required tab state. Adequate for a simple setter but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence states the action, object, unit, and a useful example with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-required-param primitive tool with no output schema, the essential call facts are present: what to set, the units, and a realistic value pair. Missing 'when' guidance is minor for this level of complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema describes only tabId; width and height are bare integers. The description compensates partially by giving units ('CSS pixels') and a concrete example ('390x844'), but it does not document ranges, constraints, or the relationship between width and height.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise action ('Set'), the resource ('the tab's viewport size'), and the unit ('CSS pixels'), with a concrete phone-size example. This clearly distinguishes it from all listed siblings such as navigate, screenshot, or tabs_activate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The '390x844 for phone' example implies mobile/responsive viewport testing, but there is no explicit statement of when to call this tool rather than alternatives or whether it should precede screenshots. No sibling provides the same function, so routing is less critical.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screenshotB

Screenshot the visible viewport as PNG. region [x0,y0,x1,y1] crops (zoom). savePath also writes the file.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdYesTab id from tabs_context
regionNo
savePathNo

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description bears full responsibility for behavioral disclosure. It mentions the side effect of savePath writing a file, which is useful, but it does not describe the return value (e.g., PNG data format) or whether the operation is non-destructive. It also lacks details on how region coordinates are interpreted (e.g., relative to viewport). This gap is significant for a tool with no annotation safety hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: two sentences. The first front-loads the core purpose, and the second efficiently explains the two optional parameters. There is zero filler or redundant phrasing, making it highly scannable for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple screenshot tool, the description covers the essential action, cropping, and file saving, but it omits critical context: the return format (e.g., does it return a base64 PNG?), coordinate system for region, and any caveats about scrolling or viewport boundaries. Since there is no output schema and no annotations, the description should be more thorough to fully equip the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds meaning to region and savePath beyond the schema: it explains region as a crop box using [x0,y0,x1,y1] and savePath as an optional file write. However, it does not fully specify the coordinate system (e.g., pixels relative to viewport) or whether cropping implies scaling ('zoom' is ambiguous). With schema coverage at only 33%, the description partially compensates but leaves important details unclear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it screenshots the visible viewport as PNG, specifying the verb (screenshot), resource (visible viewport), and output format. It distinguishes from read_page (which likely reads content) but does not explicitly name a sibling, so it's clear but not fully differentiated. It also mentions region cropping, adding scope specificity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus alternatives like read_page or resize_viewport. The description focuses on mechanics (region, savePath) but lacks context on when a screenshot is preferred over other observation tools. No exclusions or prerequisites are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scrollB

Scroll with the mouse wheel at a point (default viewport centre) or over a ref. amount = wheel ticks of 100px.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
refNoElement ref from read_page or find, e.g. ref_12 or f1:ref_3 (f1 = inside frame 1)
tabIdYesTab id from tabs_context
amountNo
directionYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It adds useful detail (default viewport centre, amount=100px ticks) but doesn't disclose side effects, reversibility, or limitations (e.g., page-specific behavior).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, dense sentence that front-loads the core action and meaning of amount, with no redundant words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter tool with no annotations or output schema, the description covers core usage but omits edge cases like precedence when both coordinates and ref are provided, and doesn't mention any return behavior. Adequate for basic calls, but not exhaustive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is low (33%), so the description compensates by explaining x/y (point), ref (over a ref), and amount (wheel ticks of 100px). Direction and tabId are either self-evident (enum) or already documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (scroll with mouse wheel) and the target (at a point or over a ref), with a note on amount units. It distinguishes from scroll_to by implying wheel-based scrolling, but doesn't explicitly contrast with any sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like scroll_to. The description provides no conditions, exclusions, or recommended scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scroll_toA

Scroll an element (by ref) into the centre of the viewport.

ParametersJSON Schema
NameRequiredDescriptionDefault
refYesElement ref from read_page or find, e.g. ref_12 or f1:ref_3 (f1 = inside frame 1)
tabIdYesTab id from tabs_context

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It clearly states the scroll behavior (centering in viewport), which is useful. However, it doesn't disclose edge cases like whether the scroll is instant/animated, what happens if the element is not found, or whether it affects scrollable containers vs the whole page.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, clear sentence that front-loads the action and target. No wasted words; every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with full schema coverage, the description is mostly complete. However, given the sibling 'scroll' tool exists, the description could clarify the distinction (element-level vs page-level scrolling) to prevent mis-selection. No output schema means return behavior is not documented, but for a scroll action this is less critical.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters. The description adds the behavioral context that the ref is used to identify the element to scroll, but doesn't add meaning beyond the schema's parameter descriptions. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('scroll') and resource ('an element by ref') with a clear target ('into the centre of the viewport'). It distinguishes from the sibling 'scroll' tool by specifying element-level scrolling via ref, though it doesn't explicitly name the sibling or contrast with viewport scrolling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: use this when you need to bring a specific element into view, as opposed to general page scrolling. However, it doesn't explicitly state when to use this vs the sibling 'scroll' tool, nor does it mention prerequisites like needing a ref from read_page or find (though the schema hints at this).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tabs_activateB

Bring a tab to the front.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdYesTab id from tabs_context

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It states the action but does not disclose side effects, such as changing the active tab, whether it returns any value, or behavior on invalid tabId. For a mutation tool, this lacks important transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no extraneous content. It is concise and to the point, achieving high efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one parameter, the description is minimal but does not explain return values (no output schema) or error handling. The agent knows what it does but not what to expect after calling, which is a notable gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides full parameter documentation: tabId is described as 'Tab id from tabs_context' with 100% coverage. The description adds no additional meaning beyond the schema, so it meets the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Bring a tab to the front' clearly states the action and resource. It distinguishes from siblings like tabs_create and tabs_close by implying activation of an existing tab, though it could be more explicit about the 'front' meaning (e.g., active tab in current window).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance. The description does not mention that this tool is for switching to an existing tab, nor does it compare with alternatives like navigate or tabs_create. The parameter schema hints that tabId comes from tabs_context, but the description itself offers no contextual direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tabs_closeC

Close a tab.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdYesTab id from tabs_context

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full burden of behavioral disclosure. 'Close a tab' implies a destructive action, but it does not state the consequences, reversibility, prerequisites, or potential failure modes. The description only states the trivial action without clarifying the actual behavioral impact, which is a significant gap for a tool that executes a destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, short sentence that gets to the point with no wasted words. For a simple, one-parameter tool, this is appropriately concise, but it is also so bare that it offers no structure beyond a statement of purpose. It earns its place, but nothing more.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (one parameter, no output schema) and the sibling context, the description is minimally adequate. The schema provides the parameter semantics and the sibling list conveys the domain (browser tabs). Yet there is no description of when the operation might fail, what the result is, or whether the tab is closed from the user's browser view. The description leans heavily on the schema and context rather than enriching them.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% coverage for the single parameter tabId, including its description and origin from tabs_context. The tool description adds no further parameter semantics; it merely repeats the action. Because the schema already handles the parameter meaning, a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Close a tab.' uses a specific verb and resource, and it is clearly distinct from sibling tools like tabs_activate, tabs_create, and tabs_context. However, it does not explicitly contrast itself with those siblings; it is simply a standalone action. That still provides enough clarity for an agent to identify the tool's core purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool relative to alternatives. There is no mention of when closing is appropriate, no prerequisites (e.g., that the tab ID must come from tabs_context, which is in the schema), and no exclusions. The agent must infer usage from the schema and sibling names alone, which is insufficient for a mutation tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tabs_contextA

List Firefox tabs (id, url, title, active). Call this first; tab ids are stable while the server runs.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It discloses the stable nature of tab ids, which is valuable state-related information, and the listed output fields clarify what the call returns. It does not mention read-only status explicitly, but the verb 'List' and the absence of mutation language make the behavior reasonably clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences deliver the resource, returned fields, and the critical usage instruction without any filler. Every part earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter list tool with no output schema, the description is complete: it names the output fields and the key behavioral fact (stable ids), which is all an agent needs to invoke it correctly and use the results with sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description does not need to compensate for undocumented inputs. The baseline of 4 applies, and no parameter explanation is required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List'), resource ('Firefox tabs'), and enumerates the returned fields (id, url, title, active). This clearly differentiates it from sibling tab management tools like tabs_create, tabs_close, and tabs_activate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear usage guidance: 'Call this first' and explains that tab ids are stable while the server runs, which signals that this tool should be used to obtain tab identifiers for later operations. It does not explicitly name alternatives or exclusion cases, but the context strongly implies its role relative to sibling tabs_* tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tabs_createA

Open a new tab, optionally navigating to a URL. Returns its tab id.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of disclosing behavior. It states the action and return value, which covers the essential side effect of creating a tab, but it does not mention whether the new tab becomes active, what happens if the URL is omitted, or error behavior on invalid URLs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, well-structured sentence states the action, optional parameter, and return value with no wasted words. The essential information is front-loaded and immediately actionable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one optional parameter and no output schema, the description is nearly complete: it states the operation, the optional input, and the return value. Minor omissions like default behavior without a URL or focus/activation behavior are understandable gaps but not critical for basic invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only the property name 'url' and type 'string' with 0% description coverage, so the description must compensate. It confirms that the URL is optional and clarifies its semantic role, but it does not explain default behavior when omitted or expected URL format. This adds some meaning but not comprehensive guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Open a new tab') and resource, clearly distinguishing it from sibling tools like tabs_activate, tabs_close, and navigate. It also states the key output ('Returns its tab id'), making the tool's purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies this tool is for creating a new tab, optionally with a URL, but it does not explicitly state when to prefer it over navigate, tabs_activate, or tabs_close. Usage context is present but exclusions and alternatives are left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

typeA

Type text with real key events into the focused element, or into ref (clicked first). clear=true selects all and deletes first. pressEnter=true submits.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoElement ref from read_page or find, e.g. ref_12 or f1:ref_3 (f1 = inside frame 1)
textYes
clearNo
tabIdYesTab id from tabs_context
pressEnterNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full transparency burden. It discloses that typing uses 'real key events', that a ref is 'clicked first', that clear does select-all-and-delete, and that pressEnter submits. It omits edge behavior like missing focus or errors, but the key behavioral traits are present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no filler. The primary behavior is front-loaded, and each clause adds a distinct piece of useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a side-effectful typing tool with five parameters, the description plus schema cover required inputs and all optional behaviors. It does not describe return values or failure modes, but those are less critical for a keyboard automation tool without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 40%, so the description must compensate. It adds meaning beyond the schema for ref ('clicked first'), clear ('selects all and deletes first'), and pressEnter ('submits'). Text and tabId are not elaborated, but text is self-evident and tabId has a schema description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Type text with real key events'), a target ('focused element' or 'ref'), and includes the side-effect flags. It does not explicitly call out sibling tools like form_input or press_key, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus alternatives such as form_input or press_key. It only explains how the target is chosen and what flags do, not when this tool should be preferred or avoided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

waitA

Wait N seconds (max 10) for the page to settle.

ParametersJSON Schema
NameRequiredDescriptionDefault
secondsYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It states that the tool waits a specified number of seconds and caps at 10, which is transparent about the basic behavior. However, it does not mention whether the wait is unconditional, whether it returns anything, or whether it interacts with page load state beyond 'settle'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that front-loads the action and the key constraint. Every part serves a purpose, and there is no unnecessary verbosity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter wait tool with no output schema, the description is largely complete: it explains the action, the parameter, and the intended context. It could mention that the wait is a fixed delay rather than a condition-based wait, but overall it gives enough to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, so the description must compensate. It does clarify that the parameter is a duration in seconds and includes the maximum value, but the parameter name 'seconds' and schema constraints already convey much of this. The description adds minimal extra semantic value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the action ('Wait N seconds') and the resource/context ('for the page to settle'), so an agent understands what the tool does. It does not explicitly differentiate from sibling tools, but the operation is distinct enough among the listed siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for the page to settle' implies when the tool should be used, but it does not provide explicit conditions, exclusions, or alternatives. An agent gets a general sense of use but no concrete guidance on when not to use it or when a different tool would be better.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 22 tool updatesv0.1.0
    • First observedclick
    • First observeddrag
    • First observedevaluate
    • First observedfind
    • First observedform_input
    • First observedget_page_text
    • First observedhover
    • First observednavigate
    • First observedpress_key
    • First observedread_console
    • First observedread_network
    • First observedread_page
    • First observedresize_viewport
    • First observedscreenshot
    • First observedscroll
    • First observedscroll_to
    • First observedtabs_activate
    • First observedtabs_close
    • First observedtabs_context
    • First observedtabs_create
    • First observedtype
    • First observedwait

TDQS

B3.4/5.0

Scored across 22 tools

Disambiguation5/5

Each tool targets a distinct browser automation concern—tab management, navigation, element discovery, input, scrolling, and diagnostics—so an agent can reliably select the right one. Even similar tools like read_page and find have clear separation: one lists elements, the other searches by selector/text. No two tools appear to do the same job.

Naming Consistency3/5

Naming is readable but mixes conventions: tabs_* forms a clear prefix group, read_* and get_* group read operations, while interaction tools use bare verbs like click, type, scroll, and hover. This category-based pattern is understandable but not a uniform verb_noun convention.

Tool Count3/5

At 22 tools, the set is comprehensive but on the heavy side, falling into the borderline range for a single server. Most tools are individually justified for browser automation, yet some input-related tools (type, form_input, press_key) could potentially be consolidated.

Completeness4/5

The surface covers tab lifecycle, navigation, element interaction, scrolling, screenshots, scripting, and diagnostics, which is strong coverage for browser automation. Minor gaps like explicit cookie management, file upload, or waiting for a specific element exist, but agents can work around them with evaluate and wait.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers