Skip to main content
Glama

firefox-use

Browser control for Claude Code, on Firefox.

Claude Code can drive Chrome. This drives Firefox, with the same tools under the same names: screenshots and clicks, typing and keys, forms, tabs, page reading, console and network logs, file uploads, and GIF recordings of a whole flow. A browser task written for one browser runs on the other.

It talks to Firefox over WebDriver BiDi, which Firefox speaks natively - no extension, no geckodriver, and no npm dependencies: the WebSocket client, the MCP server, the PNG decoder and the GIF encoder are all in src/, on Node built-ins.

you: "log into the staging site and screenshot the dashboard"
     -> read_page finds the form, form_input fills it, computer clicks and screenshots

Install

As a Claude Code plugin - the tools and the skill in one step, with the server bundled, so nothing is downloaded at install time:

claude plugin marketplace add leetomo2528/firefox-use
claude plugin install firefox-use@firefox-use

From npm, if you would rather register it yourself:

npx firefox-use install          # registers the MCP server with Claude Code
npx firefox-use install-skill    # adds the usage skill to ~/.claude/skills
npx firefox-use doctor           # checks Firefox, the profile, the port, the handshake

From source:

git clone https://github.com/leetomo2528/firefox-use.git
cd firefox-use && node bin/firefox-use install && node bin/firefox-use install-skill

Start a new Claude Code session afterwards. Firefox launches by itself on the first tool call.

Needs Firefox 129+ and Node 18+. Developed and tested on macOS; Linux runs in CI. Windows is not supported yet.

Related MCP server: Browser Tools for Claude Code

The tools

Tool

What it does

tabs_context_mcp

The tabs this session can act on. Call it first

tabs_create_mcp / tabs_close_mcp

Open and close tabs of your own

navigate

Go to a URL, back, or forward

computer

Click, double/triple/right click, hover, drag, type, keys, scroll, screenshot, zoom, wait

read_page

Accessibility tree with a ref_N on every element

find

"the search box", "the accept button" - refs by description

form_input

Set a field by ref: text, number, checkbox, select

get_page_text

Readable text, markup stripped

read_console_messages

Console output, filtered by a regex you supply

read_network_requests

Requests with status, type, timing and size

javascript_tool

Evaluate an expression in the page

browser_batch

Several calls in one round trip, stopping at the first error

resize_window

Resize the window a tab lives in

file_upload / upload_image

Fill a file input; drop a screenshot onto a page

gif_creator

Record a flow and export an annotated GIF

Plus three that are ours alone: firefox_status, firefox_restart, and firefox_dialog - Chrome's integration is finished by a modal dialog, here you can answer it.

Command line

firefox-use start [url]      # launch Firefox with BiDi enabled
firefox-use status           # what is running, which profile, which tabs
firefox-use open <url>
firefox-use shot page.png [url] [--full]
firefox-use stop | restart | doctor
firefox-use mcp              # the MCP server on stdio (what Claude Code runs)

Options: --headless, --port <n>, --profile <dir>. Environment: FIREFOX_BIN, FIREFOX_USE_PORT, FIREFOX_USE_PROFILE, FIREFOX_USE_HEADLESS=1, FIREFOX_USE_DOWNLOADS, FIREFOX_USE_UPLOAD_ROOTS, FIREFOX_USE_AUTO_RESTART=0.

How it works

Claude Code ──stdio JSON-RPC (MCP)──▶ bin/firefox-use mcp
                                        │
                    src/tools.js         registry, shared session, downloads
                    src/pages.js         navigate, tabs, javascript_tool, form_input, resize
                    src/a11y.js          read_page, find, get_page_text
                    src/computer.js      mouse, keyboard, screenshots, zoom
                    src/observers.js     console and network capture
                    src/uploads.js       file_upload, upload_image
                    src/gif.js/gifenc.js recording, PNG decode, GIF89a encode
                    src/batch.js         browser_batch
                                        │
                    src/session.js       tab group, element refs, screenshots, input
                    src/bidi.js          WebDriver BiDi commands and events
                    src/ws.js            RFC 6455 client
                                        ▼
                        firefox --remote-debugging-port 9222

Worth knowing:

  • A profile of its own at ~/.firefox-use/profile. Log in there once and it sticks; your everyday Firefox is never touched. firefox-use stop is the off switch.

  • macOS launches through open -n -a. A Firefox started as a child of the shell inherits the shell's permissions and dies with "Could not find profile folder" under macOS app-data protection; open hands the launch to launchd instead.

  • A fixed 1280x800 viewport per tab, at devicePixelRatio 1, so screenshot pixels and click coordinates share one coordinate space. resize_window changes it deliberately, per tab.

  • Refs are minted per document. After a navigation, a ref from the old page reports that it is gone rather than resolving to whatever now sits at that index.

  • One automation session per browser. The server closes its session on shutdown; firefox_restart clears one left behind by a crash.

Development

npm test                    # unit tests, no browser
npm run test:integration    # + a real headless Firefox, end to end
npm run spec                # generate the parity fixture from your Claude Code install
npm run sync-plugin         # copy the skill into the plugin and stamp versions

npm run spec writes .spec/chrome-tool-spec.json - the Chrome integration's tool definitions, read out of the Claude Code binary on your own machine. test/parity.test.js compares our schemas against it and test/prose.test.js checks that our wording is not borrowed from it. The file is never committed; those tests skip without it.

Limits

  • The Firefox window is the whole world: no OS-level input, no about: privileged pages.

  • The accessibility tree covers the top-level document; cross-origin iframes are not walked (click by coordinate, or use javascript_tool).

  • One browser instance per port.

Notices

MIT licensed. An independent project - not affiliated with Mozilla or Anthropic. Tool names and parameter shapes match Claude Code's Chrome integration on purpose, so tasks are portable; the descriptions, the skill and the documentation are written for this project. See NOTICE.

Available Tools

20 tools
browser_batchA

Push a whole run of browser tool calls through one request rather than spending a turn on each. Every entry is {name, input}, where input is the same object you would hand that tool on its own. Entries run one after another, never in parallel, and the run halts at the first failure: you get the output of everything that finished, the error from the entry that broke, and a count of the entries that never started. The list is checked before anything touches the page, so an unknown tool name, or an attempt to put browser_batch inside itself, rejects the call outright instead of half-applying it. Batch aggressively whenever you can see two or more moves ahead - open the page, click into the search box, type the query, press Return, take a screenshot. Images come back inline among the text outputs, each labelled with the entry that produced it. Two things bite here. First, any coordinate you write in these entries, and any ref you name, come from a screenshot taken BEFORE the call or a read_page that ran before it, since nothing this batch paints is visible to you while you compose it - and once an entry navigates or reloads, every earlier ref and coordinate is stale, so end the batch there and look again. Second, every entry that acts on a page has to state tabId - inside a batch no tab is inferred, so a batch that creates a tab and then drives it must repeat that new id on each later entry. Unlike the Chrome extension, firefox-use runs no per-site permission check between entries; it drives its own automation profile. A modal dialog does still freeze the page, so clear it with firefox_dialog rather than letting the rest of the batch queue up behind it.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionsYesThe ordered calls to make, one or more, run from the top down. Example: [{"name":"navigate","input":{"tabId":7,"url":"https://docs.example.org/search"}},{"name":"computer","input":{"tabId":7,"coordinate":[420,260],"action":"left_click"}},{"name":"computer","input":{"tabId":7,"text":"budget report","action":"type"}}]

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Details execution order (sequential, never parallel), failure behavior (halts at first failure, returns partial output, error, and skipped count), pre-validation (rejects unknown tools and nested batches before touching the page), stale reference caveats, mandatory tabId requirement, and differences from the Chrome extension. With no annotations provided, this description carries the full transparency burden and exceeds it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence contributes essential operational detail—no filler or repetition. It is organized into a clear flow: general behavior, validation, then two specific 'bites' and an extension difference. The structure aids comprehension despite the density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a batching tool with multiple failure modes and safety caveats, the description covers everything an agent needs: how to structure entries, what happens on failure, how validation works, tabId requirements, stale reference hazards, permission differences, and modal dialog handling. Despite no output schema, the return behavior is described sufficiently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already defines the actions array and its items, but the description adds real meaning: it explains the exact shape of each entry, provides a concrete example, clarifies that tabId must be repeated, and notes that browser_batch cannot be nested. This goes well beyond the schema's field names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('push a whole run... through one request') and resource ('browser tool calls'), and distinguishes it from spending a turn on each call. It is immediately clear what the tool accomplishes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells when to use it ('whenever you can see two or more moves ahead') and contrasts with alternatives ('rather than spending a turn on each'). It also specifies when not to use it (nesting a batch, after navigation, when a modal dialog appears), leaving no ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computerA

Drive one Firefox tab with synthetic pointer, wheel and keyboard input, and capture what it looks like. Each call performs exactly one action from the enum, against the tab named by tabId; ask tabs_context_mcp for a valid id before the first call.

  • Aim with coordinate (viewport pixels) or, preferably, with ref (an element id from read_page or find). Coordinates must come from your latest screenshot: pages scroll and re-render, and stale numbers land on whatever occupies that spot now.

  • A ref belongs to the snapshot that produced it. After a navigation or a re-render, ask read_page or find for new ids rather than reusing old ones.

  • Put the pointer in the middle of a control, not on its border. When a click appears to have done nothing, take a fresh screenshot and re-aim before repeating it.

  • Firefox specifics: this drives a profile the server owns, separate from your everyday browser, and no per-site approval step stands in the way. A native alert, confirm, prompt or beforeunload freezes the tab, and every action here fails until firefox_dialog accepts or dismisses it.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoElement handle from an earlier read_page or find call, written like "ref_1". `scroll_to` requires one and accepts nothing else. For clicks, hover, drag and scroll it stands in for `coordinate`, and it is the sturdier of the two: the element is looked up and scrolled into view at the moment the call runs. With `type` it means "click this field first, then type into it". Handles expire with the snapshot that produced them - navigate, or let the page re-render, and you need fresh ones. A handle that resolves to something hidden (a closed menu, an inactive tab panel) comes back as an error instead of a click at the page origin.
textNoPayload for two of the actions. Under `type` it is the literal string to enter, newlines included - each newline goes out as an Enter press. Under `key` it names a key or a chord, or several of them separated by spaces to be pressed in order: "Enter", "cmd+a", "ArrowDown ArrowDown Enter". The modifier names understood are ctrl, shift, alt, cmd (meta) and win (windows), attached to the key with "+"; macOS spells with cmd what Windows and Linux spell with ctrl. Chords that rescale the page - "ctrl+-", "cmd+0" and the like - are refused, since the new scale would be invisible to you; crop with the `zoom` action instead.
tabIdYesThe tab this action is performed on, from the tab group this session owns. The schema marks it required: tabs_context_mcp lists the ids that exist, tabs_create_mcp returns the id of a tab it opens, and a page opened from a tab already in the group joins it. An id from outside the group is rejected.
actionYesWhich operation to run. Every other field is read according to this choice. * `left_click`: press and release the primary button on the target. * `right_click`: secondary-button click, the usual way to raise a context menu. * `double_click`: two primary clicks in quick succession - opens an item, selects a word. * `triple_click`: three in a row, which selects the whole line or paragraph under the pointer. * `hover`: move the pointer onto the target and press nothing, to raise a tooltip, drop a menu open or trigger a :hover style. * `left_click_drag`: hold the primary button down from `start_coordinate` (or from wherever the pointer already sits) and release it on the target. * `scroll`: turn the wheel `scroll_amount` notches in `scroll_direction` over the point you give, so the scrollable box under that point is what moves. * `type`: send `text` to whatever holds keyboard focus; pass `ref` to click that field first. * `key`: press the keys or chords named in `text`, `repeat` times over. * `wait`: pause for `duration` seconds while something loads or animates. * `screenshot`: capture the visible area of the tab. * `zoom`: capture `region` alone, for reading small icons or fine print. * `scroll_to`: bring the element named by `ref` into view; it needs no coordinate.
regionNoCrop box for `zoom`, required by it and read by nothing else: [x0, y0, x1, y1] in viewport pixels, top-left corner first and bottom-right second. x1 has to exceed x0 and y1 exceed y0, and the whole box must fit inside the viewport. What comes back is that rectangle at capture resolution, which is what makes a 16-pixel icon or a line of fine print legible.
repeatNoHow many times to replay the whole sequence in `text`: a whole number from 1 to 100, and 1 when you leave it out. Only the `key` action looks at it. One call carrying repeat 20 beats twenty calls when you are walking a list with ArrowDown or emptying a field with Backspace.
durationNoSeconds the `wait` action sits idle, anywhere from 0 up to 10. Required by `wait`, read by nothing else. Fractions count (0.5 is half a second), and one call will never sit longer than 10 seconds, so wait a second time - with a screenshot in between - when a page needs more.
modifiersNoKeys held down for the length of a pointer action - the four clicks, hover and drag all honour it. Name one of ctrl, shift, alt, cmd (or meta) or win (or windows), and join several with "+" when you need more than one, as in "ctrl+shift". This is how you shift-click a range or cmd-click a link into its own tab. Optional, and unrelated to the chords the `key` action parses out of `text`.
coordinateNoTarget point as [x, y] in CSS pixels measured from the top-left corner of the viewport, not of the screen. Give this or `ref` for left_click, right_click, double_click, triple_click, hover, scroll and left_click_drag; on a drag it marks where the button comes up, the end of the movement. Take the numbers from your most recent screenshot - an older one may name a spot the page has since moved. A tab here renders at devicePixelRatio 1, so one pixel of that image is one unit here; no scaling to do. `type` ignores this field, because typing follows keyboard focus.
save_to_diskNoOnly meaningful for `screenshot` and `zoom`: set it true to also write the captured image into the firefox-use downloads folder, and the result line will carry the file path so you can attach it to a message. The picture is returned in the response either way, so leave this off while you are merely looking at the page yourself.
scroll_amountNoWheel notches per `scroll` call, from 1 to 10; omit it and you get 3. A notch is worth roughly 100 pixels of travel here, so the default moves about 300.
scroll_directionNoWhich way the wheel turns during `scroll`, and mandatory there: up, down, left or right. `down` walks further down the document, the way a real wheel would.
start_coordinateNoWhere `left_click_drag` puts the button down, as [x, y] in viewport pixels; `coordinate` or `ref` supplies the point where it lifts again. Optional - leave it out and the drag begins wherever the previous action left the pointer, which is only predictable when that action was one of yours.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description takes on the full burden of behavioral disclosure. It is remarkably transparent: it explains that hidden elements return errors, that certain chords are refused to avoid creating an invisible zoom, that screen captures are returned and optionally saved to disk, and that actions fail while a dialog is open. These details leave very little to discover at runtime.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place. It is well-structured with a clear summary, a bulleted list of actions, and detailed per-parameter explanations that are logically grouped and easy to scan. There is no redundancy or fluff, and the content is densely packed with useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the tool (13 parameters, 13 enum values, and interactions with sibling tools), the description is exceptionally complete. It explains all actions, parameter dependencies, return values (screenshots, file paths), and cross-tool prerequisites (e.g., asking tabs_context_mcp for a tab ID). Without an output schema, the description adequately covers what the agent can expect back.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 13 parameters and the description covers 100% of them with rich context. For example, it explains the lifecycle of `ref` handles, the pixel mapping of `coordinate` (CSS pixels, devicePixelRatio 1), the exact behavior of `text` under `type` vs `key`, and the semantics of `scroll_direction`. This goes far beyond the raw schema and gives the agent everything it needs to use each parameter correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Drive one Firefox tab with synthetic pointer, wheel and keyboard input, and capture what it looks like') and clearly distinguishes this tool from siblings like tabs_close_mcp, navigate, and read_page. It also hints at the screenshot capability, leaving no ambiguity about the tool's core function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides extensive, actionable guidance: how to choose actions, when to use `ref` vs `coordinate`, how to obtain a valid tab ID via tabs_context_mcp, how to interpret results, and what to do when actions fail (e.g., take a fresh screenshot). It also explains which parameters are required for which actions and gives practical tips like avoiding stale coordinates.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_uploadA

Puts one or more files from this machine into an on the page. Never try to click the input or the button that fronts it: Firefox answers a click with the operating system's own picker window, which sits outside the page where nothing here can reach it. Locate the input with read_page or find instead, then call this tool with the ref it was listed under. All three arguments - paths, ref and tabId - are required. Reading a file is gated: a path is accepted only when it resolves inside an upload root, meaning this session's download folder, the directory the server process was started in, or an absolute path named in the FIREFOX_USE_UPLOAD_ROOTS environment variable (separate several with the platform path separator). Symlinks are followed before that check, so a link parked in a root but aimed outside one is turned down as well, and the refusal lists the roots that would have worked. Chrome instead scopes uploads to whatever the user shared with the conversation; this server has no such notion. One call may carry less than 10 MB in total across every file named. Afterwards the input is read back and the result states how many files it now holds, which is how you catch an input with no multiple attribute keeping just one of several.

ParametersJSON Schema
NameRequiredDescriptionDefault
refYesThe id read_page or find printed beside the file input, in the form ref_12. A ref belongs to the snapshot it came from, so read the page again after a navigation or a re-render rather than reusing an older one.
pathsYesOne absolute path per file, with a leading ~ expanded for you. Folders are turned down; name the files themselves. Each entry has to sit, after symlinks are resolved, in an upload root - this session's download folder, the directory the server was launched from, or an entry of FIREFOX_USE_UPLOAD_ROOTS. Sizes are added up and a total of 10 MB or more fails the call before the page is touched at all.
tabIdYesThe tab holding that input. Run tabs_context_mcp when you do not already have a live id. There is no per-site approval step as in Chrome - the file lands in the page straight away, in the dedicated Firefox profile this server drives.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full responsibility and does an excellent job. It discloses critical behaviors: path restrictions based on upload roots, symlink resolution, total size limit of 10 MB, read-back verification, and the lack of per-site approval (unlike Chrome). It also warns about Firefox's native picker opening on click. This is exceptionally transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but information-dense. Every sentence carries a distinct operational point—caveats, prerequisites, constraints, verification—so the length is justified. It is structured as flowing prose rather than bullet points, which is acceptable given the density of important behavioral warnings. A slight trim could make it more scannable, but it is not redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is complex—file path restrictions, symlink behavior, size limits, ref lifecycle, and cross-browser differences—and the description covers all of it. It even explains the output (the input is read back and the result states file count) despite having no output schema. It provides enough context for an agent to use the tool correctly without needing additional external documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers all three parameters with detailed descriptions (100% coverage), so the baseline is 3. The tool description goes beyond the schema by explaining the semantic implications: ref belongs to a snapshot and becomes stale after navigation, paths cannot be folders and must resolve to upload roots, and tabId requires a live tab with no approval step. This adds meaningful context beyond the parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: putting files from the machine into a file input on the page. It distinguishes itself from sibling tools by explicitly directing users away from clicking and toward read_page/find for locating the input, and by contrasting with Chrome's file-sharing model.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit when-to-use guidance: never click the input, use read_page or find to get the ref, and call this tool with all three required arguments. It also explains when to use tabs_context_mcp for tabId and how to verify success by reading back the input's file count. It names alternatives (read_page, find, tabs_context_mcp) and conditions, leaving no ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

findA

Look up elements by describing them in ordinary language instead of by selector: "the search box at the top", "whatever button adds this to a cart", "the heading that mentions pricing". Scoring runs over the same tree read_page builds - accessible names, an element's own text, and attributes such as id, name, placeholder, title and href - with extra weight for a handful of intents (search, login, cart, menu, close, submit, dropdown, checkbox, heading). Results come back ranked, best first, each with the ref_N handle the action tools need, and at most 20 are printed; when more than that matched, the output tells you to describe the target more tightly. The search reaches 30 levels deep and the first 5000 elements walked, and it says so when it stopped short - so an empty result on a huge page means widen the net with read_page, not that the element is absent. query and tabId are both required, and handles expire with the document, so run find again after any navigation; a browser_batch step must state the id itself. As everywhere on this server the automation profile skips any per-site approval step, while an unanswered alert or confirm holds the page until firefox_dialog clears it.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesDescribe the target the way you would to a person: by its job ("login button", "the search field in the header") or by wording visible on it ("Add to cart", "organic mango"). A few concrete words beat a full sentence; filler such as "the", "please" and "click" is dropped before matching.
tabIdYesRequired: the session tab to search. Ask tabs_context_mcp or firefox_status for valid ids; only the tabs this session owns are addressable, and a browser_batch step must spell the id out.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so thoroughly. It discloses the search ranking approach, the 20-result limit, the 30-level and 5000-element traversal bounds, handle expiration, and the behavior around alerts and approval prompts. This gives the agent a solid model of side effects and limitations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but mostly purposeful. It front-loads the core purpose and then covers algorithm details, limits, and usage edge cases. A few points about tabId are repeated between the main description and the schema, but overall it is efficiently organized and each clause adds operational value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and lack of an output schema, the description is remarkably complete. It explains what results look like (ranked list with ref_N handles), how many can be returned, what happens when limits are exceeded, when results become stale, and how alerts affect execution. This gives the agent enough context to use the tool reliably without hidden surprises.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds substantial semantic detail beyond the schema: it explains how to formulate the query, that filler words are dropped, that only session-owned tabs are addressable, and how to obtain valid tab IDs. This helps the agent invoke the tool correctly in real workflows.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: looking up elements by natural language description rather than by selector. It gives concrete examples and distinguishes itself from read_page by referencing the same tree, making the tool's role unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides practical usage guidance: how to phrase queries, what to do when too many results are returned, and when to fall back to read_page on an empty result. It also notes that handles expire after navigation and that browser_batch must spell out tabId. It does not explicitly contrast with every sibling tool, but the main alternative is covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

firefox_dialogA

Accept or dismiss a native dialog (alert, confirm, prompt, beforeunload) that is blocking the page. Firefox-only: in Chrome a blocking dialog ends the automation session, here it can be cleared.

ParametersJSON Schema
NameRequiredDescriptionDefault
textNoText to enter into a prompt
tabIdNoTab showing the dialog. Use tabs_context_mcp if you do not have a valid tab id.
actionNoDefault accept

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the burden. It mentions that the dialog is blocking the page and can be cleared, but does not detail side effects such as navigation consequences after accepting or dismissing, or behavior when no dialog exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, using two sentences to convey purpose and platform-specific context. No redundant or irrelevant information is included.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for an agent to understand when the tool applies (blocking native dialogs in Firefox). Given there is no output schema and the tool's scope is narrow, the provided context is mostly sufficient, though a note about required action (accept vs dismiss) being mandatory would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers all three parameters with descriptions, giving 100% coverage. However, the descriptions are minimal (e.g., 'Default accept' for action) and do not add deeper meaning beyond the schema itself. The description text adds no additional parameter context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (accept or dismiss) and the target (native dialog such as alert, confirm, prompt, beforeunload). It also distinguishes the Firefox-specific behavior from Chrome, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context about Firefox-only behavior but does not explicitly state when to use this tool versus alternatives like javascript_tool or navigate. No clear conditions or scenarios are given for choosing this tool over others.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

firefox_restartA

Restart Firefox. Use this when the browser stops responding or reports a stale automation session. Firefox-only.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description must carry the full burden. It states that the tool restarts Firefox, which implies termination and restart, but it does not disclose potential side effects like losing open tabs, clearing session state, or whether the action is destructive. Given the nature of a restart, some behavior is implied, but explicit disclosure of consequences is missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the action ('Restart Firefox') and then adds the usage condition. There is no fluff, and every word serves a purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, parameterless tool with no output schema, the description provides enough context: what it does, when to use it, and its scope. It does not describe the return value, but that is likely irrelevant for a restart operation. The only minor gap is the lack of explicit side-effect disclosure, which is already penalized under behavioral transparency.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema is empty. According to the rubric, a baseline of 4 is appropriate when no parameters exist. The description adds no parameter information, which is fine since there are none to explain.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear, specific action: 'Restart Firefox.' It identifies the resource (Firefox) and the operation, and distinguishes it from siblings like firefox_status by specifying the action and context. It also notes 'Firefox-only,' which sets scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit usage conditions are provided: 'Use this when the browser stops responding or reports a stale automation session.' This tells the agent exactly when to invoke it, and the 'Firefox-only' qualifier warns against use on other browsers. No alternative tools are named, but the condition is clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

firefox_statusA

Report whether Firefox is running, which profile, port and viewport it uses, and which tabs are in this session's group. Firefox-only; the Chrome integration has no equivalent.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It conveys a read-only operation via the verb 'Report' and specifies the exact information returned, which implicitly signals no side effects. It does not detail potential authentication or environment constraints, but for a status tool this is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that front-loads the core purpose and appends the Firefox-only caveat. There is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-output-schema tool, the description covers the essential behavior and scope. It could mention the output format, but since the tool is a straightforward status report and the description lists what it reports, it is sufficiently complete for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4 per the rubric. The description adds no parameter-specific details (none are needed), and the empty schema requires no compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Report') and a specific resource (Firefox status), and enumerates the exact data points: running status, profile, port, viewport, and session tabs. It also explicitly distinguishes itself from the Chrome integration, making it clearly unique among siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly states this is Firefox-only and that the Chrome integration has no equivalent, which implies when to use it (when Firefox-specific status is needed) and when not to (Chrome contexts). It doesn't explicitly name alternatives, but the Firefox vs. Chrome distinction is a clear usage pointer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

form_inputA

Fill one form control that read_page has already tagged with a ref, without aiming a click at it. ref, value and tabId are all required. Text boxes, textareas, number and date fields, contenteditable regions, checkboxes, radios and dropdowns are covered; a file input is not, and goes through file_upload or upload_image, which read only from the directories this server is allowed to open. The write goes through the element's native setter and is followed by input and change events, while checkboxes and radios are toggled with a real click, so React and its relatives notice the change. A dropdown is matched against option values first, then against the labels shown to the user, ignoring case on a last pass; when nothing matches, the error lists the options that exist. Disabled controls are turned down. A ref lives only as long as the node behind it, so after a navigation, a reload, or a re-render that swaps the element out, run read_page again or this fails with an unknown-ref error.

ParametersJSON Schema
NameRequiredDescriptionDefault
refYesThe element id read_page handed out, shaped like ref_1 or ref_2. It names one live node, so take it from the most recent read of this page; older ids are usually stale.
tabIdYesThe tab holding the element; required. Ids come from tabs_context_mcp, and the ref has to belong to that same tab's document.
valueYesWhat the control should end up holding. Text-like fields take a string or a number. Checkboxes and radios take true or false (an empty string, or "false", "no", "0", "off", also reads as unchecked). A dropdown takes either the option's value attribute or the label the user sees. Arrays and objects are rejected instead of being stringified into the field.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully covers behavior: native setter, input/change events, real clicks for checkboxes, dropdown matching logic, disabled control handling, and staleness errors. This gives a complete picture of side effects and failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is detailed yet tightly organized, with each sentence conveying necessary information without redundancy. The main action and requirements are front-loaded, followed by control-specific semantics and edge cases.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers all relevant operational details: how to obtain the ref, how to pair it with a tab, how values map to controls, what happens with invalid inputs, and when errors occur (stale ref, disabled, unmatched dropdown). It is complete for an agent to use correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Even though schema coverage is 100%, the description adds significant nuance beyond the schema: it clarifies stale refs, tab ownership, accepted value types per control, and rejection of arrays/objects. This meaningfully enhances understanding beyond the raw property definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool fills a form control using a ref from read_page, and explicitly distinguishes it from clicking. It enumerates supported control types, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states required parameters and gives detailed guidance on ref freshness, tabId association, and value semantics per control type. It contrasts with file upload tools but does not explicitly enumerate all alternative tools or conditions for choosing this one over them, though the context is largely sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_page_textA

Read a page as prose: plain text, no markup, no element handles. The extractor takes the article bodies when the page has real ones, otherwise , otherwise the block with the best text-to-markup ratio, falling back to the body, and throws away navigation, banners, footers, sidebars, toolbars, select menus, iframes and anything marked hidden or aria-hidden, so a long post arrives as paragraphs rather than as the furniture around them. Reach for it when the goal is to read or quote; reach for read_page or find when you need something clickable, since nothing here carries a ref. Output stops at 200000 characters, cut on a line break, and the closing note reports how long the page actually was - only a pathological page ever hits that. tabId is required, must name a tab this session owns, and has to be spelled out on a browser_batch step. A modal alert or confirm blocks the read until firefox_dialog answers it; no site approval is asked for, since the browser runs in its own profile.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdYesRequired: the tab whose readable text you want. Valid ids come from tabs_context_mcp or firefox_status; tabs outside this session's group cannot be read, and a browser_batch step has to name the id itself.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses extraction priorities, content filtering, truncation at 200000 characters, the closing note about actual length, modal blocking until firefox_dialog answers, and that no site approval is requested due to the browser profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a clear first sentence, but it contains some stylistic redundancy and filler, such as 'so a long post arrives as paragraphs rather than as the furniture around them' and 'only a pathological page ever hits that.' These phrases add color but do not strictly earn their place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description adequately explains the output format ('plain text, no markup'), the 200000-character truncation, and the closing note reporting actual length. It also describes extraction priorities and error-relevant behavior like modal blocking, though exact error formats are not specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, tabId, already has 100% schema coverage, but the description adds meaningful constraints beyond the schema: 'Valid ids come from tabs_context_mcp or firefox_status', 'tabs outside this session's group cannot be read', and 'a browser_batch step has to name the id itself.'

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Read a page as prose: plain text, no markup, no element handles.' It also explicitly distinguishes itself from sibling tools: 'Reach for it when the goal is to read or quote; reach for read_page or find when you need something clickable, since nothing here carries a ref.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance and names alternatives: 'Reach for it when the goal is to read or quote; reach for read_page or find when you need something clickable.' It also notes important usage constraints like tab ownership, browser_batch self-naming, and modal-dialog blocking behavior.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gif_creatorA

Films a tab as a run of screenshots and encodes them into an animated GIF, with the clicks and drags of the run marked on the frames that show what they did. Four actions, and every one of them needs tabId because a recording belongs to a single tab: start_recording opens the loop and grabs a first frame right away, stop_recording ends the loop after one closing frame and keeps everything captured, export encodes the frames into a .gif on disk, clear throws them out. The loop grabs a frame about every 700 ms and halts itself once 120 are in hand; frames are shrunk to at most 800 px wide, and the delay stored in the GIF follows the cadence that was really measured, not the nominal one. Calling export while the loop still runs stops it first, and the frames outlive the export - clear them when you are done, or the next start_recording carries on the same set. Frames live in server memory only: a tab that closes mid-recording ends the loop with a note in the next reply, and any frame whose size differs from the first (a window resize partway through) is dropped at encode time and counted in the result. Where Chrome pushes the finished GIF through the browser download machinery, this server writes it into the session's download folder itself and gives you back the full path.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdYesThe tab to film. Frames are filed under this id, so stop_recording, export and clear only touch the recording that was started for the same tab. tabs_context_mcp lists the ids in this session.
actionYesWhich step to run. start_recording begins capturing this tab, or picks up a set of frames that was stopped earlier; stop_recording closes the loop while keeping the frames; export turns them into a GIF file; clear discards them and frees the memory.
optionsNoOverlay and encoding settings; only export reads them. The five overlay switches - showClickIndicators, showDragPaths, showActionLabels, showProgressBar, showWatermark - are booleans that are all on to begin with, and quality is a number from 1 to 30 that starts at 10. Anything you omit keeps that starting value.
downloadNoKept so a call written for Chrome runs here unchanged, where it routes the GIF through the browser's downloads. This server always writes the file itself and returns its path, so the flag alters nothing; pass it alongside export or leave it off.
filenameNoWhat the exported file should be called; read by export and ignored elsewhere. Left out, the server counts up through recording-1.gif, recording-2.gif and so on, stepping over names already in the folder. Any directory part is stripped, and .gif is appended when the name lacks it.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It thoroughly discloses side effects and quirks: frames live in server memory, tab close stops the loop, dropped frames are counted, the server writes the file itself (ignoring the download flag), filename auto-numbering and .gif appending, and the meaning of the quality ranges. These are far beyond typical tool descriptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is thorough but quite verbose and repetitive. It restates default values multiple times (e.g., quality default, overlay defaults) and uses long, winding sentences. The structure is logical (purpose, actions, details, parameters) but could be tightened significantly without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (4 actions, 5 parameters, nested options, multiple edge cases), the description is remarkably complete. It covers all actions, parameter effects, default behaviors, memory handling, file output, and even contrasts with Chrome's behavior. No significant aspect is left unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema already has 100% description coverage, the prose adds substantial extra meaning: default values for options, the fact that options are only read by export, the filename fallback numbering and directory stripping, and the clarifying note that the download flag is a no-op. This enriches the semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear statement of purpose: 'Films a tab as a run of screenshots and encodes them into an animated GIF.' It also names the four actions and their roles, and the tool is unique among the siblings (no other tool records or exports GIFs).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly contrast this tool with alternatives like 'computer' or 'screenshot', but the purpose is so distinct that the intended use case is obvious. It does give guidance on when to use each action (start, stop, export, clear) and mentions that 'export' is the only action that reads options.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

javascript_toolA

Evaluate JavaScript inside a page and hand back what its final expression produced. All three fields are required: action pinned to javascript_exec, text carrying the source, and tabId naming the page. The snippet runs in the page's own world, so the document, window, and whatever globals the site defined are all within reach, and anything it changes is really changed. It behaves like a console rather than a function body: top-level await is allowed, and the value of the last expression comes back by itself, so finish with the expression instead of a return statement. Results are stringified (objects as indented JSON) and cut off at 10000 characters, with a line saying how long the full text was. A throw from the page fails the call with its message and the frame that raised it. Two things script cannot do here: fill a file input, which needs file_upload or upload_image, and get past its own alert or confirm - that parks the page until firefox_dialog answers it, which in Firefox is recoverable rather than the end of the session.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesThe source to evaluate. Console semantics: top-level await works, and the last expression is the result, so write `document.title` or `await fetch('/api/status').then(r => r.json())` and skip the return; statements ahead of it run normally. It executes in the page context, with the DOM, window and page globals available and its writes sticking to the live document. Anything past 10000 characters of output is trimmed.
tabIdYesThe page to run in; required, never guessed. Use an id tabs_context_mcp listed for this session - a tab outside the group, or one already closed, is refused.
actionYesFixed string: send javascript_exec, spelled exactly that way. Any other value is turned away before the page is touched.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so the description carries full burden. It thoroughly discloses execution context (page's own world), side effects (writes stick to live document), output trimming at 10000 characters, error behavior (throw fails the call), and limitations (file input, alert/confirm). This is exceptionally transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence contributes unique information: purpose, parameter details, execution semantics, output handling, error behavior, and limitations. It is front-loaded with the primary purpose and avoids redundancy, despite being detailed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Fully complete for the tool's complexity: explains output format and truncation, error propagation, execution context, side effects, parameter requirements, and alternatives. No output schema exists, but the description covers return behavior adequately, leaving no critical gaps for an agent to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema descriptions already thoroughly cover each parameter (console semantics, tabId source, fixed action value), and the tool description adds extra meaning beyond schema—notably the explicit limitations and alternatives. Given 100% schema coverage, the description still enriches understanding, warranting above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States precisely: evaluates JavaScript in a page and returns the final expression's value. The verb 'evaluate' and resource 'JavaScript inside a page' are specific, and it clearly distinguishes from sibling tools like navigate, read_page, and form_input.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly explains when to use it (console semantics, top-level await, page context) and when not to (cannot fill file inputs or handle alerts, pointing to file_upload/upload_image and firefox_dialog as alternatives). Also directs the agent to obtain tabId from tabs_context_mcp, covering prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_console_messagesA

Hand back the console output firefox-use has buffered for one tab: log, info, warn and error entries, plus uncaught page exceptions with their stack frames. Every line arrives as an ISO timestamp, the level in brackets and the text, oldest first, and an error also shows the frame it was thrown from. Capture is lazy - the buffer only starts filling the first time you call this tool or read_network_requests in a session, so anything printed before that was never seen; reload and read again. A tab keeps its 1000 most recent entries, and the whole log is dropped when that tab moves to a different origin (a new scheme, host or port each count), so read before you send the page somewhere else; the reply says so when either of those cost you entries. tabId is the only required argument and has to name a tab this session owns - tabs_context_mcp lists them. Send a pattern nearly every time: on a talkative app an unfiltered read buries the two lines you were after.

ParametersJSON Schema
NameRequiredDescriptionDefault
clearNoSet true to empty this tab's console buffer once the reply has been assembled, so your next read shows only what arrived after this one. False by default, which means the same lines will be reported again.
limitNoCeiling on how many matching entries come back, newest kept. 100 when omitted. The header counts any matches this cap held back, so raise it once you see that it actually bit.
tabIdYesNumeric id of the tab to read, as reported by tabs_context_mcp or tabs_create_mcp. It must be one of the tabs in this session's group, and it is required - leave it out and the call errors instead of guessing which tab you meant.
patternNoA JavaScript regular expression, applied case-insensitively. An entry survives when the expression hits either its text or its level, so 'warn|error' selects by severity while 'CheckoutForm' selects by content. A malformed expression comes back as an error rather than being quietly ignored, and so does one that backtracks past the one-second matching budget - drop the nested quantifiers if that happens. Optional, but give one whenever you can narrow the search: it is the difference between the lines you need and a wall of noise.
onlyErrorsNoSet true to keep just the severe entries - console.error, console.assert, anything logged at severe or critical, and uncaught exceptions - and discard info, log and warn. Omitted it is false, and every level comes through.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description thoroughly discloses behavior: lazy capture, buffer size (1000 entries per tab), origin-change reset, output format (ISO timestamp, level, text, stack frames), behavior of the clear flag (empties buffer after reply), limit cap effect, and error handling for patterns (malformed or backtracking). It also states that read-only semantics are not pure because clear can modify state, and annotations are absent, so the description carries the full burden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Although lengthy, every sentence carries essential information—no fluff. The structure front-loads the main function, then details buffering behavior, then parameter specifics. Repetitive phrasing about buffer limits and lazy capture reinforces critical constraints without adding noise. The description is dense but well-organized and necessary given the tool's nuances.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers all necessary context an agent needs: how to trigger capture, what the output looks like, how to limit results, how to filter, error conditions, and state-changing behavior (clear). It explains edge cases like origin changes and lazy capture, and mentions the header reporting held-back entries. Nothing essential is missing for safe and effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description goes beyond the field-level comments by adding context: it restates required tabId, explains the interaction between clear and subsequent reads, clarifies that limit caps apply to matches and headers count held-back entries, details pattern matching rules (case-insensitive, hits text or level), and gives a concrete example ('warn|error'). This elevates understanding of each parameter's practical behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: returning console output (log, info, warn, error, uncaught exceptions) for a specific tab. It explicitly distinguishes from sibling read_network_requests by naming it and describing the different resource, satisfying the verb+resource requirement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides actionable usage guidance: explains the lazy capture (first call triggers buffering, so reload if needed), advises using the pattern parameter to narrow results, warns about malformed regexes causing errors, and gives clear instructions on the clear parameter for subsequent reads. It tells when to use the tool and how to get the best results.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_network_requestsA

List the HTTP traffic recorded for one tab - documents, scripts, stylesheets, images, XHR and fetch alike - as one padded row per request, oldest first, giving method, status, resource type, elapsed time, bytes transferred and URL, flagged when the request failed or is still in flight. Use it to confirm an API call actually fired, see what came back, or find the asset costing you seconds. Traffic to any host is kept, not only the page's own origin, but a tab's log is emptied as soon as that tab moves to a different origin (scheme, host or port), and only its 1000 most recent requests survive. Recording begins with your first call to this tool or read_console_messages, so a request made earlier than that does not exist here; reload and read again. tabId is required and has to be a tab in this session's group - tabs_context_mcp will give you one.

ParametersJSON Schema
NameRequiredDescriptionDefault
clearNoSet true to discard this tab's recorded requests after the reply is built, so a later read starts from empty instead of repeating what you have already looked at. Defaults to false.
limitNoLargest number of matching requests to return, most recent kept; 100 unless you say otherwise. The header reports how many matches the cap hid, which is your cue to ask for more.
tabIdYesThe tab whose requests you want, as the numeric id from tabs_context_mcp. Traffic is buffered per tab, and only tabs belonging to this session can be read.
urlPatternNoA plain substring, matched case-insensitively against the whole URL: '/api/' isolates API traffic, 'fonts.' one host's assets. Literal text, not a regular expression - dots and slashes stand for themselves. Leave it off and every recorded request comes back.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully discloses side effects: the clear parameter discards recorded requests, recording begins on first call, only session tabs are readable, and only the most recent 1000 requests are kept. Also notes the header reporting hidden matches.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence carries essential operational detail. It is well-structured, covering purpose, usage, and limitations without redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description explains the output format (rows with method, status, resource type, etc.) and the header behavior regarding capped matches. This gives the agent a clear picture of what to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Every parameter is described in the schema, and the tool description adds context: tabId is linked to tabs_context_mcp, urlPattern is explicitly a literal substring (not a regex), and limit's behavior is clarified with the default and the header cue.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states it lists HTTP traffic for a tab, enumerating the types of resources and the fields per row. It clearly differentiates from siblings like read_page or read_console_messages by focusing on network requests.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides direct use cases: confirming an API call fired, inspecting responses, and finding slow assets. Also explains key limitations (per-tab session, recording start, buffer cap) which helps the agent decide when to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_pageA

Dump a loaded page as an indented accessibility tree: one line per element carrying its role, accessible name, state, its viewport box when it has one, and a ref_N handle, which is what computer, form_input and the other element tools take as a target. Everything is listed by default, including elements scrolled out of view or hidden by CSS; pass filter "interactive" to keep only the controls. The walk stops 15 DOM levels below the body unless depth raises it, gives up after 5000 elements, and the rendered text is trimmed to 50000 characters (max_chars) on a line break so that no handle is cut in half - whenever any of those three fires, a bracketed note at the end names the limit that hit, how much of the whole you got and how to reach the rest. Handles are minted per document: after a navigation or a reload the old ones stop resolving, so read the page again instead of reusing them. tabId is required (tabs_context_mcp or firefox_status lists the ids this session owns), and a step inside browser_batch has to carry it too. Firefox specifics: the browser runs in its own automation profile, so nothing asks you to approve a site, but an open alert or confirm freezes the page until firefox_dialog answers it.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNoHow many DOM levels below the body the listing may descend, 15 when omitted. Anything further down is skipped and reported in a closing note - lower it to cut a bulky page down to its skeleton, raise it when the control you want sits deeper than the walk reached.
tabIdYesRequired: the session tab to walk. Ids come from tabs_context_mcp or firefox_status; a tab outside this session's group, or one already closed, is refused rather than guessed. State it explicitly on a browser_batch step.
filterNo"interactive" keeps only what a person can operate - links, buttons, form fields, and anything carrying a click handler, tabindex or an interactive ARIA role. The default, "all", lists every element plus the page's text runs, hidden and off-screen ones included, which is what you want when you are working out the layout rather than acting on it.
ref_idNoRead one subtree instead of the whole document: give a handle from an earlier read_page or find and you get that element together with all of its descendants. This is the usual fix when a full read overflowed - focus on the form, table or card you care about. A handle from before the last navigation is gone and comes back as an error.
max_charsNoUpper bound on the rendered output, 50000 characters when omitted. The cut falls on a line boundary and the trailing note reports how much of the total you got. Raise it when your client can swallow a big payload; otherwise prefer narrowing with depth or ref_id.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It thoroughly explains the walk limits (depth, element count, char trim), that handles are per-document and invalidate after navigation, and that alerts/confirms freeze the page. It even reports how limits are signaled via bracketed notes. This is exemplary transparency beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence adds unique value: purpose, filter usage, limits, handle invalidation, tabId requirements, and Firefox specifics are all covered. It is front-loaded with the core purpose and structured logically, with no filler or redundancy. The length is justified by the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description fully explains what the output looks like (indented accessibility tree with roles, names, boxes, ref handles, and limit notes). It covers all parameters, the scenarios for using them, and edge cases (stale handles, alert freezing). An agent can invoke the tool correctly with only this description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

While schema coverage is 100%, the description adds substantial meaning to each parameter: depth is contextualized for cutting bulky pages, filter distinguishes interactive vs layout exploration, ref_id is presented as the fix for overflow, and max_chars is tied to client capacity. The rationale and usage guidance far exceed the schema's terse property descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Dump a loaded page as an indented accessibility tree') with clear resource and scope. It differentiates itself from siblings by explaining it provides ref_N handles for other element tools, and explicitly mentions the 'interactive' filter as a contrast to the full tree. This leaves no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: it is the source of ref_N handles for form_input and other element tools, and recommends the 'interactive' filter when only controls are needed. It also specifies tabId requirements, mentions browser_batch usage, and notes Firefox-specific behavior (alerts freeze the page). This covers both usage conditions and alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resize_windowA

Change the size of the Firefox window a tab lives in, so a layout can be checked at phone, tablet or desktop width and so screenshots taken afterwards come out at that geometry. width, height and tabId are all required; there is no implicit tab here. The real window is resized when Firefox allows that; headless runs and older builds refuse that, and the tab's viewport is stretched instead, which puts the page at the same size. The reply names whichever path ran and quotes the innerWidth x innerHeight the page ended up reporting - expect it to fall short of the numbers you asked for, since toolbars come out of the window box and the browser clamps sizes it thinks are too small. Only the tab you name is affected; the rest of the group keeps rendering as it was.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdYesThe tab whose window gets resized; required, with nothing guessed for you. Take the id from tabs_context_mcp. If Firefox turns the window resize down, only this tab's viewport is changed and the other tabs are left alone.
widthYesHow wide to make the window, in pixels; must be positive. Browser chrome takes part of it, so the page area lands narrower than this - the reply says what the page actually measured.
heightYesHow tall to make the window, in pixels; must be positive. The tab strip and toolbar are subtracted from it, so the visible page is shorter; the reply quotes the innerHeight it settled on.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of disclosing behavior. It transparently explains that headless runs and older builds may refuse window resizing and instead stretch the viewport, that the reply reports innerWidth/innerHeight which may fall short due to toolbars and clamping, and that only the specified tab is affected. This is comprehensive behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose and presents necessary detail, but it contains some redundancy (e.g., repeating that tabId is required and that only the named tab is affected). The longer sentence about expected fall-short and clamps is dense but informative. Overall it is reasonably concise given the complexity, but not maximally tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description fully equips the agent to use the tool correctly: it covers input requirements, behavior under different conditions (headless, older builds), effect on other tabs, and what the reply will contain. Since there is no output schema, explicitly describing the reply format (innerWidth x innerHeight) ensures completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Though the schema already covers 100% of parameters, the description adds substantial semantic detail for each: tabId explains which tab and fallback behavior, width and height both clarify positivity requirements, that browser chrome reduces the actual page area, and that the reply reports measured dimensions. This goes well beyond the baseline for full coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Change the size of the Firefox window a tab lives in') and identifies the specific resource (a Firefox window/tab). It also explains the intended use cases (layout checking at phone/tablet/desktop widths, screenshots). This distinguishes it from sibling tools, none of which perform window resizing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context for when to use the tool (before layout checks or screenshots) and explicitly notes that tabId is required and there is no implicit tab. However, it does not explicitly contrast with alternative tools or state when not to use it, though such guidance is less critical given no sibling tool overlaps.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tabs_close_mcpA

Shut one tab out of the session tab group, addressed by id - the way you tidy up once a page has served its purpose. tabId is required, and has to be an id tabs_context_mcp reported: a tab from one of your own windows, or one that is already gone, is refused rather than quietly ignored. Ids are handed out in sequence and never reused, so a closed one stays invalid for the rest of the session. Shutting the last remaining tab empties the group and Firefox drops it; a later tabs_context_mcp that passes createIfEmpty, or any tabs_create_mcp, builds a fresh one.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabIdYesId of the tab to close, an integer taken from tabs_context_mcp. It has to be one of the tabs this session opened for itself.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Excellent transparency: it discloses that closing the last tab empties the group and Firefox drops it, and that a later createIfEmpty creates a fresh one. It also explains that already-closed or invalid tabs are refused, giving a clear behavioral contract.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is unnecessarily verbose, using metaphorical phrasing ('tidy up once a page has served its purpose') and repeating the 'tabs_context_mcp' source multiple times. While all content is relevant, it could be condensed to a more direct statement without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of an output schema, the description fully covers the tool's behavior, constraints, and edge cases (invalid IDs, last tab scenario). It provides everything an agent needs to call it correctly and anticipate outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers the parameter (tabId, integer, required), and the description adds significant meaning: it must be a valid ID from tabs_context_mcp, that IDs are never reused, and that stale IDs cause refusal. This goes well beyond the raw schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: closing a single tab by ID. The metaphorical language ('shut one tab') is still unambiguous, and the context of sibling tools (navigate, create, etc.) makes the purpose distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides concrete guidance on when to use the tool: when closing a tab, and explicitly instructs that the tabId must come from tabs_context_mcp. It also warns against invalid or stale IDs, which helps avoid misuse. However, it does not explicitly compare to siblings like tabs_create or navigate, though the action is inherently distinct.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tabs_context_mcpA

List the tabs this session owns - numeric id, current address, page title, and which one holds focus. Run it before any other browser tool, because they all address a page by an id from this list, and run it again after tabs open or close. The group only ever holds tabs this server opened in the Firefox profile it launched, so the windows you browse in yourself are neither listed nor driven, and no site has to grant automation permission first. When no group exists yet the reply says so and points at createIfEmpty. Give a new conversation a tab of its own with tabs_create_mcp rather than taking over one that is already busy, unless the user asked you to work in the tab that is open.

ParametersJSON Schema
NameRequiredDescriptionDefault
createIfEmptyNoPass true to have an empty group opened for you when the session has none: a new window with one blank tab, listed in the reply and ready to navigate. Left out it counts as false, and a session with no group is only told it has none. It does nothing once a group exists and never appends a second tab - that is what tabs_create_mcp is for.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses that the tool only lists tabs opened by this server in its Firefox profile (not user's own windows), that no site permission is needed, and the behavior when no group exists (reply says so, points to createIfEmpty). It also clarifies that createIfEmpty never appends a second tab. These are significant behavioral details beyond what a schema could convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded with the core purpose. Every sentence adds value: scope, usage timing, exclusions, fallback behavior, and guidance for new conversations. No fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema, but the description explicitly lists the returned fields (id, address, title, focus) and explains the createIfEmpty behavior. It also covers edge cases like empty groups and permission requirements. An agent has everything needed to call and interpret the result correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100% for the single parameter createIfEmpty, and the schema already explains its effect. The description reiterates the 'when no group exists' context but does not add new semantics beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'List the tabs this session owns' and enumerates the exact fields returned (numeric id, current address, page title, focus). It clearly distinguishes itself from siblings like tabs_create_mcp and navigate by stating it is a listing tool that precedes other browser actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: run it before any other browser tool and again after tabs open or close. It also explains when to prefer tabs_create_mcp for new conversations and when createIfEmpty is relevant, providing clear alternatives and exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tabs_create_mcpA

Open one more blank tab (about:blank) in the session tab group and report its id together with what the group now holds. It takes no arguments. The tab joins whichever Firefox window the group already occupies, or brings one of its own when the group is empty. Call tabs_context_mcp once first, so you know the group you are adding to, then put the id you were just given on a navigate to load something into it. Whatever you open is yours to clear away: close it with tabs_close_mcp the moment it stops being useful, and leave none of your own tabs behind when the task ends - keep one only when the user wants to look at it.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses that it creates a tab and returns data, and notes the cleanup expectation. Does not mention potential side effects like focus changes or prerequisites, but the core behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is somewhat verbose, mixing workflow instructions and cleanup policy into the tool definition. While structured and logical, it could be more concise by separating generic guidance from the tool's core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Provides enough context for an agent to use the tool correctly, including what it returns (id and group contents) and its relationship to other tools. Lacks exact output format, but that is not critical given the absence of an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters exist, and the description correctly states 'It takes no arguments.' Baseline score of 4 applies for zero parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the tool's action: open a new blank tab in the session tab group and report its id and the group's contents. Distinct from siblings like tabs_close_mcp and tabs_context_mcp.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to call tabs_context_mcp first to know the group, then use navigate on the new tab id, and close it with tabs_close_mcp when no longer needed. Provides clear when-to-use and how-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_imageA

Feeds a picture the session is already holding - a screenshot or zoom taken with the computer tool, which prints an imageId such as img_3 - into a page, so no file has to travel through you. There are two ways to aim it and you must pick exactly one: ref, an element from read_page or find, which is how you reach a file input the page keeps hidden behind styling; or coordinate, an [x, y] point in the viewport, for editors that take a dropped file and expose no input at all. Giving both is an error and so is giving neither; imageId and tabId are always needed. Down the ref path the image is written to a scratch file in this session's download folder and handed to the input, and should that element turn out not to be a file input, the tool scrolls to it and drops the image on it instead. Down the coordinate path whatever sits under the point gets a synthesized dragenter, dragover and drop, plus a paste event as a fallback when none of the three was taken up. Read the result line: it names which of drop, paste or nothing the page actually handled, so a target that quietly ignored the image cannot pass for success.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoAn element id from read_page or find, in the form ref_9. Reach for it whenever the target can be named - above all a file input the page hides - rather than aiming at pixels. Cannot be combined with coordinate, and goes stale as soon as the page navigates or re-renders.
tabIdYesThe tab that is to receive the image; always required. tabs_context_mcp lists the ids of this session's tabs if you do not have one.
imageIdYesThe id reported when the picture was captured by the computer tool (its screenshot or zoom action), for instance img_3; an image the user supplied carries one too. Only the 20 most recent survive in the session, so an id from much earlier may already be gone.
filenameNoThe name the page should see, image.png when you leave it out. A directory part is stripped off, and a bare name with no suffix gains .png, or .jpg when the capture is a JPEG.
coordinateNoA point [x, y] in viewport pixels, for a visible region that accepts dropped files without exposing an input. Take the numbers off the newest screenshot; a scroll or a layout shift since then has moved the target out from under them. Cannot be combined with ref.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full behavioral transparency. It discloses side effects in detail: writing a scratch file, scrolling and falling back to drop for non-input elements, synthesizing drag/drop/paste events, and reporting which mechanism the page actually handled via the result line.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence supplies necessary operational detail with no filler. The single paragraph flows logically from purpose to constraints, per-path behavior, and result interpretation, making it dense yet well organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given five parameters, conditional exclusivity, fallback behaviors, and a result contract, the description is complete. It tells the agent how to obtain ids, what to expect on errors, and how to validate success via the result line, leaving no operational gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema descriptions already cover all five parameters, and the prose adds critical semantics beyond them: imageId freshness ('Only the 20 most recent survive'), filename defaulting and sanitization ('directory part stripped', '.png' added), ref staleness after navigation, and coordinate invalidation after scroll or layout shift.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool's purpose precisely: 'Feeds a picture the session is already holding ... into a page'. It distinguishes the action from sibling tools by referencing the computer tool for capture and read_page/find for element references, making the verb and resource unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly contrasts the two targeting paths (ref vs coordinate), states when each is appropriate ('above all a file input the page hides' vs 'a visible region that accepts dropped files without exposing an input'), and notes constraints like 'Giving both is an error and so is giving neither'. It also points to sibling tools (tabs_context_mcp, read_page, find) for obtaining required ids.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 20 tool updatesv0.2.0
    • First observedbrowser_batch
    • First observedcomputer
    • First observedfile_upload
    • First observedfind
    • First observedfirefox_dialog
    • First observedfirefox_restart
    • First observedfirefox_status
    • First observedform_input
    • First observedget_page_text
    • First observedgif_creator
    • First observedjavascript_tool
    • First observednavigate
    • First observedread_console_messages
    • First observedread_network_requests
    • First observedread_page
    • First observedresize_window
    • First observedtabs_close_mcp
    • First observedtabs_context_mcp
    • First observedtabs_create_mcp
    • First observedupload_image

TDQS

A4/5.0

Scored across 20 tools

Disambiguation4/5

Most tools are clearly separated by role: tab management, page reading, form input, file handling, network/console logging, and Firefox maintenance. The main overlap is that tabs_context_mcp and firefox_status both report the session's tab list, and read_page, find, and get_page_text have somewhat similar page-reading purposes even though the descriptions distinguish them well.

Naming Consistency2/5

Although most names use snake_case, there is no consistent pattern across the set. Tab tools use an odd _mcp suffix, several tools are bare or generic nouns like find, computer, and gif_creator, and there is no uniform verb_noun or noun_verb convention across the rest. A new agent cannot reliably guess related tools from their names.

Tool Count3/5

20 tools is on the heavy side and somewhat at the edge where complexity starts to overwhelm. The count reflects a broad browser-automation scope, but some responsibilities are spread across multiple tools rather than being consolidated, so the set feels larger than necessary even though it is not unreasonable.

Completeness4/5

The tool surface covers a nearly complete browser workflow: tabs, navigation, reading and finding elements, form input, file and image uploads, screenshots, keyboard/pointer control, JS evaluation, console and network logging, GIF capture, and Firefox restart/dialog handling. Missing pieces like cookie/storage management or explicit screenshot download are minor and usually workable with changes or extensions.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers