Skip to main content
Glama

Give your agent a browser it can work with. Tablaze connects Codex and other MCP clients to Chromium through seventeen focused tools. Inspect a page, fill a form, extract a result, and check that the task actually succeeded.

Built with Playwright. Browser sessions stay running between calls; your MCP client supplies the reasoning. The MCP browser tools require no additional model API key. The optional standalone Agent loop uses an explicitly configured planner.

See it work

Watch the complete demo — about 2 minutes →

A real MCP SDK client completes one travel workflow: six form actions in one batch, five result checks, then a changed approval button. The stale reference is rejected with no approval or order; a fresh snapshot lets the workflow continue through an explicitly followed popup, four approval actions and three receipt checks. Finally, the recorder verifies the downloaded CSV's bytes and SHA-256, closes both owned tabs and confirms zero remaining sessions.

Real MCP demo: one approved trip, a verified CSV and zero remaining sessions

Batch the form → Verify → Stop and re-observe → Approve in the popup → Check the file

The 154-second recording presents actual tool responses and screenshots with conversational English male narration and optional captions. It contains no simulated cursor. A deterministic SDK script drives the local fixture; no model inference is involved. It is a presentation recording, not a continuous video feed from the controlled tab.

Reproduce the recording · Inspect the full trace

Related MCP server: hronaut

Measured development results

The third matched Codex run used five visible development tasks, one attempt per engine and task. Both engines passed independent business checks 5/5; Tablaze returned complete success 4/5, Browser Use 5/5. Tablaze's order task was interrupted by the old inference bridge after the order was recorded. Its failed call had no usage data, so total tokens remain unknown.

A separate order follow-up after the bridge fix completed successfully for both engines, each creating one order with no duplicate write:

Follow-up measurement

Tablaze

Browser Use

Agent reported completion

59.247 s

75.354 s

End-to-end time

59.415 s

93.706 s

Model calls, including judging

4

4 (1 judge)

Input / output tokens

59,123 / 664

70,868 / 1,002

Browser Use's default judge is included in its totals. Agent and end-to-end clocks have different starting boundaries. This single follow-up does not replace the original 4/5 result, and these small, visible samples do not establish overall superiority. Both reports retain raw results, source hashes, settings, and limitations.

A new four-task matched development smoke tested iframe input, interrupted-response order recovery, a virtual list and a visual canvas. Both engines completed all four with independent business acceptance; Tablaze's iframe sample was slower (83.799 s versus 52.515 s), while its other three whole-run samples were shorter. Each task ran once per engine, so this does not prove stable efficiency or overall superiority. The initial iframe observation gives newly attached child frames a bounded readiness wait. When the main page has no controls and exactly one visible child frame contains a form field, the first tab_open returns actionable child refs without another snapshot call; a real-Chrome regression fills and submits a delayed child form from those refs. In a six-pair matched Codex follow-up, both sides passed 6/6 with zero duplicate writes and all Tablaze runs skipped the extra frame snapshot. Tablaze's visible median whole-run time was 34.128 versus Browser Use's 57.526 s, while Agent-done time was 33.973 versus 39.778 s; Browser Use's judge is included only in whole-run time. One fixture does not establish stable speed or overall superiority. In the three-pair post-fix iframe follow-up, both engines passed every independent business check, but Tablaze's visible median whole time was 79.699 s versus Browser Use's 56.040 s. Two Tablaze runs spent extra model rounds recovering from an unnecessary text check of an input value; efficiency parity remains unproven. Action feedback now retains the operated child frame after such a failed post-check, allowing the next decision to inspect the right page without another frame-switch snapshot. In a fresh three-pair matched follow-up, both sides passed 3/3 with no duplicate writes. Tablaze's median whole-run time was 46.037 s versus 56.132 s, but its observed Agent-done time was 45.880 s versus 39.114 s. No run triggered the failed-check path, so the comparison does not measure this fix's speed effect.

The current Browser Use capability and public-issue audit tracks what is implemented, what remains missing across Agent/MCP/Harness/Pi/cloud, and which reported defects have a local reproduction or regression. Password fields now reveal only whether a value is present, while keeping its bytes redacted.

A four-task, two-seed matched Codex comparison covers forms, a virtual list, visual canvas and interrupted-response order recovery. Both engines passed 8/8 independent business checks with one correct write and no duplicates per attempt. Tablaze's visible median whole-run time was 56.331 versus Browser Use's 61.211 s, but Agent-done time was 56.177 versus 44.063 s; Browser Use's default judge runs after Agent completion. This development sample does not prove stable efficiency or overall parity.

The new first-open canvas image was tested in a separate four-pair Codex follow-up. Tablaze passed 4/4 business checks with no duplicate writes; Browser Use passed 3/4 after one attempt clicked twice despite reporting success. All four Tablaze traces skipped a separate tab_capture. On the three jointly successful seeds, median whole-run time was 41.552 versus 61.113 s and Agent-done time was 41.401 versus 42.578 s. This one visible task does not establish overall superiority.

Get started

Requires Node.js 20+, npm and Git. This is a developer preview; npm publication is pending, so install from source:

git clone https://github.com/SweetDianDian/tablaze.git
cd tablaze
npm ci
npm run build
node dist/cli.js setup

For installed Chrome, skip setup, run node dist/cli.js doctor --channel chrome, and append --channel chrome to the Codex command below. This launches a separate browser session.

On Linux, use npx playwright install --with-deps chromium in place of setup to install the browser and system dependencies.

Browser modes and diagnostics

Connect Codex

From the cloned directory, register the built server:

codex mcp add tablaze -- "$(node -p 'process.execPath')" "$PWD/dist/cli.js"
codex mcp get tablaze

Then ask Codex:

Use Tablaze to open https://example.com, read the heading and links, verify that the title contains “Example Domain”, then close the session. Report the result of the checks.

Complete Codex setup covers desktop configuration, browser selection and troubleshooting. Other MCP clients can launch the same node /absolute/path/to/tablaze/dist/cli.js command over stdio.

For workflows that need page-origin JavaScript, --page-script adds an opt-in tab_script tool. It has the page's full authority, including account data and network requests; it is unavailable with configured secrets, external CDP or a navigation policy. Scripts require a current main-frame snapshot and return a fresh one. See the page-script contract and recovery limits. The default catalog remains seventeen tools.

A separate native-Codex page-script smoke passed the independent authenticated-response judge once for both Tablaze and Browser Use CLI-MCP, with one correct write and no duplicates per arm. Codex-process samples were 198.359 s and 239.400 s respectively; browser setup differed, so this is not a controlled speed ranking or overall win.

Run a standalone task

The optional run command supports codex, anthropic, ollama, and openai-compatible planners. Always specify a model. Codex reuses the installed CLI and its existing login:

node dist/cli.js run --provider codex --model "<your-codex-model>" \
  --task "<authorized task>" --channel chrome

For a task with exactly one intended starting URL, add --direct-open-task-url to open it before the first model call. This is opt-in; --start-url <url> remains available when the caller already knows the exact page. See initialization and resume behavior.

For trusted setup before the first model call, --initial-actions ./initial-actions.json accepts an HTTP(S) starting url, then more URLs or {"click":{"name":"Open details","role":"button"}}. A click requires one uniquely named current-page control; ambiguity or truncated observation stops input. The SDK exposes the same sequence as initialActions. Each attempted action is checkpointed and never replayed automatically after an uncertain outcome. See the format and recovery rules.

In a three-pair same-Codex initial-click task, both Tablaze and pinned Browser Use passed independent business checks 3/3. Tablaze clicked before planning; Browser Use's queued click after navigation was skipped and its model completed the click later. Tablaze's median whole-run time was 68.562 versus 62.911 seconds and used one more model call per run. This visible task does not show a speed lead or overall feature parity.

In the same-source Codex task-URL follow-up, all 18 form/iframe attempts across Tablaze on/off and Browser Use passed the independent business judge once. On the form task, Tablaze's median Agent-done time was 30.378 s with direct-open versus 55.583 s without it; Browser Use took 37.647 s before its default judge. On the iframe task, Tablaze still took 56.317 s versus Browser Use's 40.495 s because all three Tablaze runs added an invalid input-as-page-text check. These visible tasks do not prove a general speed lead.

After making text and value check scopes explicit in the model-facing tool schemas, a six-pair iframe follow-up passed 6/6 per side with one correct write each. None of the six Tablaze runs repeated that invalid check; four finished in two planning calls and two used a separate verification call. Median Agent-done time was 42.162 s Tablaze versus 45.713 s Browser Use, with a 22-second Tablaze disadvantage on one seed. This remains a visible synthetic sample, not stable speed or overall superiority.

Use --output-schema ./result.schema.json to require a schema-valid JSON deliverable before the Agent can finish; the application should still check business correctness. The same schema is required on resume. See the Agent contract.

SDK callers can use createAgentControl() to pause at a safe boundary, add trusted steering, and resume with a fresh plan; queued writes are skipped and prior verification must be repeated. See live operator intervention.

The default compatible provider still requires --endpoint. Native Anthropic and Ollama use their own protocols and default endpoints; authentication, output settings, and model capabilities differ. An explicit --fallback-model can take over after transient planning failures, then remain active for that run without replaying browser actions. See provider setup and boundaries, CLI examples, and Agent verification and recovery. Anthropic/Ollama have local protocol coverage, without live inference results; the earlier comparison results above retain their original runtime hashes.

The new production Codex path also has a separate live smoke: two visible tasks passed independent business checks and completed successfully (2/2), with zero duplicate writes. This is not a new matched Browser Use comparison.

A new matched virtual-list task completed once on each engine with independent acceptance and zero duplicate writes. Under a shared 120,000-token budget, whole-run time was 86.830 s for Tablaze and 141.597 s for Browser Use, including its default judge. This one visible development task does not establish a general performance or success-rate advantage.

The delayed-control development comparison records a real Codex run that opened a menu and clicked its newly appearing option in one guarded Tablaze action batch. In the final visible matched attempt, both agents completed with one correct write and zero duplicates; Tablaze took 80.715 s and Browser Use 77.720 s. The report retains an earlier wrong-role failure and recovery. These changing-source samples do not establish a stable speed or overall advantage.

The separate cross-origin authorization-return task also passed once on each engine with one provider authorization and one app submission. Browser Use's 119.515 s whole run was slightly shorter than Tablaze's 123.719 s. This synthetic popup flow does not establish production login parity or an overall ranking.

For debugging an actual run, --record-video optionally records each owned isolated tab as a private WebM. tab_close returns finalized paths and SHA-256 digests; the standalone run report includes them after cleanup. It is off by default and adds no simulated cursor or audio. See real browser recording and privacy limits.

--record-har and --record-trace optionally save an owned session's network HAR and Playwright trace ZIP. HAR response bodies are omitted by default; trusted operators can explicitly select --har-content embed|attach and --har-mode full|minimal. Finalized private paths, sizes and SHA-256 digests appear after cleanup. Every mode can contain private request data, and body modes can include sensitive responses. See diagnostic artifacts.

Trusted operators can set viewport, screen, pixel ratio, user agent, locale, time zone, mobile behavior, touch and page permissions on owned browser contexts. --no-viewport lets desktop content follow the browser window; headed sessions also accept --window-size and --window-position. --device-preset pixel-7 and pixel-7-pro apply a coherent set of mobile emulation properties from pinned Playwright descriptors. doctor reports effective settings without echoing the user-agent string; external CDP browsers cannot be reconfigured. See browser configuration and limits. Owned contexts also accept --proxy-server with optional bypass rules and credentials read from an environment variable; the configuration guide records the tested HTTP routing path and remaining proxy validation.

A small loop, with useful controls

Capability

What it gives your agent

Warm sessions

Keep page state between calls. Temporary, isolated browser contexts by default; explicit CDP attachment for an existing profile.

Owned persistent profile

Opt into a dedicated private Chrome directory with a required profile ID on reuse and exclusive ownership. Browser-managed login state can survive a cold restart.

Compact observations

Full or incremental snapshots with element refs, text budgets and visible truncation.

Guarded actions

Check snapshot revisions and DOM targets before input. Re-observe when a target changes.

Ordered batches

Send up to 20 actions in one call, with completed, failed and skipped steps. Stops on error; earlier effects remain.

Delayed controls

After clicking an observed opener, wait for one exact-name visible button or menu item and click it in the same guarded batch. Ambiguous matches stop before input.

Explicit verification

Check URL, title, text, field values, visibility and counts. Known same-document outcomes can be checked inside tab_act after its batch, with passing evidence available to the Agent in that call.

Optional response journal

Use --capture-network to inspect bounded responses from owned pages and popups through an eighteenth MCP tool, tab_network. Off by default; not a programmable JS/CDP surface.

Optional real video

Use --record-video for private WebM recordings of owned isolated tabs, finalized with SHA-256 at session close. Recording scope and limits.

Optional navigation policy

Exact-origin allow/deny rules for HTTP(S) document requests, including redirects, frames and popups, in owned isolated browsers. No external CDP; not a network firewall.

Seventeen tools

Tool

Use it to

tab_open

Open a page and receive its first snapshot.

tab_snapshot

Observe the page, changes or a selected frame.

tab_find

Search visible text while scrolling a page or observed virtual-list container; return a fresh actionable snapshot.

tab_act

Guarded form input, explicit readonly-input/native-listbox selection, drag/drop, container scrolling, file selection and coordinate actions; optional post-action checks.

tab_click_named

Click exactly one actionable main-document control by accessible name after a fresh bounded observation; refuse ambiguous matches.

tab_verify

Test explicit page assertions and guarded form values by ref or CSS selector.

tab_extract

Read text, links or tables.

tab_capture

Capture a JPEG screenshot.

tab_list

Inspect owned sessions.

tab_close

Close a session and release its resources; return finalized WebM artifact metadata when recording was enabled.

tab_navigate

Navigate, go back/forward, or reload without losing session state.

tab_tabs

Open, switch, and close owned tabs, including popups.

tab_downloads

Inspect download status and local artifacts.

tab_dialog

Arm a one-shot accept/dismiss response to a native dialog.

tab_state

Save authentication state for explicit import into a new session.

tab_extract_structured

Extract typed fields against a JSON schema with DOM source citations.

tab_pdf

Export a private PDF artifact with source URL and SHA-256 digest.

Explore the project

Guide

Inside

Codex integration

Copyable configuration, tool arguments and troubleshooting.

Runtime reference

Reference lifetime, batch semantics, browser modes and current limits.

Typed custom tools

SDK schemas, trusted application context, browser bindings and recovery contracts.

Model providers

Codex CLI, native Anthropic/Ollama and compatible HTTP contracts, authentication and usage.

Navigation policy

Trusted CLI/SDK configuration, document-only scope, transport-loss limits and checkpoint identity.

Scoped credentials

Trusted aliases, exact-origin login fills, redaction limits and recovery.

Real browser recording

Opt-in WebM artifacts, CLI/MCP/SDK access and privacy limits.

Browser Use variants audit

Distinguish Agent, MCP, Harness, Pi and cloud comparison targets; source audit, not a benchmark.

Benchmark

Reproduce the local workload and inspect every raw sample.

Validation evidence

Real browser, SDK and Codex results, with their measured scope.

Security

Data handling, resource ownership and private reporting.

Release packaging

Build the npm tarball and source archive.

The historical 0.1.0 Ubuntu CI run on Node 20 and 22 passed its build, browser/MCP tests and package inspection. The later Ubuntu Node 20/22 run for commit c9d091a, which includes the provider adapters, also passed both jobs. That run predates the navigation-policy increment. Current development validation is recorded separately. To check a source checkout yourself, run npm test; use npm run bench for the separate local benchmark.

Contribute

Bring a reproducible browser case, improve a guide, or send a focused fix. Open an issue · Submit a pull request · Read the contribution guide

MIT licensed · Dependency credits and project provenance

The current development branch adds scoped/viewport snapshots, consistent composed-DOM reading, rich controls, file workflows, tabs, structured extraction, and a model-driven agent loop with checkpoint/resume. extractWithPlanner and tablaze-extract use a separately selected model for quote-checked, schema-valid extraction from supplied source text. tablaze run --extraction-model also exposes browser-bound extraction to the Agent. These additions have not yet been published to npm. See the one real Codex extraction check and current validation. The Browser Use comparison records remaining gaps and unmeasured acceptance criteria; it does not claim superiority.

CLI Agent and stdio MCP uploads use a trusted file list: no host files are available by default, --available-file grants specific files, and same-session downloads can be uploaded by ID. For authenticated tasks, --storage-state-file /absolute/private/state.json applies a trusted Playwright state file to each new isolated session without disclosing its path to the Agent.

Available Tools

16 tools
tab_actAct on observed elementsA
Destructive

Execute 1–20 ordered actions against a current snapshot. Optional post_checks run only after the whole batch completes, using this same snapshot's refs; they poll for asynchronous outcomes and return verification evidence in this call. A failed postcondition keeps completed effects, sets replan_required, and skips later queued tools. Ref checks fail if the document or node changes; use separate tab_verify after navigation. Use fill_secret with an observed ref and a secret alias from snapshot.available_secrets for configured credentials. Includes hover, double_click, explicit local-file upload, and click_xy in main-viewport CSS pixels after visual inspection (coordinates lack DOM identity guards). Default 30s action budget, maximum 60s; verification has a separate timeout. Stops on first action failure. A followed popup returns its snapshot and replan_required without running post_checks. Cancellation closes the session; completed effects remain.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionsYes
session_idYesSession ID returned by tab_open.
timeout_msNo
post_checksNo
snapshot_idYes
include_snapshotNo
verify_timeout_msNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations: failed postconditions keep effects and set replan_required, the batch stops on first failure, popup redirects skip post_checks, cancellation closes the session but retains effects, and click_xy coordinates lack DOM identity guards. This is exactly the behavioral context an agent needs beyond readOnlyHint/destructiveHint/openWorldHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and information-rich with no filler, and it fronts the core action-execution purpose. However, it reads as one long run-on paragraph rather than structured guidance, which slightly reduces scannability for an agent needing to quickly extract key constraints.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity—14 action variants, post_checks, snapshot refs, timeouts, popups, and cancellation—the description covers nearly all behavior an agent needs to call it correctly. It even explains effect persistence, failure semantics, and coordinate caveats. The only minor omission is a detail about include_snapshot, but that is not essential for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is very low at 14%, but the description compensates with meaningful semantics: post_checks reference the same snapshot's refs and poll for async outcomes, click_xy is in main-viewport CSS pixels, fill_secret uses an alias from snapshot.available_secrets, and action budgets/timeouts are specified. It does not explain every parameter (e.g., include_snapshot) or every action type, but the schema covers the structural constraints and the description adds the critical runtime semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Execute') and clear resource ('ordered actions against a current snapshot'), and immediately distinguishes tab_act from tab_verify by routing navigation-dependent checks elsewhere. It leaves no doubt about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says 'use separate tab_verify after navigation' and instructs using fill_secret with an observed ref and a snapshot secret alias, giving concrete when-to-use guidance. It also clarifies that post_checks are the in-call verification mechanism, with separate timeout semantics, which helps the agent choose between this and sibling verification tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tab_captureCapture pageA
Read-only

Capture the current viewport, or the full page when explicitly requested. Images can contain visible page data. Returns an image and URL metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
full_pageNo
session_idYesSession ID returned by tab_open.

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and destructiveHint. The description adds meaningful behavioral context by warning that images can contain visible page data and clarifying that full-page capture requires explicit request. This goes beyond the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler: core behavior is front-loaded, the privacy warning is valuable, and the return type is stated. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter capture tool with no output schema, the description covers the key operational details: default scope, full-page opt-in, output type, and data-visibility risk. It lacks an explicit comparison to tab_snapshot and precise output encoding, but it is otherwise sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema documents session_id but leaves full_page with only a type. The description compensates partially by indicating that full-page capture happens only when explicitly requested, implying the default is viewport-only. However, it does not name the parameter or explain its format beyond that implication.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly names the action ('Capture') and the resource ('current viewport' or 'full page'), and states the output ('image and URL metadata'). It is specific enough to distinguish it from status/navigation tools, though it does not explicitly differentiate it from the similarly named sibling tab_snapshot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to choose this tool over tab_snapshot, tab_extract, or other siblings, and no exclusions or alternative routing. The 'when explicitly requested' note is about the full_page parameter, not about tool selection, so it does not satisfy the usage-guideline dimension.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tab_closeClose sessionA
Destructive

Close one session owned by this server and release its resources. Unsaved changes in that session are lost.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesSession ID returned by tab_open.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already include destructiveHint=true, but the description adds valuable behavioral context by warning that 'Unsaved changes in that session are lost.' It also clarifies the side effect of releasing resources, which goes beyond what the annotations alone convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences, front-loading the action and then adding the critical data-loss warning. Every word contributes value, with no repetition of the title or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter destructive tool, the description covers the target resource, the resource-release effect, and the key risk of unsaved data loss. It does not describe error behavior or return value, but the simplicity of the operation and the presence of annotations make this a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema fully documents session_id as a required string 'returned by tab_open,' so schema_description_coverage is 100%. The tool description adds little parameter-specific meaning beyond reinforcing that the session is owned by this server. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Close one session owned by this server and release its resources.' This clearly distinguishes tab_close from sibling tools like tab_open, tab_state, and tab_act. The object of the operation is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The context is implied rather than explicit: closing a server-owned session and releasing resources suggests this tool is for cleanup/termination. However, it does not explicitly state when to use it versus alternatives or mention any exclusion cases, leaving some inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tab_dialogHandle next native dialogA
Destructive

Arm a one-shot accept or dismiss response before the action that opens an alert, confirm or prompt. Optional prompt_text is entered only for that dialog. Unarmed dialogs are dismissed and reported as unsupported flow; check actual outcomes before retrying.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes
session_idYesSession ID returned by tab_open.
prompt_textNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark destructive and open-world behavior; the description adds valuable context beyond that: one-shot nature, unarmed dialogs dismissed and reported as unsupported flow, and advice to check actual outcomes before retrying.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the primary purpose, then behavioral notes. Every sentence earns its place without filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the arm-before-dialog workflow, one-shot semantics, and unarmed fallback. However, with no output schema, it does not describe return values or what happens if no dialog appears, so not a perfect 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema documents only session_id; the description compensates for low coverage by explaining prompt_text is entered only for prompt dialogs and action is a single accept/dismiss choice. It does not add further session_id semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the specific verb 'Arm' and resource 'accept or dismiss response' for native dialogs, clarifies scope with 'alert, confirm or prompt'. It clearly distinguishes from sibling tools like tab_act by focusing on one-shot dialog handling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says to use it 'before the action that opens an alert, confirm or prompt', providing clear timing and context. It does not name alternatives explicitly, but the usage context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tab_downloadsInspect downloadsA
Read-only

List downloads created by this session. Supply download_id to wait up to timeout_ms for completion. Completed results include an owned local artifact path. Pending is not completed; failed downloads set isError. Files remain after closing the session.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesSession ID returned by tab_open.
timeout_msNo
download_idNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool read-only and non-destructive. The description adds valuable behavioral detail beyond annotations: completed results include an owned local path, pending is not completed, failed downloads set isError, and files persist after session close. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four compact sentences deliver all essential behavior with no repetition or filler. The main listing action is front-loaded, and each following sentence adds distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the burden of explaining what results look like. It covers completion, pending, failure, artifact paths, and file persistence. It could be slightly more explicit about the exact status representation, but an agent can reasonably invoke the tool correctly with this information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%, so the description must compensate for download_id and timeout_ms. It does: 'Supply download_id to wait up to timeout_ms for completion' explains both parameters' roles. It does not mention defaults or whether timeout_ms is ignored without download_id, but the core semantics are covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific action ('List downloads created by this session') and supports optional waiting on a particular download. It clearly identifies the resource and scope, and no sibling tool overlaps with this purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear invocation context: omit download_id to list, supply download_id to wait up to timeout_ms. It does not explicitly name alternatives or exclusions, but no sibling appears to compete, so the guidance is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tab_extractExtract page dataA
Read-only

Read bounded text, links or table data, optionally within a CSS selector. Returns truncation information when the result is limited.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindYes
selectorNo
max_itemsNo
session_idYesSession ID returned by tab_open.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this as read-only and non-destructive, and the description's 'Read' aligns with that. It adds useful behavioral context beyond annotations by disclosing that results may be bounded and that truncation information is returned when the result is limited.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler: it states the core action, the data types, the optional selector, and a key return behavior. Every phrase earns its place, though 'bounded' is slightly ambiguous.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with no output schema, the description covers the main capability and truncation behavior but omits important operational details such as how max_items applies to different kinds, what table extraction returns, and whether a default selector covers the whole page. It is adequate but has clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low (25%), but the description adds meaning by naming the kind enum values ('text, links or table data') and mentioning the optional CSS selector. It does not clearly explain the semantics of max_items beyond the vague word 'bounded', nor does it describe the output shape for each kind.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Read') and names concrete resources: bounded text, links, or table data, optionally scoped by a CSS selector. It clearly conveys what the tool does, though it does not explicitly differentiate itself from the similar sibling tab_extract_structured.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: call this when you need to read text, links, or tables from a page, possibly within a selector. However, it gives no explicit when-to-use vs alternatives guidance, nor does it mention cases where a sibling like tab_extract_structured or tab_find would be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tab_extract_structuredExtract schema-validated fieldsA
Read-only

Read named fields from observed DOM using selectors, validate strict JSON Schema draft-07, and return per-field source URL/selector/quote provenance. Supports text, attributes, current non-sensitive values, typed scalars and arrays. At most 30 fields, 20 matches per field, 100 total. Does not call a model or infer missing facts; hidden/password controls and truncated evidence are rejected.

ParametersJSON Schema
NameRequiredDescriptionDefault
fieldsYes
schemaYes
session_idYesSession ID returned by tab_open.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the read-only and non-destructive annotations, the description adds meaningful behavioral detail: hard limits (30 fields, 20 matches per field, 100 total), explicit rejection of hidden/password/truncated evidence, and the guarantee that no model or inference is used. This gives an agent a strong, accurate picture of what will and won't happen.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and efficient: the core action and return value are front-loaded, followed by supported modes, limits, and explicit non-behaviors. Every sentence conveys essential operational information without fluff or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is sufficiently complete for a 3-parameter tool with no output schema: it explains what is returned (per-field source URL, selector, quote provenance), what constraints apply, and what is rejected. An agent can reasonably decide whether to call it and understand the expected result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, so the description must compensate. It does: it clarifies that 'fields' can use text, attribute, and value modes, supports typed scalars/arrays, and imposes field-count limits. It doesn't explain the 'required' or 'multiple' flags explicitly, but the schema already enumerates those and the description's added semantics are useful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Read'), a precise resource ('named fields from observed DOM using selectors'), and the differentiating value-add: JSON Schema draft-07 validation and per-field provenance. This clearly separates it from siblings like tab_extract and tab_find by emphasizing schema validation, typed extraction, and evidence provenance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys when to use this tool: when you need schema-validated, selector-based field extraction with provenance, and not when you need model-based inference ('Does not call a model or infer missing facts'). It does not explicitly name alternative sibling tools or contrast them, but the context is clear and the exclusions (hidden/password controls, truncated evidence) are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tab_findFind text through a long or virtual listA
Destructive

Search the current frame for visible text, scrolling a page or one observed vertical scroll container in bounded steps. Use container_ref with its current snapshot_id for a virtual list. Returns a fresh viewport snapshot with actionable refs when found; found:false means no match within the searched range, not proof the whole application lacks it. This changes scroll position but does not click or submit.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
frame_idNo
session_idYesSession ID returned by tab_open.
timeout_msNo
max_scrollsNo
snapshot_idNo
container_refNoElement reference from the current snapshot.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses side effects (changes scroll position) and non-actions (does not click or submit), which aligns with destructiveHint=true and readOnlyHint=false. It also explains the bounded step search and the fresh snapshot return, adding valuable context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three focused sentences, front-loaded with the core action, followed by specific usage details, then a side-effect note. No redundancy or unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers core behavior, return format (fresh snapshot with refs), side effects, and a key interpretation caveat about found:false. Missing explicit detail on timeout_ms/max_scrolls defaults and the exact output structure, but given the annotations and the description, it is reasonably complete for a search tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 29% schema coverage, the description compensates partially by explaining the container_ref/snapshot_id relationship for virtual lists and implying max_scrolls via 'bounded steps'. However, other parameters like timeout_ms, frame_id, and text receive no additional meaning, leaving gaps for the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (search for visible text) and the resource (current frame), with specific scrolling behavior. It distinguishes itself from sibling tools like tab_extract or tab_verify by focusing on text discovery and returning a snapshot with actionable refs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance for virtual lists (use container_ref with current snapshot_id) and clarifies the semantics of found:false, which helps the agent interpret results. However, it does not explicitly name alternative tools or state when not to use this tool, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tab_listList sessionsA
Read-only

List the sessions owned by this server. Does not launch a browser or discover unrelated user tabs.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With readOnlyHint and destructiveHint already covering safety, the description adds non-obvious behavioral facts: this operation does not launch a browser and does not discover unrelated user tabs. That goes beyond the annotations and clarifies the tool's limited scope.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The positive statement comes first and the behavioral exclusions are packed into a second sentence that earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless, read-only listing tool with annotations covering safety, the description is sufficient: it states the returned resource and the key non-behaviors. No output schema exists, but the absence of parameters makes the call surface fully clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and schema coverage is 100%, so there is no parameter documentation burden for the description. The baseline of 4 applies; no additional parameter semantics are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete action and resource ('List the sessions owned by this server') and adds negative scope ('Does not launch a browser...'), making the intent unambiguous. It is not a restatement of the title or name, and it differentiates from tab-discovery tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear context: use when a server-owned session listing is needed, and it explicitly excludes browser launching and unrelated user tab discovery. It stops short of naming an alternative sibling tool (e.g., tab_tabs), so the when-not is implicit rather than pointing to a specific replacement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tab_navigateNavigate current tabA
Destructive

Navigate within an existing session while retaining cookies and tabs. Supports goto, back, forward and reload. Returns a fresh snapshot; previous references are invalidated.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNo
actionYes
session_idYesSession ID returned by tab_open.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as unsafe and destructive, and the description adds meaningful behavioral detail: cookies and tabs are retained, navigation returns a fresh snapshot, and previous references become invalidated. This goes beyond the annotation flags and clarifies the real-world side effect of using the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences carry the full purpose, supported actions, and key behavioral consequence with no filler. The most important scoping fact ('existing session') is front-loaded, and the invalidation warning is compactly placed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a navigation tool with no output schema, the description covers the essential semantics: what the tool does, which actions are supported, what is preserved, and what happens to prior snapshots. It is slightly incomplete on url semantics and route selection versus sibling tools, but it gives enough for a confident call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 33% schema description coverage, the description should compensate for undocumented parameters, especially url. It mostly restates the action enum already present in the schema and does not explain how url applies (e.g., that it is required for goto and irrelevant for back/forward/reload). The description adds little parameter-level meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Navigate') with a clear resource: an existing session's current tab, and enumerates the supported operations (goto, back, forward, reload). It is distinct from sibling tools like tab_open, which creates a session, by anchoring usage to an already-existing session.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'within an existing session' provides clear context for when this tool applies, implying it should be used after a session has been created rather than to create one. It does not explicitly name alternatives or state when not to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tab_openOpen browser sessionA
Destructive

Open an HTTP(S) page in a new session and return a compact full snapshot. Cancellation cleans up this opening attempt, including late-created pages, without closing sibling sessions or the shared browser. The browser stays warm for subsequent calls.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
storage_stateNoExplicit local storage-state file from tab_state. Restores cookies, localStorage and IndexedDB into an isolated session.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, openWorldHint=true, destructiveHint=true), the description discloses cancellation semantics, cleanup of late-created pages, session isolation, and that the browser stays warm. This is rich, non-obvious behavioral detail that helps an agent predict side effects and manage resources. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the first states the core action and output, the second explains cancellation behavior, and the third sets performance expectations. Information is front-loaded and there is no filler or repetition of schema data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description covers the main behaviors: new session, snapshot return, cancellation cleanup, session isolation, and warm browser. The 'compact full snapshot' phrase is somewhat generic and does not specify the snapshot's structure, but combined with annotations the essential context is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: storage_state is described in the schema, while url is not. The description adds some meaning to url by specifying 'HTTP(S) page' and implies session context, but it never mentions storage_state or its role. The schema description covers storage_state, so overall the agent can understand both parameters, but the tool description itself adds only partial compensation for the low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Open an HTTP(S) page'), a specific resource ('in a new session'), and what the tool returns ('a compact full snapshot'). It is clear and unambiguous. However, it does not explicitly contrast with sibling tools like tab_navigate or tab_tabs, though 'new session' implicitly points to a distinction from operating on existing tabs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear contextual guidance: this tool opens a page in a new session, and cancellation cleans up the attempt without closing sibling sessions or the shared browser. This tells the agent the tool is appropriate for isolated, temporary page opens. It lacks explicit when-not-to-use instructions or named alternatives, but the context is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tab_pdfExport page as PDFA
Destructive

Print the active page into a private local PDF artifact and return its path, SHA-256 and source URL. Files survive closing. Print layout can differ from the screen; inspect the artifact when layout matters.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNo
landscapeNo
session_idYesSession ID returned by tab_open.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnly=false, destructiveHint=true, and openWorldHint=true. Beyond those, the description adds valuable context: the artifact is private and local, files survive closing, and print layout may differ from screen layout. This gives the agent important expectations about side effects and output fidelity without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: the first states the core operation and outputs, the second notes persistence, and the third warns about layout fidelity. The most important information is front-loaded, and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description properly explains the return values: path, SHA-256, and source URL. It also covers persistence and the layout caveat. Missing details are limited to parameter defaults or behavior for format/landscape, which are partially inferable from the enum schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, with only session_id documented in the schema. The description does not mention format or landscape at all, so the agent gets no guidance on how to choose A4 vs Letter or when to set landscape. With low schema coverage, the description needed to compensate for these parameters but did not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Print') and a specific resource ('the active page ... into a private local PDF artifact'), and clearly states what is returned: path, SHA-256, and source URL. This clearly distinguishes it from sibling tools like tab_capture or tab_snapshot, which are more likely screenshots or structured captures.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the basic use case clear: export the active page to a PDF. However, it does not explicitly name alternatives or say when not to use this tool, even though siblings like tab_snapshot, tab_extract, and tab_capture exist. The guidance about print layout differing from the screen is useful but is more behavioral than usage-routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tab_snapshotObserve pageA
Read-only

Read the page and mint a new snapshot_id. Scope to one CSS root with selector or the visible viewport with viewport_only to reach controls beyond a truncated page. Changing scope resets the diff baseline. Respect truncation and frame metadata. When secrets are configured, available_secrets lists aliases available for the observed frame and top-level origin; it contains no secret values.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo
frame_idNo
selectorNo
session_idYesSession ID returned by tab_open.
text_limitNo
max_elementsNo
viewport_onlyNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool read-only, and the description adds meaningful context beyond that: changing scope resets the diff baseline, and secrets handling exposes only aliases, never values. The instruction to 'Respect truncation and frame metadata' is useful but slightly terse.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five compact sentences, with the core action front-loaded. Scope guidance, the diff-baseline caveat, truncation/frame caution, and secret-handling behavior are all relevant, and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With seven parameters, low schema coverage, and no output schema, the description must supply substantial context. It covers scope selection and diff behavior well, but leaves full/diff mode semantics and the return payload largely implicit. Adequate with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 14%, so the description carries most of the parameter burden. It adds real meaning for selector and viewport_only, and hints at frame_id via frame metadata, but it never explains the mode full/diff choices, text_limit, or max_elements. This is partial compensation, not full.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Read the page and mint a new snapshot_id', which is a specific action with a distinct output resource. It does not explicitly name a sibling alternative, but the snapshot_id concept and scope language make the tool's role clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete selection guidance: use a CSS selector for a root or viewport_only to reach controls beyond a truncated page. It also warns that changing scope resets the diff baseline, which is useful when deciding scope. It does not explicitly compare against sibling tools or state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tab_stateSave browser authentication stateA
Destructive

Save cookies, localStorage and IndexedDB to a private local artifact file for a later tab_open storage_state import. The file can contain credentials; state contents are not returned to the model. SessionStorage, extensions and open tabs are not saved.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesSession ID returned by tab_open.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description adds meaningful behavioral details: data is stored privately, may contain credentials, is not returned to the model, and explicitly excludes sessionStorage, extensions, and open tabs. This is valuable context for a write operation and does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences with no filler. The core behavior is front-loaded, and the key exclusions and security caveat are packed efficiently into the second sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with a straightforward write-side purpose, the description covers what is saved, where, the security implication, and the consumer of the saved state. No critical invocation information is missing even without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, session_id, is already fully documented in the schema as 'Session ID returned by tab_open.' The description adds no additional parameter-level detail, so the baseline of 3 applies given 100% schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Save'), specific resources ('cookies, localStorage and IndexedDB'), and the destination ('private local artifact file'). It also defines the downstream purpose ('later tab_open storage_state import'), which clearly distinguishes this tool from snapshot/extraction siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use case is clearly stated: preserve browser authentication state for a later tab_open import. It does not explicitly list alternatives or when-not-to-use conditions, but the stated purpose gives enough context for an agent to select this tool over reading/extraction tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tab_tabsManage owned tabsA
Destructive

List, open, switch or close tabs within a session. Popups stay open and appear in snapshots; switch explicitly to work in them. Only this session's pages are accessible. Closing its last tab closes the session.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNo
actionYes
tab_idNo
session_idYesSession ID returned by tab_open.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, and destructiveHint=true, so the description doesn't need to restate those. It adds valuable behavioral context: popups stay open and appear in snapshots, switching is required to work in them, and closing the last tab closes the session. This goes beyond the annotations and helps the agent understand side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each carrying distinct information: the action scope, the popup behavior, and the session lifecycle. No filler or repetition of schema details. The most important usage constraint (popups) is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-action tool with no output schema, the description covers the key behavioral nuances: popup handling, session scoping, and session closure. It doesn't describe return values or error conditions, but the action enum and parameter names give a reasonable picture. The missing parameter-to-action mapping is a minor gap, but the description is otherwise complete enough for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25% (only session_id has a description). The description does not explain the meaning of url, tab_id, or the action enum values beyond the enum itself. It implies that 'new' likely needs a url and 'switch'/'close' need a tab_id, but it doesn't explicitly map parameters to actions. Baseline 3 is appropriate because the description adds some context but doesn't fully compensate for the low schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb set ('List, open, switch or close tabs') and a specific resource ('tabs within a session'), and it distinguishes itself from sibling tools by noting that only this session's pages are accessible. It also clarifies the popup behavior, which is a unique trait not evident from the name or schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit context: popups stay open and appear in snapshots, and you must switch explicitly to work in them. It also states that only this session's pages are accessible, which implies when to use this tool versus other tab tools. However, it does not explicitly name alternative tools or say when not to use this tool, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tab_verifyVerify browser outcomeA
Read-only

Check 1–20 explicit assertions. For current input/textarea/select values, use kind:value with exactly one observed ref or CSS selector. Any ref requires the current snapshot_id; the same snapshot remains readable after act if no newer snapshot or document change replaced it. Text checks inspect visible page text, excluding raw form values. A successful click alone does not prove success.

ParametersJSON Schema
NameRequiredDescriptionDefault
checksYes
session_idYesSession ID returned by tab_open.
timeout_msNo
snapshot_idNoRequired when any value check uses ref. Use the latest snapshot_id, including one returned by tab_act.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds significant behavioral context beyond the readOnly/openWorld/destructive annotations: snapshot validity can expire after a newer snapshot or document change, text assertions operate on visible page text rather than raw form values, and a successful click does not guarantee a successful outcome. These are non-obvious semantic details that materially affect how the tool should be used.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four sentences, front-loaded with the core purpose, and every sentence carries useful operational information. It avoids repeating schema constraints and packs the most important caveats into a compact, readable form.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with nested check variants and snapshot-semantics, the description is almost complete: it covers assertion count, value-check ref/selector rules, snapshot staleness, and text-check scope. It does not describe the return shape or pass/fail format, but the tool's purpose as an assertion checker makes the outcome type fairly inferable, so the gap is minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With only 50% schema description coverage, the description compensates by clarifying key parameter behaviors: exactly one observed ref or CSS selector for value checks, the snapshot_id requirement tied to refs, and the visibility semantics of text checks. It does not enumerate all six check kinds, but the schema already provides the const values and the description adds meaning beyond the raw property definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and scope: 'Check 1–20 explicit assertions,' which clearly identifies this as a verification tool for browser outcomes. It enumerates concrete behaviors (value checks, text checks, visibility) that distinguish it from snapshot/find/act siblings. It is not a tautology and directly supports the title's 'Verify browser outcome.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives actionable when-to-use guidance: use kind:value for current form element values, use the current snapshot_id when a ref is needed, and understand that text checks exclude raw form values. It does not explicitly name sibling alternatives or say when not to use the tool, but the usage context is clear enough for an agent to invoke it correctly after a browser action.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 16 tool updatesv0.1.0
    • First observedtab_act
    • First observedtab_capture
    • First observedtab_close
    • First observedtab_dialog
    • First observedtab_downloads
    • First observedtab_extract
    • First observedtab_extract_structured
    • First observedtab_find
    • First observedtab_list
    • First observedtab_navigate
    • First observedtab_open
    • First observedtab_pdf
    • First observedtab_snapshot
    • First observedtab_state
    • First observedtab_tabs
    • First observedtab_verify

TDQS

A4.1/5.0

Scored across 16 tools

Disambiguation4/5

Most tools map cleanly to a distinct browser operation: session lifecycle, navigation, snapshots, actions, extraction, verification, downloads, dialogs, and PDF. The closest overlap is between tab_snapshot/tab_find and tab_extract/tab_extract_structured, but the descriptions clearly separate search-and-scroll from simple read and simple extraction from schema-validated extraction.

Naming Consistency5/5

Every tool uses a consistent tab_ prefix and snake_case naming, making the set highly predictable. Although a few names like tab_tabs and tab_downloads are noun-based rather than strict verb_noun, the uniform prefix convention removes ambiguity.

Tool Count4/5

At 16 tools, this is at the upper edge of a comfortable browser automation toolkit, but each tool covers a genuinely distinct capability such as dialogs, downloads, PDF, multi-tab management, and storage state. It feels slightly heavy but not bloated for the stated purpose.

Completeness4/5

The set covers the core browser automation lifecycle: session open/close, navigation, snapshots, actions, verification, extraction, capture, downloads, dialogs, PDF, and state persistence. Minor gaps such as direct cookie/localStorage editing and an explicit wait primitive can be worked around via tab_state and post_checks or tab_verify.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI agents to operate an isolated local Chromium browser through MCP, with semantic snapshots, ref-based actions, search, research, crawling, and CDP access.
    Apache 2.0
  • F
    license
    Not graded
    quality
    A
    maintenance
    Enables AI agents to control a persistent local browser with live tabs, navigation, interaction, inspection, and state management through MCP.
    6
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides a headless Chromium browser through MCP, enabling AI agents to browse JavaScript-rendered pages, search the web, capture screenshots, extract tables and data, and run stateful multi-step interactions like clicking, typing, and form submission.
    2,021,532 npm
    1
    MIT