duplex
Supports Google Gemini as a native model API and Gemini CLI as an external agent tool, enabling Gemini models to drive or assist browser automation.
Integrates with Ollama local models via an OpenAI-compatible endpoint for the built-in agent.
Supports OpenAI-compatible APIs and the native OpenAI Responses API as model providers for the built-in agent.
Duplex
One browser shared by a human and an AI — the human sees the rendered page, the AI reads the DOM and page source. Same tabs, same live session, at the same time.
English | 中文

Duplex start page (dark theme). A light theme is also included.
Demo

The AI drives: it searches Bilibili, opens the video and starts playback — the side panel streams every step.
Full demo — 3 min, bilingual subtitles, with music:

One continuous walkthrough: Act 1 — the built-in agent searches Bilibili and plays a paper-explainer video; Act 2 — Kimi browses to Claude and asks what CUDA is, then summarizes the answer; feature highlights — multi-model APIs, skills, external agents, emergency stop; and where to get it.

Controlling Microsoft Excel for the web — no dedicated spreadsheet-editing skill installed; only the stock API and a few simple automation skills 😅
A feasibility demo, not a recommended workflow. With no task-specific skill to lean on, Duplex worked with the stock API plus a few simple automation skills alone — improvising clipboard read/write channels, DOM probing and the like to push a whole grade sheet into Excel for the web. It got there, but the process wasted a lot of time and tokens. 😅
Related MCP server: Ghostlight
⚠️ Usage notes
Security boundaries (high stakes — please read). Duplex lets the AI drive your real browser session, including login state, cookies and local data. Under this architecture, a wrong AI move can touch real accounts and data (sending messages, submitting forms, modifying or deleting content), potentially with serious consequences. Don't leave the agent running unattended in environments where sensitive accounts are signed in.
Esctakeover is a last-resort human brake — it is no substitute for your own judgement about what the AI should be allowed to touch.The tooling layer is still being tuned. The goal is for the human to stay in command of the AI's tools, rather than have AI tooling and the human's own actions crowd each other out of the pipeline (contending for the same page, interrupting input, interleaving conflicting actions). The trade-offs here are still evolving — feedback on human/agent contention is welcome via Issues.
What's new in v0.2
Skills — Claude Code-compatible skill format (
SKILL.md): indexes your existing~/.claude/skillsand supports importing or authoring duplex-only skills under~/.cobrowse/skills. The built-in agent reads and follows them on demand (every script run asks for confirmation first).Multi-protocol model APIs — beyond OpenAI-compatible, native support for Anthropic Messages, OpenAI Responses and Google Gemini.
Credential import — reuse your local Codex (ChatGPT subscription) and opencode logins directly; no API key typing needed.
External agent tools — switch the panel between opencode / Codex / Claude Code / Gemini CLI / Qwen Code / custom CLIs: transcripts mirror into the panel (history + live), Codex / Claude Code sessions can be continued by replying right from the panel (the message headlessly resumes the same session), and new sessions can be started right from the panel. Codex / Claude Code also get a model picker (default follows the CLI config; one click writes the choice back to the global config, with a timestamped backup); custom tools can be added manually (session directory + start command).
Safety & robustness — global emergency-stop hotkeys (default
Esc/F2, customizable from the ⌨ button), triple-confirm deletes, task watchdog with automatic render recovery, process-tree termination, and more.
How Duplex compares
Duplex | browser-use | Browser MCP | Playwright MCP | AI browsers (Atlas / Comet) | |
Form | Desktop browser (Electron app) | Python automation framework | MCP server (browser extension) | MCP server (Microsoft) | Closed-source product |
Who uses the browser | Human and AI share the same tab and the same live session | AI only (separate automation instance) | AI drives your current Chrome | AI only (Playwright instance) | AI assistant alongside/operating |
What the AI sees | DOM outline snapshot + source + screenshots | Vision + DOM | Screenshots + a11y tree | Accessibility tree | Internal |
Human collaboration | Real-time side-by-side; | Logs afterwards | Human spectates | Human spectates | Limited intervention |
Connectable AI | Built-in models + opencode / Codex / Claude Code / Gemini / Qwen / any MCP client | Bring your own LLM | Any MCP client | Any MCP client | Official model only |
Conversation visibility | Live side-panel mirror (including external CLIs' chats and tool calls) | Logs/terminal | In the client | In the client | In-app |
Data | Fully local | Local/cloud | Local | Local | Cloud |
In one line: browser-use / Playwright MCP let the AI run a flow for you; Duplex lets the AI use the browser together with you — same tab, same session, takeover anytime.
Highlights
A real browser — tabs, address bar with search (Baidu / Bing / Google), back / forward / reload, loading state, themes (light / dark / follow system) and a wallpaper start page with a clock and search.
24 MCP tools for AI agents —
snapshotcompresses any page into a compact DOM outline with[eN]refs; the other tools cover tabs, navigation, clicking, typing, dragging, file upload, scrolling, waiting, console logs, JS evaluation and page annotations.Zero-setup bridge — a stdio MCP bridge (
mcp-bridge) auto-launches the browser on the first tool call. Works with opencode, Claude Code, or any MCP client.Live session mirror — when your AI works through opencode, its replies, reasoning and tool-call cards stream into the side panel in real time. Type in the panel to inject a message into the same session.
AI action visualization — a translucent cursor, element highlight and a status bar ("AI is clicking «…» — Esc to take over") are drawn in a Shadow-DOM overlay, so you always see what the AI is doing on the page.
Esc takeover — press
Esc(or click the status bar) to take control instantly: the running tool call is aborted, in-flight waits return early, and the AI is told the user took over.Page annotations — press
?or useannotation_modeto draw a box / circle / arrow / point on any page and attach a question. The annotation is compiled into a structured text brief (DOM outline + visible text + selectors + geometry) and sent to the AI.Built-in agent (optional) — connect any OpenAI-compatible API (DeepSeek, Kimi, Qwen, GLM, Ollama, …) and let the browser drive itself. Provider management supports one-click import from opencode.
Conversation history — built-in agent sessions are saved locally and can be reopened from the history menu.
Feature tour
1. Start page

The start page shows a clock, date and search box over a wallpaper that follows the active theme (light and dark wallpapers switch automatically with the system).
2. Built-in agent drives the browser

Ask the built-in model in the side panel — in Chinese or English. It calls tools (navigate, snapshot, …), with each call shown as a card, and reports back in the panel while you watch the page change.
3. opencode session mirror

When your AI runs through opencode, its messages and tool calls are mirrored into the panel while it drives the same visible browser. The "opencode" / "built-in" tabs at the top of the panel switch between the two ways of working.
4. Session picker

Pick which opencode session the panel is connected to. "Auto" follows the most recent conversation.
5. Model providers

Manage OpenAI-compatible providers for the built-in agent: add, edit, delete, or import providers directly from your opencode configuration. Local models via Ollama (http://localhost:11434/v1) work out of the box.
6. Conversation history

Built-in agent conversations are stored locally and can be reopened at any time.
7. AI action visualization and Esc takeover

Every AI action is drawn on the page: a cursor ring, an element highlight and a status bar. Press Esc at any moment to take the browser back — the AI stops immediately and waits for your instruction.
8. Page annotations

Draw a box (or circle / arrow / point) around anything and ask a question about it. The annotation is converted into a structured text brief for the AI, including the DOM outline of the region, visible text, selectors and geometry — so even text-only AIs can "see" what you mean.
Quick start
Installer (Windows / macOS / Linux)
Download the latest installer from Releases:
Duplex Setup x.y.z.exe(Windows),.dmg(macOS, Apple Silicon / Intel) or.AppImage(Linux).Run it and launch Duplex.
Build from source
npm install
npm run build # Electron app -> out/
npm run build:bridge # stdio MCP bridge -> dist-bridge/index.cjsIf Electron's binary download fails (e.g. behind a slow mirror), retry with:
$env:ELECTRON_MIRROR="https://npmmirror.com/mirrors/electron/"; node node_modules/electron/install.js
Run it:
npm run dev # development mode with HMR
# or
npm start # preview the production buildConnect opencode (recommended)
Install the mirror plugin. Copy
integrations/opencode/plugins/cobrowse-mirror.tsto your global opencode plugin directory (~/.config/opencode/plugins/), or into.opencode/plugins/of a project.Register the MCP server in your opencode config (
opencode.json), project-level or global — seeopencode.example.json:
{
"mcp": {
"duplex": {
"type": "local",
"command": ["node", "C:\\path\\to\\Duplex\\dist-bridge\\index.cjs"],
"enabled": true
}
}
}In opencode, ask the AI to browse: "open example.com and describe the page". The browser launches automatically on the first tool call (no need to start it manually).
In the Duplex side panel choose the opencode tab and select the session you want to follow (or keep "Auto").
Use the built-in model (no opencode required)
Panel → Built-in tab → Model settings → add an OpenAI-compatible provider (base URL, API key, model name), or click Import from opencode to reuse your existing provider configuration. For local models, point the base URL to Ollama, e.g. http://localhost:11434/v1.
MCP tools
Tool | Description |
| Tab management — human and AI share the same tabs |
| Open a URL (auto-searches if not a URL), search (Baidu default, or Bing / Google), back / forward / reload |
| Page as text — compact DOM outline; interactive elements carry |
| Raw HTML, or detailed info for elements matching a CSS selector |
| PNG screenshot of the visible page (for multimodal models) |
| Click / double-click / hover; |
| Type text (optionally submitting) and press keys, incl. combos like |
| Real drag & drop from one point/element to another |
| Native |
| Scroll the page or a specific element into view |
| Wait for time / selector / text — interruptible by Esc |
| Read page console logs (errors and warnings) |
| Enter/exit the annotation overlay (human draws a box/circle/arrow/point + question) |
| Evaluate JS in the page, returns JSON-serializable results |
How it works

The Electron main process serves a small local HTTP API on
127.0.0.1protected by a per-launch bearer token (endpoint info is written to~/.cobrowse/endpoint.json).dist-bridge/index.cjsis a stdio MCP server that proxies tool calls to that API and auto-launches the app when it is not running.The opencode plugin pushes session events (text, reasoning, tool calls) into the panel, and long-polls for messages queued in the browser (~10 ms injection latency).
Panel messages and page annotations are injected into the active opencode session with
session.promptAsync.Codex and Claude Code are mirrored from their local session transcripts (history + live tail); replies sent from the panel headlessly resume the same session (
codex exec resume/claude --resume) and the new turns stream back through the same tail.External scripts can also push messages into the session:
POST /api/chat { "text": "..." }.
Known limitations
Page internals inside iframes and shadow DOM are not covered by
snapshot/click— shadow DOM is only detected, not entered.eNrefs are invalidated by navigation; re-runsnapshotafter the page changes.Message injection targets the "active session"; with several opencode sessions the target may occasionally be ambiguous.
The mirror store holds the recent event stream in memory; it resets on browser restart.
The built-in agent is a convenience option: local / smaller models are noticeably less reliable at long tool-use chains than a full opencode setup.
Roadmap
Richer client integrations — Codex / Claude Code support has landed; more on the way:
The bridge is client-agnostic (standard stdio MCP), so the tool layer is not opencode-specific by design.
Session mirroring is in: transcripts replay into the panel live, and replies sent from the panel resume the same session headlessly (Codex
exec resume/ Claude Code--resume) — the CLI appends to its transcript and the tail streams it back into the panel.Next: richer event mapping (tool calls / reasoning cards) and more clients (Gemini CLI / Qwen Code).
Development
npm run dev # dev mode (electron-vite, HMR for the renderer)
npm run typecheck # TypeScript checks (node + web)
npm test # unit tests (vitest)
npm run smoke # end-to-end smoke test (launches the bridge and a real browser)
npm run dist # build the Windows installer (electron-builder)Debug helpers:
GET /api/debug/ui-snapshot— captures the current window to~/.cobrowse/ui-snapshot.png.POST /api/debug/panel-eval— runs JS in the panel renderer (only when the app is started withCOBROWSE_DEBUG_UI=1).Logs:
~/.cobrowse/app.log(main + renderer log lines).
License
Available Tools
24 toolsannotation_modeA
Turn the human annotation mode on/off on a tab. When on, the human can draw a box/circle/arrow (or pick an element) and attach a question; the annotation is delivered to you as a structured text message.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No | ||
| active | Yes | true = enter annotation mode, false = exit |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It usefully discloses the key behavioral trait that annotations arrive as a structured text message, but omits persistence/scope behavior (per-tab only? survives navigation?), the default state, and whether tabId is optional or defaults to the active tab.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the action is front-loaded and the second sentence explains the payoff of the mode. No filler, no repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description explains the input side (what the human will do) but not the mechanics an agent needs to call it correctly: tabId optionality, mode lifetime, and whether annotations interrupt or arrive asynchronously. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: the schema documents 'active' (true = enter, false = exit) but leaves 'tabId' unexplained. The description only hints at tabId via 'on a tab' and never clarifies whether it is optional or defaults to the current tab, so it only marginally compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('turn the human annotation mode on/off on a tab') and then describes the concrete effect of the mode — human draws shapes or picks an element and attaches a question. This is clearly distinguishable from every sibling (screenshot, click, evaluate) without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the usage context (enable it so the human can send you structured annotation messages) but never states when to enable vs. leave off, whether it should be toggled off after receiving input, or any alternative for obtaining human feedback. Guidance is inferable but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clickC
Click an element, identified by a [ref] from snapshot (e.g. "e12") or a CSS selector.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No | ||
| target | Yes | ref like "e12" or CSS selector |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full behavioral burden. It says nothing about whether the element is auto-scrolled into view, waits for actionability, how failures are reported, or whether the click targets a specific tab. For a core interaction tool this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence, front-loaded with the action and the accepted target formats, with zero filler. Nothing could be removed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and one of two parameters undocumented, the description is too thin for a foundational click tool. It omits targeting scope (tabId), actionability/wait semantics, and failure behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: the target parameter is already documented in the schema with the same 'e12 or CSS selector' phrasing the description repeats, adding no new meaning, while tabId is undocumented in both the schema and the description. The description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (click) and resource (an element), and explains how the target is identified via a ref from snapshot or a CSS selector, which implicitly links to the sibling snapshot tool. It does not explicitly differentiate from near-neighbors like dblclick or hover, but the verb itself is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use click versus alternatives such as dblclick, hover, press, or type, nor any conditions or prerequisites (e.g. element must exist first). The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_tabB
Close a tab by its id.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | Yes | Tab id from list_tabs |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the core action but says nothing about side effects, permissions, whether closing is reversible, or what happens to tab state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with zero waste. For a simple one-parameter tool, this is appropriately sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one required parameter and no output schema, so the description is minimally sufficient for invocation. However, with no annotations and no behavioral detail, it leaves gaps around side effects and error conditions for a mutating action.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter tabId is documented as 'Tab id from list_tabs'. The description merely repeats 'by its id' without adding any format, source, or constraint details beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Close a tab by its id.' This clearly distinguishes the action from siblings like new_tab, switch_tab, and list_tabs, though it does not explicitly name or contrast those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives, nor any prerequisites or conditions for closing a tab. The intended usage is only implied by the verb 'close'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dblclickB
Double-click an element by ref or CSS selector.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No | ||
| target | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and mostly doesn't. It doesn't disclose whether the element is auto-scrolled into view, what happens if the target is missing or obscured, or whether the action waits for completion.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that names the action and the accepted targeting forms with zero filler. Nothing is padded or restated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an action tool with no annotations and no output schema, the description is thin: it omits the meaning of tabId, the element-readiness/auto-scroll behavior, and failure modes. It gives just enough to attempt a call but not enough to call it reliably.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the description does partially compensate by clarifying that 'target' accepts either a ref or a CSS selector, which is real information beyond the bare string type. However, 'tabId' remains completely undefined in both the schema and the description, so the compensation is incomplete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (double-click) and resource (element), which cleanly separates it from the sibling 'click' by action semantics. It stops short of any sibling routing or scope detail, but the core action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the verb itself; the description never says when a double-click is preferred over click, hover, or drag, nor whether the element must be visible or in the viewport first. An agent can infer intent but gets no explicit conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dragB
Drag from one point/element to another (refs or CSS selectors) using real mouse events.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | Target element ref or CSS selector | |
| from | Yes | Source element ref or CSS selector | |
| tabId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. It usefully reveals the mechanism is 'real mouse events' rather than synthetic or HTML5 drag events, which materially affects whether a DnD target responds. It does not disclose side effects, visibility requirements, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that front-loads the action and endpoints with no filler. Slightly under-specified rather than verbose, so length is not the problem.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter mutation-style browser action with no annotations and no output schema, the description covers the core mechanic but omits tab targeting semantics and any statement about what a successful drag returns or how to verify it. Adequate minimum, clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% with from/to documented as refs or CSS selectors — the description repeats that same detail rather than adding new meaning. The undocumented tabId parameter is not explained anywhere, so the description neither compensates for the gap nor adds value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (drag) and the source/target resources, and clarifies both endpoints accept refs or CSS selectors. It is implicitly distinguishable from siblings like click, hover, and dblclick, though no sibling is named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the verb — an agent can infer this is the tool for drag-and-drop interactions rather than clicking. However, there is no explicit when-to-use guidance, no mention of prerequisites (e.g. elements must be visible or already located via snapshot), and no routing away from click/hover.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluateB
Run JavaScript in the page and return the JSON-serializable result. Powerful; prefer snapshot/query when possible.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No | ||
| script | Yes | JS expression or statements; use "return x" for a value |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose the return contract (JSON-serializable result) and flags the tool as "Powerful," which is genuine context, but it says nothing about side effects, page-state mutation, permissions, or execution constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short clauses, front-loaded with the capability before the caution. Nothing wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param tool with no output schema, the return behavior is covered by the description. The remaining gaps — tabId semantics and the consequences of executing arbitrary JS — are meaningful for an eval tool but not fatal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% and the description adds no parameter meaning at all — notably, tabId is undocumented in both schema and description, leaving the agent to guess which page the script runs in. The script description in the schema already covers the return-value idiom.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: run JavaScript in the page and return the serializable result. It also names the siblings it overlaps with (snapshot, query), which helps an agent separate it from read-only inspection tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"prefer snapshot/query when possible" gives a real routing hint toward the alternatives, but the condition is vague — it never says when evaluate IS the right choice, nor when it is unsafe or unnecessary. Adequate but incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_consoleB
Read recent console messages of a tab (errors/warnings/logs) collected since load.
| Name | Required | Description | Default |
|---|---|---|---|
| clear | No | Clear the buffer after reading | |
| limit | No | Max messages to return (default 50) | |
| tabId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that the buffer accumulates from page load (so pre-load messages are unavailable) and which severities are captured. It does not disclose rate limits, whether reads are non-destructive by default, or that a 'clear' flag can mutate the buffer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the verb and resource and packs the scope constraint ('since load') efficiently. Nothing is wasted, though the parenthetical list is slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with no output schema and no annotations, the description is adequate but leaves gaps: it does not state that all three parameters are optional, what happens when tabId is omitted, or that the buffer can be cleared. An agent can call it correctly but not optimally.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67%: 'clear' and 'limit' are documented in the schema, while 'tabId' has no description. The description only implies the tab scoping ('of a tab') and adds nothing about the limit or clear semantics, so it neither compensates for the gap nor adds value over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read recent console messages of a tab') and enumerates the message types (errors/warnings/logs), so the agent knows exactly what data comes back. The resource is unique among siblings (snapshot, get_html, evaluate), so no explicit sibling differentiation is strictly needed, though it is not called out.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no named alternatives among the many siblings (e.g., evaluate or snapshot for other diagnostic data). The phrase 'collected since load' hints at the data's scope but does not tell the agent when this tool is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_htmlB
Get HTML source. Without selector: cleaned body HTML. With selector: outerHTML of the first match.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No | ||
| maxChars | No | Truncate to this many chars (default 40000) | |
| selector | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. It helpfully explains the two output modes (cleaned body HTML vs. outerHTML of the first match), but leaves open what 'cleaned' means, what happens when no element matches, and any permission or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core verb+resource, then the mode-specific returns. Every word earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-required-parameter browser extractor with no output schema, the description should explain the return shape (it partly does) and which tab is targeted by tabId (it does not). It is adequate but leaves the tab-selection default and fallback behavior unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only maxChars is documented in-schema). The description compensates by defining selector semantics ('outerHTML of the first match'), but says nothing about tabId, leaving a third of the parameters undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Get HTML source') and then explains the two extraction modes, so the agent knows exactly what is returned in each case. It does not explicitly distinguish itself from close siblings like query, snapshot, or evaluate, which is the only thing keeping it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use / when-not-to-use guidance and no alternatives are named. The with/without-selector wording describes behavior rather than telling the agent which situation should select this tool over query, snapshot, or evaluate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
historyC
Go back, go forward, or reload in a tab.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No | ||
| action | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden, and it discloses almost nothing: no mention of required permissions, what happens if back/forward is unavailable, or which tab the action applies to by default. Only the three literal actions are conveyed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the action set front-loaded and no filler. It is efficient, though its brevity is partly under-specification rather than disciplined concision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and an undocumented optional tabId, the definition leaves key questions unanswered: default tab behavior, failure modes, and return value. For a navigation-mutating tool this is thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description never mentions tabId, leaving the optional target tab completely undocumented. The action enum values are self-explanatory from the schema, but the second parameter receives no explanation anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states three concrete operations (back, forward, reload) scoped to a tab, which is more specific than the bare name 'history'. However it does not differentiate from the sibling 'navigate', which also moves a tab's location.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus navigate or switch_tab, nor any precondition (e.g. whether history exists to go back to). Usage has to be inferred entirely from the enum values.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hoverA
Move the mouse over an element (ref or CSS selector) to trigger hover menus/tooltips; the pointer stays there.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No | ||
| target | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses that the pointer remains on the element ('the pointer stays there'), implying persistent post-action state, but it omits whether the tool waits for the menu/tooltip to appear, what happens if the element is missing, and any side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the action, the accepted input formats, and the resulting state are all conveyed economically.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter interaction tool with no annotations and no output schema, the description covers the core action and input format. It still leaves tabId undefined and doesn't mention waiting behavior or failure modes, so an agent lacks some context needed for robust invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does add real value by clarifying that 'target' accepts either a ref or a CSS selector, which the bare 'string' schema does not convey, but the second parameter 'tabId' is never explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('move the mouse over') plus the resource ('an element') and the outcome it enables (hover menus/tooltips), which clearly separates it from siblings like click and dblclick. It stops short of explicitly naming an alternative tool, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'to trigger hover menus/tooltips' implies when to reach for it, which is decent implied guidance. However, there are no when-not conditions, no prerequisites (e.g., element must be visible), and no named alternatives among the many interaction siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tabsA
List all open browser tabs (id, url, title, loading, active). The human sees the same tabs.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It does disclose that it returns all tabs unfiltered and that the human sees the same set (useful context), but it says nothing about ordering, permissions, or how the operation behaves. Adequate but incomplete for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero waste. The purpose is front-loaded and the field list follows immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description's enumeration of returned fields (id, url, title, loading, active) is genuinely load-bearing and covers the return contract. For a parameterless read tool the only remaining gap is ordering/behavioral detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the baseline is 4. The description's field list arguably describes the (absent) input, but there is nothing to misinterpret.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource ('List all open browser tabs') and enumerates the returned fields (id, url, title, loading, active), which lets an agent distinguish it from siblings like new_tab, close_tab, and switch_tab without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the verb 'List' and the fact that it returns all tabs; there is no explicit when-to-use guidance, no mention of the obvious alternatives (switch_tab, close_tab), and no exclusions. An agent can infer purpose but gets no routing help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
new_tabB
Open a new tab in the browser, optionally navigating to a URL.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to open; omit for a blank page |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It states the basic action but omits key traits such as whether the new tab becomes active, whether it waits for navigation, or what it returns (no output schema).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the core action and includes the optional parameter behavior without any wasted words. Perfectly sized for a simple tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with no annotations or output schema, the description covers the core action adequately. However, it lacks integration hints (e.g., return value or focus behavior) that would help an agent chain it with siblings like switch_tab, leaving a meaningful gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single 'url' parameter, and the description adds no syntax, format, or default details beyond what the schema already provides. Baseline 3 is appropriate when the schema fully documents the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Open a new tab in the browser'), which clearly distinguishes it from list_tabs, close_tab, and switch_tab. However, it does not explicitly differentiate from the sibling 'navigate', which also accepts a URL, leaving some ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by noting the URL is optional ('optionally navigating to a URL'), but gives no explicit when-to-use guidance or alternatives (e.g., when to use new_tab versus navigate). Usage context is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pressC
Press a key or combo (Enter, Escape, Tab, PageDown, Control+A, Shift+Tab, ...) in the page.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Single key or combo like "Control+A" | |
| tabId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it discloses only that the key is sent "in the page." It says nothing about whether an element must be focused, whether the target tab must be active, whether the press blocks or returns immediately, or what side effects (navigation, form submit) can result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the examples are the one addition that earns its place by pinning down the key syntax.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a browser-interaction tool with a second (tabId) parameter, no output schema, and no annotations, the description leaves out focus requirements, tab targeting, and outcome behavior. An agent could call it but could easily press into the wrong context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%. The examples ("Enter, Escape, Tab, PageDown, Control+A, Shift+Tab") genuinely clarify the format of the `key` parameter beyond the schema's terse description, but `tabId` is undocumented in both the schema and the description, leaving a required-targeting gap unaddressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Press a key or combo... in the page") with concrete key examples (Enter, Escape, Control+A, Shift+Tab) that make the action unambiguous. It does not explicitly differentiate from the adjacent `type` sibling, which is the only real gap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance at all. An agent must infer that this is for discrete keypresses vs. `type` for text entry, and nothing says what prerequisites (focused element, active tab) are required before pressing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
queryB
Query elements by CSS selector; returns details for each match (tag, id, classes, text, href, rect, visibility).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max matches (default 20) | |
| tabId | No | ||
| selector | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does disclose the return shape (tag, id, classes, text, href, rect, visibility), which is genuinely useful. It omits whether the query is non-mutating, how it behaves on zero matches, whether it waits for dynamic content, and how the limit truncates results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the action and then lists the return fields. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description usefully documents the returned fields, which is a meaningful contribution. Still missing are no-match behavior, wait semantics for dynamic DOM, and the tabId parameter's role, which leaves the agent guessing in this DOM-heavy tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% — selector and tabId have no schema descriptions. The description clarifies that selector is a CSS selector, which is the most important clarification, but tabId's meaning (active tab vs. specific tab) is left unexplained in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Query elements by CSS selector') and enumerates the returned fields, so the agent knows exactly what comes back. It does not, however, distinguish itself from nearby read-oriented siblings like snapshot, get_html, or evaluate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool rather than snapshot, get_html, or evaluate, and no exclusions or prerequisites. The agent must infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshotB
Take a PNG screenshot of the current tab. Returns an image (for vision-capable models).
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No | ||
| fullPage | No | Capture the full scrollable page (default false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the output is an image (PNG) and that vision capability is needed, which annotations would not have covered. It does not disclose default capture scope (viewport vs full page), where the image is returned (inline data vs path), or any permission/size constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and then the return type. Nothing superfluous.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, so the description does the heavy lifting on the return value, which it handles. But for a 2-parameter tool with one undocumented parameter and ambiguous overlap with 'snapshot', the definition is only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50% – 'fullPage' is documented in the schema, but 'tabId' is undocumented in both schema and description. The description mentions no parameters at all, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Take a PNG screenshot of the current tab'), so the agent knows exactly the action. However, it offers no differentiation from the sibling 'snapshot', which is exactly the tool an agent might confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to prefer this over 'snapshot', 'get_html', or 'evaluate'. The 'for vision-capable models' clause implies a condition, but it is about model capability, not usage context or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrollB
Scroll the page by dx/dy pixels (positive dy = down, positive dx = right), or scroll an element into view when selector is given.
| Name | Required | Description | Default |
|---|---|---|---|
| dx | No | Pixels to scroll horizontally (default 0) | |
| dy | No | Pixels to scroll vertically (default 600) | |
| tabId | No | ||
| selector | No | Scroll this element into view instead |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full behavioral burden, but it only discloses direction semantics and dual modes. It omits timing (instant vs. smooth), side effects on page state, failure behavior when a selector is not found, and which tab is affected. This is a significant gap for a browser automation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that packs the two modes and direction semantics without any filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple scroll tool, the description covers the primary invocation paths, and the schema supplies defaults for dx and dy. However, with no annotations and no output schema, the omission of tabId semantics and scroll behavior leaves minor gaps an agent must infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, and the description adds critical meaning beyond the schema by defining sign conventions: positive dy scrolls down and positive dx scrolls right. The selector mode is also restated usefully, though tabId remains undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Scroll') and resource ('page' or 'element'), and clarifies two modes: pixel scrolling and element-into-view. It does not explicitly differentiate itself from sibling tools like 'navigate' or 'click', but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains two internal modes (pixel scrolling vs. selector-based scrolling) but gives no guidance on when to use this tool instead of alternatives such as 'navigate', 'click', or 'snapshot'. No prerequisites, exclusions, or sibling routing are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchB
Search the web in a tab (default engine: baidu). Use for looking up information; use navigate for known URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| tabId | No | ||
| engine | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing beyond the verb. It does not say whether the search replaces the current page content in the tab, where results land (page navigation vs snapshot), or what the call returns. For a no-schema, no-annotation tool that mutates tab state, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short clauses, front-loaded with the action and immediately followed by the routing rule. No filler or restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and 0% schema description coverage across three parameters, the definition is not complete enough. An agent still cannot tell what the call returns, which tab is targeted when tabId is missing, or what side effects occur on the page.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are 3 parameters, so the description must compensate. It adds one useful fact (default engine is baidu) but leaves tabId completely unexplained and says nothing about query semantics, despite tabId being the most ambiguous parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Search the web in a tab") and explicitly distinguishes itself from the sibling navigate tool for known URLs, so an agent can separate it from the navigation tools. The default engine is named, but the "in a tab" phrasing is slightly vague about whether an existing tab context is required.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear use conditions: "Use for looking up information; use navigate for known URLs." The alternative is named and the selecting condition (known vs unknown URL) is explicit. It stops short of saying what to do when the tabId is absent or how engine choice should be made.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
select_optionA
Select an option in a native element by visible text or value. For custom dropdowns, click the trigger then click the option.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No | ||
| option | Yes | Visible option text or value (exact or substring) | |
| target | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden for a DOM-mutating action. It discloses that matching may be exact or substring via the schema, but says nothing about whether change/blur events fire, whether the element must be visible or in view, or what happens when no option matches.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero padding, with the primary native-select case front-loaded and the fallback strategy immediately after. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter, no-annotation, no-output-schema browser action, the description covers the core behavior and its main alternative. Missing are the target-parameter format, failure behavior when the option is absent, and any waiting/visibility expectation, which an agent would need for reliable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%: only 'option' is documented (exact or substring match), while 'target' and 'tabId' have no schema descriptions. The description echoes the text-or-value semantics of 'option' but adds nothing about how 'target' is specified or why 'tabId' matters, so it only partially compensates for the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Select) and resource (an option in a native <select> element) and even names the input mode ('by visible text or value'). It implicitly distinguishes itself from the sibling click tool by scoping itself to <select> elements, though it doesn't name click directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear conditional rule: use this tool for native <select> elements, but for custom dropdowns fall back to click-the-trigger-then-click-the-option. That routes the agent between this tool and siblings. It stops short of explicit when-not cases (e.g. disabled selects, hidden elements).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
snapshotA
Text snapshot of the page: compact DOM outline with [eN] refs for actionable/labeled elements. Use the refs with click/type. This is the primary way to "see" the page.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the whole burden. It discloses the return shape (compact text DOM outline with [eN] refs) and its role as the primary observation method, which is genuinely useful, but says nothing about defaults when tabId is omitted, whether hidden elements are included, or any output size limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: what it is, how to consume it, and why it matters. The key output-format detail is front-loaded in the opening clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, annotation-free, single-optional-parameter tool with no output schema, the description covers what it returns and how to use it. The only notable gap is the tabId default behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One optional parameter (tabId) with 0% schema description coverage, and the description never mentions it. 'Of the page' weakly implies the current/active page default but does not explain what tabId does or which tab is snapshotted without it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Text snapshot of the page') and characterizes the output ('compact DOM outline with [eN] refs for actionable/labeled elements'), which implicitly separates it from siblings like get_html (raw markup) and screenshot (image). An agent can tell what this produces without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear positive guidance: use the [eN] refs with click/type, and frames it as 'the primary way to "see" the page.' It lacks explicit when-not-to-use guidance (e.g., use get_html for full markup or screenshot for visual layout), so it stops short of full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
switch_tabB
Bring a tab to the foreground so the human sees it.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | Yes | Tab id from list_tabs |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses foregrounding behavior but omits whether the action requires permissions, what happens if the tab is already active or invalid, whether it changes focus across windows, and what the operation returns. This is a significant gap for a mutation/UI-control tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with zero wasted words. It states the action and its purpose efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tab-switching tool with complete parameter schema coverage, the description is minimally adequate but thin. It does not mention edge cases, permissions, or return behavior, and with no annotations or output schema, more context would improve reliability.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the schema itself documents the single parameter ('Tab id from list_tabs'). The description adds no further parameter meaning, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Bring') and resource ('a tab') with the outcome ('to the foreground'), making the tool's effect immediately clear. It does not explicitly differentiate from sibling tab tools like new_tab or close_tab, though the distinction is implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies when to use the tool ('so the human sees it') but provides no explicit when-to-use criteria, prerequisites, or alternatives to consider. Context is inferable but not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
typeC
Type text into an input/textarea/contenteditable, by ref or CSS selector. Optionally press Enter after.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| clear | No | Clear existing value first (default true) | |
| tabId | No | ||
| submit | No | Press Enter after typing | |
| target | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It notes the optional Enter press, but omits whether typing requires prior focus, how it handles a missing/unmatched target, whether it waits for the element, and what happens to existing text by default (clear=true is only in the schema).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the core action front-loaded and the optional Enter behavior appended. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter mutation tool with no annotations and no output schema, the description leaves significant gaps: focus/waiting requirements, error behavior on unmatched selectors, and the semantics of clear/tabId are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40% (text and target lack descriptions), so the description must compensate. It usefully clarifies that 'target' accepts either a ref or a CSS selector and that submit triggers Enter, but it says nothing about the 'clear' default behavior or 'tabId' scoping.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (type) and the exact targets it accepts (input/textarea/contenteditable), which clearly separates it from siblings like click, press, or hover. It does not explicitly name which sibling to use for adjacent needs (e.g., press for Enter alone), so it stops short of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says what it does but never states when to choose it over siblings such as press or evaluate, nor any prerequisites (e.g., element must exist or be focused first). An agent must infer the selection context from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
uploadC
Set files on an element (absolute local file paths).
| Name | Required | Description | Default |
|---|---|---|---|
| files | Yes | Absolute local file paths | |
| tabId | No | ||
| target | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden, and it delivers very little. It does not say whether change/input events are dispatched, what happens if the target is not a file input or is missing, whether the call waits for the input, or what a failure looks like. The only behavioral detail is that paths must be absolute, which duplicates the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single focused sentence with no padding, and the key constraint (absolute local paths) is placed inline. It is efficient, though brevity here comes at the cost of the missing behavioral and parameter detail noted above.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter mutation tool with no annotations, no output schema, and no enums, the description is too thin. It omits target semantics, event-dispatch behavior, error conditions, and required permissions or preconditions, leaving the agent to discover core invocation requirements by trial and error.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is low (33%) and the description only re-states the 'files' parameter meaning already given in the schema. The required 'target' parameter (presumably a selector or element ref) and the optional 'tabId' are left entirely undocumented in both the description and the schema, so the agent must guess how to address the element.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Set files') and the exact resource targeted ('<input type="file"> element'), which is distinct from every sibling in the browser-automation set (click, type, select_option, etc.). It is clear at a glance and does not merely restate the tool name 'upload', though it could be sharpened by naming what it is not suited for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the agent infers it should be called when a file input is present on the page. There is no explicit when-to-use, when-not-to-use, or alternative such as 'type' or 'click' for other element types, and no prerequisites like needing a resolved element reference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
waitA
Wait for time and/or page state. Provide ms, or selector (until it matches), or text (until the page contains it).
| Name | Required | Description | Default |
|---|---|---|---|
| ms | No | Wait this many milliseconds (max 30000) | |
| text | No | Wait until the page contains this text | |
| tabId | No | ||
| timeout | No | Max wait for selector/text in ms (default 10000) | |
| selector | No | Wait until this CSS selector matches |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses the three wait conditions and that selector/text waits are conditional on page state, which is useful. However, it doesn't mention default timeouts (10000ms for selector/text), the 30000ms cap on 'ms', or what happens on timeout, though some of that is in the schema parameter descriptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, under 20 words, front-loaded with the purpose and followed by the three modes. Every clause earns its place, and there is no repetition of schema details. Highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with no output schema and no annotations, the description must cover usage and behavior. It does a good job covering the three primary modes and their conditions. It falls short on timeout behavior, defaults, and the role of tabId, but given that the schema covers most parameters and this is a straightforward wait utility, it is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80%, meaning the schema already documents ms, text, timeout, and selector. The description provides the same mapping (ms→time, selector→matches, text→contains) without adding syntax, format, or behavioral details beyond the schema. The tabId parameter is undocumented in both places. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb-resource: 'Wait for time and/or page state.' It further clarifies the three modes (ms, selector, text), which distinguishes it from most siblings in this browser-automation toolbox. It doesn't explicitly name an alternative, but no sibling tool overlaps with waiting functionality, so differentiation is adequately implied.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for each mode: 'Provide ms, or selector (until it matches), or text (until the page contains it).' This tells the agent exactly which parameter to use for which scenario. It doesn't state when to avoid the tool or when another tool is preferable, but the lack of an overlapping sibling makes such guidance less critical.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
24 tool updates
v0.2.1- First observed
annotation_mode - First observed
click - First observed
close_tab - First observed
dblclick - First observed
drag - First observed
evaluate - First observed
get_console - First observed
get_html - First observed
history - First observed
hover - First observed
list_tabs - First observed
navigate - First observed
new_tab - First observed
press - First observed
query - First observed
screenshot - First observed
scroll - First observed
search - First observed
select_option - First observed
snapshot - First observed
switch_tab - First observed
type - First observed
upload - First observed
wait
TDQS
Scored across 24 tools
Tools largely have distinct purposes, but page inspection tools (snapshot, get_html, query, screenshot) overlap in their goal of revealing page content, which could occasionally cause an agent to pick a less efficient method. Interaction tools (click, dblclick, hover, drag) are well separated, and navigation/tab management tools are clearly distinct.
All names use snake_case, but the verb/noun pattern is inconsistent: some are verb_noun (list_tabs, get_html, select_option), while many are single verbs (navigate, click, type) or nouns (history, snapshot, screenshot). It remains readable, but the convention is mixed.
With 24 tools, the set feels heavy for a browser automation server. Most tools are useful, but the count lands in the borderline-heavy range (16-25) where consolidation or clearer grouping could improve usability.
The surface covers core browser automation needs: tab management, navigation, page inspection, user interactions, waits, console reading, search, and even human annotation and JavaScript evaluation. Minor gaps exist around cookies/storage, dialogs, and frame handling, but these can often be worked around with evaluate.
Maintenance
Related MCP Connectors
Live browser debugging for AI assistants — DOM, console, network via MCP.
Browser MCP for logged-in tasks. Uses your Chrome — credentials stay local. Zero-token replay.
Hyperbrowser MCP — wraps the Hyperbrowser AI-agent browsing API
- mcpOAuthcom.screenshotink
Screenshot, diff, audit and sitemap-capture any web page — 5 MCP tools for AI agents.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenancesingle-binary MCP server that gives AI agents a browser. 66 tools for navigation, form filling, data extraction, screenshots, and DOM diffing — built on pure Chrome DevTools Protocol.13MIT
- AlicenseAqualityAmaintenanceMCP server that lets AI agents drive your real Chromium browser with your existing signed-in sessions, providing visible, local, and inspectable automation for tasks like navigation, clicking, typing, and form filling.251Apache 2.0
- AlicenseNot gradedqualityAmaintenanceA zero-dependency MCP server that drives a real Chrome browser through a companion extension, enabling AI agents to automate real user sessions with trusted input events, compact accessibility-tree snapshots, and 14 tools for navigation, interaction, scripting, and inspection.280 npmMIT
- AlicenseNot gradedqualityAmaintenanceA native MCP server that gives AI assistants full Chrome control via the DevTools Protocol and an unpacked extension, including semantic element refs, real input, screenshots, and browser/tab/window management. It exposes 43 callable tools with a visible control overlay and an optional read-only security profile.2MIT