duplex
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| list_tabsA | List all open browser tabs (id, url, title, loading, active). The human sees the same tabs. |
| new_tabB | Open a new tab in the browser, optionally navigating to a URL. |
| close_tabB | Close a tab by its id. |
| switch_tabB | Bring a tab to the foreground so the human sees it. |
| navigateB | Navigate a tab (default: active tab) to a URL. |
| historyC | Go back, go forward, or reload in a tab. |
| snapshotA | Text snapshot of the page: compact DOM outline with [eN] refs for actionable/labeled elements. Use the refs with click/type. This is the primary way to "see" the page. |
| get_htmlB | Get HTML source. Without selector: cleaned body HTML. With selector: outerHTML of the first match. |
| queryB | Query elements by CSS selector; returns details for each match (tag, id, classes, text, href, rect, visibility). |
| screenshotB | Take a PNG screenshot of the current tab. Returns an image (for vision-capable models). |
| clickC | Click an element, identified by a [ref] from snapshot (e.g. "e12") or a CSS selector. |
| typeC | Type text into an input/textarea/contenteditable, by ref or CSS selector. Optionally press Enter after. |
| pressC | Press a key or combo (Enter, Escape, Tab, PageDown, Control+A, Shift+Tab, ...) in the page. |
| scrollB | Scroll the page by dx/dy pixels (positive dy = down, positive dx = right), or scroll an element into view when selector is given. |
| hoverA | Move the mouse over an element (ref or CSS selector) to trigger hover menus/tooltips; the pointer stays there. |
| dblclickB | Double-click an element by ref or CSS selector. |
| dragB | Drag from one point/element to another (refs or CSS selectors) using real mouse events. |
| select_optionA | Select an option in a native element by visible text or value. For custom dropdowns, click the trigger then click the option. |
| uploadC | Set files on an element (absolute local file paths). |
| waitA | Wait for time and/or page state. Provide ms, or selector (until it matches), or text (until the page contains it). |
| get_consoleB | Read recent console messages of a tab (errors/warnings/logs) collected since load. |
| searchB | Search the web in a tab (default engine: baidu). Use for looking up information; use navigate for known URLs. |
| annotation_modeA | Turn the human annotation mode on/off on a tab. When on, the human can draw a box/circle/arrow (or pick an element) and attach a question; the annotation is delivered to you as a structured text message. |
| evaluateB | Run JavaScript in the page and return the JSON-serializable result. Powerful; prefer snapshot/query when possible. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 24 tools
Tools largely have distinct purposes, but page inspection tools (snapshot, get_html, query, screenshot) overlap in their goal of revealing page content, which could occasionally cause an agent to pick a less efficient method. Interaction tools (click, dblclick, hover, drag) are well separated, and navigation/tab management tools are clearly distinct.
All names use snake_case, but the verb/noun pattern is inconsistent: some are verb_noun (list_tabs, get_html, select_option), while many are single verbs (navigate, click, type) or nouns (history, snapshot, screenshot). It remains readable, but the convention is mixed.
With 24 tools, the set feels heavy for a browser automation server. Most tools are useful, but the count lands in the borderline-heavy range (16-25) where consolidation or clearer grouping could improve usability.
The surface covers core browser automation needs: tab management, navigation, page inspection, user interactions, waits, console reading, search, and even human annotation and JavaScript evaluation. Minor gaps exist around cookies/storage, dialogs, and frame handling, but these can often be worked around with evaluate.