superpowers-chrome
A single use_browser tool controls a persistent Chrome browser via CDP with automatic page capture.
Navigation:
navigate,back,forward,new_tab,close_tab,list_tabs,switch_tabInteraction:
click,type,hover,double_click,right_click,drag_drop,mouse_move,scroll,keyboard_press,select,file_uploadExtraction:
extract(text/html/markdown),attr,eval,screenshotWaiting:
await_element,await_textwith configurable timeoutDialogs: handle alerts/confirms/prompts/auth/device choosers via
dialog::accept/dialog::dismissselectorsViewport:
set_viewport,get_viewport,clear_viewportConsole:
enable_console_logging,get_console_messages,clear_console_messagesBrowser lifecycle:
show_browser,hide_browser,browser_mode,kill_chrome,restart_chromeProfiles/state:
set_profile,get_profile,clear_cookiesSecret handling:
set_attrwrites onlydata-sen-secret/data-sen-nonce; token-shaped pages suppress captures and redact outputAuto-capture: DOM actions save screenshot, markdown, HTML, and console log to a session directory
Help:
action='help'returns full per-action documentation
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@superpowers-chromenavigate to example.com and extract the content"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Superpowers Chrome - Claude Code Plugin
Direct browser control via Chrome DevTools Protocol. Two modes available:
Skill Mode - CLI tool for Claude Code agents (
browsingskill)MCP Mode - Ultra-lightweight MCP server for any MCP client
Features
Zero dependencies - Built-in WebSocket, no npm install needed
Idiotproof API - Tab index syntax (
0,1,2) instead of WebSocket URLsPlatform-agnostic -
chrome-ws startworks on macOS, Linux, Windows17 commands covering all browser automation needs
Complete documentation with real-world examples
Related MCP server: chrome-devtools-mcp
Installation
/plugin marketplace add obra/superpowers-marketplace
/plugin install superpowers-chrome@superpowers-marketplaceQuick Start
# Find your plugin installation path (varies by marketplace and version)
# Common locations:
# ~/.claude/plugins/cache/superpowers-marketplace/superpowers-chrome/<version>/skills/browsing
# ~/.claude/plugins/cache/superpowers-chrome/skills/browsing
cd ~/.claude/plugins/cache/superpowers-marketplace/superpowers-chrome/*/skills/browsing
./chrome-ws start # Launch Chrome
./chrome-ws new "https://example.com" # Create tab
./chrome-ws navigate 0 "https://google.com"
./chrome-ws fill 0 "textarea[name=q]" "test"
./chrome-ws click 0 "button[name=btnK]"Port allocation: Chrome gets a dynamically allocated port (range 9222-12111) to avoid conflicts. Port assignment is persisted per profile in ~/.cache/superpowers/browser-profiles/{name}.meta.json. Override with --port=N flag or CHROME_WS_PORT env var. Multiple profiles can run in parallel on different ports.
Parallel MCPs on one host (3.0+): the bridge auto-disambiguates the default profile. The first MCP claims superpowers-chrome:9222, the next silently falls through to superpowers-chrome-2:9223, then -3:9224, etc., each driving its own Chrome with its own profile dir. To intentionally share a Chrome between processes (e.g., a chrome-ws CLI session + a Claude MCP attaching to it), set a fixed profile via CHROME_WS_PROFILE=name (env var) or call {action: "set_profile", payload: "name"} at runtime — explicit profiles share rather than disambiguate.
Windows tip: The tooling defaults to 127.0.0.1 for DevTools traffic. Override via CHROME_WS_HOST / CHROME_WS_PORT or --port=N if you forward Chrome elsewhere.
Linux/WSL2 tip: For headed mode (visible browser), the MCP server needs the DISPLAY environment variable. If show_browser doesn't work, configure "env": {"DISPLAY": ":0"} in your MCP server config. See mcp/README.md for details. Running as root or inside a container is detected automatically and disables Chrome's sandbox; on a headless box add CHROME_EXTRA_ARGS="--headless=new --disable-gpu".
Custom Chrome flags: Set CHROME_EXTRA_ARGS to a whitespace-separated list of flags that will be appended to the Chrome command line on launch. Useful for headless containers that need software WebGL:
CHROME_EXTRA_ARGS="--use-gl=angle --use-angle=swiftshader-webgl --enable-unsafe-swiftshader"Windows Verification (November 7, 2025)
node skills/browsing/chrome-ws startlaunched Chrome with remote debugging enabled on a fresh Windows 11 Pro install.node skills/browsing/chrome-ws tabsandnode skills/browsing/chrome-ws navigate 0 https://example.comconfirmed CLI control with the IPv4 default binding.codex exec -c "mcp_servers.superpowers-chrome.enabled=true" "List Chrome tabs via MCP to verify the Windows override patch."listed the Example Domain tab through the MCP server, demonstrating that the overrides also work through Codex.
Commands
Setup:
start(auto-detects platform)Tab management:
tabs,new,closeNavigation:
navigate,wait-for,wait-textInteraction:
click,fill,selectExtraction:
eval,extract,attr,htmlExport:
screenshot,markdownRaw protocol:
raw(full CDP access)
Dialog Handling
Pages that open JavaScript dialogs (alert, confirm, prompt, beforeunload), WebUSB/Bluetooth/Serial/HID device choosers, HTTP basic-auth challenges, or permission prompts (camera, microphone, notifications, geolocation, clipboard) no longer wedge the connection. The dialog is surfaced as a synthetic page response and the agent interacts with it using the existing click and type actions against a small dialog::* selector grammar.
What an agent sees
While a dialog is open, any page-targeted action (extract, screenshot, eval, attr, click <real-selector>, etc.) returns a clear refusal with the dialog content and instructions:
Page is behind a dialog. Handle dialog::accept or dialog::dismiss first.
# Dialog: confirm
Tab origin: https://example.com
> Are you sure you want to leave?
Buttons:
- dialog::accept (OK)
- dialog::dismiss (Cancel)
To interact:
click selector="dialog::accept"
click selector="dialog::dismiss"Browser-targeted actions (list_tabs, new_tab, close_tab, etc.) pass through unaffected.
Selector grammar
Selector | Purpose |
| OK / Grant / Provide credentials, depending on dialog kind |
| Cancel / Deny |
| Stage prompt text; commit on |
| Pick a device in the chooser (USB, BT, Serial, HID) |
| Basic-auth credentials |
Worked example
# 1. Page on load: alert('Saved!')
extract payload=text
# → refused with synthetic dialog markdown
# 2. Dismiss
click selector="dialog::accept"
# 3. Page is interactive again
extract payload=text
# → returns the page textPermission prompts (getUserMedia, Notification.requestPermission, geolocation, clipboard) are caught by a document_start JS-API shim and surfaced through the same flow.
See docs/superpowers/specs/2026-05-13-dialog-handling-design.md for the full design.
MCP Server Mode
Ultra-lightweight MCP server with a single use_browser tool. Perfect for minimal context usage with automatic page captures.
Installation Options
Option 1: NPX from GitHub (Recommended)
{
"mcpServers": {
"chrome": {
"command": "npx",
"args": [
"github:obra/superpowers-chrome"
]
}
}
}Option 1b: NPX with Headless Mode
{
"mcpServers": {
"chrome": {
"command": "npx",
"args": [
"github:obra/superpowers-chrome",
"--headless"
]
}
}
}Option 2: Git Clone + Local Path (Current)
git clone https://github.com/obra/superpowers-chrome.git
cd superpowers-chrome/mcp && npm install && npm run build{
"mcpServers": {
"chrome": {
"command": "node",
"args": [
"/path/to/superpowers-chrome/mcp/dist/index.js"
]
}
}
}Auto-Capture Features
DOM-changing actions (navigate, click, type, select, eval) automatically capture:
Page HTML: Full rendered DOM state
Page Markdown: Structured content extraction
Screenshot: Visual page state
DOM Summary: Token-efficient page structure
Session Organization: Time-ordered captures in temp directory
Pages showing credential-shaped content are not captured; see Credential-shaped pages.
Response format:
→ https://example.com (capture #001)
Size: 1200×765
Snapshot: /tmp/chrome-session-123/001-navigate-456/
Resources: page.html, page.md, screenshot.png, console-log.txt
DOM:
Example Domain
Interactive: 0 buttons, 0 inputs, 1 links
Layout: bodyCredential-shaped pages
When a page shows credential-shaped content, auto-capture writes no files for that action and the response carries only metadata (URL, size, element counts, layout) plus a ⚠️ Page shows credential-shaped content; auto-capture and DOM output suppressed. line. No markdown, headings, title, or DOM diff is returned. This covers Slack tokens (xox[abposr]-, xoxe./xoxe-, xapp-), GitHub tokens (ghp_, gho_, ghu_, ghs_, ghr_, github_pat_), 1Password service-account tokens (ops_eyJ…) and Secret Keys (A3-…), otpauth:// URIs carrying a secret=, and any page containing an element with the data-sen-secret attribute (for secrets with no distinctive shape, like backup codes; see below for what marking does not cover). The check covers the HTML, the rendered markdown, the DOM summary, the page's rendered text (innerText, which joins a token split across inline spans), open shadow roots, and live input/textarea values. It runs after the screenshot, and any match deletes every artifact already written for that action, so a token revealed while the capture runs doesn't survive in the PNG. It can't see inside closed shadow roots or cross-origin iframes.
In addition, every use_browser result and error has credential-shaped substrings replaced with [REDACTED credential-shaped], and screenshot refuses on such a page.
The data-sen-secret marker: an accident guard, not a boundary
Mark an element with the data-sen-secret attribute (via set_attr) when it holds a secret with no distinctive shape — a bare base32 TOTP seed, a backup code — that the credential-shape check above can't recognize. When to mark: as soon as the action that revealed the secret returns, before any other action on that page, then capture the value with the credential broker. Marking only affects actions after it: the capture files the revealing action already wrote (typically its NNN-navigate.html/.md or NNN-click.html/.md) still hold the value, and marking does not delete them. Once marked:
evalrefuses outright, page-wide, if ANY element on the page currently carriesdata-sen-secret— checked against the live DOM immediately before the expression runs. This is a point check at the moment of the call: it does not remember that the page was ever marked, and it does not re-check after the expression finishes.extractandattrrefuse when the resolved element or any ancestor of it (through shadow-root hosts) is marked; where a read doesn't require touching the marked value at all (extract's HTML/markdown forms), they instead read off a clone imported into a fresh, inert document with marked descendants stripped, rather than refusing the whole page over one marked corner of it. Screenshot and auto-capture use the same live scan, extended to open shadow roots and same-originiframe/frame/object/embed, and write nothing to disk for a marked page from then on.This is not a security boundary.
evalruns arbitrary JavaScript in the exact same JS realm as the marked element — it can read the value directly,fetch()it to another origin, stash it inwindow.nameor storage for another tab, or erase the marker withremoveAttributeas its own last step, and no in-page check can close any of that. Same-originfetch, another browser tab, and page JavaScript that moves the value out from under the marked element can all still reach it regardless of marking. Don't mark an element and then runevalon that page expecting the secret to stay contained — after marking, use the credential broker for the value andset_attrfor writes.Native dialog (
alert/confirm/prompt/beforeunload) text keeps only the credential-shape redaction and file suppression described above; a marked-but-shapeless value in dialog text is not specially caught, for the same reasonevalisn't a boundary.Not covered either: capture files written before the mark (see "When to mark" above), and markup inside an
<iframe srcdoc="…">attribute.extract's HTML form andattr srcdocreturn that attribute text verbatim, so a marked element written intosrcdoccomes back with it, even though the live scan sees the loaded frame andevalrefuses.
set_attr is a separate, write-only action exempt from all of the above: it takes no caller JavaScript (values travel as CDP call arguments, never as code text), and can write exactly two attribute names — data-sen-nonce (a broker nonce, deliberately not any other data-*/aria-* name, since page JS and frameworks routinely wire arbitrary ones to behavior) and data-sen-secret itself, the write path for the marker. Writing data-sen-nonce resolves to the single first-VISIBLE match, the same way extract/click/type resolve a selector, and refuses if that element is already marked. Writing data-sen-secret instead marks every element the selector matches, hidden duplicates included, and can never remove or weaken an existing mark — re-marking an already-marked element is a no-op, not a refusal. set_attr's different outcomes (ok / no element matched / refused: target element is marked) are themselves a limited prefix oracle over what's on the page for a caller who can vary the selector and watch which one comes back; this is a known, accepted limitation, not something the guard closes.
Set SUPERPOWERS_CHROME_ALLOW_CREDENTIAL_CAPTURE=1 to restore the old, unguarded behavior when debugging your own browser.
Separately, whatever auto-capture does write to disk (.html, .md, and -diff.txt) is redacted by both attribute and value. A plain password or 6-digit code has no shape the credential-shape check above recognizes, so it wouldn't otherwise be caught when a page's own JavaScript mirrors it into an attribute, a hidden input, or visible text.
Which fields count:
input[type="password"](case-insensitive, and remembered on that same element after a "show password" toggle switches its type totext), any field whoseautocompletecontainscurrent-password,new-password,one-time-code,cc-number, orcc-csc(case-insensitive substring match, soautocomplete="section-2fa one-time-code"still matches), anydata-sen-secret-marked element, and anyinput/textareawhose live.valueexactly matches one of its own OTHER attributes (a self-mirror, found by comparing values, not by a fixed attribute-name list — this is how a field with no recognized type or autocomplete, like Google's 2-step verificationtotpPinfield mirroring the typed code intodata-initial-value, still gets caught). Self-mirror detection skips inputs nobody types into (submit,button,reset,checkbox,radio,hidden,image) and does not compare against attributes that name or label a field (type,name,id,aria-label,title,placeholder,for,class), so a submit button whose value matches itsnameoraria-label, or a radio whose value matches itsid, is not treated as a secret. The rest of thecc-*family (cc-exp-month,cc-exp-year,cc-name,cc-type, ...) is not included.Attribute stripping: an inert clone of the page (
document.implementation.createHTMLDocument+importNode, nevercloneNodeon the live document, so no resource re-fetch and no handler re-fires) hasvalueand everydata-*/aria-*attribute removed from fields matched by type/autocomplete/marker. This applies at any value length. A field found only by self-mirror has no selector to re-match on the clone, so its attribute NAME is not stripped this way — only its value, by the redaction pass below (floor applies).Value redaction: the live
.valueof each of those fields (what was typed, not just thevalueattribute), plus its HTML-entity-escaped forms (an attribute value escapes& " < >and U+00A0 as ; text content escapes& < >and U+00A0), is replaced with[REDACTED]wherever it appears in the.html/.md/-diff.txtoutput — any attribute, on any element, not just the field's ownvalue. This catches a mirror onto a different element, a different attribute (title,aria-*,data-*), or echoed text. Only values of at least 4 characters are redacted this way, or at least 3 forcc-csc(most card security codes are three digits). Dialog captures (alert/confirm/prompttext written to.md/.html) are redacted with the values from the most recent capture in this MCP session (which may have been on another tab), since a page can't be read while a dialog is open; a value typed in the same action that opened the dialog was never seen by a capture and is not redacted there.Side effect: value redaction is a plain string replace over the serialized capture, not a structural edit. A secret that also occurs as ordinary page text or markup (a password of
password, a 4-digit code matching a year) replaces those occurrences too, for exampletype="[REDACTED]". That degrades the artifact but leaks nothing. The same happens when self-mirror detection flags a non-secret field: a text field, search box, or textarea whose value equals one of its own attributes outside the excluded list above (a page that syncs a search query intodata-query, or a prefilled field with a matchingdata-defaultoraria-description) has that value replaced across the capture if it is at least 4 characters.Known gaps: these still write the secret to disk. (1) Split one-digit OTP boxes that a page joins into an aggregate hidden input: each digit is below the length floor, so the aggregate is written in clear. (2) A field cleared on submit whose mirror remains: the field no longer has a value to redact, so the mirror is written in clear. (3) A "show password" toggle that swaps in a new element instead of changing
type: the new element was never a password input, so its value is not redacted. (4) A self-mirrored value shorter than the length floor: the field is still found (self-mirror detection has no floor of its own), but its value is too short to substring-redact, so the mirror is written in clear. Screenshots (.png) are pixels, not scrubbed text, so a visible typed value (e.g. atype="text"one-time code) can show up in a screenshot.
The live page is never touched by any of this: the fields the browser actually uses, and the page's own image/resource loads, are unaffected.
Usage
{
"action": "navigate",
"payload": "https://example.com"
}Get help: {"action": "help"} - Returns complete documentation
See mcp/README.md for complete documentation.
When to Use
Use Skill Mode when:
Working with Claude Code agents
Need full CLI control with 17 commands
Use MCP Mode when:
Using Claude Desktop or other MCP clients
Want minimal context usage (single tool)
Use Playwright MCP when:
Need fresh browser instances
Complex automation with screenshots/PDFs
Prefer higher-level abstractions
Documentation
SKILL.md - Complete skill guide
EXAMPLES.md - Real-world examples
chrome-ws README - Tool documentation
License
MIT
Available Tools
1 tooluse_browserA
Control persistent Chrome browser with automatic page capture.
Every DOM action (navigate, click, type, select, eval) auto-captures to the session dir:
{prefix}.png — viewport screenshot
{prefix}.md — page content as structured markdown
{prefix}.html — full rendered DOM
{prefix}-console.txt — browser console messages
Prefer reading these files to using 'extract' or 'screenshot' whenever possible. Actions on a page showing token-shaped content (Slack/GitHub/1Password tokens, otpauth:// URIs) or any data-sen-secret element write no capture files, token-shaped values are redacted from all output, and eval refuses outright while any data-sen-secret element is on the page — an accident guard against a cooperating agent's own reads, not a security boundary. A secret with no distinctive shape (bare TOTP seed, backup code) is captured like any other text until it is marked: mark it with 'set_attr' name 'data-sen-secret' as soon as the action that revealed it returns, before any other action on that page. Capture files written before the mark keep the value and are not deleted, so don't read them back. 'set_attr' is a write-only setter for exactly 'data-sen-nonce' or 'data-sen-secret' that is exempt from that eval restriction (see 'help' for details) — use it, not eval, to write onto the page while a secret is present.
Schema: 4 parameters — action, selector (CSS/XPath or null), payload (string or object), timeout (ms). selector targets a DOM element (null/omit for navigation, eval, tab management, etc.). payload is a string for simple actions (navigate=URL, type=text, eval=JS, keyboard_press=key). payload is an object for structured actions (set_viewport={width,height}, drag_drop={target}, etc.) — a JSON-encoded string of the same object works too. Tabs are tracked as sticky state; use switch_tab to change the active tab. Use action='help' for full per-action payload shapes.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | Action to perform. action='help' lists all actions with payload shapes. | |
| payload | No | Extra data for the action. Both a plain object and an equivalent JSON-encoded string are accepted for every structured shape below (e.g. set_viewport accepts either {width:390,height:844} or '{"width":390,"height":844}'). Literal string for code/free-text actions, taken as-is and never JSON-parsed even if it happens to look like JSON (eval=JS source, type=literal text, await_text=literal text to wait for, select=literal option value). String or object for simple cases (navigate=URL, set_profile=name, new_tab=URL, attr=attribute name or {selector,attr}). Structured object (or its JSON-string equivalent) for the rest: set_viewport={width,height,mobile?} (no bare-string form), keyboard_press=key string or {key,modifiers:{shift?,ctrl?,alt?,meta?}}, extract=format string or {format:'text'|'html'|'markdown',selector?}, screenshot=path string or {path,fullpage?,selector?}, scroll=direction string or {deltaX?,deltaY?,selector?}, drag_drop=target selector string, {x,y} target coords, or {source,target}, mouse_move={x,y,steps?,fromX?,fromY?} (no bare-string form), file_upload=path string, JSON array-of-paths string, or {selector,files:[...]}, get_console_messages={since:epochMs} or a bare epoch-ms timestamp, switch_tab=tab index/url-substring/title-substring). See action='help' for per-action payload shapes. | |
| timeout | No | Timeout in ms for await_element / await_text actions. | |
| selector | No | CSS or XPath selector — what to act on. Null/omitted for actions that don't target an element (navigate, eval, list_tabs, etc.). XPath must start with / or //. dialog::accept and dialog::dismiss are special selectors for handling open dialogs. | |
| tab_index | No | Legacy: behaves like switch_tab. Sets the active tab to this index before running the action. Prefer the switch_tab action. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover only the generic hints (readOnly=false, openWorld=true, idempotent=false, destructive=false); the description adds far more: every DOM action auto-writes .png/.md/.html/-console.txt artifacts, secret-shaped content suppresses captures and redacts output, eval refuses outright when a data-sen-secret element is present, set_attr is a write-only setter exempt from that restriction, and pre-mark capture files are retained and must not be re-read. This is exactly the behavioral context annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded (purpose → capture artifacts → secret guardrails → parameters) and every block carries information an agent needs. The secret-handling paragraph is dense and reads as run-on prose, which costs some clarity, but little of the text is pure filler given the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 38-action tool with five parameters and no output schema, the description covers the safety profile, capture side effects, the sticky-tab model, and defers per-action payload shapes to action='help' — a legitimate escape hatch rather than an omission. The only residual gap is that the destructive-looking actions (kill_chrome, clear_cookies, restart_chrome) are not called out against the destructive=false hint.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes beyond it: it explains selector nullability by action class, distinguishes literal-string payloads (never JSON-parsed) from structured object payloads, gives concrete shape examples (set_viewport={width,height}, drag_drop={target}, keyboard_press={key,modifiers}), and clarifies that a JSON-encoded string is accepted for every structured shape. That is real added meaning over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource: 'Control persistent Chrome browser with automatic page capture.' The automatic-capture behavior is a distinguishing characteristic stated up front, and the enumerated action set (navigate, click, type, eval) confirms the scope. An agent immediately knows what class of tool this is.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing advice — 'Prefer reading these files to using extract or screenshot whenever possible' — and names the alternative actions to avoid. It also directs the agent to action='help' for payload shapes and instructs 'use it [set_attr], not eval' under a specific condition. What's missing is higher-level when/when-not framing across the many actions, but the actionable preferences are present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v3.0.8- Changed
use_browser1 field changed- changed
Input schema / properties / action / enumPrevious value: -[ - "navigate", - "back", - "forward", - "click", - "type", - "extract", - "screenshot", - "eval", - "select", - "attr", - "await_element", - "await_text", - "new_tab", - "close_tab", - "list_tabs", - "switch_tab", - "show_browser", - "hide_browser", - "browser_mode", - "set_profile", - "get_profile", - "help", - "hover", - "drag_drop", - "mouse_move", - "scroll", - "double_click", - "right_click", - "file_upload", - "keyboard_press", - "set_viewport", - "clear_viewport", - "get_viewport", - "clear_cookies", - "enable_console_logging", - "get_console_messages", - "clear_console_messages", - "kill_chrome", - "restart_chrome" -]New value: +[ + "navigate", + "back", + "forward", + "click", + "type", + "extract", + "screenshot", + "eval", + "select", + "attr", + "set_attr", + "await_element", + "await_text", + "new_tab", + "close_tab", + "list_tabs", + "switch_tab", + "show_browser", + "hide_browser", + "browser_mode", + "set_profile", + "get_profile", + "help", + "hover", + "drag_drop", + "mouse_move", + "scroll", + "double_click", + "right_click", + "file_upload", + "keyboard_press", + "set_viewport", + "clear_viewport", + "get_viewport", + "clear_cookies", + "enable_console_logging", + "get_console_messages", + "clear_console_messages", + "kill_chrome", + "restart_chrome" +]
1 tool update
v3.0.4- Changed
use_browser1 field changed- changed
Input schema / properties / payload / descriptionPrevious value: -"Extra data for the action. String for simple cases (navigate=URL, type=text, eval=JS, keyboard_press=key, set_profile=name, new_tab=URL). Object for structured cases (set_viewport={width,height,mobile?}, keyboard_press={key,modifiers:{shift?,ctrl?,alt?,meta?}}, extract={format:'text'|'html'|'markdown'}, screenshot={path?,fullpage?}, scroll={deltaX?,deltaY?} or direction string, drag_drop={x,y} or selector string for target, mouse_move={x,y,steps?,fromX?,fromY?}, file_upload={files:[...]}, get_console_messages={since:epochMs}, await_text=text string or {text,timeout?}, switch_tab=tab index/url-substring/title-substring). See action='help' for per-action payload shapes."New value: +"Extra data for the action. Both a plain object and an equivalent JSON-encoded string are accepted for every structured shape below (e.g. set_viewport accepts either {width:390,height:844} or '{\"width\":390,\"height\":844}'). Literal string for code/free-text actions, taken as-is and never JSON-parsed even if it happens to look like JSON (eval=JS source, type=literal text, await_text=literal text to wait for, select=literal option value). String or object for simple cases (navigate=URL, set_profile=name, new_tab=URL, attr=attribute name or {selector,attr}). Structured object (or its JSON-string equivalent) for the rest: set_viewport={width,height,mobile?} (no bare-string form), keyboard_press=key string or {key,modifiers:{shift?,ctrl?,alt?,meta?}}, extract=format string or {format:'text'|'html'|'markdown',selector?}, screenshot=path string or {path,fullpage?,selector?}, scroll=direction string or {deltaX?,deltaY?,selector?}, drag_drop=target selector string, {x,y} target coords, or {source,target}, mouse_move={x,y,steps?,fromX?,fromY?} (no bare-string form), file_upload=path string, JSON array-of-paths string, or {selector,files:[...]}, get_console_messages={since:epochMs} or a bare epoch-ms timestamp, switch_tab=tab index/url-substring/title-substring). See action='help' for per-action payload shapes."
1 tool update
v3.0.1- First observed
use_browser
TDQS
Scored across 1 tool
There is only a single tool, so there is no possibility of confusing it with another tool on this server. All routing between operations happens through the documented 'action' parameter rather than through tool selection, which removes cross-tool misselection entirely.
The lone name 'use_browser' follows a clean verb_noun convention. With only one tool there is no pattern to break, so consistency is trivially satisfied, though the naming dimension is largely uninformative here.
A single tool for the whole browser-control domain is on the thin side of appropriate, pushing a large multi-action surface (navigate, click, type, select, eval, tabs, captures) into one dispatcher. It works as a pattern but makes discovery and schema validation harder than a small set of focused tools.
The action set appears to cover the core browser lifecycle: navigation, DOM interaction, eval, tab management, and automatic capture artifacts, plus a help action for discovering payload shapes. Minor gaps (explicit wait-for-condition, network interception, download/upload handling) are largely workable via eval.
Maintenance
Related MCP Connectors
Hosted real Google Chrome MCP with per-user persistent state. Navigate, click, type, screenshot.
Browserless MCP — wraps the Browserless headless-Chromium REST API (browserless.io)
Browser MCP for logged-in tasks. Uses your Chrome — credentials stay local. Zero-token replay.
Live browser debugging for AI assistants — DOM, console, network via MCP.
Related MCP Servers
- AlicenseAqualityAmaintenanceControls a running Chrome/Chromium browser via the Chrome DevTools Protocol, enabling navigation, JavaScript evaluation, tab management, and raw CDP commands through MCP tools.5MIT
- AlicenseNot gradedqualityBmaintenanceEnables coding agents to control and inspect a live Chrome browser via MCP, providing Chrome DevTools capabilities for automation, debugging, and performance analysis.1,731,154 npmApache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables AI coding agents to control and inspect a live Chrome browser via MCP, including performance tracing, network and console debugging, screenshots, and reliable Puppeteer-based automation.1,731,154 npmApache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables MCP clients to control Chrome/Edge browsers via CDP, including page navigation, element interaction, snapshots, network debugging, script execution, and performance tracing.AGPL 3.0