Skip to main content
Glama
obra

superpowers-chrome

by obra

Superpowers Chrome - Claude Code Plugin

Direct browser control via Chrome DevTools Protocol. Two modes available:

  1. Skill Mode - CLI tool for Claude Code agents (browsing skill)

  2. MCP Mode - Ultra-lightweight MCP server for any MCP client

Features

  • Zero dependencies - Built-in WebSocket, no npm install needed

  • Idiotproof API - Tab index syntax (0, 1, 2) instead of WebSocket URLs

  • Platform-agnostic - chrome-ws start works on macOS, Linux, Windows

  • 17 commands covering all browser automation needs

  • Complete documentation with real-world examples

Related MCP server: chrome-devtools-mcp

Installation

/plugin marketplace add obra/superpowers-marketplace
/plugin install superpowers-chrome@superpowers-marketplace

Quick Start

# Find your plugin installation path (varies by marketplace and version)
# Common locations:
#   ~/.claude/plugins/cache/superpowers-marketplace/superpowers-chrome/<version>/skills/browsing
#   ~/.claude/plugins/cache/superpowers-chrome/skills/browsing

cd ~/.claude/plugins/cache/superpowers-marketplace/superpowers-chrome/*/skills/browsing
./chrome-ws start                        # Launch Chrome
./chrome-ws new "https://example.com"   # Create tab
./chrome-ws navigate 0 "https://google.com"
./chrome-ws fill 0 "textarea[name=q]" "test"
./chrome-ws click 0 "button[name=btnK]"

Port allocation: Chrome gets a dynamically allocated port (range 9222-12111) to avoid conflicts. Port assignment is persisted per profile in ~/.cache/superpowers/browser-profiles/{name}.meta.json. Override with --port=N flag or CHROME_WS_PORT env var. Multiple profiles can run in parallel on different ports.

Parallel MCPs on one host (3.0+): the bridge auto-disambiguates the default profile. The first MCP claims superpowers-chrome:9222, the next silently falls through to superpowers-chrome-2:9223, then -3:9224, etc., each driving its own Chrome with its own profile dir. To intentionally share a Chrome between processes (e.g., a chrome-ws CLI session + a Claude MCP attaching to it), set a fixed profile via CHROME_WS_PROFILE=name (env var) or call {action: "set_profile", payload: "name"} at runtime — explicit profiles share rather than disambiguate.

Windows tip: The tooling defaults to 127.0.0.1 for DevTools traffic. Override via CHROME_WS_HOST / CHROME_WS_PORT or --port=N if you forward Chrome elsewhere.

Linux/WSL2 tip: For headed mode (visible browser), the MCP server needs the DISPLAY environment variable. If show_browser doesn't work, configure "env": {"DISPLAY": ":0"} in your MCP server config. See mcp/README.md for details. Running as root or inside a container is detected automatically and disables Chrome's sandbox; on a headless box add CHROME_EXTRA_ARGS="--headless=new --disable-gpu".

Custom Chrome flags: Set CHROME_EXTRA_ARGS to a whitespace-separated list of flags that will be appended to the Chrome command line on launch. Useful for headless containers that need software WebGL:

CHROME_EXTRA_ARGS="--use-gl=angle --use-angle=swiftshader-webgl --enable-unsafe-swiftshader"

Windows Verification (November 7, 2025)

  • node skills/browsing/chrome-ws start launched Chrome with remote debugging enabled on a fresh Windows 11 Pro install.

  • node skills/browsing/chrome-ws tabs and node skills/browsing/chrome-ws navigate 0 https://example.com confirmed CLI control with the IPv4 default binding.

  • codex exec -c "mcp_servers.superpowers-chrome.enabled=true" "List Chrome tabs via MCP to verify the Windows override patch." listed the Example Domain tab through the MCP server, demonstrating that the overrides also work through Codex.

Commands

  • Setup: start (auto-detects platform)

  • Tab management: tabs, new, close

  • Navigation: navigate, wait-for, wait-text

  • Interaction: click, fill, select

  • Extraction: eval, extract, attr, html

  • Export: screenshot, markdown

  • Raw protocol: raw (full CDP access)

Dialog Handling

Pages that open JavaScript dialogs (alert, confirm, prompt, beforeunload), WebUSB/Bluetooth/Serial/HID device choosers, HTTP basic-auth challenges, or permission prompts (camera, microphone, notifications, geolocation, clipboard) no longer wedge the connection. The dialog is surfaced as a synthetic page response and the agent interacts with it using the existing click and type actions against a small dialog::* selector grammar.

What an agent sees

While a dialog is open, any page-targeted action (extract, screenshot, eval, attr, click <real-selector>, etc.) returns a clear refusal with the dialog content and instructions:

Page is behind a dialog. Handle dialog::accept or dialog::dismiss first.

# Dialog: confirm
Tab origin: https://example.com

> Are you sure you want to leave?

Buttons:
  - dialog::accept   (OK)
  - dialog::dismiss  (Cancel)

To interact:
  click selector="dialog::accept"
  click selector="dialog::dismiss"

Browser-targeted actions (list_tabs, new_tab, close_tab, etc.) pass through unaffected.

Selector grammar

Selector

Purpose

click dialog::accept

OK / Grant / Provide credentials, depending on dialog kind

click dialog::dismiss

Cancel / Deny

type dialog::prompt <value>

Stage prompt text; commit on dialog::accept

click dialog::device[id="…"]

Pick a device in the chooser (USB, BT, Serial, HID)

type dialog::username <value> / type dialog::password <value>

Basic-auth credentials

Worked example

# 1. Page on load: alert('Saved!')
extract payload=text
# → refused with synthetic dialog markdown

# 2. Dismiss
click selector="dialog::accept"

# 3. Page is interactive again
extract payload=text
# → returns the page text

Permission prompts (getUserMedia, Notification.requestPermission, geolocation, clipboard) are caught by a document_start JS-API shim and surfaced through the same flow.

See docs/superpowers/specs/2026-05-13-dialog-handling-design.md for the full design.

MCP Server Mode

Ultra-lightweight MCP server with a single use_browser tool. Perfect for minimal context usage with automatic page captures.

Installation Options

Option 1: NPX from GitHub (Recommended)

{
  "mcpServers": {
    "chrome": {
      "command": "npx",
      "args": [
        "github:obra/superpowers-chrome"
      ]
    }
  }
}

Option 1b: NPX with Headless Mode

{
  "mcpServers": {
    "chrome": {
      "command": "npx",
      "args": [
        "github:obra/superpowers-chrome",
        "--headless"
      ]
    }
  }
}

Option 2: Git Clone + Local Path (Current)

git clone https://github.com/obra/superpowers-chrome.git
cd superpowers-chrome/mcp && npm install && npm run build
{
  "mcpServers": {
    "chrome": {
      "command": "node",
      "args": [
        "/path/to/superpowers-chrome/mcp/dist/index.js"
      ]
    }
  }
}

Auto-Capture Features

DOM-changing actions (navigate, click, type, select, eval) automatically capture:

  • Page HTML: Full rendered DOM state

  • Page Markdown: Structured content extraction

  • Screenshot: Visual page state

  • DOM Summary: Token-efficient page structure

  • Session Organization: Time-ordered captures in temp directory

Pages showing credential-shaped content are not captured; see Credential-shaped pages.

Response format:

→ https://example.com (capture #001)
Size: 1200×765
Snapshot: /tmp/chrome-session-123/001-navigate-456/
Resources: page.html, page.md, screenshot.png, console-log.txt
DOM:
  Example Domain
  Interactive: 0 buttons, 0 inputs, 1 links
  Layout: body

Credential-shaped pages

When a page shows credential-shaped content, auto-capture writes no files for that action and the response carries only metadata (URL, size, element counts, layout) plus a ⚠️ Page shows credential-shaped content; auto-capture and DOM output suppressed. line. No markdown, headings, title, or DOM diff is returned. This covers Slack tokens (xox[abposr]-, xoxe./xoxe-, xapp-), GitHub tokens (ghp_, gho_, ghu_, ghs_, ghr_, github_pat_), 1Password service-account tokens (ops_eyJ…) and Secret Keys (A3-…), otpauth:// URIs carrying a secret=, and any page containing an element with the data-sen-secret attribute (for secrets with no distinctive shape, like backup codes). The check covers the HTML, the rendered markdown, the DOM summary, the page's rendered text (innerText, which joins a token split across inline spans), open shadow roots, and live input/textarea values. It runs after the screenshot, and any match deletes every artifact already written for that action, so a token revealed while the capture runs doesn't survive in the PNG. It can't see inside closed shadow roots or cross-origin iframes.

In addition, every use_browser result and error has credential-shaped substrings replaced with [REDACTED credential-shaped] (this is how extract and eval output is handled), and screenshot refuses on such a page. Set SUPERPOWERS_CHROME_ALLOW_CREDENTIAL_CAPTURE=1 to restore the old behavior when debugging your own browser.

Usage

{
  "action": "navigate",
  "payload": "https://example.com"
}

Get help: {"action": "help"} - Returns complete documentation

See mcp/README.md for complete documentation.

When to Use

Use Skill Mode when:

  • Working with Claude Code agents

  • Need full CLI control with 17 commands

Use MCP Mode when:

  • Using Claude Desktop or other MCP clients

  • Want minimal context usage (single tool)

Use Playwright MCP when:

  • Need fresh browser instances

  • Complex automation with screenshots/PDFs

  • Prefer higher-level abstractions

Documentation

License

MIT

Available Tools

1 tool
use_browserA

Control persistent Chrome browser with automatic page capture.

Every DOM action (navigate, click, type, select, eval) auto-captures to the session dir:

  • {prefix}.png — viewport screenshot

  • {prefix}.md — page content as structured markdown

  • {prefix}.html — full rendered DOM

  • {prefix}-console.txt — browser console messages

Prefer reading these files to using 'extract' or 'screenshot' whenever possible. Pages showing credential-shaped content (tokens, 2FA seeds) are never captured, and such values are redacted from all output.

Schema: 4 parameters — action, selector (CSS/XPath or null), payload (string or object), timeout (ms). selector targets a DOM element (null/omit for navigation, eval, tab management, etc.). payload is a string for simple actions (navigate=URL, type=text, eval=JS, keyboard_press=key). payload is an object for structured actions (set_viewport={width,height}, drag_drop={target}, etc.) — a JSON-encoded string of the same object works too. Tabs are tracked as sticky state; use switch_tab to change the active tab. Use action='help' for full per-action payload shapes.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesAction to perform. action='help' lists all actions with payload shapes.
payloadNoExtra data for the action. Both a plain object and an equivalent JSON-encoded string are accepted for every structured shape below (e.g. set_viewport accepts either {width:390,height:844} or '{"width":390,"height":844}'). Literal string for code/free-text actions, taken as-is and never JSON-parsed even if it happens to look like JSON (eval=JS source, type=literal text, await_text=literal text to wait for, select=literal option value). String or object for simple cases (navigate=URL, set_profile=name, new_tab=URL, attr=attribute name or {selector,attr}). Structured object (or its JSON-string equivalent) for the rest: set_viewport={width,height,mobile?} (no bare-string form), keyboard_press=key string or {key,modifiers:{shift?,ctrl?,alt?,meta?}}, extract=format string or {format:'text'|'html'|'markdown',selector?}, screenshot=path string or {path,fullpage?,selector?}, scroll=direction string or {deltaX?,deltaY?,selector?}, drag_drop=target selector string, {x,y} target coords, or {source,target}, mouse_move={x,y,steps?,fromX?,fromY?} (no bare-string form), file_upload=path string, JSON array-of-paths string, or {selector,files:[...]}, get_console_messages={since:epochMs} or a bare epoch-ms timestamp, switch_tab=tab index/url-substring/title-substring). See action='help' for per-action payload shapes.
timeoutNoTimeout in ms for await_element / await_text actions.
selectorNoCSS or XPath selector — what to act on. Null/omitted for actions that don't target an element (navigate, eval, list_tabs, etc.). XPath must start with / or //. dialog::accept and dialog::dismiss are special selectors for handling open dialogs.
tab_indexNoLegacy: behaves like switch_tab. Sets the active tab to this index before running the action. Prefer the switch_tab action.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover safety flags (readOnly=false, openWorld=true, idempotent=false). The description goes well beyond them: it enumerates the four artifacts written per action (.png/.md/.html/-console.txt), discloses the credential redaction/never-capture policy, and notes tabs are sticky state. That is exactly the side-effect context an agent needs before invoking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose, then the capture artifacts in a scannable list before the parameter prose. Some of the payload-form text duplicates the schema and could be trimmed, but every section is readable and earns most of its space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, 39 enum actions, and destructive-capable actions like kill_chrome/clear_cookies, the description does well by pointing to action='help' and explaining capture/redaction. It stops short of warning about irreversible actions or persistence across sessions, which would fully round it out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents action, payload, selector, and timeout forms. The description largely restates those same payload shapes rather than adding meaning; the only genuinely additive note is the sticky-tab pointer to switch_tab. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Control persistent Chrome browser') and immediately names the distinguishing capability: automatic page capture on every DOM action. An agent knows exactly what class of tool this is without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit preference rule — read the auto-captured files rather than calling 'extract' or 'screenshot' — and routes to action='help' for payload shapes. It lacks guidance on when not to use the tool or on the sticky-tab workflow's tradeoffs, but the core selection heuristic is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev3.0.4
    • Changeduse_browser1 field changed
      • changedInput schema / properties / payload / description
        Previous value: -"Extra data for the action. String for simple cases (navigate=URL, type=text, eval=JS, keyboard_press=key, set_profile=name, new_tab=URL). Object for structured cases (set_viewport={width,height,mobile?}, keyboard_press={key,modifiers:{shift?,ctrl?,alt?,meta?}}, extract={format:'text'|'html'|'markdown'}, screenshot={path?,fullpage?}, scroll={deltaX?,deltaY?} or direction string, drag_drop={x,y} or selector string for target, mouse_move={x,y,steps?,fromX?,fromY?}, file_upload={files:[...]}, get_console_messages={since:epochMs}, await_text=text string or {text,timeout?}, switch_tab=tab index/url-substring/title-substring). See action='help' for per-action payload shapes."New value: +"Extra data for the action. Both a plain object and an equivalent JSON-encoded string are accepted for every structured shape below (e.g. set_viewport accepts either {width:390,height:844} or '{\"width\":390,\"height\":844}'). Literal string for code/free-text actions, taken as-is and never JSON-parsed even if it happens to look like JSON (eval=JS source, type=literal text, await_text=literal text to wait for, select=literal option value). String or object for simple cases (navigate=URL, set_profile=name, new_tab=URL, attr=attribute name or {selector,attr}). Structured object (or its JSON-string equivalent) for the rest: set_viewport={width,height,mobile?} (no bare-string form), keyboard_press=key string or {key,modifiers:{shift?,ctrl?,alt?,meta?}}, extract=format string or {format:'text'|'html'|'markdown',selector?}, screenshot=path string or {path,fullpage?,selector?}, scroll=direction string or {deltaX?,deltaY?,selector?}, drag_drop=target selector string, {x,y} target coords, or {source,target}, mouse_move={x,y,steps?,fromX?,fromY?} (no bare-string form), file_upload=path string, JSON array-of-paths string, or {selector,files:[...]}, get_console_messages={since:epochMs} or a bare epoch-ms timestamp, switch_tab=tab index/url-substring/title-substring). See action='help' for per-action payload shapes."
  2. 1 tool updatev3.0.1
    • First observeduse_browser

TDQS

A4.3/5.0

Scored across 1 tool

Disambiguation5/5

There is only one tool, so there is no risk of selecting the wrong tool. The action parameter internally routes behavior, but tool-level disambiguation is trivially perfect.

Naming Consistency5/5

The single tool name 'use_browser' follows a clear verb_noun snake_case pattern. With only one name, there is no inconsistency to evaluate.

Tool Count3/5

A single tool is borderline thin for the broad browser-automation domain, even though it multiplexes many actions via the action parameter. It avoids tool sprawl but may be under-split for discoverability.

Completeness4/5

The tool covers navigation, interaction, evaluation, tab management, viewport, drag/drop, and keyboard input, plus automatic capture and redaction. A few niche browser operations may still be missing, but core lifecycle coverage is strong.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers