OpenChrome
# OpenChrome
OpenChrome is a browser automation MCP server for controlling a real Chrome
browser from Claude Code, Codex CLI, OpenCode, or any MCP client.
It ships as a Node CLI plus MCP runtime. Desktop apps, browser extensions,
native-host installers, deployment templates, and release artifact builders are
outside this repository surface.
## Install
Requires Node.js 20 or newer.
```bash
npm install -g openchrome-mcp
openchrome setup --client codex
```
For Claude Code:
```bash
openchrome setup --client claude
```
Restart the MCP host after setup so it reloads the generated configuration.
## Run
```bash
openchrome serve --auto-launch --auto-elect --minimal
```
Manual Codex CLI configuration:
```bash
openchrome config --client codex
```
Add the printed `[mcp_servers.openchrome]` block to `~/.codex/config.toml`.
## CLI
OpenChrome can call its MCP tools directly from the shell:
```bash
oc run navigate --arg url=https://example.com
oc run read_page --arg mode=dom --json
oc navigate https://example.com
oc click ref_5
```
Playbooks run deterministic YAML scenarios:
```bash
oc playbook run scenario.yaml --vars url=https://iana.org --out report.md
```
## HTTP Mode
Run a long-lived MCP HTTP daemon when multiple clients should share one managed
Chrome owner:
```bash
openchrome serve --http 3100 --auth-token <token> --idle-timeout 30m
curl -s http://127.0.0.1:3100/health
```
Independent stdio clients should use separate `--port` and `--user-data-dir`
profiles, or connect through broker mode with `--auto-elect`.
Both stdio and HTTP serve the stateless MCP `2026-07-28` revision alongside
legacy `initialize`-based clients; see [`docs/mcp-2026-07-28.md`](docs/mcp-2026-07-28.md).
## Capabilities
- Real Chrome control through CDP.
- Navigation, clicks, typing, screenshots, DOM reads, accessibility reads, and
natural-language element lookup.
- Parallel tab/session workflows with broker-safe profile ownership.
- Compact page serialization for lower-token agent loops.
- Outcome contracts, evidence bundles, diffs, and diagnostics.
- Optional pilot-tier recovery and skill runtime behind `--pilot`.
Full tool catalogue: [`docs/agent/capability-map.md`](docs/agent/capability-map.md).
## Documentation
| Topic | Link |
| --- | --- |
| Architecture | [`docs/architecture.md`](docs/architecture.md) |
| Getting started | [`docs/getting-started.md`](docs/getting-started.md) |
| CLI | [`docs/cli.md`](docs/cli.md) |
| Playbooks | [`docs/cli/playbook.md`](docs/cli/playbook.md) |
| MCP topologies | [`docs/mcp/topologies.md`](docs/mcp/topologies.md) |
| MCP 2026-07-28 support | [`docs/mcp-2026-07-28.md`](docs/mcp-2026-07-28.md) |
| HTTP daemon | [`docs/getting-started/http-daemon.md`](docs/getting-started/http-daemon.md) |
| Security model | [`SECURITY.md`](SECURITY.md) |
| Repository structure | [`docs/dev/project-structure.md`](docs/dev/project-structure.md) |
## Development
```bash
git clone https://github.com/shaun0927/openchrome.git
cd openchrome
npm install
npm run build
npm test
```
Useful checks:
```bash
npm run lint
npm run lint:repo-structure
npm run lint:tier
npm run docs:capability-map:check
```
## License
MIT
TDQS
Scored across 122 tools
Multiple overlapping tool families create real confusion: three separate task/run systems (oc_task_*, oc_task_run_*, oc_run_*) have nearly identical start/get/list/cancel verbs, and several page-reading tools (read_page, inspect, page_content, extract_data, oc_observe) have fuzzy boundaries. Element location is split across find, query_dom, vision_find, oc_query, and oc_observe, while interact/act/computer require careful reading to distinguish. Detailed descriptions help, but the sheer number of near-duplicate surfaces will cause misselection.
The oc_ prefixed tools follow a reasonably consistent verb_noun pattern (oc_task_run_start, oc_session_snapshot), but the legacy non-prefixed tools mix conventions wildly: some are verb_noun (read_page, fill_form), others noun_verb (page_screenshot, page_reload), and several are bare nouns (computer, network, storage, cookies, memory). The confusingly similar oc_task_run_* vs oc_task_* prefixes and odd names like javascript_tool and batch_paginate further erode consistency.
122 tools is far beyond any reasonable scope for a browser automation server. There are multiple redundant subsystems (three task/run ledgers, at least five page-reading tools, three performance analyzers, two network capture tools) that inflate the surface without adding proportionate capability. This is a textbook case of tool sprawl that will overwhelm agents and add selection latency.
The browser automation domain is very thoroughly covered — navigation, reading, interaction, forms, network, performance, crawling, screenshots, and workflows all have extensive support with no obvious dead ends or missing core operations. If anything, completeness is over-achieved at the cost of coherence; the sprawl means completeness is high but at the expense of the other dimensions.