chrome-mcp
# chrome-mcp
[English](./README.md) | [简体中文](./README.zh-CN.md)
A [Model Context Protocol](https://modelcontextprotocol.io) server for browser automation, powered by [DrissionPage](https://github.com/g1879/DrissionPage).
Lets MCP clients (Claude Desktop, Claude Code, Cursor, etc.) drive a real Chromium browser through a minimal 8-tool surface — all sharing one "current tab" pointer.
## Why
- **Zero learning curve for agents.** Facade tools (`execute_js`, `run_cdp`, `navigate`, `capture`) speak only universal knowledge — JavaScript, the CDP protocol, URLs. No niche-library syntax required for the 90% cases.
- **Full power underneath.** `run_drission_code` is the complete-capability base: execute DrissionPage Python code in-process with `page` / `browser` / `context` (persistent dict) / `switch_tab()` / `tabs()` / `relaunch()` injected — loops, waits, multi-tab orchestration, anything one tool call can't express.
- **One pointer, always in sync.** Switch tabs via DrissionPage code or `Target.createTarget`/`Target.activateTarget` CDP commands — every tool follows. (Call tools serially per session; the shared pointer is not concurrency-safe.)
- **Multi-instance safe.** Per-process isolation (atomic port allocation via `socket.bind` + PID/UUID private user-data dirs) — spawn N servers, zero conflicts.
- **Self-documenting.** `get_manual` returns a compact built-in cheat sheet (DrissionPage locator syntax, error-prone CDP recipes, capture workflow, launch options) — agents fetch it on demand, zero token cost otherwise.
## Installation
```bash
uv tool install chrome-mcp
# or
uvx chrome-mcp
# or
pip install chrome-mcp
```
Requires Python ≥ 3.10 and a local Chromium-based browser.
## Usage
Add to your MCP client config (e.g. `claude_desktop_config.json`). Two equivalent forms:
```json
{
"mcpServers": {
"chrome": {
"command": "chrome-mcp"
}
}
}
```
Or run on the fly with `uvx` (no install needed):
```json
{
"mcpServers": {
"chrome": {
"command": "uvx",
"args": ["chrome-mcp"]
}
}
}
```
A typical agent flow:
```
navigate("https://httpbin.org/get") → returns tab list
execute_js("return document.body.innerText") → page data as JSON
capture(action="start", url_filter="api/")
... trigger requests ...
capture(action="get") → overview inline,
full bodies in $TMPDIR/chrome-mcp-captures/*.json
```
### Tools
| Tool | Description |
|---|---|
| `navigate` | Navigate in current tab, or open a new tab (`new_tab`) and switch the pointer to it. Returns the full tab list |
| `execute_js` | Run JavaScript in the current tab (must `return` the result). `preset: "dom_tree"` outputs a DOM structure tree without writing the template |
| `run_cdp` | Raw CDP passthrough to the current tab. `Target.createTarget` / `activateTarget` auto-switch the pointer — no `attachToTarget` needed |
| `run_drission_code` | **Full-capability base**: execute DrissionPage Python code with injected `page`/`browser`/`context`/`switch_tab()`/`tabs()`/`relaunch()`/`get_browser()` |
| `capture` | Network capture in one tool: `action` = `start` / `get` / `stop`. Overview returned inline; full data (with bodies) saved to `$TMPDIR/chrome-mcp-captures/capture_*.json` |
| `get_manual` | Return the built-in authoring manual (locator syntax, CDP recipes, capture workflow, launch options) |
| `get_browser` | **Instance management**: no args = ensure an instance is ready (starts one if absent); `cdp` = attach to an already-running browser (login state preserved; hard error if unreachable — never falls back to launching); any launch option = rebuild with new params (destructive, same semantics as `relaunch`) |
| `close_browser` | Close the browser instance |
### Launch options
Customize browser startup via command-line args:
```json
{
"mcpServers": {
"chrome": {
"command": "chrome-mcp",
"args": ["--headless", "--proxy", "http://127.0.0.1:7890", "--arg", "--lang=zh-CN"]
}
}
}
```
| Arg | Description |
|---|---|
| `--headless` | Run browser headless |
| `--proxy URL` | Proxy server |
| `--user-agent UA` | Custom User-Agent |
| `--user-data PATH` | User data dir to reuse login state. ⚠️ Conflicts if your system Chrome is using the same profile |
| `--browser-path PATH` | Path to a specific browser binary |
| `--incognito` | Incognito mode |
| `--no-imgs` | Don't load images. ⚠️ May be ignored by recent Chrome versions — reliable alternative: `run_cdp("Network.setBlockedURLs", {"urls": ["*.png", "*.jpg"]})` |
| `--arg ARG` | Pass through any Chrome flag (repeatable) |
Options can also be changed mid-session (destructive, restarts the browser) from inside `run_drission_code`:
```python
return relaunch(headless=True, user_agent="Mozilla/5.0 ...")
```
The `get_browser` tool exposes the same options (plus `cdp`) at the tool level — e.g. `get_browser(cdp="127.0.0.1:9222")` takes over your logged-in browser; `get_browser(user_data=...)` rebuilds with a given profile.
## Testing
8 scenario suites run against the real MCP stdio protocol (each spawns an independent server + headless browser — itself a multi-instance concurrency test):
```bash
uv run python tests/run_all.py # full regression
uv run python tests/run_all.py --only smoke,c
```
Covers: data extraction (DOM tree/pagination/iframes), form interaction (DP actions + CDP input sequences), network capture (filters/timing traps/body fidelity), multi-tab pointer consistency, CDP capabilities (screenshots/emulation/cookies), lifecycle (relaunch/kill-recovery), and real-site scenarios (httpbin/TLS).
## License
[GPL-3.0-only](./LICENSE)
TDQS
Scored across 7 tools
run_cdp, execute_js, and run_drission_code all provide code/protocol execution on the current tab, creating overlap that could confuse an agent. get_manual and close_browser are distinct, but the three execution tools lack clear boundary guidance.
Mostly consistent snake_case verb_noun style (run_cdp, get_manual, close_browser, run_drission_code, execute_js), with 'capture' and 'navigate' as minor single-verb exceptions.
Seven tools are well-scoped for a Chrome automation server, covering navigation, JS execution, CDP, network capture, manual lookup, and browser lifecycle without bloat.
Covers the core lifecycle: launch/close, navigate/tabs, JS/CDP execution, network capture, and a manual. Minor gaps exist around explicit tab management (though navigate and run_drission_code provide switch_tab/tabs), and no dedicated element-click/type tools, but coverage is strong for the stated purpose.