chrome-mcp
You can automate a real Chromium browser via MCP, with tools for navigation, JavaScript execution, CDP commands, DrissionPage Python scripting, network capture, and documentation.
Navigate to URLs in the current tab or open/switch to a new tab (with timeout and background options).
Execute JavaScript in the current tab, including a
dom_treepreset for DOM structure output.Send raw Chrome DevTools Protocol (CDP) commands to the current tab;
Target.createTarget/activateTargetauto-switch the active tab.Run arbitrary DrissionPage Python code in-process with injected
page,browser,context,switch_tab(),tabs(),relaunch(), enabling loops, waits, and multi-tab orchestration.Capture network traffic:
start/get/stop, with HTTP method and URL keyword filters; overview inline, full bodies saved to JSON.Retrieve a built-in manual covering locator syntax, CDP recipes, capture workflow, and launch options.
Close the browser instance.
(Per README) Manage browser instances with
get_browser(ensure ready, attach via CDP, relaunch with options) and customize launch via CLI args like--headless,--proxy,--user-data.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@chrome-mcpOpen https://example.com and return the page title"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
chrome-mcp
A Model Context Protocol server for browser automation, powered by DrissionPage.
Lets MCP clients (Claude Desktop, Claude Code, Cursor, etc.) drive a real Chromium browser through a minimal 8-tool surface — all sharing one "current tab" pointer.
Why
Zero learning curve for agents. Facade tools (
execute_js,run_cdp,navigate,capture) speak only universal knowledge — JavaScript, the CDP protocol, URLs. No niche-library syntax required for the 90% cases.Full power underneath.
run_drission_codeis the complete-capability base: execute DrissionPage Python code in-process withpage/browser/context(persistent dict) /switch_tab()/tabs()/relaunch()injected — loops, waits, multi-tab orchestration, anything one tool call can't express.One pointer, always in sync. Switch tabs via DrissionPage code or
Target.createTarget/Target.activateTargetCDP commands — every tool follows. (Call tools serially per session; the shared pointer is not concurrency-safe.)Multi-instance safe. Per-process isolation (atomic port allocation via
socket.bind+ PID/UUID private user-data dirs) — spawn N servers, zero conflicts.Self-documenting.
get_manualreturns a compact built-in cheat sheet (DrissionPage locator syntax, error-prone CDP recipes, capture workflow, launch options) — agents fetch it on demand, zero token cost otherwise.
Related MCP server: Real Browser MCP
Installation
uv tool install chrome-mcp
# or
uvx chrome-mcp
# or
pip install chrome-mcpRequires Python ≥ 3.10 and a local Chromium-based browser.
Usage
Add to your MCP client config (e.g. claude_desktop_config.json). Two equivalent forms:
{
"mcpServers": {
"chrome": {
"command": "chrome-mcp"
}
}
}Or run on the fly with uvx (no install needed):
{
"mcpServers": {
"chrome": {
"command": "uvx",
"args": ["chrome-mcp"]
}
}
}A typical agent flow:
navigate("https://httpbin.org/get") → returns tab list
execute_js("return document.body.innerText") → page data as JSON
capture(action="start", url_filter="api/")
... trigger requests ...
capture(action="get") → overview inline,
full bodies in $TMPDIR/chrome-mcp-captures/*.jsonTools
Tool | Description |
| Navigate in current tab, or open a new tab ( |
| Run JavaScript in the current tab (must |
| Raw CDP passthrough to the current tab. |
| Full-capability base: execute DrissionPage Python code with injected |
| Network capture in one tool: |
| Return the built-in authoring manual (locator syntax, CDP recipes, capture workflow, launch options) |
| Instance management: no args = ensure an instance is ready (starts one if absent); |
| Close the browser instance |
Launch options
Customize browser startup via command-line args:
{
"mcpServers": {
"chrome": {
"command": "chrome-mcp",
"args": ["--headless", "--proxy", "http://127.0.0.1:7890", "--arg", "--lang=zh-CN"]
}
}
}Arg | Description |
| Run browser headless |
| Proxy server |
| Custom User-Agent |
| User data dir to reuse login state. ⚠️ Conflicts if your system Chrome is using the same profile |
| Path to a specific browser binary |
| Incognito mode |
| Don't load images. ⚠️ May be ignored by recent Chrome versions — reliable alternative: |
| Pass through any Chrome flag (repeatable) |
Options can also be changed mid-session (destructive, restarts the browser) from inside run_drission_code:
return relaunch(headless=True, user_agent="Mozilla/5.0 ...")The get_browser tool exposes the same options (plus cdp) at the tool level — e.g. get_browser(cdp="127.0.0.1:9222") takes over your logged-in browser; get_browser(user_data=...) rebuilds with a given profile.
Testing
8 scenario suites run against the real MCP stdio protocol (each spawns an independent server + headless browser — itself a multi-instance concurrency test):
uv run python tests/run_all.py # full regression
uv run python tests/run_all.py --only smoke,cCovers: data extraction (DOM tree/pagination/iframes), form interaction (DP actions + CDP input sequences), network capture (filters/timing traps/body fidelity), multi-tab pointer consistency, CDP capabilities (screenshots/emulation/cookies), lifecycle (relaunch/kill-recovery), and real-site scenarios (httpbin/TLS).
License
Available Tools
7 toolscaptureA
网络抓包三合一。action='start':在当前 tab 启动监听(⚠️ 必须在触发请求之前调用;监听挂载在该 tab,不随指针迁移);action='get':拉取结果(概览直接返回,完整数据落盘 JSON 供 read);action='stop':停止并清空。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | get:最多返回条数(0=全部) | |
| action | Yes | start=启动监听 / get=拉取结果 / stop=停止 | |
| method | No | start:HTTP 方法过滤,多个用逗号分隔(如 'GET,PUT')。不传时 DrissionPage 缺省只抓 GET/POST | |
| url_filter | No | start:URL 关键词过滤(空=全部) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses important behavioral traits: the start-before-request timing constraint, tab-bound monitoring that does not migrate, that 'get' returns an inline overview while full data lands in a JSON file for 'read', and that 'stop' clears the capture. These are meaningful operational details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with a summary ('三合一') followed by per-action groupings, each earning its place. No redundant sentences beyond the heading. Dense but efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-mode capture tool with no output schema and no annotations, the description covers the critical timing, scoping, and result-handling details an agent needs. It could mention permissions or how the JSON file path is obtained, but the core operational picture is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters including the enum and method filter defaults. The description adds only a brief parenthetical about the overview/JSON-file split for 'get' at the top level, which is marginal beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb+resource: network capture with three modes (start/get/stop) broken out explicitly. It is distinguishable from siblings like run_cdp or execute_js, though the exact relationship to browser-attached siblings is not explained. Purpose is clear without being fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear sequencing guidance: 'start' must be called before triggering the request, and monitoring is bound to the current tab rather than following the pointer. It does not name alternatives or when-not-to-use, but the timing and scoping instructions are actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_browserC
关闭浏览器实例
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the action (close) but doesn't describe what happens—e.g., whether it gracefully shuts down the browser, kills processes, affects other tools, or has side effects like losing unsaved data. This is a significant gap for a potentially destructive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence ('关闭浏览器实例') that directly states the purpose without waste. It's appropriately sized for a simple tool, though it could be slightly more informative without losing conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's potential complexity (closing a browser could involve cleanup, state changes, or errors) and lack of annotations and output schema, the description is incomplete. It doesn't address behavioral aspects, return values, or error conditions, leaving gaps for an AI agent to understand its full context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, and schema description coverage is 100%, so no parameter documentation is needed. The description doesn't add parameter semantics, but that's acceptable here. Baseline is 4 for zero parameters, as there's nothing to compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description '关闭浏览器实例' (close browser instance) states a clear action (close) on a resource (browser instance), but it's vague about scope—it doesn't specify whether this closes all browser windows, a specific instance, or just the current tab. It distinguishes from siblings like 'navigate_to_page' or 'get_browser_status' by being a termination action, but lacks precision.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. For example, it doesn't mention if this should be used after completing tasks, to free resources, or as part of cleanup, nor does it reference sibling tools like 'launch_chrome_manually' for context. The description alone offers no usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
execute_jsA
在当前 tab 执行 JavaScript 代码。⚠️ 必须使用 return 语句返回结果。preset='dom_tree' 时 code 改为传 JSON 参数对象 {"max_depth":5,"include_text":true,"max_text_length":50,"selector":null},直接输出 DOM 结构树。
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | 要执行的 JavaScript 代码(必须用 return 返回结果!),或 preset='dom_tree' 时的 JSON 参数 | |
| preset | No | 预置脚本模板:dom_tree=DOM 结构树(免写模板 JS) | |
| timeout | No | 本次 JS 执行超时(秒,可选)。默认 150 秒。 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the critical behavioral rule that the code must use return to produce a result, plus the default 150s timeout. It does not cover error behavior, sandbox/permission constraints, or what happens when the script throws, leaving meaningful gaps for a code-execution tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact and front-loaded: the core action comes first, then the return-value warning, then the preset override. No filler sentences, though the preset paragraph is dense and could be slightly tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description must explain the return contract, which it partially does via the return-statement warning. However, it is silent on error handling, result serialization, and side effects, which matter for an arbitrary JS execution tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and each parameter is already documented, so baseline would be 3. The description adds value beyond the schema by showing the exact JSON argument shape for preset='dom_tree' with concrete keys (max_depth, include_text, max_text_length, selector), clarifying how the overloaded 'code' parameter behaves.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: execute JavaScript code in the current tab, and clarifies a second mode (preset='dom_tree') that changes what 'code' means. It does not name or contrast with siblings like run_cdp or run_drission_code, so an agent must infer which execution tool to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives the condition for the dom_tree preset (pass JSON params instead of code), which is useful usage guidance, but it never says when to prefer this tool over run_cdp or run_drission_code, nor any exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_manualB
返回内置编写手册:定位语法、元素动作、等待、tab 管理、抓包、CDP 高频配方、启动定制等速查。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose the essential behavior — a zero-parameter read that returns reference content, therefore inherently non-destructive — but says nothing about return format, length, or whether the content is static or context-dependent. That omission is notable for a tool whose entire value is the payload.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that leads with what is returned before the topic enumeration. Every clause earns its place by telling the agent what content to expect. Slightly list-heavy, but no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema retrieval tool, the description supplies the one thing the agent needs: what knowledge it will obtain. The lack of any hint about return size or format is the main residual gap, but the topic list makes the tool's purpose sufficiently clear to call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the 4 baseline applies. The description makes clear it is invoked without arguments, matching the empty object schema with 100% coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: it returns a built-in authoring manual, and enumerates the topics covered (locator syntax, element actions, waits, tab management, capture, CDP recipes, launch customization). This clearly distinguishes it from siblings like run_cdp or run_drission_code, which execute rather than document. It stops short of a 5 only because the manual's scope beyond the listed topics is undefined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no reference to the sibling tools whose DSL it presumably documents. The agent must infer on its own that this is a reference to consult before calling run_drission_code or run_cdp. Nothing states when-not to call it or whether it is cheap to call repeatedly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_cdpA
对当前 tab 透传 Chrome DevTools Protocol 命令。method 如 'Page.captureScreenshot',params 为参数对象。命令始终作用于当前 tab,无需 attachToTarget(Target.createTarget/activateTarget 会自动切换操作 tab)。详见 get_manual。
| Name | Required | Description | Default |
|---|---|---|---|
| method | Yes | CDP 方法名,如 'Page.navigate'、'Runtime.evaluate' | |
| params | No | CDP 命令参数对象(可选) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose a non-obvious side effect: Target.createTarget/activateTarget silently change which tab subsequent commands act on. That is exactly the kind of behavioral trait annotations would otherwise cover. It still omits return shape, error behavior, and any auth/permission constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the core action and followed by the tab-scoping caveat and a pointer to get_manual. No filler, though the parenthetical about tab switching is dense and could be split for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a raw passthrough tool with no output schema and no annotations, the description covers what it does, its scoping constraint, the auto tab-switch side effect, and where to get deeper detail (get_manual). Return values and failure modes are undefined but are inherently method-dependent, so this is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description's example ('Page.captureScreenshot') and its note that params is an optional argument object largely restate what the schema already documents for method and params, adding little new semantic content.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: passing through Chrome DevTools Protocol commands to the current tab, with a concrete example method. It implicitly distinguishes itself from siblings like execute_js, capture, and run_drission_code by being the raw-CDP escape hatch, though it never names an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives useful operating context (commands always act on the current tab, no attachToTarget needed, createTarget/activateTarget auto-switch the operating tab) and defers to get_manual. However, it never says when to prefer this over execute_js, capture, or run_drission_code, which is the main selection decision an agent faces.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_drission_codeB
执行 DrissionPage Python 代码(server 进程内,完整能力底座)。预置变量:page=当前tab、browser=浏览器对象、context=跨调用持久字典、DOM_TREE_JS=DOM树JS模板。预置函数:switch_tab(tab_id)、tabs()、relaunch(**opts)。⚠️ 必须用 return 返回结果。
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | 要执行的 Python 代码(必须用 return 返回结果!) | |
| timeout | No | 整体兜底超时(秒,可选)。默认 120 秒。 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose meaningful runtime behavior: it runs server-side, exposes preset page/browser/context variables, and must return results. However, for arbitrary code execution it omits any warning about side effects, permissions, or failure behavior, which is a notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense block, front-loaded with the purpose before environment details, with no filler sentences. The ⚠️ return requirement is correctly emphasized, though the run-on layout slightly reduces scanability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a code-execution tool with no output schema and no annotations, the description supplies the essential execution context: environment, preset variables, preset functions, and the return convention. It is largely complete, missing only error/edge-case behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (code, timeout) are already documented, establishing the baseline of 3. The preset variables and functions listed in the description enrich the code-writing context but do not clarify the parameters themselves beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('执行 DrissionPage Python 代码') plus scope ('server 进程内'), which clearly separates it from the JS-oriented execute_js sibling. It never explicitly names an alternative, so it lands at clear-but-undifferentiated rather than a full 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use/when-not guidance and no mention of alternatives like execute_js, run_cdp, or navigate. The phrase '完整能力底座' only hints that this is the general-purpose escape hatch; the agent must infer selection criteria on its own.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.2.1- First observed
capture - First observed
close_browser - First observed
execute_js - First observed
get_manual - First observed
navigate - First observed
run_cdp - First observed
run_drission_code
TDQS
Scored across 7 tools
run_cdp, execute_js, and run_drission_code all provide code/protocol execution on the current tab, creating overlap that could confuse an agent. get_manual and close_browser are distinct, but the three execution tools lack clear boundary guidance.
Mostly consistent snake_case verb_noun style (run_cdp, get_manual, close_browser, run_drission_code, execute_js), with 'capture' and 'navigate' as minor single-verb exceptions.
Seven tools are well-scoped for a Chrome automation server, covering navigation, JS execution, CDP, network capture, manual lookup, and browser lifecycle without bloat.
Covers the core lifecycle: launch/close, navigate/tabs, JS/CDP execution, network capture, and a manual. Minor gaps exist around explicit tab management (though navigate and run_drission_code provide switch_tab/tabs), and no dedicated element-click/type tools, but coverage is strong for the stated purpose.
Maintenance
Related MCP Connectors
Hosted real Google Chrome MCP with per-user persistent state. Navigate, click, type, screenshot.
Browserless MCP — wraps the Browserless headless-Chromium REST API (browserless.io)
Access Kernel's cloud-based browsers and app actions via MCP (remote HTTP + OAuth).
- TabfleetOAuthcom.tabfleet
Launch, inspect, control, and share isolated cloud browsers for your agents.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables MCP clients to drive a real, logged-in Chrome browser for web automation tasks like navigation, clicking, typing, and screenshotting.01MIT
- AlicenseNot gradedqualityDmaintenanceEnables browser automation over MCP using a real Chrome browser with existing profile, supporting real tabs, downloads, cookies, and RPA workflows.120 npmMIT
- FlicenseNot gradedqualityDmaintenanceEnables MCP clients to control a real local browser window for web automation tasks such as clicking, typing, scrolling, and taking screenshots.7 npm-
- AlicenseAqualityBmaintenanceEnables MCP clients to control a ChromiumFish browser for web automation, including page management, navigation, interaction, and content extraction.243MIT