Skip to main content
Glama

chrome-mcp

English | 简体中文

A Model Context Protocol server for browser automation, powered by DrissionPage.

Lets MCP clients (Claude Desktop, Claude Code, Cursor, etc.) drive a real Chromium browser through a minimal 8-tool surface — all sharing one "current tab" pointer.

Why

  • Zero learning curve for agents. Facade tools (execute_js, run_cdp, navigate, capture) speak only universal knowledge — JavaScript, the CDP protocol, URLs. No niche-library syntax required for the 90% cases.

  • Full power underneath. run_drission_code is the complete-capability base: execute DrissionPage Python code in-process with page / browser / context (persistent dict) / switch_tab() / tabs() / relaunch() injected — loops, waits, multi-tab orchestration, anything one tool call can't express.

  • One pointer, always in sync. Switch tabs via DrissionPage code or Target.createTarget/Target.activateTarget CDP commands — every tool follows. (Call tools serially per session; the shared pointer is not concurrency-safe.)

  • Multi-instance safe. Per-process isolation (atomic port allocation via socket.bind + PID/UUID private user-data dirs) — spawn N servers, zero conflicts.

  • Self-documenting. get_manual returns a compact built-in cheat sheet (DrissionPage locator syntax, error-prone CDP recipes, capture workflow, launch options) — agents fetch it on demand, zero token cost otherwise.

Related MCP server: Real Browser MCP

Installation

uv tool install chrome-mcp
# or
uvx chrome-mcp
# or
pip install chrome-mcp

Requires Python ≥ 3.10 and a local Chromium-based browser.

Usage

Add to your MCP client config (e.g. claude_desktop_config.json). Two equivalent forms:

{
  "mcpServers": {
    "chrome": {
      "command": "chrome-mcp"
    }
  }
}

Or run on the fly with uvx (no install needed):

{
  "mcpServers": {
    "chrome": {
      "command": "uvx",
      "args": ["chrome-mcp"]
    }
  }
}

A typical agent flow:

navigate("https://httpbin.org/get")                  → returns tab list
execute_js("return document.body.innerText")         → page data as JSON
capture(action="start", url_filter="api/")
... trigger requests ...
capture(action="get")                                → overview inline,
                                                       full bodies in $TMPDIR/chrome-mcp-captures/*.json

Tools

Tool

Description

navigate

Navigate in current tab, or open a new tab (new_tab) and switch the pointer to it. Returns the full tab list

execute_js

Run JavaScript in the current tab (must return the result). preset: "dom_tree" outputs a DOM structure tree without writing the template

run_cdp

Raw CDP passthrough to the current tab. Target.createTarget / activateTarget auto-switch the pointer — no attachToTarget needed

run_drission_code

Full-capability base: execute DrissionPage Python code with injected page/browser/context/switch_tab()/tabs()/relaunch()/get_browser()

capture

Network capture in one tool: action = start / get / stop. Overview returned inline; full data (with bodies) saved to $TMPDIR/chrome-mcp-captures/capture_*.json

get_manual

Return the built-in authoring manual (locator syntax, CDP recipes, capture workflow, launch options)

get_browser

Instance management: no args = ensure an instance is ready (starts one if absent); cdp = attach to an already-running browser (login state preserved; hard error if unreachable — never falls back to launching); any launch option = rebuild with new params (destructive, same semantics as relaunch)

close_browser

Close the browser instance

Launch options

Customize browser startup via command-line args:

{
  "mcpServers": {
    "chrome": {
      "command": "chrome-mcp",
      "args": ["--headless", "--proxy", "http://127.0.0.1:7890", "--arg", "--lang=zh-CN"]
    }
  }
}

Arg

Description

--headless

Run browser headless

--proxy URL

Proxy server

--user-agent UA

Custom User-Agent

--user-data PATH

User data dir to reuse login state. ⚠️ Conflicts if your system Chrome is using the same profile

--browser-path PATH

Path to a specific browser binary

--incognito

Incognito mode

--no-imgs

Don't load images. ⚠️ May be ignored by recent Chrome versions — reliable alternative: run_cdp("Network.setBlockedURLs", {"urls": ["*.png", "*.jpg"]})

--arg ARG

Pass through any Chrome flag (repeatable)

Options can also be changed mid-session (destructive, restarts the browser) from inside run_drission_code:

return relaunch(headless=True, user_agent="Mozilla/5.0 ...")

The get_browser tool exposes the same options (plus cdp) at the tool level — e.g. get_browser(cdp="127.0.0.1:9222") takes over your logged-in browser; get_browser(user_data=...) rebuilds with a given profile.

Testing

8 scenario suites run against the real MCP stdio protocol (each spawns an independent server + headless browser — itself a multi-instance concurrency test):

uv run python tests/run_all.py            # full regression
uv run python tests/run_all.py --only smoke,c

Covers: data extraction (DOM tree/pagination/iframes), form interaction (DP actions + CDP input sequences), network capture (filters/timing traps/body fidelity), multi-tab pointer consistency, CDP capabilities (screenshots/emulation/cookies), lifecycle (relaunch/kill-recovery), and real-site scenarios (httpbin/TLS).

License

GPL-3.0-only

Available Tools

7 tools
captureA

网络抓包三合一。action='start':在当前 tab 启动监听(⚠️ 必须在触发请求之前调用;监听挂载在该 tab,不随指针迁移);action='get':拉取结果(概览直接返回,完整数据落盘 JSON 供 read);action='stop':停止并清空。

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoget:最多返回条数(0=全部)
actionYesstart=启动监听 / get=拉取结果 / stop=停止
methodNostart:HTTP 方法过滤,多个用逗号分隔(如 'GET,PUT')。不传时 DrissionPage 缺省只抓 GET/POST
url_filterNostart:URL 关键词过滤(空=全部)

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses important behavioral traits: the start-before-request timing constraint, tab-bound monitoring that does not migrate, that 'get' returns an inline overview while full data lands in a JSON file for 'read', and that 'stop' clears the capture. These are meaningful operational details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with a summary ('三合一') followed by per-action groupings, each earning its place. No redundant sentences beyond the heading. Dense but efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-mode capture tool with no output schema and no annotations, the description covers the critical timing, scoping, and result-handling details an agent needs. It could mention permissions or how the JSON file path is obtained, but the core operational picture is complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters including the enum and method filter defaults. The description adds only a brief parenthetical about the overview/JSON-file split for 'get' at the top level, which is marginal beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb+resource: network capture with three modes (start/get/stop) broken out explicitly. It is distinguishable from siblings like run_cdp or execute_js, though the exact relationship to browser-attached siblings is not explained. Purpose is clear without being fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear sequencing guidance: 'start' must be called before triggering the request, and monitoring is bound to the current tab rather than following the pointer. It does not name alternatives or when-not-to-use, but the timing and scoping instructions are actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

close_browserC

关闭浏览器实例

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the action (close) but doesn't describe what happens—e.g., whether it gracefully shuts down the browser, kills processes, affects other tools, or has side effects like losing unsaved data. This is a significant gap for a potentially destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence ('关闭浏览器实例') that directly states the purpose without waste. It's appropriately sized for a simple tool, though it could be slightly more informative without losing conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's potential complexity (closing a browser could involve cleanup, state changes, or errors) and lack of annotations and output schema, the description is incomplete. It doesn't address behavioral aspects, return values, or error conditions, leaving gaps for an AI agent to understand its full context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has 0 parameters, and schema description coverage is 100%, so no parameter documentation is needed. The description doesn't add parameter semantics, but that's acceptable here. Baseline is 4 for zero parameters, as there's nothing to compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description '关闭浏览器实例' (close browser instance) states a clear action (close) on a resource (browser instance), but it's vague about scope—it doesn't specify whether this closes all browser windows, a specific instance, or just the current tab. It distinguishes from siblings like 'navigate_to_page' or 'get_browser_status' by being a termination action, but lacks precision.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. For example, it doesn't mention if this should be used after completing tasks, to free resources, or as part of cleanup, nor does it reference sibling tools like 'launch_chrome_manually' for context. The description alone offers no usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

execute_jsA

在当前 tab 执行 JavaScript 代码。⚠️ 必须使用 return 语句返回结果。preset='dom_tree' 时 code 改为传 JSON 参数对象 {"max_depth":5,"include_text":true,"max_text_length":50,"selector":null},直接输出 DOM 结构树。

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes要执行的 JavaScript 代码(必须用 return 返回结果!),或 preset='dom_tree' 时的 JSON 参数
presetNo预置脚本模板:dom_tree=DOM 结构树(免写模板 JS)
timeoutNo本次 JS 执行超时(秒,可选)。默认 150 秒。

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the critical behavioral rule that the code must use return to produce a result, plus the default 150s timeout. It does not cover error behavior, sandbox/permission constraints, or what happens when the script throws, leaving meaningful gaps for a code-execution tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Compact and front-loaded: the core action comes first, then the return-value warning, then the preset override. No filler sentences, though the preset paragraph is dense and could be slightly tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description must explain the return contract, which it partially does via the return-statement warning. However, it is silent on error handling, result serialization, and side effects, which matter for an arbitrary JS execution tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and each parameter is already documented, so baseline would be 3. The description adds value beyond the schema by showing the exact JSON argument shape for preset='dom_tree' with concrete keys (max_depth, include_text, max_text_length, selector), clarifying how the overloaded 'code' parameter behaves.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: execute JavaScript code in the current tab, and clarifies a second mode (preset='dom_tree') that changes what 'code' means. It does not name or contrast with siblings like run_cdp or run_drission_code, so an agent must infer which execution tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives the condition for the dom_tree preset (pass JSON params instead of code), which is useful usage guidance, but it never says when to prefer this tool over run_cdp or run_drission_code, nor any exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_manualB

返回内置编写手册:定位语法、元素动作、等待、tab 管理、抓包、CDP 高频配方、启动定制等速查。

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose the essential behavior — a zero-parameter read that returns reference content, therefore inherently non-destructive — but says nothing about return format, length, or whether the content is static or context-dependent. That omission is notable for a tool whose entire value is the payload.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that leads with what is returned before the topic enumeration. Every clause earns its place by telling the agent what content to expect. Slightly list-heavy, but no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-output-schema retrieval tool, the description supplies the one thing the agent needs: what knowledge it will obtain. The lack of any hint about return size or format is the main residual gap, but the topic list makes the tool's purpose sufficiently clear to call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the 4 baseline applies. The description makes clear it is invoked without arguments, matching the empty object schema with 100% coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: it returns a built-in authoring manual, and enumerates the topics covered (locator syntax, element actions, waits, tab management, capture, CDP recipes, launch customization). This clearly distinguishes it from siblings like run_cdp or run_drission_code, which execute rather than document. It stops short of a 5 only because the manual's scope beyond the listed topics is undefined.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no reference to the sibling tools whose DSL it presumably documents. The agent must infer on its own that this is a reference to consult before calling run_drission_code or run_cdp. Nothing states when-not to call it or whether it is cheap to call repeatedly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_cdpA

对当前 tab 透传 Chrome DevTools Protocol 命令。method 如 'Page.captureScreenshot',params 为参数对象。命令始终作用于当前 tab,无需 attachToTarget(Target.createTarget/activateTarget 会自动切换操作 tab)。详见 get_manual。

ParametersJSON Schema
NameRequiredDescriptionDefault
methodYesCDP 方法名,如 'Page.navigate'、'Runtime.evaluate'
paramsNoCDP 命令参数对象(可选)

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose a non-obvious side effect: Target.createTarget/activateTarget silently change which tab subsequent commands act on. That is exactly the kind of behavioral trait annotations would otherwise cover. It still omits return shape, error behavior, and any auth/permission constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, front-loaded with the core action and followed by the tab-scoping caveat and a pointer to get_manual. No filler, though the parenthetical about tab switching is dense and could be split for readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a raw passthrough tool with no output schema and no annotations, the description covers what it does, its scoping constraint, the auto tab-switch side effect, and where to get deeper detail (get_manual). Return values and failure modes are undefined but are inherently method-dependent, so this is close to complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description's example ('Page.captureScreenshot') and its note that params is an optional argument object largely restate what the schema already documents for method and params, adding little new semantic content.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: passing through Chrome DevTools Protocol commands to the current tab, with a concrete example method. It implicitly distinguishes itself from siblings like execute_js, capture, and run_drission_code by being the raw-CDP escape hatch, though it never names an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives useful operating context (commands always act on the current tab, no attachToTarget needed, createTarget/activateTarget auto-switch the operating tab) and defers to get_manual. However, it never says when to prefer this over execute_js, capture, or run_drission_code, which is the main selection decision an agent faces.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_drission_codeB

执行 DrissionPage Python 代码(server 进程内,完整能力底座)。预置变量:page=当前tab、browser=浏览器对象、context=跨调用持久字典、DOM_TREE_JS=DOM树JS模板。预置函数:switch_tab(tab_id)、tabs()、relaunch(**opts)。⚠️ 必须用 return 返回结果。

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes要执行的 Python 代码(必须用 return 返回结果!)
timeoutNo整体兜底超时(秒,可选)。默认 120 秒。

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose meaningful runtime behavior: it runs server-side, exposes preset page/browser/context variables, and must return results. However, for arbitrary code execution it omits any warning about side effects, permissions, or failure behavior, which is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense block, front-loaded with the purpose before environment details, with no filler sentences. The ⚠️ return requirement is correctly emphasized, though the run-on layout slightly reduces scanability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a code-execution tool with no output schema and no annotations, the description supplies the essential execution context: environment, preset variables, preset functions, and the return convention. It is largely complete, missing only error/edge-case behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (code, timeout) are already documented, establishing the baseline of 3. The preset variables and functions listed in the description enrich the code-writing context but do not clarify the parameters themselves beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('执行 DrissionPage Python 代码') plus scope ('server 进程内'), which clearly separates it from the JS-oriented execute_js sibling. It never explicitly names an alternative, so it lands at clear-but-undifferentiated rather than a full 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use/when-not guidance and no mention of alternatives like execute_js, run_cdp, or navigate. The phrase '完整能力底座' only hints that this is the general-purpose escape hatch; the agent must infer selection criteria on its own.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv0.2.1
    • First observedcapture
    • First observedclose_browser
    • First observedexecute_js
    • First observedget_manual
    • First observednavigate
    • First observedrun_cdp
    • First observedrun_drission_code

TDQS

B3.4/5.0

Scored across 7 tools

Disambiguation3/5

run_cdp, execute_js, and run_drission_code all provide code/protocol execution on the current tab, creating overlap that could confuse an agent. get_manual and close_browser are distinct, but the three execution tools lack clear boundary guidance.

Naming Consistency4/5

Mostly consistent snake_case verb_noun style (run_cdp, get_manual, close_browser, run_drission_code, execute_js), with 'capture' and 'navigate' as minor single-verb exceptions.

Tool Count5/5

Seven tools are well-scoped for a Chrome automation server, covering navigation, JS execution, CDP, network capture, manual lookup, and browser lifecycle without bloat.

Completeness4/5

Covers the core lifecycle: launch/close, navigate/tabs, JS/CDP execution, network capture, and a manual. Minor gaps exist around explicit tab management (though navigate and run_drission_code provide switch_tab/tabs), and no dedicated element-click/type tools, but coverage is strong for the stated purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables MCP clients to drive a real, logged-in Chrome browser for web automation tasks like navigation, clicking, typing, and screenshotting.
    0
    1
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables browser automation over MCP using a real Chrome browser with existing profile, supporting real tabs, downloads, cookies, and RPA workflows.
    120 npm
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables MCP clients to control a real local browser window for web automation tasks such as clicking, typing, scrolling, and taking screenshots.
    7 npm
    -