chatgpt-remote-browser-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@chatgpt-remote-browser-mcplist the tabs I've shared from Chrome and screenshot the current one"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ChatGPT Remote Browser MCP
An MCP server that exposes OpenClaw's Chrome-extension browser controls while preserving the extension's Selected Tabs access boundary.
The server is designed for a topology in which ChatGPT reaches a private MCP process through an outbound Secure MCP Tunnel, OpenClaw runs on a gateway host, and Chrome may run on a different machine with the OpenClaw extension.
Security model
Only tabs currently shared by the OpenClaw extension are published.
Raw Chrome/OpenClaw target IDs are replaced with opaque, short-lived handles.
The current tab ACL is checked again before every operation.
Revoked, expired, unknown, and closed handles fail with
ACCESS_DENIED.Read-only operations may retry a small set of transient target-resolution errors. Mutating operations are never automatically retried.
There is no generic shell, gateway RPC, Relay, or raw CDP passthrough.
Upload paths are restricted to OpenClaw-owned inbound/upload locations.
Generated file responses are capped at 20 MiB and temporary output is removed.
native_browser includes powerful operations such as JavaScript evaluation,
navigation, uploads, downloads, and tab lifecycle changes. Use MCP client
confirmations and share only tabs that are safe to automate.
See Security for the complete trust model.
Related MCP server: SecureFoxMCPServer
Architecture
ChatGPT / MCP client
|
| outbound/private MCP transport
v
chatgpt-remote-browser-mcp (stdio)
|
| fixed OpenClaw browser CLI calls
v
OpenClaw Gateway -> Chrome extension -> Selected Tabs ACLSee Architecture for details.
Tools
Compatibility tools:
Viewer:
list_shared_tabs,snapshot,screenshotOperator:
click,type,press,scroll,navigate
Native high-level tool:
native_browseractions:tabs,open,focus,close,label,snapshot,screenshot,navigate,console,requests,errors,text,emulate,pdf,download,waitfordownload,upload,dialog,highlight,responsebody, andactactkinds:batch,click,clickCoords,type,press,hover,scrollIntoView,drag,select,fill,resize,wait,evaluate, andclose
Internal extension transport operations such as attach/detach are deliberately encapsulated. Local profile lifecycle and Gateway administration are outside the MCP surface.
Requirements
Node.js 22 or newer
A working OpenClaw installation and browser profile
OpenClaw Chrome extension paired with the Gateway
Optional: an OpenAI Secure MCP Tunnel for private ChatGPT connectivity
Install and run
npm ci
npm run check
npm test
npm install -g .
openclaw-browser-mcpThe MCP transport is stdio. Configuration is provided through environment variables:
OPENCLAW_BIN: OpenClaw executable; default/opt/homebrew/bin/openclawOPENCLAW_BROWSER_PROFILE: browser profile; defaultchromeOPENCLAW_BROWSER_TIMEOUT_MS: backend timeout; default30000
Private tunnel deployment
Generic deployment templates are in deploy/. They intentionally
contain placeholders rather than machine paths, tunnel identifiers, or
credentials. Render them locally into protected runtime configuration; never
commit the rendered files.
See Deployment for setup, health checks, operation, and rollback.
Tests
npm run check
npm test
npm run test:e2e
npm run test:remote
npm run test:acl
npm audit --omit=devUnit/schema/security suite: 22 cases
Local managed-browser E2E
Self-contained remote Chrome-extension E2E using a disposable tab
Interactive ACL-revocation E2E
The remote E2E creates, shares, exercises, closes, and post-validates its own tab. It does not require a manually prepared test tab. See Testing.
Status
Version 0.2.2 has been validated with:
22/22 unit, schema, dispatch, and security tests
local browser E2E
remote Chrome-extension E2E across all compatibility tools and 16 native scenarios
immediate
ACCESS_DENIEDafter ACL revocation and after tab closurea private ChatGPT-to-MCP tunnel invocation
zero known production dependency vulnerabilities at validation time
License
Available Tools
9 toolsclickClick in shared browser tabA
Click an element in a currently shared tab. This changes browser or external-site state.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| button | No | left | |
| double | No | ||
| handle | Yes | ||
| modifiers | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explicitly states that clicking changes browser or external-site state, which is a side effect disclosure. This aligns with the annotations that indicate it is not read-only, and adds specificity about the nature of the change.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, consisting of two short sentences. It is well-structured, clearly stating the action and its side effect without superfluous information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description gives the core purpose and side effect, but lacks details about parameters, expected outcomes, or error conditions. Given the minimal schema coverage, it leaves out necessary context for an agent to fully understand how to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not explain any of the five parameters (ref, button, double, handle, modifiers). Since the schema provides no descriptions, the tool description fails to convey the meaning or usage of these parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (click) and a specific resource (element in a currently shared tab). It clearly distinguishes from sibling tools like type, scroll, and navigate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a condition "in a currently shared tab" that guides when to use the tool. However, it does not explicitly compare with alternatives, but the condition implies it is for shared tab contexts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
native_browserNative OpenClaw browser controlCDestructive
Full high-level OpenClaw Chrome-extension browser surface for Selected Tabs. Use list_shared_tabs first and pass its opaque handle. Supports tab lifecycle, DOM interaction, evaluation, diagnostics, emulation, files and batch actions. Raw shell and raw target IDs are never accepted.
| Name | Required | Description | Default |
|---|---|---|---|
| x | No | ||
| y | No | ||
| fn | No | ||
| key | No | ||
| ref | No | ||
| url | No | ||
| kind | No | ||
| text | No | ||
| urls | No | ||
| clear | No | ||
| depth | No | ||
| frame | No | ||
| label | No | ||
| level | No | ||
| limit | No | ||
| paths | No | ||
| width | No | ||
| accept | No | ||
| action | Yes | ||
| button | No | ||
| device | No | ||
| endRef | No | ||
| fields | No | ||
| filter | No | ||
| handle | No | ||
| height | No | ||
| labels | No | ||
| locale | No | ||
| origin | No | ||
| slowly | No | ||
| submit | No | ||
| timeMs | No | ||
| values | No | ||
| actions | No | ||
| compact | No | ||
| delayMs | No | ||
| dismiss | No | ||
| element | No | ||
| headers | No | ||
| offline | No | ||
| accuracy | No | ||
| dialogId | No | ||
| fullPage | No | ||
| inputRef | No | ||
| latitude | No | ||
| maxChars | No | ||
| selector | No | ||
| startRef | No | ||
| textGone | No | ||
| efficient | No | ||
| imageType | No | ||
| loadState | No | ||
| longitude | No | ||
| modifiers | No | ||
| timeoutMs | No | ||
| promptText | No | ||
| timezoneId | No | ||
| colorScheme | No | ||
| doubleClick | No | ||
| interactive | No | ||
| responseUrl | No | ||
| stopOnError | No | ||
| downloadName | No | ||
| snapshotFormat | No | ||
| clearGeolocation | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructive/open-world behavior and the description adds capability context such as tab lifecycle, files, and batch actions. No contradiction exists, but concrete side effects (e.g., closing tabs, downloading files, dismissing dialogs) are not spelled out.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and readable, with the key prerequisite and constraint front-loaded. It avoids boilerplate, though the phrase 'Full high-level' is slightly redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 65 parameters, nested objects, and no output schema, the description provides only broad categories and leaves the agent without enough information to construct valid calls or interpret results. No return value or error behavior is mentioned.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description does not explain any of the 65 parameters, their relationships, or required combinations. Even the referenced 'opaque handle' is not explicitly mapped to the handle parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies this as the high-level OpenClaw browser control surface for tabs and lists the supported capability categories (tab lifecycle, DOM interaction, evaluation, etc.). It also differentiates it from raw shell/target ID access, though it doesn't name sibling tools explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It instructs the agent to call list_shared_tabs first and pass the opaque handle, and warns that raw shell/target IDs are not accepted. However, it does not explain when to prefer this tool over the sibling actions like navigate, click, or type.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pressPress key in shared browser tabB
Press a keyboard key in a currently shared tab. This can change browser or external-site state.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| handle | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses that pressing a key can change browser or external-site state, which goes beyond the annotations (readOnlyHint=false, destructiveHint=false). It also implies a precondition that the tab must be currently shared, offering useful behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured, with two short sentences that convey the essential action and side effect without unnecessary detail. It is efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a simple tool with no output schema, but it lacks usage guidance and parameter context. It mentions the shared-tab prerequisite and potential state changes, but does not explain how the handle identifies the tab or what keys are accepted, leaving some gaps for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not mention the parameters 'handle' or 'key' at all, providing no insight into what values they expect or how they relate to the action. Since schema coverage is 0% and the description does not compensate, parameter semantics are completely lacking.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (press a keyboard key) and the target (a currently shared tab). It also distinguishes from sibling tools like type, click, and scroll by specifying 'keyboard key' and 'currently shared tab', making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool over alternatives such as type or click. It neither explains the difference between pressing a single key and entering text, nor mentions any prerequisites beyond the tab being shared, leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshotScreenshot shared browser tabCRead-only
Capture an image of a currently shared tab.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | png | |
| handle | Yes | ||
| labels | No | ||
| fullPage | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry the safety profile (readOnlyHint=true, destructiveHint=false), and the description is consistent with them—no contradiction. The description adds only the 'currently shared' context but does not disclose output format, what the captured image contains, or edge-case behavior when the tab is unavailable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with no waste, which is good, but it borders on under-specification rather than disciplined conciseness. There is room to add value without bloating the text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 parameters (1 required), zero parameter descriptions, and no output schema, a one-line description is inadequate. The agent lacks information on the return value (where/how the image is returned), the meaning of labels and fullPage, and the behavior when the handle references an unshared tab.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it mentions none of the four parameters. An agent gets no explanation of what 'labels' does, what 'fullPage' changes, or how 'type' and 'handle' interact, leaving significant ambiguity that the schema alone cannot resolve.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Capture') and resource ('an image of a currently shared tab'), so the core action is clear. However, there is a sibling tool named 'snapshot' whose purpose could overlap, and the description does nothing to differentiate screenshot from it, leaving an agent to guess which to use.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, no prerequisites stated, and no exclusions. The phrase 'currently shared tab' implies a precondition but does not explicitly state it as a requirement or describe what happens if the tab is not shared.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrollScroll shared browser tabC
Scroll a currently shared tab by one or more pages.
| Name | Required | Description | Default |
|---|---|---|---|
| pages | No | ||
| handle | Yes | ||
| direction | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations indicate non-read-only and non-destructive, which aligns with scrolling. However, the description does not disclose any side effects (e.g., state changes, animation, or limitation when the tab is not shared). It adds minimal behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single clear sentence with no unnecessary words. It is concise and to the point, effectively communicating the core action in minimal space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is too sparse to be complete. It omits parameter details, expected behavior (e.g., what happens if the tab is not shared or if the end of the page is reached), and any output or return value. With no output schema, the description should provide more context but does not.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 3 parameters with 0% description coverage. The description only implicitly references 'pages' with 'one or more pages' but leaves 'handle' and 'direction' entirely unexplained. No effort is made to clarify parameter meaning or constraints, so the description fails to compensate for the schema gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'scroll' and the object 'a currently shared tab', specifying the action of moving by pages. It distinguishes from navigation, typing, and clicking, though it could be slightly more explicit about the page increment semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like navigate, type, or click. The description does not mention prerequisites (e.g., tab must be shared) or situations where scrolling is appropriate, leaving the agent to infer context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
snapshotRead shared browser tabARead-only
Read the accessible page structure and text from a currently shared tab.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| format | No | ai | |
| handle | Yes | ||
| interactive | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description's 'Read' is consistent with the readOnlyHint=true and destructiveHint=false annotations. It adds useful context that the tab must be 'currently shared,' but it does not mention potential return details or any parameter-dependent behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence with no redundant wording. It front-loads the core purpose and is easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is too sparse to cover the tool's parameter semantics; it does not explain the format enum values, the limit bounds, or the interactive flag, leaving important gaps for correct usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description provides no explanation of handle, limit, format, or interactive. None of the parameters' meanings, defaults, or constraints are conveyed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Read'), the target ('accessible page structure and text'), and the specific scope ('currently shared tab'). This distinguishes it from visual tools like screenshot and navigation tools like navigate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for extracting text/structure from the shared tab, but it does not explicitly explain when to prefer this tool over siblings such as screenshot or list_shared_tabs. The guidance is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
typeType in shared browser tabC
Type text into an element in a currently shared tab. This changes browser or external-site state.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| text | Yes | ||
| handle | Yes | ||
| slowly | No | ||
| submit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description notes that typing 'changes browser or external-site state,' which adds some context, but the annotations already indicate readOnlyHint=false and destructiveHint=false. No additional behavioral details are provided about optional parameters like slowly or submit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise and well-structured, with no redundant information or excessive detail. It efficiently conveys the core functionality.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lacks essential context for correct usage, including parameter explanations, when to use it instead of alternatives, and potential side effects beyond a generic state-change warning. An agent would need to infer much of the behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain any of the five parameters. While handle, ref, and text are somewhat self-explanatory, slowly and submit are ambiguous without further detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Type text into an element') and the context ('in a currently shared tab'), providing a specific verb and resource. It is distinct from sibling tools like click and press, though it does not explicitly name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as click or press. It only mentions a side-effect, not usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.2.2- First observed
click - First observed
list_shared_tabs - First observed
native_browser - First observed
navigate - First observed
press - First observed
screenshot - First observed
scroll - First observed
snapshot - First observed
type
TDQS
Scored across 9 tools
native_browser is a full browser surface that overlaps with nearly every specific tool (navigate, click, type, scroll, screenshot, etc.), making tool selection ambiguous. The specific tools are individually distinct, but the presence of an umbrella tool blurs the boundary between simple actions and the high-level API.
Most tools follow a consistent imperative verb pattern in snake_case (navigate, click, type, scroll, press, snapshot, screenshot, list_shared_tabs). native_browser breaks the pattern by being a noun rather than a verb, but the overall convention is still readable and predictable.
Nine tools is a reasonable size for a browser-control server and fits the well-scoped 3-15 range. The count is not excessive, though native_browser does add some redundancy.
The set covers the core browser interaction lifecycle: navigation, clicking, typing, scrolling, key presses, reading state, and screenshots. native_browser extends coverage to tab lifecycle, diagnostics, emulation, files, and batch actions, so the domain appears well covered.
Maintenance
Related MCP Connectors
Browser MCP for logged-in tasks. Uses your Chrome — credentials stay local. Zero-token replay.
Hosted real Google Chrome MCP with per-user persistent state. Navigate, click, type, screenshot.
Access Kernel's cloud-based browsers and app actions via MCP (remote HTTP + OAuth).
- TabfleetOAuthcom.tabfleet
Launch, inspect, control, and share isolated cloud browsers for your agents.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables MCP-compatible agents to securely control the user's already authenticated Chrome browser via explicit tab authorization and DOM-based actions.1Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables secure browser control and automation through MCP, adding domain validation to tab-specific operations while preserving navigation, content, window, history, bookmark, and request-monitoring tools.MIT
- FlicenseNot gradedqualityCmaintenanceEnables MCP clients to control real Chrome tabs in existing profiles, supporting navigation, interaction, screenshots, and console/network inspection across multiple Chrome instances.-
- AlicenseNot gradedqualityBmaintenanceEnables local control of existing Chrome/Chromium browser tabs through MCP, including tab management, navigation, content reading, screenshots, and page interaction.Apache 2.0