Browser Controller
Provides the companion Chrome extension distributed via the Chrome Web Store, enabling AI agents to control the user's real browser with existing sessions, cookies, and login states.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Browser Controllergo to github.com and click the notifications bell"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
What this project solves
You ship a fix. Your agent says "done, please verify." You alt-tab to Chrome, navigate to the page, log in, click around, find the bug.
Your agent just wrote the code. It could also verify it. It already has your browser open right there. It just can't see it.
Now it can. Browser Controller gives any MCP-compatible AI agent (Cursor, Claude Desktop, Windsurf, …) direct control of the browser you already have open — your real sessions, your logins, your cookies. No headless browser, no fresh profile, no re-authentication.
What's new in v2 (multi-client)
v1 was a single-client server: one agent, one active tab. If you switched tabs or ran two agents, they stepped on each other. v2 is a multi-client daemon with explicit tab targeting:
Multiple agents at once. Cursor can drive tab 10 while Claude drives tab 11 — both through one shared daemon, neither blocking the other.
Tab targeting, not "the active tab." Every action names a
tabId. Move your mouse, switch tabs, watch YouTube — the agent keeps working on the tab you told it to. It never hijacks the page you're reading.Per-tab isolation. Element refs, console logs, and network buffers are scoped per tab. A ref from tab 10 can never click something in tab 20.
Per-tab concurrency. Two actions on the same tab serialize (no races); actions on different tabs run in parallel.
Tab locking. An agent can claim a tab so others queue behind it instead of racing (
browser_tabs { action: "lock" }).Authenticated WebSocket. The extension↔daemon connection requires a token, so no other local process can silently drive your browser.
No debugger banner.
browser_evaluatenow runs in the page's MAIN world viachrome.scripting— no yellow "this tab is being debugged" banner.
Related MCP server: ChromeWire MCP
How it works
Three pieces, all on your machine. Nothing leaves localhost.
Agent (Cursor / Claude / Windsurf) ── other agents connect too ──┐
│ stdio (MCP protocol) │
▼ ▼
┌─────────────────────────────┐ ┌─────────────────────────────────────────┐
│ thin MCP client │ │ thin MCP client │
│ (npx browser-controller) │ │ (npx browser-controller) │
│ - speaks MCP over stdio │ │ - spawns daemon if not running │
│ - forwards calls to daemon │ │ - gets its own sessionId │
└──────────────┬──────────────┘ └────────────────────┬───────────────────┘
│ local IPC socket (AF_UNIX / named pipe, token-auth) │
▼ ▼
┌──────────────────────────────────────────────────────────────────────────┐
│ DAEMON (single long-running process, owns port 7225) │
│ - multiplexes N clients → 1 extension │
│ - tags every call with the client's sessionId │
│ - enforces per-tab mutex + tab-lock coordination │
└──────────────────────────────┬───────────────────────────────────────────┘
│ WebSocket ws://127.0.0.1:7225?token=…
▼
┌──────────────────────────────────────────────────────────────────────────┐
│ Chrome Extension (Manifest V3 service worker) │
│ - resolves the target tabId (never "the active tab" implicitly) │
│ - serializes same-tab actions, parallelizes cross-tab actions │
│ - executes click/type/snapshot/evaluate against the named tab │
└──────────────────────────────────────────────────────────────────────────┘Key idea: the first time any agent runs, npx browser-controller spawns a background daemon that owns port 7225 and the extension connection. Every subsequent agent (even from a different MCP client) connects to that same daemon over a local IPC socket and gets its own sessionId. The extension sees one stable connection and routes each call to the exact tab the caller specified.
Quick Start
Two parts: the MCP server (runs on your machine, talks to your AI agent) and the Chrome extension (sits in your browser, executes commands).
1. Add the MCP server
Cursor (one click):
Or add manually in Cursor Settings > MCP > "Add new MCP server":
{
"mcpServers": {
"real-browser": {
"command": "npx",
"args": ["-y", "browser-controller"]
}
}
}By default the daemon names each connection after its parent IDE ("Cursor",
"Claude", …). To override — e.g. when several agents share one IDE, or to label
them by project — pass --agent <name> in the args. It takes priority over
every auto-detection:
{
"mcpServers": {
"real-browser": {
"command": "npx",
"args": ["-y", "browser-controller", "--agent", "My Project Agent"]
}
}
}The name appears in the popup's Connected Agents list. (You can also set the
MCP_AGENT_NAME env var — equivalent.) Reconnecting with the same name replaces
the old entry, so IDE restarts don't pile up duplicates.
Claude Desktop: Edit ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows). Add the same JSON block.
Windsurf: Settings > MCP. Same config.
Multiple agents at once (v2): just add the server in each client. They all share the one daemon automatically — no extra config.
{
"mcpServers": {
"real-browser": { "command": "npx", "args": ["-y", "browser-controller"] }
}
}Any MCP-compatible client works.
2. Install the Chrome extension
Or load from source:
git clone https://github.com/compnew2006/browser-controller.gitOpen
chrome://extensionsand enable Developer mode (toggle in the top right)Click Load unpacked and select the
extension/folder from the cloned repo
Click the Browser Controller icon in your toolbar.
Green dot = connected. Gray = waiting for the daemon.
3. Pair the extension with the daemon's auth token
The daemon generates a secret token on first run. The extension must present it to connect.
Run any browser command once (e.g. ask your agent to "list my browser tabs") — this starts the daemon and creates the token at
~/.browser-controller/token.json.Read it:
cat ~/.browser-controller/token.json(on Windows:%USERPROFILE%\.browser-controller\token.json)Click the extension icon → paste the token into the Auth Token field.
Green dot = you're connected. Your agent can now see your browser.
The token prevents any other local process from opening a WebSocket and driving your authenticated browser sessions. It's optional in the sense that the daemon will still start without it being set in the extension, but the extension won't connect until it matches.
How to use it (steps)
The v2 model is tab-first: the agent always says which tab to act on. It never assumes "the active tab."
Basic workflow
List tabs to get a
tabId:browser_tabs { action: "list" } → [{ id: 15, url: "...", title: "...", active: true, lockedBy: null }, ...]Snapshot that tab to see its structure and get element refs:
browser_snapshot { tabId: 15 } → { tree: [ { ref: "e3", role: "button", name: "Sign in" }, ... ] }Refs are valid only for this tabId. If you navigate or the DOM changes, re-snapshot. New elements since the last snapshot are tagged
isNew: true— after an action opens an overlay/dropdown, the agent can focus on just those instead of re-reading the whole tree.Interact using the ref and the same tabId:
browser_click { tabId: 15, ref: "e3" } browser_type { tabId: 15, ref: "e5", text: "hello@example.com" } browser_press_key { tabId: 15, key: "Enter" }If a ref is stale but the element still exists, it's found automatically via a robust selector + text/role scan (response carries
via: "fallback"). If the element was scrolled away entirely (virtualized feeds), the response carriesfreshRefs: [...]with a fresh snapshot inline — retry with one of those new refs in the same step, no separate snapshot needed.Verify — snapshot or read text again after the action.
Multi-agent coordination (two agents, two tabs)
Agent A lists tabs, picks tab 10, optionally locks it:
browser_tabs { action: "lock", tabId: 10 }Agent B lists tabs, picks tab 11, locks it:
browser_tabs { action: "lock", tabId: 11 }Both work in parallel. Each agent's calls serialize against its own tab; the two tabs never interfere.
When done:
browser_tabs { action: "unlock", tabId: 10 }.
The popup is your control panel for all of this:
Open Tabs — every open tab with a per-tab dropdown to pin it to a specific agent, or unpin it.
Connected Agents — each connected agent with its name, session id, uptime, and a ✕ to disconnect it immediately (clears a zombie that the heartbeat hasn't reaped yet).
Tab Locks + Unlock All — the current lock map, with a one-click release if an agent crashed mid-lock.
Things to know
Forgot
tabId? You'll get a clear error:tabId is required. Call browser_tabs list first.Protected pages (
chrome://, the Web Store, devtools) can't be scripted — you'll getCannot access protected page (chrome://...)instead of a silent hang.browser_navigateis the one tool wheretabIdis optional (defaults to the active tab) — but for multi-agent safety, pass it explicitly.browser_evaluatenow runs in the page's MAIN world (no debugger banner, CSP-safe). It's powerful but non-idempotent — it won't be auto-retried on timeout.Scrolling virtualized feeds (Facebook/Instagram/Twitter):
browser_scrollreturnsrefsMayBeStale: truebecause those sites recycle DOM nodes. Re-snapshot before your next interaction.Duplicate elements: when several elements share text+role (e.g. 3 "Like" buttons), the fallback resolver picks the correct one by ordinal (
nth), not just the first match.
🧠 Teach Your Agent
The agent can use all 22 tools out of the box, but it works better when it knows the tab-first workflow. Run one command to install the rules:
npx browser-controller --setup cursorThis installs:
~/.cursor/rules/browser-controller.mdc— the tab-targeting workflow, dropdown handling, when to lock tabs~/.cursor/commands/check-browser.md— adds/check-browserto your Cursor chat
After that, type /check-browser in any chat. Or just say "check the result in my browser" and the agent knows what to do.
npx browser-controller --setup claudeAdds an AGENTS.md to your project root. Claude Code auto-discovers it.
See agent-config/ for manual installation or to customize the rules.
What It Can Do
22 tools. Every page-interaction tool takes a tabId (the one exception is browser_navigate, where it's optional).
See
Tool | What it does |
| Accessibility tree with element refs. Compact mode (default) returns only interactive elements. Traverses shadow DOM + iframes. |
| Capture a tab as an image (activates the tab first to capture) |
| Extract raw text from page or element |
| Query elements by natural language |
Interact
Tool | What it does |
| Click by ref or CSS selector |
| Click by visible text. Works through React portals and overlays |
| Type into inputs and contenteditable fields |
| Key combos (Enter, Escape, Ctrl+A) |
| Scroll pages and virtual containers |
| Trigger tooltips and dropdowns |
| Pick from native |
| Wait for elements to appear or disappear |
| Fill multiple form fields in one call |
| Drag element-to-element (uses CDP for reliability) |
| Upload files through |
Navigate
Tool | What it does |
| Go to a URL in a tab ( |
| List / create / close / focus / lock / unlock tabs |
Debug & Advanced
Tool | What it does |
| Console output (log, warn, error) — per-tab, capped at 200 entries |
| XHR/fetch requests with status codes — per-tab |
| Run JavaScript in the page's MAIN world (no banner, CSP-safe) |
| Handle alert/confirm/prompt dialogs |
| Run a self-contained JS action object via CDP |
How Others Compare
Browser Controller | Playwright MCP | Chrome DevTools MCP | |
Uses your existing browser | Yes | No, launches new | Partial, needs debug port |
Sessions and cookies | Already there | Fresh profile | Manual setup |
Works behind corporate SSO | Yes | No | Depends |
Multiple agents, multiple tabs | Yes (v2) | No | No |
Tab-targeting (won't hijack active tab) | Yes (v2) | N/A | No |
Authenticated local connection | Yes (v2) | N/A | No |
Setup | Extension + MCP config | Headless browser | Chrome with |
Configuration
Env var | Default | What it does |
|
| WebSocket port the daemon uses for the extension connection |
| (unset) | Set to |
Daemon state files
The daemon keeps everything in ~/.browser-controller/ (Windows: %USERPROFILE%\.browser-controller\):
File | Purpose |
| Auth token the extension must present (mode |
| The IPC socket thin clients connect to (AF_UNIX on mac/linux; named pipe on Windows) |
| Daemon metadata (pid, port, start time) — used to detect a running daemon |
| Daemon stdout/stderr when spawned by a client |
To fully reset: stop your MCP clients, delete the folder, and the next run recreates it with a fresh token.
Reliability
The daemon is auto-spawned the first time any client runs and left running detached.
Connection drops use exponential backoff (1s → 30s), ping/pong health checks every 10s.
Per-tool timeouts (5s for clicks, 60s for navigation).
Read-only tools (snapshot, screenshot, text, find, console, network) are retried on timeout; side-effecting tools (click, type, navigate, evaluate) are never retried — a click can't fire twice.
If a previous daemon crashed while holding port 7225, the new one finds and evicts the stale PID (cross-platform:
lsofon mac/linux,Get-NetTCPConnectionon Windows).
Run two daemons on different ports by setting WS_PORT per client:
{
"mcpServers": {
"browser-work": {
"command": "npx", "args": ["-y", "browser-controller"]
},
"browser-personal": {
"command": "npx", "args": ["-y", "browser-controller"],
"env": { "WS_PORT": "9333" }
}
}
}Update the port in each extension popup to match.
Everything stays on your machine. The extension connects to the daemon via an authenticated WebSocket on localhost; MCP clients connect to the daemon via a local IPC socket. No cloud, no proxy, nothing leaves your browser.
browser-controller/
├── mcp-server/ MCP server (npm package, TypeScript)
│ └── src/
│ ├── daemon.ts Single multi-client daemon (owns WS :7225)
│ ├── daemon-config.ts IPC protocol, paths, auth token, idempotency set
│ ├── index.ts Thin stdio MCP client (spawns daemon, multiplexes)
│ ├── bridge.ts Extension WS server + cross-platform port probe
│ └── tools/ One file per tool (22), registry pattern
├── extension/ Chrome extension (Manifest V3, plain JS)
│ ├── background.js Service worker: tab resolution, per-tab mutex, locks
│ ├── lib/tab-concurrency.js Pure, unit-tested mutex + lock primitives
│ ├── content.js Console capture
│ └── popup/ Status, Open Tabs (pin/unpin), Connected Agents (disconnect), tab-lock viewer
├── agent-config/ Pre-built configs for Cursor + Claude Code
│ ├── cursor/ Rules and commands
│ ├── skills/ Browser automation skill
│ └── setup.mjs One-command installer
└── tests/ Bridge + registry + concurrency tests (45 passing)Stack: TypeScript (strict) · MCP SDK · WebSocket · Chrome Extension Manifest V3 · Vitest
git clone https://github.com/compnew2006/browser-controller.git
cd browser-controller
npm install
npm run build
npm testCommand | What it does |
| Compile TypeScript → |
| Watch mode |
| Run the full test suite (45 tests) |
| Type check without emitting |
| Install Cursor rule + command |
The test suite covers the WebSocket bridge (including token-auth rejection), the tool registry, and the per-tab concurrency primitives (same-tab serialization + cross-tab parallelism). The daemon's IPC auth and the WS token enforcement are also smoke-tested via the build artifacts in mcp-server/dist/.
FAQ
That's the whole point. The extension runs inside your actual Chrome — same cookies, same sessions, same local storage. No re-authentication needed.
No. The MCP clients, the daemon, and the extension all talk over localhost (IPC socket + WebSocket). Nothing leaves your machine. There's no analytics, no telemetry, no cloud component. Privacy policy.
Any MCP-compatible client. Cursor, Claude Desktop, Claude Code, Windsurf, Cline, and anything else that speaks the MCP protocol. v2 lets several of them run at once against the same daemon.
Yes — that's the main v2 improvement. Each agent connects to the shared daemon, gets its own sessionId, and targets a specific tabId. Actions on the same tab serialize through a per-tab mutex; actions on different tabs run in parallel. Optionally an agent can lock a tab to claim exclusive access; other agents queue behind the lock rather than failing.
It can't — not silently. Every page-interaction tool requires a tabId, and if it's missing you get a clear tabId is required error. The agent can never accidentally act on the tab you happen to be looking at. (The one exception is browser_navigate without a tabId, which uses the active tab — but for multi-agent use you should always pass tabId.)
Without it, any local process on your machine could open a WebSocket to port 7225 and drive your authenticated browser sessions (your bank, your email, your company SSO). The daemon generates a secret token in ~/.browser-controller/token.json and the extension must present it to connect.
They launch a new browser instance from scratch — no state, no cookies, no sessions. You have to replay the full login flow every time. This connects to the browser you already have open with everything already loaded.
Contributing
Bug reports, feature requests, and PRs welcome. Open an issue first for larger changes.
Author
noiemany — GitHub
License
MIT © noiemany
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/compnew2006/browser-controller'
If you have feedback or need assistance with the MCP directory API, please join our Discord server