agent-browser-extension
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@agent-browser-extensionclaim the tab showing my dashboard and take a screenshot"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
agent-browser-extension
A standalone Chrome extension and local bridge for autonomous AI browser tasks. Works with MCP-compatible agents and direct HTTP clients. Any compatible agent host can connect; no particular app or model provider is required.
Agents discover existing tabs, claim the ones relevant to their task, or create background tabs in their own named groups. There is no required tab picker. The extension uses the signed-in profile where you install it.
What it does
Existing tab discovery, competing-task ownership, background tab creation and task groups.
Accessibility snapshots with element references and semantic locators by role/name, text, label, placeholder, or test ID.
Rendered page text and current form state; idempotent checkbox, radio and switch controls; click, fill, caret typing, native dropdown selection, keyboard, hover, pointer drag, scrolling and browser history.
Frame discovery and scoped actions across same-origin, cross-origin and nested frames, plus selectors inside open shadow roots.
PNG/JPEG viewport and bounded full-page screenshots with image dimensions and CSS geometry for visual interaction; bounded JavaScript extraction.
File uploads, task-attributed downloads and JavaScript dialogs.
Immediate popups correlated with agent actions, including
target="_blank", middle-clicked links andwindow.open, join the task's ownership and cleanup.Task cleanup: close temporary agent tabs, preserve deliverables and handoffs, and release existing user tabs.
Visible connection/task status, Show controls and Stop controls in the extension.
Authenticated local transport, reconnect without replaying actions, and independent MCP client sessions.
Concurrent work across tabs, ordered commands within each tab, and session-specific cleanup barriers.
Website approvals in the extension: ask first, allow for a task, always allow, block, and revoke. Debugging has a separate grant.
Observable waits for URL, text, document readiness, visibility and enabled state.
Bounded console/error and network metadata diagnostics, plus viewport, color-scheme and reduced-motion emulation that resets on release.
Document performance traces, performance counters, and element layout/computed-style inspection, with bounded capture and chunked exports. CPU profiling reports availability explicitly; normal Chrome extensions cannot assume Profiler access.
Developer preview: releases, support and bug reports, and privacy and data handling.
See the capability and tool reference for all 39 tools, targeting examples, and limits. See security, contributing, and source release preparation for the project boundaries and development workflow.
Claimed tabs show a ๐ค title prefix. Agents can call browser_activity with tabId and activity (researching ๐, editing โ๏ธ, testing ๐งช, waiting โณ, or active for the robot alone). Change this at meaningful work phases; waiting means waiting for user input, not an automatic pause. Page title changes and navigation keep the marker while the tab is owned, and release, finish, Stop, access revocation and bridge disconnect remove it while preserving the latest page title. No extra host permissions are needed. The marker uses the existing debugger connection and is best effort: unavailable pages return an indicator warning; an unexpected debugger/worker failure expires the marker after 60 seconds when page timers run. Reconnection restores markers for still-owned, permitted tabs.
Background tabs are ordinary Chrome tabs. Collapsing a group keeps the workspace quiet; it does not make the browser invisible. Chrome displays its own debugger attachment indicator.
Related MCP server: antigravity-browser-bridge
Local setup
Requires Node.js 22.12+ (22.x, 24.x or 26+) and Chrome 125+. The current distribution is GitHub source with an unpacked Chrome extension. Chrome Web Store publication is deferred.
git clone https://github.com/dbalders/agent-browser-extension.git
cd agent-browser-extension
npm ci
npm run build
node dist/cli.js serveKeep the bridge process running. In another terminal:
node dist/cli.js pair
node dist/cli.js setupOpen
chrome://extensions, enable Developer mode, choose Load unpacked, and select this project'sdist/extensionfolder.Open the extension's connection screen. Enter the Port and Connection code printed by
pair, then click Connect.Run
node dist/cli.js statusto verify the browser connection.Add the MCP configuration printed by
setupto your agent. It uses absolute paths to your Node executable and the built CLI.
Pairing connects the extension. Website access starts in Ask first mode, including after upgrading a build that did not have website controls. Approve an exact website origin in the popup or connection screen, either for the requesting task or persistently. A denied command returns SITE_ACCESS_REQUIRED and is not replayed when approved; the agent must retry it. Agents can inspect pending requests with browser_access but cannot grant access. Individual tasks never require manual tab selection.
serve --port <port> changes the loopback port. --home <directory> or AGENT_BROWSER_HOME changes the connection directory; use the same directory for the bridge, MCP client and any HTTP adapter. On POSIX systems, this directory and its connection file must be private (700 and 600).
The default connection file is ~/.agent-browser-extension/connection.json. If you used an earlier development build, stop its bridge and pair this release again, or pass --home explicitly to both bridge and clients to reuse your chosen private connection directory. Treat the connection code as a credential: possession allows control of the connected profile through the local bridge. Do not commit it or paste it into an agent conversation. Normal bridge restarts preserve pairing. Use Disconnect in the extension to stop tasks and disable automatic reconnection.
Agent workflow
browser_status
browser_start {name: "Research suppliers"}
browser_tabs {query: "supplier"}
browser_use {tabId: ...} OR browser_open {url: "https://example.com"}
browser_snapshot {tabId: ...}
browser_click / browser_fill / browser_press / ...
browser_keep {tabId: ..., disposition: "deliverable"}
browser_finishbrowser_open defaults to background. browser_show brings a tab forward for handoff. browser_finish closes only unretained tabs created by that session; pre-existing tabs stay open. Explicit browser_close can close a claimed tab, so agents should use it only when authorized. MCP sessions are isolated and ordinary client shutdown attempts cleanup. If an agent crashes or is force-killed, use the extension's Stop control to clean up its task.
Use browser_frames to discover embedded documents, then pass a frameId to browser_snapshot or a targeted action. Fresh element references remember their frame. CSS selectors search the selected document and its open shadow roots; ambiguous matches are rejected. Use browser_select for native dropdowns, browser_type to insert text at the caret, and browser_fill to replace a field.
Actions also accept a locator instead of ref or selector, for example {role: "button", name: "Save"} or {label: "Email"}. Role/name uses Chrome's accessibility tree. Text, label, placeholder, and test-ID criteria are convenience matching for visible elements; they do not implement the full accessible-name algorithm. All target forms require a unique match. Use CSS selectors for waits involving attached/detached elements, since semantic locators find visible elements.
Use browser_read to check rendered text and current form values without writing JavaScript. Password and hidden input values are omitted. browser_check sets a requested checked state without toggling an already correct control. browser_wait can combine URL, document readiness, rendered text and target conditions; its conditions must all match.
browser_console and browser_network begin observing when the debugger first attaches to that tab. They have bounded buffers and incremental cursors; they cannot recover earlier activity. Network records omit URL credentials, query strings, fragments, headers and bodies. URL paths and console text may still contain private page data. browser_emulate changes viewport size, color scheme or reduced motion on the claimed tab and restores defaults on release/finish; it does not change user agent or provide touch emulation.
browser_profile, browser_performance and browser_inspect require a separate debugging grant. CPU profiling depends on the Chrome build: the local Chrome for Testing 148 smoke captured real samples, while a previously tested Chrome 153 build rejected Profiler commands and returned PROFILE_UNAVAILABLE. On a browser that permits profiling, captures cover the selected renderer's V8 isolate and can include other contexts sharing it. Performance traces contain the selected document's PerformanceObserver entries, not browser/GPU traces. Captures stop automatically within 30 seconds; read the returned artifact in chunks before releasing the tab. See the debugging reference for formats and scope.
For canvas or other visual controls, inspect browser_screenshot before using x/y with click, hover or drag. These coordinates are top-level viewport CSS pixels, which may differ from image pixels. Map viewport image points using viewport.width / image.width and viewport.height / image.height. For full-page images, use page.width / image.width and page.height / image.height, then subtract viewport.pageX/pageY; scroll and capture again if the target is outside the viewport. Both drag endpoints must be visible.
Page text is untrusted input. Agents must follow the user's task, respect ownership errors and avoid treating website content as instructions. Browser control includes the user's authenticated sessions and should only be exposed to trusted agents.
Architecture
Any MCP agent โโ stdio MCP CLI โโโ
โโโ authenticated loopback HTTP bridge
Any HTTP client โโโ โ
authenticated WebSocket
โ
Chrome MV3 extension
tabs + groups + CDPThe bridge listens only on 127.0.0.1. HTTP requests require a bearer token; ordinary website origins are rejected. The extension's WebSocket handshake requires a valid token and Chrome-compatible extension origin. Only one extension connection is accepted at a time. Commands carry independent session IDs, request IDs and deadlines; the effective deadline is the earlier of the caller and bridge timeouts. Queued commands that have expired are rejected. A timeout or disconnect does not establish whether a dispatched action completed: inspect the resulting page before deciding whether to repeat it. Reconnect never replays dispatched commands. The extension tracks tab ownership in session storage and refuses competing ownership claims.
There is no hosted service, model dependency, T3-specific protocol or bundled proprietary browser implementation. The code uses public Chrome extension and Chrome DevTools Protocol APIs. No raw arbitrary CDP passthrough is exposed.
Direct clients can import BrowserClient or use authenticated GET /status, GET /tools, and POST /command. The generated dist/catalog.json describes the 39 agent tools. A command body is {sessionId, operation, args, deadlineMs?}; session IDs must be unique to the trusted agent task. See src/protocol.ts for operations. Direct integrations own their session lifecycle and must call session.end when finished. A session ID is not a separate credential: clients sharing a bridge token are mutually trusted.
Both MCP and extension scheduling allow different tabs to proceed concurrently. Commands on one tab stay ordered. Start/finish and group changes form barriers for their session; opening another tab can proceed during a page wait. Short Chrome focus/group mutations serialize to keep ownership and focus coherent. Stop revokes queued work immediately. In-flight page side effects cannot be undone, and Chrome can throttle background tabs or serialize JavaScript in a shared renderer.
MCP screenshot results include actual image content alongside geometry metadata. Agent hosts must forward those images to a vision-capable model to support visual interaction; HTTP adapters are maintained separately from this extension.
Deployment branding
Forks can build and publish their own branded extension from the same source:
BROWSER_DISPLAY_NAME="Example Browser" \
BROWSER_ICON_PATH="./branding/logo.png" \
BROWSER_BRAND_COLOR="#4353d8" \
BROWSER_PRIVACY_URL="https://example.com/privacy" \
npm run buildPocket agent is the default logo and mascot. The extension includes transparent 16, 32, 48 and 128-pixel icons and uses the same mark in the popup and connection page. The selected artwork and export details are in brand assets.
Supply your own 128ร128 PNG at the icon path. The build changes the extension name, description, popup, connection page, toolbar title, icon, accent color and privacy link. A custom logo replaces the default mascot files in the package. All overrides are optional and affect only the packaged extension; a plain npm run build restores the upstream brand.
See branded builds for forks for the settings, PowerShell commands, and extension ZIP packaging. Forks can pursue their own store publication while upstream store work is deferred. The common local bridge and MCP protocol stay compatible. The independent upstream project remains dbalders/agent-browser-extension.
Current limits
Navigation and ordinary tab discovery are limited to HTTP/HTTPS. Captured agent popups can also be used while at
about:blank. No incognito, Chrome internal pages or local file navigation.Frame pointer actions support nested and positively scaled frames. Rotated, skewed, flipped or perspective-transformed frames fail safely. Selectors do not traverse closed shadow roots.
Drag uses pointer events for controls such as sliders and resize handles; native HTML5 drag-and-drop behavior is not guaranteed. Full-page captures have size limits; use viewport captures and scrolling for large pages.
Popup attribution uses a bounded window around an agent action. Delayed or uncorrelated popups remain unowned and are not closed by task cleanup. Chrome may briefly show a popup before focus is restored.
One connected Chrome profile per bridge. Separate profiles need separate bridges, ports and configuration directories.
The bridge and Chrome must run on the same machine. A remote agent server does not automatically control Chrome on the user's client computer.
Browser restarts clear extension session ownership; bridge reconnects preserve it while Chrome remains running. Force-killed agent processes can leave tasks requiring Stop.
This is an initial implementation, not a claim of feature parity with commercial browser agents. Complex websites, anti-automation challenges and enterprise browser policies may require user intervention.
Screenshot delivery depends on the client's image support and the selected model's vision capability. Local integration verification does not establish compatibility with every provider or agent surface.
Text extraction, semantic searches, screenshots, diagnostics and waits have explicit size/time bounds. Inspect truncation and dropped-record fields; narrow the target or read incremental results when needed.
Diagnostics provide observations, not request interception, network replay, full response bodies or an exported HAR. Viewport/media emulation is not a physical-device test.
Website controls gate agent operations, not a website's own scripts, redirects or network requests. Cross-origin embedded documents need their own grants; whole-tab screenshots/input conservatively require all discovered frame origins. Tab discovery still exposes titles and URLs. Use an isolated browser profile when evaluating untrusted agents.
CPU profiling is browser-dependent; unsupported builds fail explicitly with
PROFILE_UNAVAILABLE. On a compatible browser, profiles exclude workers and out-of-process frames. Document traces exclude JavaScript stacks, browser internals, GPU events and request payloads. This is not a complete DevTools replacement.
Development and verification
npm run check
npm test
npm run build
npm run check:release
npx playwright install chromium
npm run test:browser
npm run package:sourceThe real-browser smoke uses an isolated Chrome for Testing profile and local fixtures. It exercises the extension's actual connection and permission UI, discovery, concurrent MCP tasks, ownership, page targeting, form state, input, browser history, uploads/downloads, screenshots, diagnostics/emulation, trace exports, explicit CPU capability reporting, popups, shadow roots, frames, reconnection, cleanup and user Stop. It removes the disposable profile afterward. To use an existing Chrome for Testing binary, set CHROME_EXECUTABLE to its absolute executable path. It never attaches to your personal profile. The presence of a test case is not a claim that the latest candidate has passed it; retain the actual run output and source-archive digest for a release. The comparison benchmark provides shared tasks and machine-checked outcomes for separate agent runs.
Unit and integration tests cover bridge authentication, hostile origins, malformed responses, request deadlines, disconnects, ownership isolation, stale references, revocation, private credential-file handling and MCP lifecycle behavior. Generated builds, browser profiles and credentials are excluded from Git.
The standalone extension is tested with an isolated Chrome profile and MCP client. Compatibility with a specific agent application or model provider requires separate integration testing. See the 0.1.0 release notes for the verified platform scope.
The source packager builds an allowlisted archive and a per-file hash/dependency-license manifest in test-results/release/. It excludes local runtime data, dependencies, generated output and repository history, and does not execute Git commands or publish anything. See the release process for clean extraction verification and the remaining publication decisions. Hosted CI and fresh release evidence are separate from local preparation.
Licensing
Apache-2.0. Copyright 2026 David Balderston and contributors. Third-party packages retain their own licenses. This independent project is not affiliated with or endorsed by OpenAI or Google.
This server cannot be deployed
Maintenance
Related MCP Connectors
Browser MCP for logged-in tasks. Uses your Chrome โ credentials stay local. Zero-token replay.
Hosted real Google Chrome MCP with per-user persistent state. Navigate, click, type, screenshot.
Browserless MCP โ wraps the Browserless headless-Chromium REST API (browserless.io)
Access Kernel's cloud-based browsers and app actions via MCP (remote HTTP + OAuth).
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables MCP-compatible AI agents to securely control Chrome tabs via a local bridge, with explicit per-tab sharing and human approval for consequential actions.MIT
- AlicenseNot gradedqualityBmaintenanceEnables MCP-compatible agents to securely control the user's already authenticated Chrome browser via explicit tab authorization and DOM-based actions.1Apache 2.0
- AlicenseCqualityAmaintenanceMCP server for agent interaction with user-approved Google Chrome tabs, enabling tab listing/control, navigation, clicking, filling, snapshots, and screenshots via a local authenticated bridge.1002MIT
- AlicenseNot gradedqualityBmaintenanceEnables remote AI agents to control a local, authenticated desktop browser over MCP through an outbound WebSocket bridge and Chrome extension, supporting tab management, navigation, page reading, clicking, typing, scrolling, screenshots, and JavaScript evaluation.1MIT