Skip to main content
Glama
Wecko-ai

browser-relay

Official
by Wecko-ai

browser-relay

Drives the user's real, logged-in Chrome from AI agent sessions without ever handing another browser client the user's cookies (which gets sessions invalidated). An MV3 extension lives inside real Chrome and executes commands sent over a WebSocket from a local relay daemon.

Architecture

Claude session -> MCP (HTTP :9277/mcp) -> relay daemon -> WebSocket -> MV3 extension -> real Chrome tabs
  • extension/ - the Chrome MV3 extension.

  • server/relay.js - the relay daemon: MCP streamable-HTTP server + WebSocket server. Node >= 18, single dependency (ws).

  • deploy/ai.wecko.browser-relay.plist - launchd LaunchAgent template (macOS); install.sh fills in your node and repo paths.

  • install.sh - installs deps + the LaunchAgent and prints the two remaining manual steps.

  • test/ - MCP smoke test (e2e.sh), a direct tool-call helper (call.js), and a mock extension for testing the relay without Chrome.

The extension connects out to ws://127.0.0.1:9277/ws as a client (it does not run a server itself), sends a hello on open, and answers every command the relay sends with exactly one ok:true/ok:false reply.

Install

/bin/bash install.sh

This copies the LaunchAgent, (re)starts it, and waits for http://127.0.0.1:9277/health to answer. It then prints two steps you do by hand:

  1. Register the MCP server:

    claude mcp add --transport http browserx http://127.0.0.1:9277/mcp -s user
  2. Load the extension: chrome://extensions -> enable Developer mode -> Load unpacked -> select extension/.

Tools (relay commands)

command

args

result

list_tabs

{}

{tabs:[...]}

new_tab

{url?, active?}

{id,url}

close_tab

{tab_id}

{closed:true}

activate_tab

{tab_id}

{id,title,url}

navigate

{url, tab_id?}

{id,url,title}

back

{tab_id?}

{id,url}

forward

{tab_id?}

{id,url}

snapshot

{tab_id?, max_elements?, include_text?}

{tab_id,title,url,frames,elements,truncated?,text_excerpt?}

click

{ref, tab_id?}

{clicked:ref}

type_text

{ref, text, submit?, clear?}

{typed:ref}

evaluate

{js, tab_id?, all_frames?}

{result, truncated?} (capped ~6KB)

screenshot

{tab_id?, format?}

{mime,data_base64,width?,height?} (1280px max, jpeg default)

wait_for

{text, tab_id?, timeout_ms?}

{found:true,waited_ms}

tab_id is optional everywhere it appears - it falls back to the active tab of the last focused normal window.

snapshot is the main way an agent finds things to act on: it walks every frame of the page (iframes included) for interactive elements, tags each one with a stable data-bx-id="fN:eM" ref (per-frame prefix), and returns a compact list plus frame metadata. click, type_text and wait_for search all frames, so admin apps rendered in iframes (OVH, Shopify) work directly without evaluate fallbacks.

Cost guardrails (v1.1.0)

  • snapshot defaults to 150 elements per frame (hard cap 250) and the text_excerpt is opt-in via include_text.

  • evaluate output is capped at ~6KB (strings truncated, arrays/objects limited) and prefers compact JSON.

  • screenshot is downscaled to max 1280px, jpeg quality 75 by default.

  • click fires exactly one native el.click() (v1 fired 2 click events).

  • wait_for polls every 1.5s across all frames.

The relay enforces a backstop cap on every tool result, so no single call can inflate the model context.

Known limitations

  • Synthetic input is not trusted input. click and type_text dispatch real DOM events, but they are not OS-level input, so some sites (payment iframes, bot-detection-heavy forms) may reject or flag them.

  • Screenshots require the tab to be active. chrome.tabs.captureVisibleTab only captures the active tab of a window; screenshot briefly activates the target tab if it isn't already, then restores the previous one.

  • evaluate can be blocked by a page's CSP. Strict script-src policies can reject the injected eval; when that happens the error is surfaced verbatim rather than swallowed. A future version should switch to chrome.userScripts.execute in the USER_SCRIPT world, which is exempt from page CSP.

Changelog

  • v1.1.0 - cost/token optimizations: smaller snapshots, capped evaluate output, downscaled jpeg screenshots, single-click fix, iframe support for snapshot/click/type/wait.

  • v1.0.0 - initial release.

  • The MV3 service worker gets killed by Chrome. It is normal for background.js to be torn down and revived repeatedly; the bx-keepalive alarm (every ~24s) and the reconnect-on-close logic are what keep the WebSocket connection coming back. A short gap in connectivity right after Chrome starts or the SW is revived is expected.

  • Google (and similar) may challenge with reCAPTCHA or a security check. If that happens, the agent must stop and ask the human to clear it in the real browser window rather than trying to script around it.

  • Same-origin iframes are not walked by snapshot in v1.

Troubleshooting

  • curl 127.0.0.1:9277/health - confirms the relay daemon is up.

  • tail -f relay.log - relay daemon stdout/stderr (from the LaunchAgent).

  • Reload the extension from chrome://extensions if it stops responding.

  • Inspect the service worker console: chrome://extensions -> browserx relay -> "service worker" link, to see WebSocket connect/reconnect logs and any errors from injected scripts.