Skip to main content
Glama
Wecko-ai

browser-relay

Official
by Wecko-ai
README.md
# browser-relay

Drives the user's real, logged-in Chrome from AI agent sessions without ever
handing another browser client the user's cookies (which gets sessions
invalidated). An MV3 extension lives inside real Chrome and executes
commands sent over a WebSocket from a local relay daemon.

## Architecture

```
Claude session -> MCP (HTTP :9277/mcp) -> relay daemon -> WebSocket -> MV3 extension -> real Chrome tabs
```

- `extension/` - the Chrome MV3 extension.
- `server/relay.js` - the relay daemon: MCP streamable-HTTP server + WebSocket server. Node >= 18, single dependency (`ws`).
- `deploy/ai.wecko.browser-relay.plist` - launchd LaunchAgent template (macOS); `install.sh` fills in your node and repo paths.
- `install.sh` - installs deps + the LaunchAgent and prints the two remaining manual steps.
- `test/` - MCP smoke test (`e2e.sh`), a direct tool-call helper (`call.js`), and a mock extension for testing the relay without Chrome.

The extension connects out to `ws://127.0.0.1:9277/ws` as a client (it does
not run a server itself), sends a `hello` on open, and answers every command
the relay sends with exactly one `ok:true`/`ok:false` reply.

## Install

```
/bin/bash install.sh
```

This copies the LaunchAgent, (re)starts it, and waits for
`http://127.0.0.1:9277/health` to answer. It then prints two steps you do by
hand:

1. Register the MCP server:
   ```
   claude mcp add --transport http browserx http://127.0.0.1:9277/mcp -s user
   ```
2. Load the extension: `chrome://extensions` -> enable Developer mode ->
   Load unpacked -> select `extension/`.

## Tools (relay commands)

| command | args | result |
|---|---|---|
| `list_tabs` | `{}` | `{tabs:[...]}` |
| `new_tab` | `{url?, active?}` | `{id,url}` |
| `close_tab` | `{tab_id}` | `{closed:true}` |
| `activate_tab` | `{tab_id}` | `{id,title,url}` |
| `navigate` | `{url, tab_id?}` | `{id,url,title}` |
| `back` | `{tab_id?}` | `{id,url}` |
| `forward` | `{tab_id?}` | `{id,url}` |
| `snapshot` | `{tab_id?, max_elements?, include_text?}` | `{tab_id,title,url,frames,elements,truncated?,text_excerpt?}` |
| `click` | `{ref, tab_id?}` | `{clicked:ref}` |
| `type_text` | `{ref, text, submit?, clear?}` | `{typed:ref}` |
| `evaluate` | `{js, tab_id?, all_frames?}` | `{result, truncated?}` (capped ~6KB) |
| `screenshot` | `{tab_id?, format?}` | `{mime,data_base64,width?,height?}` (1280px max, jpeg default) |
| `wait_for` | `{text, tab_id?, timeout_ms?}` | `{found:true,waited_ms}` |

`tab_id` is optional everywhere it appears - it falls back to the active tab
of the last focused normal window.

`snapshot` is the main way an agent finds things to act on: it walks every
frame of the page (iframes included) for interactive elements, tags each one
with a stable `data-bx-id="fN:eM"` ref (per-frame prefix), and returns a
compact list plus frame metadata. `click`, `type_text` and `wait_for` search
all frames, so admin apps rendered in iframes (OVH, Shopify) work directly
without `evaluate` fallbacks.

## Cost guardrails (v1.1.0)

- `snapshot` defaults to 150 elements per frame (hard cap 250) and the
  `text_excerpt` is opt-in via `include_text`.
- `evaluate` output is capped at ~6KB (strings truncated, arrays/objects
  limited) and prefers compact JSON.
- `screenshot` is downscaled to max 1280px, jpeg quality 75 by default.
- `click` fires exactly one native `el.click()` (v1 fired 2 click events).
- `wait_for` polls every 1.5s across all frames.

The relay enforces a backstop cap on every tool result, so no single call can
inflate the model context.

## Known limitations

- **Synthetic input is not trusted input.** `click` and `type_text` dispatch
  real DOM events, but they are not OS-level input, so some sites (payment
  iframes, bot-detection-heavy forms) may reject or flag them.
- **Screenshots require the tab to be active.** `chrome.tabs.captureVisibleTab`
  only captures the active tab of a window; `screenshot` briefly activates
  the target tab if it isn't already, then restores the previous one.
- **`evaluate` can be blocked by a page's CSP.** Strict `script-src` policies
  can reject the injected eval; when that happens the error is surfaced
  verbatim rather than swallowed. A future version should switch to
  `chrome.userScripts.execute` in the `USER_SCRIPT` world, which is exempt
  from page CSP.

## Changelog

- **v1.1.0** - cost/token optimizations: smaller snapshots, capped `evaluate`
  output, downscaled jpeg screenshots, single-click fix, iframe support for
  snapshot/click/type/wait.
- **v1.0.0** - initial release.
- **The MV3 service worker gets killed by Chrome.** It is normal for
  `background.js` to be torn down and revived repeatedly; the `bx-keepalive`
  alarm (every ~24s) and the reconnect-on-close logic are what keep the
  WebSocket connection coming back. A short gap in connectivity right after
  Chrome starts or the SW is revived is expected.
- **Google (and similar) may challenge with reCAPTCHA or a security check.**
  If that happens, the agent must stop and ask the human to clear it in the
  real browser window rather than trying to script around it.
- Same-origin iframes are not walked by `snapshot` in v1.

## Troubleshooting

- `curl 127.0.0.1:9277/health` - confirms the relay daemon is up.
- `tail -f relay.log` - relay daemon stdout/stderr (from the LaunchAgent).
- Reload the extension from `chrome://extensions` if it stops responding.
- Inspect the service worker console: `chrome://extensions` -> browserx
  relay -> "service worker" link, to see WebSocket connect/reconnect logs
  and any errors from injected scripts.