Skip to main content
Glama

🌊 Selenium Flow

One browser, many calls. Drive a real Chrome or Firefox on Selenium Grid from an agent over MCP β€” or from anything else over plain HTTP. Same actions, same server, one browser that stays exactly where you left it. 🧭

πŸ§ͺ Test πŸ›‘οΈ Quality πŸ“Έ Image Builder πŸ“– Wiki License: MIT Docker FastMCP


The whole idea, in one breath

Open a browser once. It stays alive β€” same page, same cookies, same scroll position β€” while an agent works a task one call at a time, or an n8n workflow steps node to node. Nothing relaunches, nothing logs in twice.

   agent  ──── MCP  /mcp ─────▢  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                                 β”‚ selenium-flow β”‚ ─────▢ β”‚ Selenium Grid β”‚ ──▢ 🌐
workflow  ──── HTTP /browser ──▢ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                  holds no browser          the browser
                                                            lives here

This server holds no browser. The browser lives on the Grid, and with WORKSPACE_STORE=redis the record naming it does too β€” so the server can restart or scale to zero without anyone losing a tab. With the default in-memory store, a restart forgets which browser was whose. πŸͺ„


Related MCP server: Browser MCP

πŸ”€ Two surfaces, one implementation

Surface

For

Path

MCP over Streamable HTTP

agents and MCP clients

/mcp

JSON over HTTP

n8n HTTP nodes, curl, scripts, anything

/browser/*

Every action is a tool and an endpoint, one to one, and a test fails the build if that stops being true. They differ in one place: screenshot hands MCP an image block a vision model can see, and HTTP a base64 payload.

GET /health and GET /openapi.yaml need no credentials.


🧰 Every action, both ways

Seventeen actions, each a tool and an endpoint with identical parameters β€” a request body is the tool's schema, with nothing added and nothing taken away.

Name your workspace on either surface: ?workspace=<name> or an X-Workspace header. There is no session id anywhere. See Workspaces.

Action

Endpoint

open_session

POST /browser

Start a browser β€” chrome or firefox πŸš€

navigate

POST /browser/navigate

Go to a URL 🧭

interact

POST /browser/interact/{action}

Click, double-click, right-click, hover, scroll to πŸ–±οΈ

drag

POST /browser/drag

Drag an element onto another, or by an offset 🀏

write

POST /browser/write

Type into a field ⌨️

press_key

POST /browser/press-key

Press a named key β€” tab, enter, arrows 🎹

extract

POST /browser/extract

Read text and HTML off the page πŸ“–

outline

POST /browser/outline

What is on the page: a checked selector each, and what works πŸ—ΊοΈ

assert

POST /browser/assert

JavaScript that must come back true, or the call fails βœ…

screenshot

POST /browser/screenshot

Capture a PNG, viewport or full page πŸ“Έ

print

POST /browser/print

Keep the page as a PDF or as HTML πŸ“„

execute_script

POST /browser/script

Run JavaScript β€” the escape hatch πŸ§ͺ

frame

POST /browser/frame

Enter and leave an iframe πŸ–ΌοΈ

dialog

POST /browser/dialog

Answer a native alert, confirm or prompt πŸ’¬

resize

POST /browser/resize

Change the window at any time πŸ“

upload_file

POST /browser/upload

Attach a file to a file input πŸ“Ž

end_browser

DELETE /browser

Give the slot back, keep the workspace 🧹

Every parameter, every return field and the traps worth knowing are one page per action in the wiki β€” generated from openapi.yaml, which is itself generated from the live tool schemas, so it cannot drift from the server. The same schemas are served at GET /openapi.yaml.

🦊 Chrome or Firefox

open_session(browser="firefox") and you are on Firefox; leave it out and you are on Chrome. Every other action behaves identically on both β€” both are plain W3C WebDriver, so only session creation differs.

🧠 A workspace is not a browser

The workspace outlives the sessions it holds β€” a session is the live browser open in it, one at a time. When the Grid reaps an idle one, or an operator ends one, the workspace keeps the browser choice, the window and the page it was on β€” so recovery is one call with no arguments:

open_session()      # same browser, same window, back where you were

Workspaces expire on WORKSPACE_TTL, slid forward on every use. Nothing else removes one. More in the wiki.

🧭 About that url parameter

click, write, press_key, extract, screenshot and execute_script all take an optional url, and it is not an assertion. If the browser is elsewhere it goes there first, so you can jump straight to a page instead of clicking a path to it.


πŸͺ Workspaces

Every workspace is named by whoever calls, on both surfaces. Say who you are and you get the browser that belongs to that name β€” there is no session id in the contract at all, so there is nothing to keep and nothing to pass.

How to name it

Looks like

Use it when

A URL parameter

…/mcp?workspace=research-bot

the usual case: one credential, each caller named in its own URL

A header

X-Workspace: research-bot

an operator pins one workspace to one credential, or a client only sends approved headers (Claude.ai custom connectors)

stdio

nothing to do

one process serves one client, and it is named stdio

The old ?session= and X-Session-Key are refused with a 400 naming the new spelling.

Sending both is a 400, not a contest one wins: two names is two ideas about who is calling, and quietly picking one hides that from whoever wired it up. Naming nothing is a 400 too, with a message saying how β€” except on the flow library, which falls back to the shared global one that everyone reads and nobody writes.

open_session always comes first. Nothing opens a browser implicitly, because that is the only place its browser, window size and timeouts can be chosen.

Call again with the same name β€” after a reconnect, a client restart, a week later β€” and you are back on the same browser at the same page.

⚠️ n8n opens a new MCP transport per tool call, so nothing the transport negotiates is ever the same twice. Name the workspace in the URL and it simply works.

πŸ”‘ A workspace name is an address, not a secret: the bearer token is the credential, so anyone holding it who knows your workspace's name can drive your browser.

Browser lifetime is the Grid's (SE_NODE_SESSION_TIMEOUT, 300s here); how long a workspace is remembered is WORKSPACE_TTL, slid forward on every call. Nothing runs a cleanup loop.

πŸ“ Where am I?

workspace://current reports the workspace name, which browser it is running, the page it is on, whether a session is open at all, whether you are inside a frame, and the window size. Reading it never opens a browser.

Everything there is to read is a resource like this one, and a client whose model cannot read resources β€” VS Code Copilot, or anything declaring ?resources=off β€” gets two tools instead: list_resources and read_resource(uri).

πŸ“– It teaches you how to use it

The server ships an Agent Skill: the strategic half tool descriptions cannot hold β€” how to name a workspace, why extract beats screenshot by orders of magnitude, how to reach a page in one call, and what a timeout usually means.

SKILL.md is a thin index; each reference is its own resource, so an agent loads only what its task needs:

Resource

Holds

skill://selenium-flow/SKILL.md

the index: naming your workspace, the three rules, where next

.../references/WORKSPACES.md

naming a workspace, sharing one, recovering a dead browser

.../references/READING_PAGES.md

extract vs script vs screenshot, and durable XPath

.../references/INTERACTION.md

forms, clicks, keys, scrolling, waiting

.../references/TROUBLESHOOTING.md

timeouts, dead sessions, blank captures

.../references/CONFIGURATION.md

setting the server up, which env var to change

skill://selenium-flow/_manifest

the file listing, with sizes and hashes

Those URIs are FastMCP's convention, served by its own SkillProvider, so list_skills and download_skill work here with no special casing.

A client that implements the MCP Skills extension also finds it through skills/list and skills/get, and can verify every file it reads.

A client that cannot read resources reads the same URIs with read_resource. MCP_SKILL=false turns the skill off.


πŸ” Flows β€” do it once, run it forever

Drive a form once, save the steps under a name, and every run after that is one call:

run_flow(name="sign-up", params={"email": "a@example.com"})

A step is just a tool call, validated against the live tool schemas when it is saved β€” so a flow that could not run is refused before it starts. It runs in whatever browser you already hold, which means the same flow checks Chrome and then Firefox without an edit. Each workspace keeps its own library, beside a shared one called global that every workspace can run and none can change. Set DATA_DIR to turn them on; workspaces live under DATA_DIR/workspaces/<name>/.

πŸ” Secrets β€” typed, never shown

Mount credentials as a directory per secret and a file per key β€” exactly how Kubernetes already mounts a Secret β€” and point SECRETS_DIRS at it. An agent sees the names and keys, never a value, and binds one where the value would go:

write(selector={"css": "#password"}, secret={"name": "nextcloud", "key": "password"})

The server types it; it never passes through the model, the transcript or a log. A secret can be pinned to the sites it may be used on, and is refused anywhere else.

πŸ—‚ Files, Screenshots, Recordings and Downloads

Four sections, addressed by path, that never merge into one list:

Section

Holds

Cleared

Downloads

whatever the site downloaded β€” the Grid's own store

dies with the browser

Screenshots

every screenshot, from the moment it is taken

kept, or cleared in bulk

Recordings

the video of a browser opened with record=true

kept, or cleared in bulk

Files

anything keep_file'd, and every print

one at a time, by an operator

keep_file(uri) moves a screenshot or recording into Files, or copies a download there before the browser ends it. upload_file(file=uri) puts a file from any of the four into a page's file input. No agent tool clears or deletes anything β€” that is an operator action in the Admin UI below.

Read it as

URI

a resource

workspace://files β€” Files, plus the three folders below

a resource

workspace://files/{name} β€” one kept file, as bytes (a video: its entry)

a resource

workspace://files/screenshots, .../screenshots/{name}

a resource

workspace://files/recordings, .../recordings/{name} β€” an entry, never the video

a resource

workspace://files/downloads, .../downloads/{name}

JSON over HTTP

GET /files, GET /files/screenshots, GET /files/recordings, GET /files/downloads

keep one, over HTTP

PUT /files/{folder}/{name}/kept β€” screenshots, recordings or downloads

Every entry carries a link to hand someone β€” signed and time-limited when the server has a token, a plain path when authentication is off β€” because an <img> tag cannot send an Authorization header. screenshot(save=false) opts out when a capture is not worth keeping.


🎬 Recording

open_session(record=true) films the browser's whole life β€” from that call until the browser ends or the Grid reaps it β€” and the video appears under workspace://files/recordings (and in the admin page's Recordings row) shortly after. Like screenshots, recordings stay until they are kept into Files or cleared.

Selenium records; you deliver; selenium-flow files. Every docker-selenium node image from 4.45.0-20260606 contains the recorder and starts it for a session that asks with se:recordVideo. The Grid has no endpoint to download a recording, so how the file reaches this server is yours to choose. The one rule:

Make the Grid's recordings arrive in RECORDING_DIR (default $DATA_DIR/recordings), and set RECORDING_ENABLED=true.

Both sides need write access to that directory: the Grid's side writes, this server moves the finished file out once it has sat unchanged for RECORDING_SETTLE seconds, because a transport may still be finishing with it (rclone checks an upload after writing it, and uploads one that vanished again). Keep SE_NODE_MAX_SESSIONS=1 on recording nodes, or recordings share one screen.

  • A shared volume. Mount the same directory at the nodes' /videos and at RECORDING_DIR here. Nothing else to configure.

  • rclone, to anything. The recorder uploads each finished file with rclone: set SE_UPLOAD_DESTINATION_PREFIX and an RCLONE_CONFIG_<REMOTE>_* remote on the nodes β€” a WebDAV such as Nextcloud whose folder is the same storage as RECORDING_DIR, S3, SFTP. Leaving --inplace out of SE_UPLOAD_OPTS makes rclone write *.partial and rename, which this server waits for.

  • A local folder for stdio or compose: bind-mount it into the Grid container at /videos. In docker-compose.yaml this is opt-in, because Docker creates a missing folder as root and this image runs as 65534: run mkdir -p data/recordings && chmod -R 777 data once, then uncomment DATA_DIR, RECORDING_ENABLED, both volumes and the Grid's SE_VIDEO_EVENT_DRIVEN=false and SE_VIDEO_RECORD_STANDALONE=true, which a standalone container may need to start its recorder (not yet verified). If a standalone image's built-in recorder does not start, docker-selenium's separate selenium/video sidecar, sharing /videos, does the same job (not yet verified).

A recording is matched to its session by the Grid's session id anywhere in its path below RECORDING_DIR (folders or file name), so any path and prefix the transport adds is fine.

What the recorder must be set to. Keep SE_VIDEO_FILE_NAME=auto (the default): a fixed name gives every session the same file, which cannot be matched. And keep either SE_VIDEO_FILE_NAME_SUFFIX=true (the default) or SE_VIDEO_SESSION_SUBFOLDER=true, so the Grid session id is in the path; with both off a custom se:videoName carries no id, the recording is made, and it is never matched. One that never arrives is given up on after RECORDING_WAIT seconds with a warning in the log.


πŸ–₯ Admin UI

GET / β€” your workspaces, marked with the browser each is running. Open one for its Files tab β€” Downloads, Screenshots, Recordings and Files, each cleared the way that fits it β€” and its Flows tab. Click a file to view it in place; End quits a stale browser and gives its Grid slot back, rather than waiting out the Grid's idle timeout β€” it ends the session and keeps the workspace.

Workspaces, not Grid sessions: browsers somebody else put on the Grid are not listed. Nothing on the MCP surface lists workspaces at all β€” a client sees its own and nothing else. More in the wiki.

The list pushes its own updates over Server-Sent Events β€” no refresh button, and no polling per tab: one loop serves every page. Rows carry the name their caller claimed; a browser this server has no record of is labelled as somebody else's rather than passed off as ours.

There are no accounts: the sign-in box asks for the server's token, since anyone holding it can already drive every browser through the API. It is kept in sessionStorage, so it does not outlive the tab.

The Grid's own console is a second tab, framed same-origin. Put both behind one host β€” the Grid at /, this server under a path β€” and its live view works with no cross-origin exception.


🧩 MCP Apps

Hosts implementing the MCP Apps extension β€” Claude, ChatGPT, VS Code, Goose β€” render a tool result as UI rather than JSON. show(uri) draws any resource the server serves: your workspace, its files, site data, flows, secrets and the skill.

The views are the app's own; the files grid, lightbox and browser marks are shared with the admin UI, not copied. Degradation is the point: one server, the client's capabilities pick the rendering.

The client can

It gets

render apps

show, and the view inline

read resources

workspace://files, flow://flows, and the file as bytes

neither

read_resource / list_resources, the same JSON, and links anything can open

Apps get a deny-by-default CSP with no network, so PUBLIC_BASE_URL is also what admits this server's images to the frame. MCP_APPS=false turns it off.


βš™οΈ Configuration

Every setting can be set in a YAML config file, in the environment, or on the command line, and a later one wins: default < config < env < args. A setting's file path is its name: redis.host is REDIS_HOST and --redis-host.

# /etc/selenium-flow/config.yaml
grid:
  url: http://selenium-hub:4444
redis:
  host: redis.data
secrets:
  entries:
    nextcloud:
      allowed_urls: [https://nextcloud.example.com]
      keys:
        password: {env: NEXTCLOUD_PASSWORD}

Point at the file with --config-file or CONFIG_FILE. Prefer file: or env: for a secret's keys β€” value: writes the value into the config file itself.

Every setting, in all three spellings: the wiki's Configuration page.

πŸ’Ύ The data directory

DATA_DIR turns on flows, kept files and recordings; unset, they are off.

$DATA_DIR/
β”œβ”€β”€ workspaces/<name>/{flows,files,screenshots,recordings}
└── recordings/          # RECORDING_DIR's default: where the Grid's videos arrive

Setting

Default

RECORDING_ENABLED

false

open_session(record=true) works; needs DATA_DIR

RECORDING_DIR

$DATA_DIR/recordings

the inbox the Grid's recordings arrive in

RECORDING_WAIT

600

seconds after a browser ends to wait for its video

RECORDING_WATCH

auto

events, poll, or auto (polls on a network filesystem)

RECORDING_POLL

1000

milliseconds between looks, when polling

RECORDING_SETTLE

10

seconds a finished video sits unchanged before it is filed; 0 files it at once

Upgrading: a :main build's $DATA_DIR/sessions/ moves itself to workspaces/ at the first boot; both present stops the boot until you merge them. Roll the upgrade out with Recreate (or scale to 0): an old pod still writing sessions/ makes the next boot stop on "both exist". With Redis, REDIS_PREFIX=selenium-flow:session: keeps the old records. From a released version (FLOW_DATA_DIR), workspaces sat at the top of the directory, and the server refuses to boot until they move β€” once:

cd "$DATA_DIR" && ls             # the old workspace folders, and recordings/ if any
mkdir workspaces && mv <each workspace folder> workspaces/

Leave recordings/ where it is, then set DATA_DIR in place of FLOW_DATA_DIR. workspaces and recordings are reserved names: an old workspace called either is refused at boot until you rename it (say workspaces-old) and move it into workspaces/.

Session defaults cascade

The browser, the window size and the two timeouts resolve in order of increasing specificity:

server default (config: session.*)  <  client default (?width= / X-Window-Width)  <  open_session argument

A bad client default is ignored and logged; a bad session.* in the config, env or args stops the boot, and an explicit browser argument fails loudly. The browser is stored with the workspace, so a reaped one reopens as the same browser.

πŸ” Auth

Setting AUTH_TOKEN turns on auth for both surfaces at once. Clients send it the usual way:

Authorization: Bearer <token>

The HTTP endpoints also accept the bare token as the Authorization value, for clients that can't express a scheme. /health is always open.

It is also the signing key for file links and the event stream, so rotating it revokes those too. In the cluster External Secrets generates it.

Behind an OIDC gateway, /mcp also takes your identity provider's JWT. Set these beside AUTH_TOKEN, and the server checks the signature, audience, expiry and role itself:

OIDC_ISSUER=https://auth.example.com/realms/example
OIDC_AUDIENCE=https://mcp.example.com
OIDC_JWKS_URI=https://auth.example.com/realms/example/protocol/openid-connect/certs
OIDC_ROLES=mcp

The admin UI can sign in with the same issuer, for a person holding an admin role β€” register a public client for it and add:

OIDC_CLIENT_ID=selenium-flow-admin
OIDC_ADMIN_ROLES=admin

The token keeps working beside both, and stays the admin's.


πŸš€ Running it

docker compose up --build

That starts the server and a standalone Grid for it to drive, with auth off:

curl -X POST localhost:8000/browser \
  -H 'X-Workspace: demo' -H 'Content-Type: application/json' \
  -d '{"url":"https://example.com","width":1280,"height":800}'

Point an MCP client at http://localhost:8000/mcp, open localhost:8000 β€” and watch the browser work live at localhost:7900, the Grid's noVNC view. πŸ‘€


πŸ›  Contributing

Setup, tests, how the OpenAPI spec is generated and how the embedded skill is packaged: CONTRIBUTING.md. Design rules worth reading before changing behaviour: AGENTS.md.


πŸ”— References


πŸ“œ Licence

MIT. See LICENSE.

Not affiliated with, endorsed by, or sponsored by the Selenium project. "Selenium" is a trademark of the Software Freedom Conservancy, used here only to identify the software this server drives.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables MCP clients to drive a real, logged-in Chrome browser for web automation tasks like navigation, clicking, typing, and screenshotting.
    2 npm
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables MCP clients to control a real local browser window for web automation tasks such as clicking, typing, scrolling, and taking screenshots.
    16 npm
    -
  • A
    license
    A
    quality
    C
    maintenance
    Enables AI agents and MCP clients to automate web browsers via Selenium WebDriver, supporting Chrome, Firefox, and Edge in headless or visible mode with tools for navigation, interaction, content extraction, screenshots, and scripting.
    21
    24 npm
    MIT
  • A
    license
    B
    quality
    C
    maintenance
    Enables MCP clients to drive a real Chromium browser for automation, including navigation, JavaScript execution, CDP commands, network capture, and multi-tab control.
    7
    GPL 3.0