Selenium MCP
Allows driving a real browser through Selenium Grid, providing tools for session management, navigation, clicking, typing, pressing keys, extracting element content, executing JavaScript, and taking screenshots.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Selenium MCPOpen example.com, search for 'MCP', and screenshot the results."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
π Selenium Flow
One browser, many calls. Drive a real Chrome or Firefox on Selenium Grid from an agent over MCP β or from anything else over plain HTTP. Same actions, same server, one browser that stays exactly where you left it. π§
The whole idea, in one breath
Open a browser once. It stays alive β same page, same cookies, same scroll position β while an agent works a task one call at a time, or an n8n workflow steps node to node. Nothing relaunches, nothing logs in twice.
agent ββββ MCP /mcp ββββββΆ βββββββββββββββββ βββββββββββββββββ
β selenium-flow β ββββββΆ β Selenium Grid β βββΆ π
workflow ββββ HTTP /browser βββΆ βββββββββββββββββ βββββββββββββββββ
holds no browser the browser
lives hereThis server holds no browser. The browser lives on the Grid, and with WORKSPACE_STORE=redis the record naming it does too β so the server can restart or scale to zero without anyone losing a tab. With the default in-memory store, a restart forgets which browser was whose. πͺ
Related MCP server: Browser MCP
π Two surfaces, one implementation
Surface | For | Path |
MCP over Streamable HTTP | agents and MCP clients |
|
JSON over HTTP | n8n HTTP nodes, curl, scripts, anything |
|
Every action is a tool and an endpoint, one to one, and a test fails the build if that stops being true. They differ in one place: screenshot hands MCP an image block a vision model can see, and HTTP a base64 payload.
GET /health and GET /openapi.yaml need no credentials.
π§° Every action, both ways
Seventeen actions, each a tool and an endpoint with identical parameters β a request body is the tool's schema, with nothing added and nothing taken away.
Name your workspace on either surface:
?workspace=<name>or anX-Workspaceheader. There is no session id anywhere. See Workspaces.
Action | Endpoint | |
| Start a browser β | |
| Go to a URL π§ | |
| Click, double-click, right-click, hover, scroll to π±οΈ | |
| Drag an element onto another, or by an offset π€ | |
| Type into a field β¨οΈ | |
| Press a named key β | |
| Read text and HTML off the page π | |
| What is on the page: a checked selector each, and what works πΊοΈ | |
| JavaScript that must come back true, or the call fails β | |
| Capture a PNG, viewport or full page πΈ | |
| Keep the page as a PDF or as HTML π | |
| Run JavaScript β the escape hatch π§ͺ | |
| Enter and leave an iframe πΌοΈ | |
| Answer a native alert, confirm or prompt π¬ | |
| Change the window at any time π | |
| Attach a file to a file input π | |
| Give the slot back, keep the workspace π§Ή |
Every parameter, every return field and the traps worth knowing are one page per action in the wiki β generated from openapi.yaml, which is itself generated from the live tool schemas, so it cannot drift from the server. The same schemas are served at GET /openapi.yaml.
π¦ Chrome or Firefox
open_session(browser="firefox") and you are on Firefox; leave it out and you are on Chrome. Every other action behaves identically on both β both are plain W3C WebDriver, so only session creation differs.
π§ A workspace is not a browser
The workspace outlives the sessions it holds β a session is the live browser open in it, one at a time. When the Grid reaps an idle one, or an operator ends one, the workspace keeps the browser choice, the window and the page it was on β so recovery is one call with no arguments:
open_session() # same browser, same window, back where you wereWorkspaces expire on WORKSPACE_TTL, slid forward on every use. Nothing else removes one. More in the wiki.
π§ About that url parameter
click, write, press_key, extract, screenshot and execute_script all take an optional url, and it is not an assertion. If the browser is elsewhere it goes there first, so you can jump straight to a page instead of clicking a path to it.
πͺ Workspaces
Every workspace is named by whoever calls, on both surfaces. Say who you are and you get the browser that belongs to that name β there is no session id in the contract at all, so there is nothing to keep and nothing to pass.
How to name it | Looks like | Use it when |
A URL parameter |
| the usual case: one credential, each caller named in its own URL |
A header |
| an operator pins one workspace to one credential, or a client only sends approved headers (Claude.ai custom connectors) |
stdio | nothing to do | one process serves one client, and it is named |
The old ?session= and X-Session-Key are refused with a 400 naming the new spelling.
Sending both is a 400, not a contest one wins: two names is two ideas about who is calling, and quietly picking one hides that from whoever wired it up. Naming nothing is a 400 too, with a message saying how β except on the flow library, which falls back to the shared global one that everyone reads and nobody writes.
open_session always comes first. Nothing opens a browser implicitly, because that is the only place its browser, window size and timeouts can be chosen.
Call again with the same name β after a reconnect, a client restart, a week later β and you are back on the same browser at the same page.
β οΈ n8n opens a new MCP transport per tool call, so nothing the transport negotiates is ever the same twice. Name the workspace in the URL and it simply works.
π A workspace name is an address, not a secret: the bearer token is the credential, so anyone holding it who knows your workspace's name can drive your browser.
Browser lifetime is the Grid's (SE_NODE_SESSION_TIMEOUT, 300s here); how long a workspace is remembered is WORKSPACE_TTL, slid forward on every call. Nothing runs a cleanup loop.
π Where am I?
workspace://current reports the workspace name, which browser it is running, the page it is on, whether a session is open at all, whether you are inside a frame, and the window size. Reading it never opens a browser.
Everything there is to read is a resource like this one, and a client whose model cannot read resources β VS Code Copilot, or anything declaring ?resources=off β gets two tools instead: list_resources and read_resource(uri).
π It teaches you how to use it
The server ships an Agent Skill: the strategic half tool descriptions cannot
hold β how to name a workspace, why extract beats screenshot by orders
of magnitude, how to reach a page in one call, and what a timeout usually means.
SKILL.md is a thin index; each reference is its own resource, so an agent loads
only what its task needs:
Resource | Holds |
| the index: naming your workspace, the three rules, where next |
| naming a workspace, sharing one, recovering a dead browser |
| extract vs script vs screenshot, and durable XPath |
| forms, clicks, keys, scrolling, waiting |
| timeouts, dead sessions, blank captures |
| setting the server up, which env var to change |
| the file listing, with sizes and hashes |
Those URIs are FastMCP's convention, served by its own SkillProvider, so
list_skills and download_skill work here with no special casing.
A client that implements the MCP Skills extension also finds it through
skills/list and skills/get, and can verify every file it reads.
A client that cannot read resources reads the same URIs with read_resource.
MCP_SKILL=false turns the skill off.
π Flows β do it once, run it forever
Drive a form once, save the steps under a name, and every run after that is one call:
run_flow(name="sign-up", params={"email": "a@example.com"})A step is just a tool call, validated against the live tool schemas when it is saved β so a flow that could not run is refused before it starts. It runs in whatever browser you already hold, which means the same flow checks Chrome and then Firefox without an edit. Each workspace keeps its own library, beside a shared one called global that every workspace can run and none can change. Set DATA_DIR to turn them on; workspaces live under DATA_DIR/workspaces/<name>/.
π Secrets β typed, never shown
Mount credentials as a directory per secret and a file per key β exactly how Kubernetes already mounts a Secret β and point SECRETS_DIRS at it. An agent sees the names and keys, never a value, and binds one where the value would go:
write(selector={"css": "#password"}, secret={"name": "nextcloud", "key": "password"})The server types it; it never passes through the model, the transcript or a log. A secret can be pinned to the sites it may be used on, and is refused anywhere else.
π Files, Screenshots, Recordings and Downloads
Four sections, addressed by path, that never merge into one list:
Section | Holds | Cleared |
Downloads | whatever the site downloaded β the Grid's own store | dies with the browser |
Screenshots | every | kept, or cleared in bulk |
Recordings | the video of a browser opened with | kept, or cleared in bulk |
Files | anything | one at a time, by an operator |
keep_file(uri) moves a screenshot or recording into Files, or copies a download
there before the browser ends it. upload_file(file=uri) puts a file from any of
the four into a page's file input. No agent tool clears or deletes anything β that is an
operator action in the Admin UI below.
Read it as | URI |
a resource |
|
a resource |
|
a resource |
|
a resource |
|
a resource |
|
JSON over HTTP |
|
keep one, over HTTP |
|
Every entry carries a link to hand someone β signed and time-limited when the
server has a token, a plain path when authentication is off β because an
<img> tag cannot send an Authorization header. screenshot(save=false)
opts out when a capture is not worth keeping.
π¬ Recording
open_session(record=true) films the browser's whole life β from that call until
the browser ends or the Grid reaps it β and the video appears under
workspace://files/recordings (and in the admin page's Recordings row) shortly
after. Like screenshots, recordings stay until they are kept into Files or
cleared.
Selenium records; you deliver; selenium-flow files. Every docker-selenium
node image from 4.45.0-20260606 contains the recorder and starts it for a
session that asks with se:recordVideo. The Grid has no endpoint to download a
recording, so how the file reaches this server is yours to choose. The one rule:
Make the Grid's recordings arrive in
RECORDING_DIR(default$DATA_DIR/recordings), and setRECORDING_ENABLED=true.
Both sides need write access to that directory: the Grid's side writes, this
server moves the finished file out once it has sat unchanged for
RECORDING_SETTLE seconds, because a transport may still be finishing with it
(rclone checks an upload after writing it, and uploads one that vanished
again). Keep SE_NODE_MAX_SESSIONS=1 on recording
nodes, or recordings share one screen.
A shared volume. Mount the same directory at the nodes'
/videosand atRECORDING_DIRhere. Nothing else to configure.rclone, to anything. The recorder uploads each finished file with rclone: set
SE_UPLOAD_DESTINATION_PREFIXand anRCLONE_CONFIG_<REMOTE>_*remote on the nodes β a WebDAV such as Nextcloud whose folder is the same storage asRECORDING_DIR, S3, SFTP. Leaving--inplaceout ofSE_UPLOAD_OPTSmakes rclone write*.partialand rename, which this server waits for.A local folder for stdio or compose: bind-mount it into the Grid container at
/videos. Indocker-compose.yamlthis is opt-in, because Docker creates a missing folder as root and this image runs as 65534: runmkdir -p data/recordings && chmod -R 777 dataonce, then uncommentDATA_DIR,RECORDING_ENABLED, both volumes and the Grid'sSE_VIDEO_EVENT_DRIVEN=falseandSE_VIDEO_RECORD_STANDALONE=true, which a standalone container may need to start its recorder (not yet verified). If a standalone image's built-in recorder does not start, docker-selenium's separateselenium/videosidecar, sharing/videos, does the same job (not yet verified).
A recording is matched to its session by the Grid's session id anywhere in its
path below RECORDING_DIR (folders or file name), so any path and prefix the
transport adds is fine.
What the recorder must be set to. Keep SE_VIDEO_FILE_NAME=auto (the
default): a fixed name gives every session the same file, which cannot be
matched. And keep either SE_VIDEO_FILE_NAME_SUFFIX=true (the default) or
SE_VIDEO_SESSION_SUBFOLDER=true, so the Grid session id is in the path;
with both off a custom se:videoName carries no id, the recording is made, and
it is never matched. One that never arrives is
given up on after RECORDING_WAIT seconds with a warning in the log.
π₯ Admin UI
GET / β your workspaces, marked with the browser each is running. Open one for its Files tab β Downloads, Screenshots, Recordings and Files, each cleared the way that fits it β and its Flows tab. Click a file to view it in place; End quits a stale browser and gives its Grid slot back, rather than waiting out the Grid's idle timeout β it ends the session and keeps the workspace.
Workspaces, not Grid sessions: browsers somebody else put on the Grid are not listed. Nothing on the MCP surface lists workspaces at all β a client sees its own and nothing else. More in the wiki.
The list pushes its own updates over Server-Sent Events β no refresh button, and no polling per tab: one loop serves every page. Rows carry the name their caller claimed; a browser this server has no record of is labelled as somebody else's rather than passed off as ours.
There are no accounts: the sign-in box asks for the server's token, since anyone holding it can already drive every browser through the API. It is kept in sessionStorage, so it does not outlive the tab.
The Grid's own console is a second tab, framed same-origin. Put both behind one host β the Grid at /, this server under a path β and its live view works with no cross-origin exception.
π§© MCP Apps
Hosts implementing the MCP Apps extension β Claude, ChatGPT, VS Code, Goose β render a tool result as UI rather than JSON. show(uri) draws any resource the server serves: your workspace, its files, site data, flows, secrets and the skill.
The views are the app's own; the files grid, lightbox and browser marks are shared with the admin UI, not copied. Degradation is the point: one server, the client's capabilities pick the rendering.
The client can | It gets |
render apps |
|
read resources |
|
neither |
|
Apps get a deny-by-default CSP with no network, so PUBLIC_BASE_URL is also what admits this server's images to the frame. MCP_APPS=false turns it off.
βοΈ Configuration
Every setting can be set in a YAML config file, in the environment, or on the command line, and a later one wins: default < config < env < args. A setting's file path is its name: redis.host is REDIS_HOST and --redis-host.
# /etc/selenium-flow/config.yaml
grid:
url: http://selenium-hub:4444
redis:
host: redis.data
secrets:
entries:
nextcloud:
allowed_urls: [https://nextcloud.example.com]
keys:
password: {env: NEXTCLOUD_PASSWORD}Point at the file with --config-file or CONFIG_FILE. Prefer file: or env: for a secret's keys β value: writes the value into the config file itself.
Every setting, in all three spellings: the wiki's Configuration page.
πΎ The data directory
DATA_DIR turns on flows, kept files and recordings; unset, they are off.
$DATA_DIR/
βββ workspaces/<name>/{flows,files,screenshots,recordings}
βββ recordings/ # RECORDING_DIR's default: where the Grid's videos arriveSetting | Default | |
|
|
|
|
| the inbox the Grid's recordings arrive in |
|
| seconds after a browser ends to wait for its video |
|
|
|
|
| milliseconds between looks, when polling |
|
| seconds a finished video sits unchanged before it is filed; |
Upgrading: a :main build's $DATA_DIR/sessions/ moves itself to workspaces/ at the first boot; both present stops the boot until you merge them. Roll the upgrade out with Recreate (or scale to 0): an old pod still writing sessions/ makes the next boot stop on "both exist". With Redis, REDIS_PREFIX=selenium-flow:session: keeps the old records. From a released version (FLOW_DATA_DIR), workspaces sat at the top of the directory, and the server refuses to boot until they move β once:
cd "$DATA_DIR" && ls # the old workspace folders, and recordings/ if any
mkdir workspaces && mv <each workspace folder> workspaces/Leave recordings/ where it is, then set DATA_DIR in place of FLOW_DATA_DIR. workspaces and recordings are reserved names: an old workspace called either is refused at boot until you rename it (say workspaces-old) and move it into workspaces/.
Session defaults cascade
The browser, the window size and the two timeouts resolve in order of increasing specificity:
server default (config: session.*) < client default (?width= / X-Window-Width) < open_session argumentA bad client default is ignored and logged; a bad session.* in the config, env or args stops the boot, and an explicit browser argument fails loudly. The browser is stored with the workspace, so a reaped one reopens as the same browser.
π Auth
Setting AUTH_TOKEN turns on auth for both surfaces at once. Clients send it the usual way:
Authorization: Bearer <token>The HTTP endpoints also accept the bare token as the Authorization value, for clients that can't express a scheme. /health is always open.
It is also the signing key for file links and the event stream, so rotating it revokes those too. In the cluster External Secrets generates it.
Behind an OIDC gateway, /mcp also takes your identity provider's JWT. Set these
beside AUTH_TOKEN, and the server checks the signature, audience, expiry and role itself:
OIDC_ISSUER=https://auth.example.com/realms/example
OIDC_AUDIENCE=https://mcp.example.com
OIDC_JWKS_URI=https://auth.example.com/realms/example/protocol/openid-connect/certs
OIDC_ROLES=mcpThe admin UI can sign in with the same issuer, for a person holding an admin role β register a public client for it and add:
OIDC_CLIENT_ID=selenium-flow-admin
OIDC_ADMIN_ROLES=adminThe token keeps working beside both, and stays the admin's.
π Running it
docker compose up --buildThat starts the server and a standalone Grid for it to drive, with auth off:
curl -X POST localhost:8000/browser \
-H 'X-Workspace: demo' -H 'Content-Type: application/json' \
-d '{"url":"https://example.com","width":1280,"height":800}'Point an MCP client at http://localhost:8000/mcp, open localhost:8000 β and watch the browser work live at localhost:7900, the Grid's noVNC view. π
π Contributing
Setup, tests, how the OpenAPI spec is generated and how the embedded skill is packaged: CONTRIBUTING.md. Design rules worth reading before changing behaviour: AGENTS.md.
π References
π Licence
MIT. See LICENSE.
Not affiliated with, endorsed by, or sponsored by the Selenium project. "Selenium" is a trademark of the Software Freedom Conservancy, used here only to identify the software this server drives.
This server cannot be deployed
Maintenance
Related MCP Connectors
Hosted real Google Chrome MCP with per-user persistent state. Navigate, click, type, screenshot.
Stealth web browser for agents: search, fetch, click, download and type in persistent MCP sessions.
- TabfleetOAuthcom.tabfleet
Launch, inspect, control, and share isolated cloud browsers for your agents.
Undetectable cloud browser sessions for AI agents and scrapers. Navigate, extract, click, captcha.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables MCP clients to drive a real, logged-in Chrome browser for web automation tasks like navigation, clicking, typing, and screenshotting.2 npmMIT
- FlicenseNot gradedqualityDmaintenanceEnables MCP clients to control a real local browser window for web automation tasks such as clicking, typing, scrolling, and taking screenshots.16 npm-
- AlicenseAqualityCmaintenanceEnables AI agents and MCP clients to automate web browsers via Selenium WebDriver, supporting Chrome, Firefox, and Edge in headless or visible mode with tools for navigation, interaction, content extraction, screenshots, and scripting.2124 npmMIT
- AlicenseBqualityCmaintenanceEnables MCP clients to drive a real Chromium browser for automation, including navigation, JavaScript execution, CDP commands, network capture, and multi-tab control.7GPL 3.0