webcontrol-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@webcontrol-mcpGo to staging.example.com, log in as test@example.com, and screenshot the dashboard."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
webcontrol-mcp
Let Claude control any website — open pages, click, fill forms, grab text and screenshots. Faster than playwright-mcp.
I'm a systems engineer and I built this to let Claude drive my test sites: spin up a staging server, point the model at it, and let it click through the flows I would otherwise check by hand. It speaks plain Model Context Protocol over stdio, so it drops straight into Claude Desktop, Cursor, or any MCP client.
Why another browser MCP?
Most browser-control MCPs spin up heavyweight automation stacks with long cold-start times. webcontrol-mcp is a single small Node process with five focused tools, lazy-loads Playwright, and starts in well under a second. On Windows, the optional native helper (below) makes screenshots up to 5x faster.
Related MCP server: WebOperator MCP Server
Quickstart
npx webcontrol-mcpOr install globally:
npm install -g webcontrol-mcp
webcontrol-mcpFirst run on a fresh machine: install the Chromium browser once with npx playwright install chromium.
Tools
Tool | Description |
| Open a URL and wait for it to load. Returns the page title. |
| Click an element by CSS selector. |
| Type text into an input field. |
| Capture the current page as PNG. |
| Read visible text from the page or a selected element. |
Claude Desktop config
Add to claude_desktop_config.json:
{
"mcpServers": {
"webcontrol": {
"command": "npx",
"args": ["-y", "webcontrol-mcp"]
}
}
}Cursor config
Add to .cursor/mcp.json:
{
"mcpServers": {
"webcontrol": {
"command": "npx",
"args": ["-y", "webcontrol-mcp"]
}
}
}Optional: native helper (Windows)
On Windows you can install a small native capture driver that speeds up screenshots up to 5x:
Download
helper-setup.exefrom the v1.0.0 release.Run it. If Windows SmartScreen shows a warning (the binary isn't broadly distributed yet), choose More info → Run anyway.
Restart
webcontrol-mcp.
The helper is entirely optional — the MCP server works fine without it (pure-Node screenshots are just slower). macOS and Linux users can ignore it completely.
Troubleshooting
Playwright is not installed— runnpm installin the package directory, or install globally.Could not launch Chromium— runnpx playwright install chromiumonce to download the browser.Port/firewall prompts — the server drives a local headless Chromium; no inbound ports are needed.
License
Available Tools
5 toolsclickB
Click an element on the current page, located by a CSS selector.
| Name | Required | Description | Default |
|---|---|---|---|
| selector | Yes | CSS selector of the element to click |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It says nothing about what happens after the click (navigation, waits, page load), whether the element is scrolled into view, or what error occurs if the selector matches nothing – all material for a browser-automation click.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no waste. The action and its targeting mechanism are stated immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter interaction tool with full schema coverage, the definition is minimally adequate, but it omits the post-click behavior and failure modes that an agent would need to sequence follow-up calls correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single parameter is already documented as a CSS selector. The description merely restates that, adding no format or edge-case detail beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (click) plus resource (an element on the current page) and the locating mechanism (CSS selector). It is clearly distinguishable from fill/get_text/screenshot in practice, but it never names or contrasts those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus fill, get_text, or other interaction tools, and no prerequisites or exclusions. The phrase 'on the current page' weakly implies a page must already be open, but nothing is stated outright.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fillC
Type text into an input field on the current page.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Text to type into the field | |
| selector | Yes | CSS selector of the input element |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not say whether existing text is cleared or appended, whether input/keyboard events are dispatched (critical for React-style frameworks), whether the element must be focused first, or what happens if the selector matches nothing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or redundancy. It is efficient, though extremely sparse given the behavioral details that are omitted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with full schema coverage and no output schema, the core action is conveyed. However, with no annotations and no output description, behavioral details such as event dispatch and handling of pre-existing content are left entirely unstated, leaving real gaps for a mutation-style UI tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both selector and text are already documented in the schema. The description adds no extra semantics such as selector syntax expectations or text escaping/formatting behavior, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (type text) and resource (an input field on the current page), which is enough to distinguish it from siblings like click, get_text, and screenshot. It stops short of explicitly contrasting itself with any named alternative, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, nor any prerequisites such as needing the page open or the element visible. The agent must infer all usage context from the tool name and siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_textB
Get the visible text content of the page, or of an element matching a CSS selector.
| Name | Required | Description | Default |
|---|---|---|---|
| selector | No | Optional CSS selector to scope the text extraction |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full behavioral burden. It discloses one useful trait, 'visible' text (implying hidden/script content is excluded), but says nothing about return format, error behavior when a selector matches nothing, or whether it waits for page load. Substantial behavioral gaps remain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with the core action first and the optional scoping variant second. No filler and every clause carries meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool with no annotations and no output schema, the description covers the essential purpose. It is not fully complete, though: the return shape and behavior on no-match or non-visible content are unaddressed, which an agent might need.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single selector parameter is documented in the schema as an optional CSS selector to scope extraction. The description reinforces this ('element matching a CSS selector') but adds no syntax or format detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Get the visible text content of the page' plus an element-scoped variant. It is clearly distinguishable from siblings like click, fill, and screenshot. It stops short of explicitly contrasting itself with any sibling, so it does not reach the top band.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'or of an element matching a CSS selector' implies the two usage modes (whole-page vs scoped), so the alternative paths are implied rather than stated. There is no explicit guidance on when to prefer this over screenshot or open_page, and no exclusions. Adequate but leaves inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_pageA
Open a URL in the controlled browser and wait for it to load. Returns the page title.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to open, e.g. https://example.com |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does add real behavioral context: it blocks until load completes and returns the page title. It omits important traits for a browser navigation tool, such as failure/timeout behavior and whether the current page state is replaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action, with no filler. Every clause (browser context, wait-for-load, return value) earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with no output schema, the definition is nearly sufficient: it names the action, the wait semantics, and the return value. The remaining gap is error/timeout and navigation-side-effect behavior, which matters slightly given the absence of annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema description coverage, including a concrete URL example, so the schema already does the work. The description adds nothing about URL form or validation beyond what the schema provides; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Open) and resource (a URL in the controlled browser), plus the blocking behavior ('wait for it to load'). Sibling tools (click, fill, screenshot, get_text) are action-specific and clearly distinct, so an agent can tell them apart, though the description never explicitly positions itself relative to them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: 'in the controlled browser' signals this is the navigation entry point preceding the other browser tools. There is no explicit statement of when to use this vs. the siblings or any prerequisite, so the guidance stays at the implied level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
screenshotB
Take a PNG screenshot of the current page. Returns the image as base64.
| Name | Required | Description | Default |
|---|---|---|---|
| fullPage | No | Capture the full scrollable page (default false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It helpfully discloses the return format (base64 PNG), which is real value, but omits whether a page must already be loaded, output size limits, or how fullPage changes behavior. Partial disclosure, not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero waste, with the action stated first and the return format second. Nothing could be trimmed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, no-output-schema tool, the description covers the essential action and return format, and the schema covers the parameter. It could note the fullPage capability or page-loaded precondition, but it is nearly complete for this tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single fullPage parameter is fully documented in the schema, so the baseline is 3. The description adds nothing about fullPage or viewport/scroll behavior beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Take) and resource (PNG screenshot of the current page), so an agent knows exactly what the tool produces. It does not explicitly differentiate itself from siblings like get_text (which also reads page content), but the visual-capture purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to prefer a screenshot over get_text or when a full-page capture is warranted. The phrase 'current page' implies a page must be open, but no prerequisites, alternatives, or exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v1.0.0- First observed
click - First observed
fill - First observed
get_text - First observed
open_page - First observed
screenshot
TDQS
Scored across 5 tools
Each tool maps to a single, clearly distinct browser action (navigate, click, type, capture, read), with no semantic overlap between them. An agent can unambiguously pick the right tool for any given step.
Most tools follow a verb_noun/verb-object pattern (open_page, get_text) or a standard single-action verb (click, fill), which is readable and predictable. The lone noun-form 'screenshot' is a minor deviation from the dominant verb style but still self-explanatory.
Five tools is a reasonable, well-scoped surface for a minimal browser controller and each tool earns its place. It leans slightly thin, since common interactions like navigation or key presses have no dedicated tool.
The core read/interact loop (open, click, fill, screenshot, read text) is covered, but notable gaps exist: no back/forward/reload navigation, no key press (e.g. Enter/Submit), no select/dropdown handling, scroll, or hover. These omissions will force agents to work around common form and multi-page flows.
Maintenance
Related MCP Connectors
Headless browser primitives for AI agents when sites need real JS rendering.
Stealth web browser for agents: search, fetch, click, download and type in persistent MCP sessions.
Provides cloud browser automation capabilities using Stagehand and Browserbase, enabling LLMs to i…
Stealth web automation for AI agents. Login, signup, navigate, screenshot.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceProvides browser automation capabilities through Claude Desktop and other MCP clients, enabling navigation, screenshot capture, content extraction, and interactive control.-
- AlicenseNot gradedqualityCmaintenanceEnables MCP clients to drive a live Chrome or Brave browser over stdio, with tools for navigation, clicking, typing, screenshots, and executing automation goals.3MIT

orchords-web-pilotofficial
FlicenseAqualityCmaintenanceEnables coding agents to drive a real Playwright-backed browser session, navigating, observing, interacting, and capturing proof on web pages via MCP tools across stdio and Streamable HTTP transports.152-- AlicenseAqualityAmaintenanceEnables assistants such as Claude Code, Codex, and Gemini CLI to drive a real browser over MCP — navigating, clicking, typing, and reading live pages from plain-English instructions. Supports optional proxy routing, persistent profiles for surviving logins, and reproducible browser identities via seeds.1631,816MIT