Browser Control
Allows controlling a Firefox browser via the Browser Control extension, providing browser automation capabilities through a shared daemon and WebSocket connection.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Browser Controlopen https://news.ycombinator.com and summarize the top 3 stories"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
browser-control-daemon
One daemon owns the Browser Control extension's WebSocket port (8089). Every Claude Code session runs a thin stdio MCP shim that talks to the daemon over a private unix socket, and starts the daemon in the background if it isn't running. Sessions are no longer limited by the number of ports the extension is configured for.
claude session 1 ─ shim ─┐
claude session 2 ─ shim ─┼─ unix socket (0600) ─ daemon ─ ws://127.0.0.1:8089 ─ extension
claude session N ─ shim ─┘Setup
npm run build
claude mcp add browser-control -e EXTENSION_SECRET=<secret from extension options> -- node $PWD/dist/server.jsExtension: configure the single port 8089 (remove the extra ports / pool wrapper).
Related MCP server: Konnect
Notes
EXTENSION_SECRET/EXTENSION_PORTare read by the daemon at spawn time. To change them:npm run stop; the next tool call respawns it.Logs:
~/.local/state/browser-control-daemon/daemon.log.Daemon requests use its own UUID correlation ids, so concurrent sessions can't collide.
Tools are a copy of upstream
mcp-server/server.ts; only theBrowserAPIimport differs.Tests:
npm test(needs unix-socket creation, so not inside the Claude Code sandbox).Stop the old per-session upstream server / port pool first: the daemon exits if 8089 is taken.
Attribution
src/server.ts and src/common/ are derived from
eyalzh/browser-control-mcp (MIT, see LICENSE).
Use it with that project's Firefox extension.
Available Tools
9 toolscapture-tab-screenshotA
Capture a screenshot of the visible area of a browser tab, by tab ID. The user must authorize each tab by clicking the extension's toolbar button while that tab is open. If the tab is not authorized, this tool returns an error explaining what to ask the user to do; relay that request to the user and retry afterwards. Authorization ends when the tab navigates or closes. Capturing brings the tab to the foreground momentarily.
| Name | Required | Description | Default |
|---|---|---|---|
| scale | No | Image scale relative to CSS pixels, lower values produce smaller images | |
| tabId | Yes | ||
| format | No | Use png only when exact pixel fidelity matters, as it is much larger | jpeg |
| quality | No | JPEG quality, ignored for png |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the per-tab user authorization requirement, that authorization expires on navigation or close, the error-and-retry protocol, and the side effect that capturing brings the tab to the foreground. This is exactly the behavioral context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, each load-bearing, with the core action front-loaded and the authorization caveat immediately following. No padding or redundant restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description would ideally state what is returned (image data and its format). Otherwise it thoroughly covers prerequisites, error handling, and side effects for a 4-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so scale, format, and quality are already documented in the schema. The description only touches tabId implicitly via 'by tab ID' and adds no syntax or semantics beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Capture') and resource ('screenshot of the visible area of a browser tab') with the scoping qualifier ('by tab ID'). An agent can distinguish this from get-tab-web-content or find-highlight-in-browser-tab without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the prerequisite (each tab must be authorized by the user via the toolbar button) and the retry flow after an authorization error. It does not name sibling alternatives such as get-tab-web-content for text extraction, so it stops short of full when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close-browser-tabsC
Close tabs in the user's browser by tab IDs
| Name | Required | Description | Default |
|---|---|---|---|
| tabIds | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not disclose that closing tabs is irreversible, whether it requires specific permissions, or what happens with invalid IDs; it merely restates the mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. Its brevity is appropriate for conciseness, though it borders on under-specification for a destructive operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and a required parameter at 0% schema coverage, the description is too thin. It should clarify irreversibility, ID sourcing, and error behavior to be complete for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. It mentions 'by tab IDs', which adds meaning to the tabIds parameter, but it does not explain the array nature, number type, or where to obtain the IDs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (close) and resource (browser tabs), and specifies the mechanism (by tab IDs). It does not explicitly differentiate itself from siblings like get-list-of-open-tabs or open-browser-tab, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, nor any prerequisites such as needing tab IDs from get-list-of-open-tabs. Usage is only implied by the verb 'close'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find-highlight-in-browser-tabC
Find and highlight text in a browser tab (use a query phrase that exists in the web content)
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | Yes | ||
| queryPhrase | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral-disclosure burden. It states the action but omits key traits such as whether highlighting is persistent, whether it scrolls or focuses the tab, what happens with zero or multiple matches, and whether any permissions are needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted language. The parenthetical adds a useful constraint efficiently. It is concise, though its brevity contributes to the completeness gaps elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a browser-tab tool with two required parameters, no annotations, no output schema, and 0% schema description coverage, the description is too thin. It does not explain the tabId parameter, highlight behavior, side effects, or expected return, leaving an agent under-informed for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two required parameters. The description loosely covers 'queryPhrase' by saying it should exist in the web content, but it does not mention 'tabId' at all and adds no matching or tab-reference semantics beyond what is already implied by the parameter name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Find and highlight text in a browser tab.' This is clear and agent-actionable. However, it does not distinguish the tool from siblings such as get-tab-web-content or capture-tab-screenshot, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical 'use a query phrase that exists in the web content' gives a useful precondition for the query phrase, implying the tool is for locating text already present. But there is no explicit guidance on when to choose this tool over alternatives like get-tab-web-content, and no exclusions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get-list-of-open-tabsC
Get the list of open tabs in the user's browser. Use offset and limit parameters for pagination when there are many tabs.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of tabs to return (default: 100, max: 500) | |
| offset | No | Starting index for pagination (0-based, must be >= 0) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden, yet it discloses nothing about ordering of tabs, whether it reflects the active window vs all windows, whether it requires browser permissions, or what each returned tab contains. It only implies read-only behavior through the verb 'Get', leaving meaningful behavioral traits unreported.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler or redundancy. The pagination note follows the purpose statement in a sensible order, though the second sentence is arguably schema-restating rather than value-adding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a listing tool with no output schema and no annotations, the description should at least sketch the return shape (what fields a tab has) and ordering behavior. It omits both, so an agent knows the tool's purpose but not what it will actually receive back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both limit and offset are fully documented in the schema (defaults, max, 0-based semantics). The description merely restates that they control pagination, adding no format or edge-case detail beyond the structured fields; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Get the list of open tabs in the user's browser'), which tells an agent exactly what the tool returns. It does not, however, differentiate itself from adjacent siblings such as get-recent-browser-history or get-tab-web-content, so the sibling boundary is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage guidance is 'Use offset and limit parameters for pagination when there are many tabs,' which is a mechanical hint about the parameters already in the schema. There is no statement of when this tool is preferable to get-recent-browser-history or get-tab-web-content, and no prerequisites or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get-recent-browser-historyB
Get the list of recent browser history (to get all, don't use searchQuery)
| Name | Required | Description | Default |
|---|---|---|---|
| searchQuery | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not disclose what 'recent' means (time window, count limit), the return shape, ordering, or any permission requirements. Only the filtering hint is given.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no waste. The parenthetical is compact, though slightly awkwardly placed at the end.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description leaves important details unspecified: what a history entry contains, how many results return, and how recency is bounded. Adequate for selection but not for confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single searchQuery parameter, so the description must compensate. It only indicates that omitting searchQuery returns everything, which conveys the filter's effect but not its matching semantics, format, or scope.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource: 'Get the list of recent browser history'. An agent can tell it retrieves history rather than tab state. However, it does not differentiate itself from the sibling tab-oriented tools or clarify scope of 'recent'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical 'to get all, don't use searchQuery' implies the searchQuery parameter filters results, which is implicit usage guidance. There is no statement of when to prefer this tool over siblings or any prerequisites/exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get-tab-web-contentA
Get the full text content of the webpage and the list of links in the webpage, by tab ID. Use "offset" only for larger documents when the first call was truncated and if you require more content in order to assist the user.
| Name | Required | Description | Default |
|---|---|---|---|
| tabId | Yes | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the return payload (text + links) and the truncation/offset paging behavior, which is real behavioral context. It stops short of stating read-only nature, permission needs, or how truncation is signalled, so it is adequate but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with what the tool returns before the conditional offset guidance. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, no-annotation, no-output-schema tool, the description covers the return payload, the scoping key, and the paging escape hatch. What is missing is any statement of safety/permission posture, a minor gap given the low complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It fully explains the non-obvious 'offset' parameter (paging after truncation), which is valuable, but gives no semantics for 'tabId' beyond its self-evident name. Partial compensation over two parameters lands at the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: returns the full text content plus the link list of a webpage, keyed by tab ID. This clearly distinguishes it from siblings like capture-tab-screenshot (visual) and find-highlight-in-browser-tab (search), so an agent can pick it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit guidance for one parameter ('Use offset only for larger documents when the first call was truncated'), which is genuine when-to-use advice. However, it never says when to reach for this tool versus alternatives such as find-highlight-in-browser-tab or capture-tab-screenshot, leaving the core selection decision implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
group-browser-tabsC
Organize opened browser tabs in a new tab group
| Name | Required | Description | Default |
|---|---|---|---|
| tabIds | Yes | ||
| groupColor | No | grey | |
| groupTitle | No | New Group | |
| isCollapsed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not disclose that this is a mutating operation, whether tabs leave their current position or window, whether the group persists across sessions, what happens to already-grouped tabs, or any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero wasted words. It is efficiently structured, though its brevity borders on under-specification rather than true economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter mutation tool with no annotations, no output schema, and 0% schema description coverage, the description is far too thin. An agent cannot determine side effects, parameter behavior, or expected results from this one line.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters, so the description must compensate — and it does not. tabIds, groupColor (a 9-value enum), groupTitle, and isCollapsed are all undocumented; the vague reference to a 'new tab group' hints at groupTitle/groupColor but adds no real meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb plus resource: 'Organize opened browser tabs in a new tab group.' The phrase 'new tab group' signals the creation of a grouping, which loosely separates it from reorder-browser-tabs or close-browser-tabs, but the description never explicitly contrasts itself with those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no preconditions, and no mention of alternatives such as reorder-browser-tabs. The agent is left to infer that this is the tool for grouping, with no rules about when grouping is preferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open-browser-tabB
Open a new tab in the user's browser (useful when the user asks to open a website)
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It states a new tab is opened but omits whether the tab becomes active/focused, whether it requires the browser to be running, what is returned, or any permission/blocking behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short sentence with the action front-loaded; the parenthetical adds usage value rather than filler. Efficient overall.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no output schema, the description covers the basic purpose adequately. However, with zero annotation coverage it leaves the tab's resulting state and return value unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the single url parameter is undocumented in both schema and description. The phrase 'when the user asks to open a website' only loosely implies a URL, adding no format, validation, or syntax detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Open a new tab in the user's browser.' This clearly contrasts with siblings like close-browser-tabs and get-list-of-open-tabs without naming them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical gives an explicit trigger condition ('useful when the user asks to open a website'), which is clear usage context. It does not mention exclusions or alternatives (e.g., reusing an existing tab), so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reorder-browser-tabsC
Change the order of open browser tabs
| Name | Required | Description | Default |
|---|---|---|---|
| tabOrder | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it says almost nothing: not whether the reorder is destructive to pinned/grouped state, whether it requires a specific browser context, whether it fails, or how the result is reported. Only the implicit 'this mutates tab order' is conveyed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no waste, which is good structurally, but the brevity here reflects under-specification rather than discipline given the complete absence of parameter and behavior detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations, no output schema, and an undocumented required array parameter, the description is far too thin. An agent cannot reliably construct a valid tabOrder value from the definition alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description adds zero meaning for 'tabOrder'. It is an array of numbers with no indication whether the values are tab indices, tab IDs, or positions, nor whether the array must be a complete permutation of open tabs or only a subset.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Change the order of open browser tabs'), which clearly separates it from close-browser-tabs, open-browser-tab, and get-list-of-open-tabs. It does not, however, distinguish itself from group-browser-tabs, which also rearranges tabs, so sibling differentiation is incomplete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus group-browser-tabs or close-browser-tabs, no prerequisites, and no statement of what state the tabs must be in first. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
capture-tab-screenshot - First observed
close-browser-tabs - First observed
find-highlight-in-browser-tab - First observed
get-list-of-open-tabs - First observed
get-recent-browser-history - First observed
get-tab-web-content - First observed
group-browser-tabs - First observed
open-browser-tab - First observed
reorder-browser-tabs
TDQS
Scored across 9 tools
Each tool has a clearly distinct purpose: tab lifecycle (open, close, reorder, group), tab/history listing, content extraction, highlighting, and screenshots. Overlap is minimal, and even the two read tools (get-tab-web-content vs capture-tab-screenshot) differ by output modality. An agent can reliably select the correct tool.
All tool names use consistent kebab-case with verb-first phrasing (open-, close-, get-, reorder-, find-, group-, capture-). Resource nouns are predictable (browser-tab, tabs, history, web-content). No mixed conventions or naming outliers.
Nine tools is well-scoped for a browser-control server and sits comfortably in the ideal 3-15 range. Each tool maps to a distinct operation, avoiding both bloat and thinness.
The surface covers core tab management, history, content extraction, highlighting, and screenshots. Minor gaps like navigating an existing tab or refreshing a page are absent, but agents can approximate these via opening a new tab or re-reading content.
Maintenance
Related MCP Connectors
- TabfleetOAuthcom.tabfleet
Launch, inspect, control, and share isolated cloud browsers for your agents.
Use your own Mac from ChatGPT, Claude or Codex: files, commands, documents, and a browser.
Live browser debugging for AI assistants — DOM, console, network via MCP.
Securely control computers you explicitly pair through files, terminals, processes, screenshots, desktop UI/input, clipboard, browser automation, diagnostics, and document tools.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA server implementation that enables controlling web browsers programmatically through Claude's desktop application, providing comprehensive Selenium WebDriver operations for browser automation with Chrome and Firefox support.3MIT
- AlicenseNot gradedqualityAmaintenanceBridges AI agents to a real browser using a persistent daemon and Chrome extension for driving actual login sessions, cookies, and tabs.MIT
- AlicenseAqualityDmaintenanceA bridge to launch Chrome on Windows from WSL2 and proxy Chrome DevTools MCP tools. It enables browser automation via Claude Code in WSL2 by managing Chrome's lifecycle and exposing debugging capabilities.212 npmMIT
- MIT