Skip to main content
Glama
v76ADrR
by v76ADrR

grok-browser-mcp

MCP server that attaches to your already-running Brave over Chrome DevTools Protocol (CDP). Grok CLI can then snapshot, click, type, and screenshot whatever tab you have open — same window, same logins.

It does not launch a fresh empty browser.

Requirements

  • Linux (Brave flags file ~/.config/brave-flags.conf)

  • Brave

  • Node.js 18+

  • Grok CLI

Related MCP server: chrome-pilot-mcp

Install

git clone https://github.com/v76ADrR/grok-browser-mcp.git
cd grok-browser-mcp
npm install

Enable CDP on Brave (once)

The Arch/CachyOS Brave wrapper reads ~/.config/brave-flags.conf. Put:

--remote-debugging-port=9222
--remote-debugging-address=127.0.0.1
--remote-allow-origins=*

Restart Brave. Tabs usually restore. Confirm:

curl -s http://127.0.0.1:9222/json/version

Or call the MCP tool browser_prepare (this restarts Brave).

CDP is localhost only. Any process on this machine can control the browser while those flags are set. Remove them and restart Brave when you are done.

Grok CLI config

~/.grok/config.toml:

[mcp_servers.browser]
command = "/usr/bin/node"
args = ["/ABS/PATH/TO/grok-browser-mcp/src/server.mjs"]
enabled = true
startup_timeout_sec = 20
tool_timeout_sec = 180

Reload: /mcps then r, or a new Grok session.

Tools

Tool

Purpose

browser_status

Brave PIDs + CDP health

browser_prepare

Write flags; restart Brave if CDP is down

browser_tabs

List / focus tabs

browser_snapshot

Clickable elements + refs

browser_find

Filter snapshot by text

browser_click

Click by ref / selector / text

browser_type / browser_press

Keyboard

browser_evaluate

JS in the page

browser_screenshot

JPEG of the current tab

browser_navigate / browser_scroll / browser_wait

Navigation

Never call Playwright browser.close() on this connection — it would quit the user's Brave.

License

MIT

Available Tools

14 tools
browser_clickB

Click an element. Prefer ref from the latest browser_snapshot / browser_find. Falls back to CSS selector or visible text.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoRef like e12 from the last snapshot
textNoVisible text to click
exactNoExact text match (default false)
selectorNoCSS selector

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It does reveal one real behavioral trait: element resolution follows a ref-first, then CSS-selector/visible-text fallback chain. It says nothing about failure behavior when no element matches, auto-waiting, scrolling into view, or navigation side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core action front-loaded and the targeting preference immediately after; no filler. It is perhaps overly terse for a tool with failure modes, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description should be doing more work: post-click behavior (navigation waits, errors when the target is missing, whether text matching is case-sensitive beyond the exact flag) is unaddressed. The reliance on the fully-documented schema keeps it minimally viable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented in the schema; baseline is 3. The description adds only the precedence among those parameters (ref over selector/text), and the ref-source hint largely duplicates the schema's "Ref like e12 from the last snapshot".

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("Click an element") that is immediately distinguishable from siblings like browser_type, browser_press, and browser_evaluate. It does not explicitly contrast with those siblings, but the action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit preference among the tool's own targeting options ("Prefer ref from the latest browser_snapshot / browser_find") and names two siblings as the source of that ref. However, it never says when to click vs. use browser_press, browser_type, or browser_evaluate, so tool-selection guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_desktop_screenshotA

Screenshot the GNOME desktop (Brave window as the user sees it), even if CDP is down. Uses org.gnome.Shell.Screenshot.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It usefully discloses the underlying mechanism ('Uses org.gnome.Shell.Screenshot') and the CDP-independence, which implies a GNOME/display requirement, but says nothing about permissions, return format, or failure modes on non-GNOME hosts.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and scope, then the resilience condition and implementation detail. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with no output schema, the description covers what it captures, when to prefer it, and its backing API. It omits the return shape (image payload vs. file path) and environment prerequisites, which matter more here since there is no output schema to fall back on.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4 and there are no parameter semantics to add. The description correctly implies it is parameterless.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Screenshot the GNOME desktop'), clarifies the capture surface ('Brave window as the user sees it'), and implicitly distinguishes itself from the CDP-based sibling browser_screenshot by noting it works 'even if CDP is down'. An agent can tell this apart from browser_screenshot without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear condition for use ('even if CDP is down'), which routes the agent here when the CDP path is unavailable. It stops short of explicitly naming browser_screenshot as the alternative or stating when this tool should NOT be used.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_evaluateC

Run JavaScript in the current tab (page.evaluate). Return value is JSON-stringified. Use for iCloud DOM inspection.

ParametersJSON Schema
NameRequiredDescriptionDefault
expressionYesJS expression or IIFE, e.g. document.title or (() => {...})()

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden. It usefully notes the return value is JSON-stringified, but says nothing about the fact that arbitrary JS can mutate page state, read sensitive data, or fail/throw — significant for an execution tool. Mutation and safety characteristics are undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with no filler; the core action leads. Slightly terse given the tool's risk profile, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description does cover the return format (JSON-stringified), which is the right thing to add. However, for a tool executing arbitrary JavaScript with no annotations, the absence of any safety or side-effect context leaves the agent under-informed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema itself already gives examples ('document.title or (() => {...})()'), so the parameter is fully documented structurally. The description adds no extra meaning about the expression parameter, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Run JavaScript in the current tab') and even names the underlying API (page.evaluate), so the agent knows exactly what happens. It doesn't explicitly differentiate from siblings like browser_snapshot or browser_find, but no sibling does JS evaluation, so confusion risk is low.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only guidance is 'Use for iCloud DOM inspection,' which is a narrow, context-specific hint rather than a when-to-use rule. It gives no conditions for preferring this over browser_snapshot, browser_find, or browser_click, and no when-not-to-use exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_findA

Find visible clickable elements whose text contains query. Returns refs for browser_click.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesCase-insensitive substring to search for

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose the scope filter ('visible clickable elements') and the return shape (refs), which is useful. But it is silent on no-match behavior, whether refs go stale, and other operational traits an agent would need when the find fails.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed sentences with zero filler; the purpose is front-loaded and the workflow hint follows immediately. Nothing could be removed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and no annotations, so the description is the only carrier of behavior. It covers purpose and return type adequately for a simple read-only lookup, but leaves the failure/empty-result case and ref lifetime undocumented, which matters for a tool whose output feeds browser_click.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and there is only one parameter, so the schema already documents the case-insensitive substring semantics. The description echoes that meaning by specifying matching against element text rather than attributes, which adds slight value but does not go beyond the structured field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Find) and a precisely bounded resource (visible clickable elements whose text contains query), and distinguishes itself from siblings by naming browser_click as the downstream consumer. An agent can tell exactly what this returns without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The sentence 'Returns refs for browser_click' establishes clear context: this is the lookup step before clicking. It does not, however, state exclusions or compare against alternatives like browser_snapshot, which also surfaces page elements, so the agent must infer which locator tool applies.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_navigateC

Navigate the current tab to a URL.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesAbsolute URL

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full behavioral burden. It does disclose one useful trait — navigation targets the *current* tab rather than opening a new one — but it says nothing about whether the call blocks until load completes, what happens on failure or invalid URLs, or any prerequisite browser state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the verb and scope, and nothing superfluous. There is no filler to remove.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a browser-automation tool with no annotations and no output schema, this is thin: it doesn't cover prerequisites (e.g., browser_prepare), load-wait semantics, or error behavior. An agent can call it, but cannot predict how it behaves in a workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'url' parameter is documented as an absolute URI in the schema, so the schema already does the work. The description adds no syntax, scheme, or format guidance beyond that, which is the expected baseline for a fully covered single-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Navigate) and resource (the current tab, to a URL), so the agent knows exactly what the call does. It implicitly distinguishes itself from tab-management siblings by saying 'current tab', but it never names an alternative sibling to make the boundary explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use or when-not-to-use guidance is given. With siblings like browser_tabs, browser_prepare, and browser_snapshot available, the agent gets no signal about when navigating is the right move versus opening a tab or preparing the browser first.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_prepareA

Write Brave CDP flags and optionally restart Brave so Playwright can attach. Restarts the user's Brave (tabs usually restore). Use restart=true when CDP is down.

ParametersJSON Schema
NameRequiredDescriptionDefault
restartNoIf true, restart Brave even if CDP is already up. If omitted, restart only when CDP is down.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the key side effect ('Restarts the user's Brave') and a mitigation ('tabs usually restore'), which are important behavioral traits. It still omits details like whether flags are persistent, whether unsaved data may be lost, or whether repeated calls are safe.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three short sentences with no wasted wording. Purpose is front-loaded, followed by a behavioral warning and then a parameter usage note, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter setup tool with no output schema, the description covers purpose, primary side effect, and parameter usage adequately. It lacks explicit guidance on when to choose this tool over browser_status or other preparatory alternatives, but otherwise gives enough context to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single restart parameter is fully documented in the schema. The description's 'Use restart=true when CDP is down' largely repeats the schema's own explanation, adding no meaningful semantic detail beyond what structured data already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Write Brave CDP flags' and 'restart Brave so Playwright can attach.' It clearly identifies a setup/preparation tool, distinguishing it implicitly from the browser_* action siblings, though it does not explicitly name or compare against any sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'so Playwright can attach' gives implied context for when to call the tool, and 'Use restart=true when CDP is down' provides a condition for the parameter. However, there is no explicit guidance on when to use this tool versus browser_status or other siblings, and no when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_pressB

Press a key (Playwright key name), e.g. Enter, Tab, Escape, Control+a, ArrowDown.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYesKey name

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the key-name format (Playwright, modifier syntax such as Control+a), which is useful, but says nothing about where the key is dispatched (focused element vs. page), whether it waits, or what happens if no element is focused.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that leads with the action and then supplies the format and examples. Nothing is wasted and no filler is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter tool with no output schema, the description is minimally adequate but leaves an important gap: it never says which element receives the key press or what state the browser must be in. With no annotations to compensate, that omission keeps it at 'minimum viable'.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With one parameter at 100% schema coverage and no enums, the baseline is 3. The description goes beyond the schema's terse 'Key name' by naming the accepted convention (Playwright key names) and giving concrete examples like 'Control+a' and 'ArrowDown', which materially clarifies syntax.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Press a key') and clarifies the input domain as Playwright key names. It is clearly distinguishable from siblings like browser_type and browser_click, though it never explicitly names or contrasts with those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus browser_type, browser_click, or browser_find. The key examples hint at the keyboard-interaction domain but do not state prerequisites such as requiring a focused element or a prior navigation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_screenshotB

JPEG screenshot of the current tab. Returns an image plus a saved file path.

ParametersJSON Schema
NameRequiredDescriptionDefault
fullPageNoCapture full scrollable page (default false)

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations and no output schema, so the description carries the full behavioral burden. It usefully discloses the return shape (image plus saved file path), but says nothing about what determines the capture region (viewport vs full page), whether the page state is affected, or any permission/side-effect profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero padding; the core identification leads and the return-value note follows, so the reader gets the essentials immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-optional-param tool with no output schema, the description adequately covers what comes back and what is captured. However, it leaves the reader without the usage context needed to choose between it and browser_desktop_screenshot, which is a real gap given the crowded screenshot sibling set.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with a single optional boolean, so the schema already explains fullPage and its default. The description adds no parameter detail, which is acceptable at this coverage level but not additive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('JPEG screenshot of the current tab'), which is enough to separate it from the sibling browser_desktop_screenshot (desktop capture) without opening a schema. It stops short of naming that sibling explicitly, so the differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites (e.g. a tab must be prepared/active), and no pointer to browser_desktop_screenshot as the alternative for desktop captures. The agent must guess which of the two screenshot tools applies.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_scrollC

Scroll the current tab with a mouse wheel delta.

ParametersJSON Schema
NameRequiredDescriptionDefault
dxNo
dyNoVertical delta pixels (default 600, negative = up)

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It states the mechanism (mouse wheel delta) and the target (current tab), but says nothing about whether scrolling is async, whether it triggers lazy-loading, what happens at scroll boundaries, or the return behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the action front-loaded and no wasted clauses. It is terse, though arguably under-specified rather than tightly concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity two-parameter tool with no output schema, the description is minimally adequate, but the undocumented dx parameter and absent behavioral notes (async, defaults, return) leave real gaps for an agent invoking it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: dy is documented with default and sign convention, but dx is undocumented in both the schema and the description. The phrase 'mouse wheel delta' gestures at the dx/dy pair but adds no units, defaults, or sign conventions to compensate for the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource ('Scroll the current tab') with an implementation detail ('mouse wheel delta'), clearly distinguishable from siblings like browser_click or browser_navigate. It stops short of explicitly contrasting with any sibling, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no preconditions, and no mention of alternatives (e.g., browser_press for keyboard paging). 'Current tab' implies the target but not when scrolling is the right call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_snapshotA

Accessibility / clickable-element snapshot of the current tab, with refs (e1, e2, …) for browser_click. Filter with query for large lists (e.g. iCloud contacts).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax elements to return (default 250)
queryNoCase-insensitive substring filter on element text
offsetNoSkip this many matches

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses the return shape (labelled refs) and the downstream consumer (browser_click), but says nothing about pagination behavior, snapshot size limits, or what happens on elements without refs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with what the tool returns and followed by the practical filtering tip. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-parameter read tool with no output schema, the description covers the essential purpose, output shape, and the key tuning knob. Minor gaps remain around pagination (limit/offset interplay) and behavior on very large pages.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds a use-case hint for query ('large lists, e.g. iCloud contacts'), but adds nothing about limit or offset semantics beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Accessibility / clickable-element snapshot of the current tab') and clarifies what the output is (refs e1, e2, …). The 'accessibility / clickable-element' framing implicitly separates it from browser_screenshot and browser_find, so an agent can route without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives usage context for the query parameter ('for large lists (e.g. iCloud contacts)') and tells the agent refs are consumed by browser_click, but never states when to prefer this over browser_screenshot, browser_find, or browser_evaluate. Usage is implied rather than contrasted with alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_statusA

Check whether Brave is running and whether Chrome DevTools Protocol is reachable on 127.0.0.1:9222.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses the exact target (Brave process and CDP endpoint 127.0.0.1:9222), which is useful behavioral context, but it omits return format, timeout/error behavior, and whether the check has any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence where every clause adds needed specificity (browser, condition, address). There is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and the description covers the diagnostic purpose, but with no annotations and no output schema it does not say what the status result looks like (e.g., boolean/status object) or what happens if the check fails. For an agent to interpret the call, that is a real gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero input parameters, so the baseline is 4. The description adds no parameter meaning because none exist, and nothing is missing in that dimension.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Check) and two precise conditions (whether Brave is running and whether CDP is reachable on a specific address). This clearly distinguishes it from the action-oriented sibling tools such as browser_navigate or browser_click.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the diagnostic nature — an agent would call it to verify the browser connection before using other browser_* tools — but the description does not explicitly state when to use it, when not to use it, or name alternatives such as browser_prepare.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_tabsB

List open Brave tabs (via CDP). Optionally focus one by index, URL substring, or title substring.

ParametersJSON Schema
NameRequiredDescriptionDefault
indexNo0-based tab index from the list
urlIncludesNoFocus the first tab whose URL contains this
titleIncludesNoFocus the first tab whose title contains this

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It hints that 'focus' mutates browser state (changing the active tab) but does not state what happens if no tab matches the substring, whether the window is raised, or that listing is read-only while focusing is a side-effecting action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero waste, front-loading the core list behavior before the optional focus behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 3-optional-param tool with no output schema this is mostly adequate: it says what it returns (tabs) and how focus selects one. The gaps are the lack of annotations explaining side effects and no statement of behavior when a match is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameters are already fully documented. The description merely restates the same three selectors (index, URL substring, title substring) without adding format or matching semantics beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List open Brave tabs') plus the secondary action (focus one) and the mechanism (via CDP). None of the 13 siblings list tabs, so it is distinguishable, but it never explicitly contrasts itself with a related tool, keeping it just below a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Optionally focus one by index, URL substring, or title substring' implies when the focus parameters apply, but there is no guidance on when to call this versus e.g. browser_navigate/browser_prepare, nor any exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_typeC

Type into a field (ref/selector) or the focused element.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNo
textYesText to type
clearNoClear existing value first (default true)
selectorNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it does not state whether typing triggers submission, whether value is cleared, whether permission/auth is required, or what happens on failure. The only behavioral hint is the focused-element fallback, which is thin for a mutation-style interaction tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the target options are packed into a compact clause. Brevity is appropriate but borders on under-specification rather than being genuinely concise-and-complete.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 4 parameters (50% documented), the description is too thin for an interaction tool. It never explains the return value, failure behavior, or how the targeting parameters relate, so an agent is left guessing on key points.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%, with text and clear documented in the schema. The description partially compensates by naming the ref/selector targeting modes, but it leaves the clear default and the ref-vs-selector distinction to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (type) and target resource (a field via ref/selector, or the focused element), which is clearer than most siblings. It does not explicitly differentiate itself from browser_click or browser_press, but an agent can still tell what the tool does at a glance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites (e.g. whether the element must already be focused), and no comparison to alternatives like browser_press or browser_click. The targeting modes are hinted at but the agent gets no rule for choosing ref vs selector vs focused element.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

browser_waitC

Wait for a number of milliseconds, or until text/selector appears.

ParametersJSON Schema
NameRequiredDescriptionDefault
msNoMilliseconds to wait (default 1000)
textNo
selectorNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not say whether waiting blocks execution, what happens if the text/selector never appears, whether there is a timeout or error, or which condition wins when both ms and a condition are supplied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence that front-loads the core action and covers both operating modes without filler. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and three optional parameters, the definition is too thin. Missing timeout/error behavior and parameter precedence leaves real gaps for correct invocation of a synchronization primitive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%; only 'ms' has a description with its default. The description partially compensates by explaining that 'text' and 'selector' are conditions to wait for, but adds no detail on matching semantics (exact vs substring, CSS vs XPath) or precedence between the three parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Wait) and both modes of operation: a fixed duration or a condition (text/selector appears). It is immediately distinguishable from siblings like browser_click or browser_snapshot, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to prefer this over alternatives such as browser_snapshot or browser_find for detecting page state, nor any prerequisites or context. The agent must infer from the name alone that this is for synchronization.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 14 tool updatesv1.0.0
    • First observedbrowser_click
    • First observedbrowser_desktop_screenshot
    • First observedbrowser_evaluate
    • First observedbrowser_find
    • First observedbrowser_navigate
    • First observedbrowser_prepare
    • First observedbrowser_press
    • First observedbrowser_screenshot
    • First observedbrowser_scroll
    • First observedbrowser_snapshot
    • First observedbrowser_status
    • First observedbrowser_tabs
    • First observedbrowser_type
    • First observedbrowser_wait

TDQS

A3.6/5.0

Scored across 14 tools

Disambiguation4/5

Most tools have clearly distinct actions and targets, but browser_snapshot and browser_find overlap because both return clickable-element refs, and browser_status/browser_prepare both concern Brave/CDP availability. The descriptions help mitigate this, so boundaries are mostly clear.

Naming Consistency5/5

All tools use the browser_ prefix followed by snake_case verbs or nouns, giving a highly predictable and readable pattern. browser_desktop_screenshot is slightly longer but still consistent with the schema.

Tool Count5/5

With 14 tools, the set sits comfortably in the expected range for a browser automation server. Each tool maps to a recognizable action or control, and the count does not feel bloated or thin.

Completeness4/5

The surface covers navigation, input, scrolling, waiting, JavaScript evaluation, snapshots, screenshots, tab listing/focusing, and CDP setup. Minor gaps exist around explicit new/close tab, back/forward/reload, and file upload or select-option handling, but browser_evaluate and browser_navigate can often work around them.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers