Skip to main content
Glama

safari-mcp

An MCP server that drives your real, logged-in Safari — your cookies, your sessions, the tab you already have open — with a local model gating every write.

list_tabs · read_page · find_elements        reads, never gated
open_tab · navigate · click · fill · run_js  writes, gated twice

Why not the obvious options

Route

Controls Safari?

Your logins?

Claude in Chrome extension

no — Chromium only

✅

Computer use

view-only on browsers

✅ but read-only

safaridriver / WebDriver

✅

❌ clean profile, signed into nothing

AppleScript do JavaScript

✅

✅

safaridriver deliberately launches an isolated automation profile, which is the right security default and also useless if you wanted to fill in a form behind your own account. AppleScript's do JavaScript runs inside the tab you are already sitting in. That is the whole idea, and it is also why this server takes safety seriously.

Related MCP server: Chrome MCP Server

Requirements

  • macOS with Safari

  • Safari Settings → Develop → Allow JavaScript from Apple Events (this is not the "Allow remote automation" checkbox above it, which is WebDriver)

  • Automation permission for whatever runs the server (macOS prompts once)

  • Ollama with a small model pulled, for the safety gate

  • Python 3.11+

Install

git clone git@github.com:lionel509/Safari-MCP.git
cd Safari-MCP
UV_PROJECT_ENVIRONMENT=~/.venvs/safari-mcp uv sync
ollama pull qwen3.5:2b

The virtualenv deliberately lives outside the project directory — the checkout doubles as a folder inside an Obsidian vault, and Obsidian should never index a venv.

Register it with your MCP host:

{
  "mcpServers": {
    "safari-mcp": {
      "type": "stdio",
      "command": "/Users/you/.venvs/safari-mcp/bin/safari-mcp",
      "args": [],
      "env": {}
    }
  }
}

Safety

Writes pass two layers. Reads pass neither — reading a page is never blocked.

Layer 1 — denylist. URL patterns in ~/.config/safari-mcp/config.json. Deterministic, checked first, and not overridable by layer 2. Ships defaulting to webmail, authentication pages, banks and brokerages, and checkout confirmation URLs.

Layer 2 — local gate model. A small model, through Ollama, sees a compact description of the proposed action — the verb, the page URL and title, and the target element's own text — and answers ALLOW or BLOCK. It never sees the whole page: that is a privacy decision as much as a latency one, since the page is your logged-in session. It is local for the same reason.

The property that makes layer 2 safe to feed page text at all: the model can only add refusals, never remove them. A hostile page will try to talk its way through the gate. The best an injection can achieve is ALLOW — which is exactly what layer 1 has already decided it permits. Nothing is gained.

The gate fails closed. If Ollama is unreachable or the verdict is unclear, the action is refused.

Choosing a gate model

Measured on an M2 Pro. The first pass used ten obvious cases, which every local model passed — a non-result, since a test everything passes cannot rank anything. The set was widened to 34 (17 allow / 17 block, including six prompt-injection dialects), and qwen3:0.6b, qwen3:1.7b and qwen3:4b were eliminated there:

Model

34-case

Avg

qwen3:0.6b

5/10

1.97s

Fails open — allowed a bank transfer and an injection

qwen3:1.7b

5/10

0.38s

Fails closed, but refuses "Add to Cart"

qwen3.5:2b

34/34

0.81s

qwen3.5:4b

34/34

1.57s

The default

gemma4:e4b

34/34

1.62s

No gain over 4b for 4× the weights

The interesting result came from ten harder cases where the danger is implicit — an innocuous label on a dangerous page, like BUTTON "✓" on /messages/draft or BUTTON "Continue" on /checkout/payment, paired with innocuous twins (BUTTON "Continue" on a Wikipedia article). There, qwen3.5:2b let the send through and qwen3.5:4b over-refused "Next" on a Google Form.

Then the description was enriched — aria-label, the enclosing form's action, the enclosing dialog's heading — and both models went 10/10, qwen3.5:2b included, at 0.95s.

The binding constraint on gate accuracy is how much signal reaches it, not how large the model reading it is. A BUTTON "✓" is unjudgeable at any parameter count; BUTTON "✓" | aria-label="Send message" | in form action="/messages/send" is trivially judgeable at 2B. Spend effort on _describe, not on a bigger model.

This is also why the gate stays local. A hosted frontier model would ship your page URLs, titles and click targets to a third party on every action, for no measurable accuracy gain over a 4B running on your own machine.

qwen3.5:4b is the default rather than the faster 2b purely on the safety asymmetry: pre-enrichment it was the one that never missed a danger, and a missed danger costs more than a second of latency. Set guard_model to qwen3.5:2b if you want the speed.

Configuration

~/.config/safari-mcp/config.json, all keys optional:

Key

Default

Meaning

denylist

see config.py

Regexes; any match blocks all writes on that URL

guard_enabled

true

Layer 2 on/off. Layer 1 stays live regardless

guard_model

qwen3.5:2b

Any Ollama model

ollama_host

http://127.0.0.1:11434

guard_timeout

20.0

Seconds

load_timeout

20.0

Seconds to wait for readyState === "complete"

Environment variables override the file: SAFARI_MCP_GUARD, SAFARI_MCP_GUARD_MODEL, SAFARI_MCP_OLLAMA_HOST, SAFARI_MCP_GUARD_TIMEOUT, SAFARI_MCP_LOAD_TIMEOUT, and SAFARI_MCP_DENY_EXTRA (comma-separated extra patterns).

Notes from building it

Three things bite when you drive Safari this way, and the server absorbs all three:

  • Parentheses in a URL make Safari's AppleScript URL setter fail silently — the tab simply stays where it was. navigate percent-encodes them.

  • "Front document" drifts. A redirect, or the person switching tabs, moves it under you. Tabs are addressed by index or URL substring and re-resolved on every call.

  • There is no load event, so navigate and open_tab poll readyState and return the URL actually landed on, which is how you notice a redirect.

And one that only shows up in forms: Google Forms ignores element.value = x. It is built on Closure, which listens for events. fill goes through the element's native setter and dispatches input and change with bubbles: true, which is also what React needs.

Tab pid looks like a stable handle and is not — same-origin tabs share a WebContent process, so two Gmail tabs report the same pid.

License

MIT

Available Tools

8 tools
clickA

Click an element. Check it with find_elements first.

Passes through the denylist and then the local safety gate, either of which may refuse — the refusal explains itself and is worth relaying verbatim rather than retrying.

Args: tab: Tab index within window 1, or a substring of its URL or title. selector: CSS selector for the element. index: Which match to click when the selector hits several. Default 0.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabYes
indexNo
selectorYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden and does meaningful work: it reveals that clicks pass through a denylist then a local safety gate that may refuse, and that refusals are self-explanatory. It does not mention side effects like navigation or form submission triggered by the click, but the safety-gate disclosure is genuinely useful added context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first sentence, followed by a one-line prerequisite, a compact safety note, and a tight Args block. Each sentence carries information and nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter tool with an output schema and no enums or nested objects, this is near-complete: all parameters, the prerequisite workflow, and refusal handling are covered. It only omits minor behavioral details, such as whether the click waits for navigation or other side effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it fully does: tab is explained as an index or URL/title substring, selector as a CSS selector, and index as the match selector with default 0. Every parameter receives meaning beyond its bare type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Click an element', a specific verb and resource, and the sibling list (find_elements, fill, navigate, run_js) makes the distinction immediate. 'Check it with find_elements first' reinforces that this is the action tool rather than the lookup tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly instructs the agent to verify the element with find_elements before clicking and to relay refusals verbatim rather than retrying. However, it does not state when NOT to use click in favor of fill, run_js, or navigate, so it stops short of full exclusion routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fillA

Type a value into an input, textarea, or contenteditable element.

Sets the value through the element's native setter and then dispatches input and change with bubbles: true. That matters: Google Forms, React and other frameworks listen for those events and ignore a bare assignment to .value, so a naive fill looks correct on screen and submits empty.

Args: tab: Tab index within window 1, or a substring of its URL or title. selector: CSS selector for the field. value: The text to enter. index: Which match to fill when the selector hits several. Default 0.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabYes
indexNo
valueYes
selectorYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses a critical behavioral trait: it uses the native setter and dispatches input/change events with bubbles: true. This goes beyond what annotations provide (none) and explains why the tool behaves differently from a simple value assignment. It also notes the consequence of not doing this (forms submit empty). However, it doesn't mention potential side effects like triggering validation or focus changes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured. It opens with a one-sentence summary, then explains the important behavioral detail about event dispatch, and finally lists parameters in a clear Args block. Every sentence earns its place, and the critical information about native setters and events is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for a form-filling tool. It covers the key behavioral nuance (event dispatch), explains all parameters, and the output schema exists so return values don't need explanation. It could mention edge cases like non-interactive elements or disabled fields, but these are minor gaps given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the 'tab' parameter as 'Tab index within window 1, or a substring of its URL or title,' which adds meaning beyond the schema's bare 'Tab' title. It also explains 'index' as 'Which match to fill when the selector hits several. Default 0.' The 'value' and 'selector' parameters are self-explanatory from their names, but the description adds context for the two less obvious ones.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Type a value into an input, textarea, or contenteditable element.' It uses a specific verb ('fill') and resource (form fields), and distinguishes itself from siblings like click and run_js by focusing on setting values with native setters and dispatching events.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains why this tool is needed over a naive assignment, mentioning frameworks like React and Google Forms that require input/change events. It doesn't explicitly name alternatives, but the context about event dispatch implies when to use this tool versus simpler approaches. It could be improved by explicitly stating when not to use it (e.g., for clicking buttons, use click).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_elementsA

Inspect the elements a CSS selector matches, before acting on them.

Use this to check that a selector hits what you think it does, and to read the index you will pass to click or fill. Returns JSON with each match's tag, id, class, visible text, href, name, type and current value.

Args: tab: Tab index within window 1, or a substring of its URL or title. selector: The CSS selector to inspect. limit: Most elements to describe. Default 20.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabYes
limitNo
selectorYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses that the tool returns JSON with per-match fields (tag, id, class, visible text, href, name, type, current value) and a default limit of 20. It could mention whether it has side effects, but as an inspection tool, its non-destructive nature is reasonably implied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded with the core purpose, followed by a brief usage note and a concise Args section. It avoids fluff, though the Args section slightly overlaps with schema titles.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity and the descriptive return-value explanation, the description is largely complete. It could benefit from noting that it matches all elements matching the selector and does not alter state, but overall it provides the needed context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description compensates. It explains tab as a tab index or substring of URL/title, selector as a CSS selector, and limit as max elements with default 20. This adds meaningful guidance beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: inspect elements matched by a CSS selector. It differentiates itself by emphasizing pre-action validation and reading the index for click or fill, distinguishing it from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use: check that a selector hits the right elements and read the index to pass to click or fill. It does not name alternative tools, but the context is clear enough for an agent to decide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_tabsA

List every open Safari tab, with the handle you use to address it.

Call this first. Every other tool takes a tab argument, and this shows you what is addressable: a 1-based index within window 1, or any distinctive substring of a tab's URL or title.

Returns one line per tab as [w<window>:t<index>] <title> — <url>.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden itself. It discloses the addressing scheme (1-based index within window 1 or URL/title substring) and the exact one-line return format. It doesn't spell out side-effect-freedom or edge cases, but 'List' and the output format make the read-only behavior clear enough.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, focused paragraphs with no filler. The main purpose is front-loaded, and each sentence adds a distinct piece of information: purpose, when to call it, handle format, and output format.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter list tool with an output schema, the description is complete: it states scope, return format, and how the result should be used downstream. No critical information is missing for an agent to select and call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, so the empty schema is already complete. The description adds useful context about what the returned handle means, which is the closest analog to parameter semantics for this tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('List every open Safari tab') and immediately explains what makes this tool distinct: it returns the handle needed by other tab-based tools. It differentiates from siblings by framing itself as the discovery/entry point.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit directive ('Call this first') and explains why, stating that other tools take a tab argument and need the handle this tool provides. This tells an agent when to invoke it relative to every sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_tabA

Open a URL in a new Safari tab and wait for it to finish loading.

Prefer this over navigate when starting a task: it leaves whatever the person was already reading untouched. Returns the new tab's handle and its final URL, which may differ from the one you asked for if the site redirected.

Args: url: The URL to open. Parentheses and spaces are encoded for you.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must cover behavioral aspects. It discloses that the tool waits for the page to finish loading, which is a significant behavioral trait, and explains that the final URL may differ due to redirects, adding valuable context. However, it doesn't mention potential side effects like affecting the tab list or requiring network permissions, which could be relevant but not essential.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured: the first line states the primary action, followed by usage guidance, then a note about the return value. Each sentence serves a purpose, and the argument documentation is minimal but adequate. No fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a single parameter, a clear output schema (returning tab handle and final URL), and no annotations, the description covers the essential use case and key behavior. It could go into more detail about error handling (e.g., invalid URLs) or if simultaneous tabs are supported, but for a tool of this simplicity, it is sufficiently complete for an agent to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is effectively 0% (no description in the schema), so the description must compensate. It explains that parentheses and spaces in the URL are encoded automatically, which goes beyond the schema's minimal type definition. It doesn't elaborate on the URL format (e.g., must be fully qualified), but the provided handling of special characters is useful and relevant.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool opens a URL in a new Safari tab and waits for loading, with a specific verb ('Open'), resource ('URL'), and context ('new Safari tab'). It also distinguishes itself from the sibling 'navigate' by noting that it preserves the user's current page, avoiding ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly advises to prefer this tool over 'navigate' when starting a task, citing the benefit of leaving the user's current reading untouched. This provides clear when-to-use guidance and names the alternative, leaving no inference needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_pageA

Read the visible text of a tab, or of one element inside it.

This is the workhorse — prefer it over screenshots. It returns rendered text (innerText), so it reflects what a person would actually see rather than raw markup.

Args: tab: Tab index within window 1, or a substring of its URL or title. selector: Optional CSS selector. Omit it to read the whole page. max_chars: Truncate beyond this many characters. Default 8000.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabYes
selectorNo
max_charsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden. It discloses that the tool returns rendered innerText reflecting what a person sees rather than raw markup, which is valuable behavioral context beyond the tool's name. It does not discuss side effects or failure modes, but the read-only nature is strongly implied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized: a one-line summary, a concise rationale for preferring this tool, then a compact Args block. Every sentence earns its place, and the most decision-relevant guidance ('prefer it over screenshots') is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

All three parameters are explained sufficiently for correct invocation, and the output schema likely covers return details. Minor gaps remain—what happens when selector matches multiple elements, and how read_page compares to find_elements—but the core usage context is complete for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description fully compensates. It explains tab as an index within window 1 or a URL/title substring, defines selector as optional CSS that defaults to the whole page, and clarifies max_chars truncation with its 8000 default. This adds meaning far beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('read') and resource ('visible text of a tab, or of one element inside it'), and explicitly contrasts itself with screenshots and raw markup ('rendered text (innerText)... rather than raw markup'). This makes the tool's purpose unmistakable and differentiates it from siblings like find_elements and run_js.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear invocation guidance: prefer this over screenshots, omit selector to read the whole page, and use selector to target a specific element. It does not explicitly say when not to use it or directly name sibling alternatives like find_elements, but the context is sufficient for an agent to route correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_jsA

Run arbitrary JavaScript in a tab. The escape hatch; prefer the others.

Your code runs as a function body, so return the value you want back. The result is JSON-encoded, so return plain data rather than DOM nodes.

Args: tab: Tab index within window 1, or a substring of its URL or title. code: JavaScript to execute. Use return to produce a result.

ParametersJSON Schema
NameRequiredDescriptionDefault
tabYes
codeYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It usefully explains that code runs as a function body, requires `return`, and that results are JSON-encoded (so return plain data, not DOM nodes). It does not mention error handling or side effects, but it discloses the key execution semantics an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is tightly organized: a one-line purpose, a short usage hint, two behavioral sentences, and two bullet-style parameter explanations. Every sentence earns its place and the most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an arbitrary-JS tool with no annotations and a minimal schema, the description covers tab targeting, code execution, and return-style expectations. It could add notes on error behavior or async handling, but it contains enough to call the tool correctly in typical cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description fully compensates. It explains `tab` as a tab index within window 1 or a URL/title substring, and `code` as JavaScript to execute with `return` for results. Both parameters gain meaning beyond their bare string type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Run arbitrary JavaScript in a tab') and immediately differentiates itself from siblings with 'The escape hatch; prefer the others.' An agent can tell exactly what it does and how it relates to the other tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear usage context by labeling itself an escape hatch and advising to prefer the other tools, which implies using it only when the dedicated siblings do not cover the need. It does not name specific alternative conditions, but the guidance is clear enough for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv0.1.0
    • First observedclick
    • First observedfill
    • First observedfind_elements
    • First observedlist_tabs
    • First observednavigate
    • First observedopen_tab
    • First observedread_page
    • First observedrun_js

TDQS

A4.4/5.0

Scored across 8 tools

Disambiguation5/5

Each tool targets a distinct action: listing tabs, reading content, inspecting elements, opening/navigating tabs, clicking, filling, and executing JS. Even overlapping actions like open_tab vs. navigate are clearly separated by whether they open a new tab or reuse an existing one. There is no ambiguity in selecting the right tool for a task.

Naming Consistency4/5

Most names follow a verb_noun pattern (list_tabs, read_page, find_elements, open_tab, run_js), but navigate, click, and fill are bare verbs without an object. The names are still predictable and all lowercase snake_case, but the pattern is not perfectly uniform.

Tool Count5/5

Eight tools is a well-scoped set for browser automation. Each tool has a clear purpose and no redundant utilities; the count feels appropriate for the domain without being sparse or bloated.

Completeness4/5

The surface covers the core browser automation lifecycle: inspect (list_tabs, read_page, find_elements), act (click, fill), navigate (open_tab, navigate), and an escape hatch (run_js). Missing operations like back/forward, close tab, or direct URL getter are minor and can be worked around with run_js or navigate.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    Provides AI assistants with Safari browser automation and developer tools access, enabling LLMs to control Safari, access console logs, monitor network activity, and perform browser automation tasks.
    13
    9 npm
    33
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Transforms Chrome browser into an AI-controlled automation tool that allows AI assistants like Claude to access browser functionality, enabling complex automation, content analysis, and semantic search while preserving your existing browser environment.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to control and automate your Chrome browser directly, leveraging existing login states and configurations for tasks like content analysis, semantic search across tabs, screenshots, network monitoring, and interactive operations.
    10
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Lets AI assistants control your real Chrome browser to perform web tasks like reading pages, taking screenshots, clicking, and typing, using your existing logged-in sessions.
    133
    MIT