VisionCLI
Captures clean screenshots from Android devices via adb for visual analysis.
Provides fast Brave browser control via AppleScript, including tab management, navigation, reading pages, clicking, typing, and scrolling.
Reads project .env files to discover database connection settings for read-only MySQL, Postgres, and SQLite queries.
Integrates with Gemini CLI and the Gemini API via API-key auth, enabling voice and typed questions with access to VisionCLI's vision tools.
Captures clean screenshots from iOS Simulator for visual analysis.
Opens file:line references in JetBrains IDEs and detects JetBrains project windows for context.
Integrates with macOS system features for screen capture, microphone/speech input, menu bar presence, automation, and AppleScript-based control.
Supports read-only MySQL schema inspection and SELECT/SHOW/DESCRIBE/EXPLAIN/WITH queries using connection details from the project's .env.
Maps PhpStorm project windows to their local project folders so questions are answered about the correct project.
Supports read-only SQLite schema inspection and SELECT/SHOW/DESCRIBE/EXPLAIN/WITH queries using connection details from the project's .env.
Interacts with terminal apps such as Warp to type questions into running Claude/Gemini CLI sessions and receive answers.
Uses Xcode for iOS Simulator screenshots and can open file:line references in Xcode.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@VisionCLIwhat is this error on my screen?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
"What is this file for?" Claude looks at the window under your mouse, finds the real file in your project and answers with
file:linereferences.Hold β₯ Space and ask out loud. A loader rides on your cursor, and the answer appears in a bubble above it, read aloud.
Works with your Claude Code session. It types into the
claudealready running for that project, or opens one, so you can type and talk in the same conversation.Points at things. Claude draws glowing boxes on your screen around what it's talking about.
Fast browser control. "Open the Stripe tab", "click sign in", "pause the video" run in under a second (via TypeSafe Jev), with a visible AI cursor.
Knows your data. Read-only database questions ("how are payments and orders related?") using your project's
.env.
Install
Requirements
macOS 13 or later (Apple Silicon or Intel)
Claude Code (
claudeCLI), signed in, and/or Gemini CLI (gemini) with a Gemini API key (see Gemini CLI)Node.js 18 or later
1. Download the app
Download
VisionCLI-<version>-macos.dmgfrom the latest release.Open it and drag VisionCLI into Applications.
The app isn't notarized by Apple yet, so the first time:
Right-click VisionCLI β Open β Open, or
run
xattr -dr com.apple.quarantine /Applications/VisionCLI.appin Terminal.
2. First launch
The π icon appears in the menu bar and a welcome screen offers to set up the CLIs it finds. That registers VisionCLI's tools for every project: claude mcp add visioncli --scope user β¦ for Claude Code, and gemini mcp add -s user --trust visioncli β¦ for Gemini CLI. You can also do it later from π β Settings β Set Up Claude Code / Set Up Gemini CLI.
macOS will ask for three permissions. VisionCLI needs all of them:
Permission | Why |
Microphone and Speech Recognition | to hear your question while you hold β₯ Space |
Screen & System Audio Recording for VisionCLI | to see the window you're asking about in voice mode. After allowing it, quit and reopen VisionCLI. |
Screen & System Audio Recording for your terminal app (iTerm, Terminal, Warpβ¦) | when you type to Claude Code, the screenshot is taken by the terminal running |
Automation (iTerm, Chrome/Brave) | asked the first time it types into your terminal or controls the browser |
If something doesn't work, check System Settings β Privacy & Security, and π β Settings β Open Log.
3. Optional extras (auto-detected)
Add | Get | How |
Whisper | Better transcription of names and jargon (~0.7s) |
|
Kokoro | Natural voices for spoken answers | install |
TypeSafe Jev | Instant browser commands | add your key via π β Settings β API Keys, or set |
Browser control | Click, scroll and read pages | in Chrome or Brave: View β Developer β Allow JavaScript from Apple Events |
Related MCP server: PeepIt MCP
Gemini CLI
VisionCLI works with Gemini CLI as well as Claude Code:
Typing in
gemini: after Set Up Gemini CLI, the visioncli tools (look, browser, database, device screenshotsβ¦) are available in every Gemini CLI project.Voice with Gemini: choose π β Assistant β Gemini CLI.
Gemini Terminal mode types your question into the
geminirunning for the project (or opens one), attaches screenshots as@file, and shows the answer in both the terminal and the bubble.Bubble Only mode runs
gemini -p β¦ -o stream-json, resuming the same session each time.
Pointing at a terminal that runs
claudeorgeminiuses that CLI, whatever the Assistant setting.
Sign-in: Google no longer serves Gemini CLI on a personal Google sign-in (the free "Code Assist for individuals" tier says to move to Antigravity). Use a Gemini API key from aistudio.google.com/apikey: put it in π β Settings β API Keys ("gemini": {"apiKey": "..."}) or set GEMINI_API_KEY before visioncli voice. VisionCLI then runs gemini with API-key auth for its own sessions (a per-run settings override), without changing your Gemini settings. Paid Code Assist (Standard/Enterprise) sign-ins keep working as-is.
Limits: in Bubble Only mode Gemini can't ask for approval, so it can only read (code, screen, browser, database); for edits and commands use Gemini Terminal mode. Highlight boxes, history, speech and Jev work the same with either assistant.
Use it
Voice (anywhere)
Keys | What happens |
Hold β₯ Space, speak, let go | Ask about what's under the mouse (or tap to start, tap again to send) |
Hold β₯ β§ Space | Ask, and drag a box around exactly what you mean |
β₯ P | Pin the answer: scroll, select text, click |
β₯ O | Open the first |
β₯ β© / esc | Approve or deny an edit or command (background mode) |
esc | Dismiss, or stop the answer |
Which project? By default VisionCLI follows the window under your mouse. PhpStorm myshop β payments maps to ~/PhpstormProjects/myshop, and a terminal running claude maps to its folder. With several projects open, pick one in π β Project. The menu bar shows the current project, with π when you've chosen one.
Which session? In Terminal mode (the default), voice questions are typed into the claude (or gemini, see Gemini CLI) already running for that project in iTerm, or a new one is opened. You can type there too; it's one conversation, and answers show in both the terminal and the bubble. Bubble Only mode uses a hidden, leaner session instead. Switch in π β Answer In.
In Claude Code (typing)
After Set Up Claude Code, just talk about what's on screen in any project:
"what is this file for?", "why is this view misaligned in the simulator?"
"let me select an area", "read the page I have open and summarize it"
"how are the orders and payments tables related?"
Slash commands: /mcp__visioncli__see <question>, /mcp__visioncli__point <question>.
Tool | |
| Screenshot the window behind the terminal, the one under the mouse, a named app, or a display; reports the pointer position |
| Pick a window, drag a box, or use the clipboard image |
| Clean iOS Simulator (needs Xcode) or Android ( |
| Tables, columns, foreign keys; single read-only SELECT/SHOW/DESCRIBE/EXPLAIN/WITH queries in a read-only transaction (MySQL, Postgres, SQLite from |
| Fast Chrome/Brave control via AppleScript |
| Start the voice overlay |
Menu
π Project (Auto, or pick from running claude/gemini sessions, open editors and recent projects) Β· Assistant (Claude Code or Gemini CLI) Β· Answer In (terminal, or bubble only) Β· Model (Sonnet or Opus, for Claude) Β· Voice & Audio (speak answers, speed up to 2Γ, voice, speech recognition, microphone, speaker) Β· History Β· Copy Last Answer Β· New Conversation Β· Settings (Set Up Claude Code / Gemini CLI, Start at Login, instant browser commands, AI cursor, API keys, logs).
API keys
π β Settings β API Keys opens ~/Library/Application Support/visioncli/integrations.json (readable only by you):
{
"typesafe": { "apiKey": "..." },
"gemini": { "apiKey": "..." }
}typesafe: instant browser commands with Jev. Without it, every request goes to the assistant.gemini: lets VisionCLI run Gemini CLI with API-key auth.
TYPESAFE_API_KEY / GEMINI_API_KEY in your shell are copied here by visioncli voice if the file has none.
Privacy
Screenshots and questions go to Claude through your own Claude Code login.
With Jev enabled, quick commands send your spoken words, open tab titles and URLs, and visible link/button text to api.typesafe.ai. Turn it off in Settings.
Whisper and Kokoro run locally. History is stored in
~/Library/Application Support/visioncli/.
Build from source
git clone https://github.com/thethirdsourcers/visioncli.git && cd visioncli
npm install # builds the MCP server (dist/)
node dist/index.js install # register with Claude Code (and Gemini CLI, if installed) for all projects
npm run build # build VisionCLI.app and install it into /Applications
node dist/index.js voice # start voice mode (under launchd: restarts after a crash)npm run setup-signing(optional, once) creates a local code-signing identity so macOS keeps permissions across rebuilds.npm run releasebuilds the downloadable universal.dmgand.zipintorelease/. Pushing av*tag does the same on GitHub Actions and attaches them to a release.
Layout: src/ is the MCP server and CLI (TypeScript). overlay/Sources/ is the menu-bar app (Swift). scripts/ has the build, release and signing scripts.
Troubleshooting
Problem | Fix |
"VisionCLI can't be opened" | Right-click β Open, or |
Claude sees only the wallpaper | Voice: enable Screen Recording for VisionCLI. Typing in Claude Code: enable it for your terminal app. Then quit and reopen that app. |
Tools missing in Claude Code / Gemini CLI | π β Settings β Set Up Claude Code / Set Up Gemini CLI (again after moving the app or switching Node versions), then restart |
Gemini: "no longer supportedβ¦" or "API key not valid" | Add a valid key from aistudio.google.com/apikey in π β Settings β API Keys |
Answers only in the terminal, or none at all | π β Settings β Open Log; check |
Microphone errors | Pick a specific mic in π β Voice & Audio β Microphone |
It crashed | It restarts automatically; details are in π β Settings β Open Crash Log |
License
MIT Β© The Third Sourcers
Available Tools
17 toolsbrowser_clickClick an elementB
Click an element by id from browser_elements (or by its visible text).
| Name | Required | Description | Default |
|---|---|---|---|
| target | Yes | Element id like e12, or visible text |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It says nothing about what happens after the click (navigation, new tab, page mutation), whether it waits for load, scroll-into-view, or failure behavior on missing elements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the essential action first and zero filler. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-param tool this is close to adequate, but with no annotations and no output schema the description omits any post-click consequence, leaving the agent unsure whether to re-read the page or expect a return value.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single param is fully documented in the schema ('Element id like e12, or visible text'). The description merely restates the same id-or-text duality, adding no syntax or format detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (click) and resource (element), and names the sibling browser_elements as the source of the id. However, it does not contrast with other interaction siblings like browser_type, so differentiation is only partial.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(or by its visible text)' implies a usage style, and referencing browser_elements hints at a prerequisite workflow. But there is no explicit when-to-use/when-not guidance and no alternatives named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_elementsList clickable elementsARead-only
Visible links, buttons and form fields on the active page with ids (e1, e2, β¦) for browser_click / browser_type.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already tells the agent this is a safe read, and the description adds two meaningful scope constraints beyond the annotation: only *visible* elements are returned, and only those on the *active* page. It does not mention whether off-screen elements appear or how ids are invalidated after navigation, but the visibility/active-page caveats are real added value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence states what is listed, where it comes from, what the ids look like, and which tools consume them. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read tool with no output schema, the description covers the essentials: scope (active page, visible only), return shape (e1, e2 ids), and downstream use. Minor unknowns remain (behavior after scroll or navigation, whether the list is paginated), but nothing blocks correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4; there is nothing for the description to disambiguate. The mention of the returned id scheme is a bonus rather than a parameter explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb+resource (visible links, buttons, form fields) scoped to the active page, and the id format (e1, e2, β¦) makes clear what it returns. It differentiates itself from content-oriented siblings like browser_read by naming the click/type consumers, though it never explicitly contrasts with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
By stating the ids exist 'for browser_click / browser_type', it implies the workflow: call this first to obtain element handles, then act. The context is clear but there is no explicit when-not guidance or alternative named (e.g., a text-extraction tool for reading rather than interacting).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_openOpen a URLB
Open a URL in the browser, in a new tab (default) or the current one.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| new_tab | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does very little with it. It does not disclose what happens to an already-open page, whether the call blocks until load, whether it steals window focus, or what it returns; only the tab default is stated, which is really a parameter fact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words, and the tab-behavior detail is placed immediately after the core action. It is efficient, though the brevity is partly what leaves the behavioral and sibling-differentiation gaps.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, this is roughly the minimum needed to invoke it. It is still missing the distinction from browser_navigate and any indication of what the call produces or how it affects the current page state.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so both parameters are undocumented in structured data. The description compensates for new_tab by naming the default and the alternative ('new tab (default) or the current one'), but it adds nothing about url beyond the obvious, so it only partially closes the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Open a URL in the browser', which is unambiguous on its own. However, it does not differentiate itself from the sibling browser_navigate, which plausibly also opens/navigates URLs, so an agent cannot tell how the two differ from this text alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the invocation choice by explaining that a new tab is the default and the current tab is the alternative, which is a conditional usage hint. It names no sibling alternative (e.g. browser_navigate) and gives no when-not-to-use guidance, so guidance remains implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_readRead the current pageARead-only
Title, URL, selected text and readable text of the active tab (capped). Faster and more exact than a screenshot.
| Name | Required | Description | Default |
|---|---|---|---|
| max_chars | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already covers the safety profile, so the description's extra value is the disclosed return contents and the '(capped)' truncation warning, which is genuine behavioral context an annotation cannot carry. It does not explain what happens to content beyond the cap (silent truncation vs. error) or how the cap is set, leaving one meaningful gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences: the return contents are front-loaded, and the comparison to screenshots follows as supporting rationale. No filler and nothing redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the returned fields, which is the right coverage. However, for a tool whose only parameter governs truncation, the missing explanation of the cap and its default leaves the definition incomplete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single max_chars parameter has 0% schema description coverage (no description, only default 12000 / min 500 / max 50000), so the description must compensate. '(capped)' gestures at truncation but never names the parameter, its unit, its default, or that the cap is configurable β an agent cannot connect the behavior to the knob.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description enumerates exactly what the tool returns (title, URL, selected text, readable text of the active tab), which identifies the resource and scope clearly alongside the verb in the name/title. It also differentiates against one sibling (screenshot) with 'faster and more exact than a screenshot', though it never names or contrasts against browser_elements, look, or the other page-reading alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Faster and more exact than a screenshot' implies a when-to-use preference over screenshot tools, so usage is implied rather than stated. There is no explicit condition, no when-not guidance (e.g. use browser_elements for structured DOM), and no mention of the active-tab prerequisite that selects this over browser_tabs or browser_switch_tab.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_scrollScroll the pageC
Scroll the active page.
| Name | Required | Description | Default |
|---|---|---|---|
| where | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not disclose how far a single scroll moves (viewport, page, pixel amount), whether it waits for lazy content, or whether it ever fails or returns anything. For a browser-interaction tool with zero structured behavioral hints, this is a substantial gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with no wasted words, but it is terse to the point of under-specification rather than genuinely concise. Brevity here costs information rather than saving it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and 0% parameter coverage, the description should compensate but does not. It omits scroll magnitude, return behavior, and any interaction with dynamic content, so an agent cannot predict the effect of invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single 'where' parameter has 0% schema description coverage, and the description adds nothing about it. The enum values (down/up/top/bottom) are largely self-explanatory, but the description does not clarify semantics such as whether 'down' scrolls by a viewport or a small increment.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (scroll) and resource (the active page), which is clear enough to separate it from siblings like browser_click or browser_read. It does not, however, differentiate itself from other navigation-type siblings or explain the scope of 'active page' vs. tabs/windows.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to scroll versus using browser_navigate, browser_read, or select_region. There are no prerequisites or exclusions given, leaving the agent to infer usage entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_switch_tabSwitch tabB
Bring a tab (id from browser_tabs) to the front.
| Name | Required | Description | Default |
|---|---|---|---|
| tab_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure, and it says almost nothing: no mention of what happens on an invalid tab_id, whether the OS window is focused, whether the tab is restored or only made active, or whether it errors. Only the bare effect ('to the front') is stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the action and the parameter provenance are both packed in without waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (one required param, no output schema, no annotations), so a short description is defensible. Still, for a state-changing view operation it should at least say what a success looks like and what occurs when the id no longer exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the schema gives only a bare string with no format, so the description nominally does add value by naming browser_tabs as the source of the tab_id. But it does not describe the expected id format or whether ids are stable across sessions, leaving a real gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Bring a tab ... to the front') and clarifies the identifier provenance ('id from browser_tabs'), which separates it from the sibling browser_tabs that lists tabs. It does not, however, contrast itself with other navigation siblings like browser_navigate or browser_open.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the agent must infer it should call this after identifying a tab via browser_tabs. There is no explicit when-to-use, no conditions under which switching is needed, and no note about behavior when the tab is already active or the id is stale.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_tabsList browser tabsARead-only
Open tabs in the front-most Brave/Chrome (ids like t1.3 = window 1, tab 3), marking the active one.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already tells the agent this is a safe read. The description adds real value beyond the annotation by disclosing the id format ('t1.3 = window 1, tab 3') and that the active tab is marked, which shapes how the agent parses results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the resource and scope, with the id-format detail packed into a parenthetical rather than a separate sentence. Slightly dense but nothing wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-param, non-mutating tool with no output schema, the description covers what is returned (tab list with ids) and the active-tab marker. It stops short of describing ordering or how many windows are covered, but that is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the schema imposes no semantic burden and the baseline is 4. The description correctly avoids restating a nonexistent parameter list.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: it lists open tabs in the front-most Brave/Chrome window. The scope qualifier 'front-most' implicitly separates it from list_windows and other browser_* siblings, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no mention of browser_switch_tab as the natural follow-up, and no prerequisite stated. The agent must infer that this is the discovery step before acting on a tab.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browser_typeType into a fieldB
Set the value of a form field (id from browser_elements), optionally submitting the form.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| submit | No | ||
| target | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses one side effect ('optionally submitting the form'), but not whether typing clears the existing value, fires input/change events, waits for navigation, or what happens on a failed submit. Critical mutation behavior is undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with the core action front-loaded and the sibling reference and optional behavior tucked into a parenthesis. No filler, nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with zero annotation coverage, no output schema, and 0% parameter documentation, the description leaves significant gaps: field-clearing behavior, event firing, submit side effects, and failure handling are all unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains 'target' (an id sourced from browser_elements) and 'submit' (submitting the form), covering two of three params, but says nothing about whether 'text' appends to or replaces existing content, or any format expectations.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Set the value') and resource ('form field') and links the target to browser_elements, which distinguishes it from browser_click and browser_read. It stops short of explicitly contrasting itself with sibling tools, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(id from browser_elements)' implies a workflow prerequisite β call browser_elements first to obtain the target id β but no when-to-use/when-not guidance or alternatives are given. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clipboard_imageRead image from clipboardBRead-only
Return the image currently on the user's clipboard (e.g. a screenshot taken with Cmd+Ctrl+Shift+4).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already tells the agent this is a safe, non-mutating read, so the description isn't carrying the safety burden. It adds the useful context that the source is typically a screenshot, but says nothing about what happens when the clipboard is empty or non-image, which is the main behavioral risk for a zero-param read.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the action, the resource, and a concrete example with zero filler. Nothing could be trimmed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations about return format, the description leaves the agent unaware of how the image is delivered and how failure (no image on clipboard) is signalled. Adequate for a trivial zero-param tool, but the primary edge case an agent will hit is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline of 4 applies. No additional parameter meaning could be added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (return) and resource (image on the user's clipboard), so the agent knows exactly what the tool fetches. However, it gives no signal distinguishing it from sibling image-capture tools such as device_screenshot or select_region.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical example (a Cmd+Ctrl+Shift+4 screenshot) implies the context where the clipboard holds an image, but there is no explicit when-to-use guidance and no mention of alternatives like device_screenshot for capturing an image directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
db_queryRun a read-only SQL queryARead-only
Run one SELECT/SHOW/DESCRIBE/EXPLAIN/WITH query against the project's database inside a read-only transaction. Results are capped at 50 rows; use LIMIT.
| Name | Required | Description | Default |
|---|---|---|---|
| sql | Yes | A single read-only SQL statement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true; the description adds substantial behavior the annotation grid cannot express: execution inside a read-only transaction, exactly one statement, an enumerated allowlist of statement types, and a hard 50-row result cap with the LIMIT workaround. The row cap is the kind of silent-truncation trait an agent must know before calling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first front-loads the operation and its constraints, the second front-loads the operational gotcha (row cap) with the remedy. No filler or restatement of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one parameter and no output schema, the description covers the essentials: allowed statements, single-statement limit, read-only scope, and result cap. It omits error/timeout behavior and the result's column shape, which are minor gaps for a tool this simple.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter is already documented as a read-only SQL statement, so the baseline is 3. The description adds genuine meaning by enumerating the accepted statement prefixes (SELECT/SHOW/DESCRIBE/EXPLAIN/WITH) and the one-statement constraint, going beyond the schema's generic phrasing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (run) and resource (SQL query against the project's database) and enumerates exactly which statement classes are accepted. An agent can immediately distinguish this data-query tool from the sibling db_schema introspection tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operating context: one statement per call, read-only only, and 'use LIMIT' because output is capped. It does not, however, explicitly route the agent to db_schema when it needs to discover tables/columns rather than query data, so the when-not guidance is incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
db_schemaDatabase schemaARead-only
Tables, columns and foreign keys of the project's database (connection read from the project's .env: Laravel DB_* or DATABASE_URL; MySQL, Postgres, SQLite). Pass a table to narrow it. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| table | No | Only this table |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered (and 'Read-only' largely repeats it). The description adds real value beyond the annotations by disclosing that the connection is read from the project's .env (Laravel DB_* or DATABASE_URL) and which engines are supported (MySQL, Postgres, SQLite), which tells the agent where the data comes from.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short clauses, front-loaded with the resource, then the connection detail, then the narrowing option and the read-only guarantee. No filler, and a reader can stop after the first clause and still know what the tool does.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must describe the return shape, and it does name the returned elements (tables, columns, foreign keys). It could go slightly further on how results are structured or whether foreign keys are nested, but for a single-parameter introspection tool with annotations already covering safety this is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the lone parameter is documented as 'Only this table', so the baseline is 3. The description improves on that by clarifying the omitted-parameter default ('Pass a table to narrow it' implies all tables are returned otherwise), which the schema does not state.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (tables, columns, foreign keys of the project's database) with the verb implied as introspection, so an agent knows exactly what it retrieves. It does not distinguish itself from the sibling db_query tool, which is the most likely source of confusion, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Pass a table to narrow it' implies the exploratory/overview use case, and 'Read-only' signals it is for inspection rather than mutation. However, it never says when to choose this over db_query or when not to use it, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
device_screenshotScreenshot a simulator or deviceARead-only
Full-resolution screenshot of the booted iOS Simulator (platform ios) or a connected Android device/emulator (platform android), without window chrome. Use it to explain or debug app UI, then find the view in code.
| Name | Required | Description | Default |
|---|---|---|---|
| platform | No | ios |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true already telling the agent this is a safe read, the description adds real behavioural context: full resolution, no window chrome, and the implicit prerequisite that a simulator must be booted or a device connected. It stops short of saying how the image is returned (inline vs file path), which would be the remaining useful disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core action and scope, then the usage hint. Every clause earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with no output schema, the description covers target, resolution, chrome behaviour and usage intent. The only real gap is how the capture is delivered back (inline image vs saved path), which the absent output schema leaves to the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the schema only lists the enum values, but the description meaningfully maps them: platform ios = booted iOS Simulator, platform android = connected Android device/emulator. That is genuine semantic value beyond the bare enum, though it omits the default ('ios').
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (full-resolution screenshot) and its exact resource (booted iOS Simulator or connected Android device/emulator), plus the 'without window chrome' scope qualifier. An agent can distinguish this from browser/window capture siblings such as look or list_windows without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use it to explain or debug app UI, then find the view in code' gives clear positive context for when to reach for this tool. It does not name alternatives or state exclusions (e.g., when to prefer browser_* or look instead), so it stops short of full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_windowsList visible windowsARead-only
List on-screen app windows (front-most first) with ids, app names and titles. Use before look when unsure which window the user means.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds genuinely useful behavior beyond that: the result ordering (front-most first) and the fields returned. It does not mention staleness or whether minimized/off-screen windows are excluded, so it falls just short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no waste: the first defines output and ordering, the second routes to the correct usage. The key scoping/ordering detail is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema present, the description compensates by naming the returned fields and their order. For a zero-parameter read-only listing, nothing an agent needs in order to invoke or interpret it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4; there is no argument semantics to explain and the schema is trivially complete. The description correctly implies a no-argument listing rather than suggesting any filtering options.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (on-screen app windows) and names the returned fields (ids, app names, titles) plus ordering (front-most first). It also differentiates from the sibling `look`, so an agent can distinguish it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly prescribes when to call it: 'Use before `look` when unsure which window the user means.' This names the alternative tool and the exact condition that selects this one, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lookLook at the screenARead-only
Take a screenshot of what the user is looking at and return it as an image. By default captures the front-most window that is not the terminal (target "window"). Use target "screen" for a whole display. Optionally pick a window by app name/title or window id.
| Name | Required | Description | Default |
|---|---|---|---|
| app | No | App name or window-title substring, e.g. "Xcode", "Simulator", "Chrome" | |
| target | No | Capture a single window or an entire display | window |
| display | No | Display number for target screen (1 = main) | |
| window_id | No | Exact window id from list_windows | |
| under_cursor | No | Capture the window under the mouse pointer instead of the front-most one (falls back if it's the terminal) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so safety is covered. The description adds real behavioral context beyond that: the default excludes the terminal window, the screen target captures a whole display, and windows can be selected by app/title/id. It stops short of mentioning fallback edge cases or multi-display behavior in detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the core action before the capture-target nuances. Little waste, though the target explanation mildly overlaps the schema's own description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description states the return is an image, covering the output form. For a 5-param, all-optional read tool with a safe annotation, this is largely complete; only deeper fallback/multi-monitor nuances are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so each parameter is already documented in the schema. The description reinforces default behavior ('target "window"' by default) and the screen alternative, but adds no syntax or format detail beyond what the schema already provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Take a screenshot of what the user is looking at and return it as an image.' This clearly identifies a screen-capture operation. It does not, however, differentiate itself from siblings like device_screenshot or select_region, which also capture visual content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the default capture behavior (front-most non-terminal window) and how to switch to a whole display via target 'screen', which guides invocation. But it offers no guidance on when to prefer this tool over siblings such as device_screenshot or select_region, leaving alternatives to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
select_regionLet the user select a regionARead-only
Show the native crosshair so the user can drag a rectangle around the exact part of the screen they mean (press Space to pick a window instead, Esc to cancel). Blocks until they finish. Tell the user to make the selection.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so safety is covered. The description adds genuinely useful behavior beyond that: the call blocks until the user finishes, and the Space/Esc escape hatches. It does not say what is returned when the selection completes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with zero waste. The core action is front-loaded, with the modifier keys and blocking behavior following in priority order.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, non-destructive interaction tool with no output schema, the blocking behavior, cancel path, and user instruction cover nearly everything an agent needs. The one gap is what the call yields on success (coordinates, image, or confirmation).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is nothing for the description to clarify, and it does not need to compensate for any undocumented inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (show native crosshair) and the exact outcome (user drags a rectangle around the part of the screen they mean). It also distinguishes itself from the window-selection path by noting Space picks a window instead, which separates it from siblings like device_screenshot and list_windows.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear in-context guidance: use this to let the user define a region, press Space for window selection, Esc to cancel, and instruct the user to make the selection. It does not explicitly contrast when to prefer this over device_screenshot or look, so it falls short of full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_voiceStart voice modeA
Launch the voice overlay for this project: the user holds Option+Space, asks a question out loud, and the answer appears in a bubble above their mouse pointer.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does add genuine context: it is a UI overlay, the user drives it via Option+Space, and output surfaces as a bubble near the pointer. However it omits the single most important behavioral fact for an agent β whether this call returns immediately or blocks awaiting the spoken question and answer β and says nothing about how to terminate voice mode.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action, with no redundant or filler content. Every clause conveys distinct information about what launching the overlay means in practice.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with no output schema and no annotations, the description covers purpose and interaction flow adequately. The remaining gap is the synchronous/asynchronous behavior of the call, which an agent would need in order to sequence it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline of 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Launch') and resource ('the voice overlay for this project') and describes the resulting interaction model. It is unmistakably distinct from every sibling tool, all of which are browser/window/screenshot operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by explaining the voice interaction flow, but never states when an agent should choose this tool versus alternatives, nor any preconditions (e.g., whether voice mode must be enabled first or whether an overlay is already running).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
17 tool updates
v0.4.1- First observed
browser_click - First observed
browser_elements - First observed
browser_navigate - First observed
browser_open - First observed
browser_read - First observed
browser_scroll - First observed
browser_switch_tab - First observed
browser_tabs - First observed
browser_type - First observed
clipboard_image - First observed
db_query - First observed
db_schema - First observed
device_screenshot - First observed
list_windows - First observed
look - First observed
select_region - First observed
start_voice
TDQS
Scored across 17 tools
Tools cluster into clear domains (browser control, screen/device capture, database, voice), and the capture tools each target a distinct source (screen, simulator/device, clipboard, user selection). Minor overlap between browser_open (new tab) and browser_navigate (active tab), and between browser_switch_tab and browser_tabs, but descriptions disambiguate them.
Browser tools share a consistent browser_ prefix and db tools a db_ prefix, mostly verb_noun. The vision tools deviate slightly with a bare verb ('look') alongside verb_noun forms (list_windows, select_region, device_screenshot), but all remain snake_case and readable.
17 tools is slightly heavy but justified: browser control legitimately needs ~8 tools, capture needs several sources, database needs schema+query, plus voice. No tool feels redundant or out of scope.
Covers browser navigation, page reading, element interaction, multiple capture sources, DB introspection, and voice. Gaps are minor and mostly by design (read-only DB, no explicit back/forward history navigation), so agents can work around them.
Maintenance
Related MCP Connectors
Use your Mac, Windows or Linux computer from ChatGPT, Claude or Codex: files, commands, documents.
Mac & Windows: let ChatGPT, Claude & Cursor use your email, calendar, iMessage, Teams, files. Free.
Screenshot, PDF and HTML-to-image rendering API so Claude and Cursor can see any web page.
Screenshot, PDF and HTML-to-image rendering API so Claude and Cursor can see any web page.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables Claude to see and interact with any macOS application using natural language commands. Perfect for testing Mac applications, UI automation, and app development with AI assistance.33MIT
- AlicenseAqualityDmaintenanceEnables AI agents to capture and analyze screenshots of macOS applications, windows, or the entire screen using local (Ollama) or cloud-based AI vision models, with non-intrusive, fast screen capture via Apple's ScreenCaptureKit.312 npm2MIT
- AlicenseNot gradedqualityBmaintenanceScreenshot and diagram tool for AI agents. Capture and annotate screenshots to show Claude what you mean β or let the agent render Mermaid diagrams and open them for visual review. Approve, annotate, or request changes with text feedback. Built-in review mode with structured responses. CLI and MCP server for Claude Code, Cursor, Windsurf, Cline. macOS, open source, free.12 npm327MIT
- AlicenseNot gradedqualityCmaintenanceProvides AI coding assistants with privacy-first, synchronized screen context including cursor position, active window metadata, and visual verification to aid in building and debugging graphical interfaces.MIT