Skip to main content
Glama

UI-Venus MCP

A remote-first, cross-platform Computer-Use MCP server for AI agents — unified GUI automation across Windows, Linux, macOS, Android, iOS and browsers. Its primary job is not the local machine: it drives remote, awkwardly-accessed targets (SSH-only intranet hosts, no-internet boxes, locked-down sessions) with zero installs on the target side. Structured APIs first (UIA / AX / AT-SPI / UIAutomator / XCUITest / DOM), vision grounding as the fallback, with UI-Venus-2-9B (official protocol, W8A8 build verified) as the default — and pluggable — vision provider.

Remote targets are the main path: one command onboards a new Windows host (scripts/win-remote/bootstrap-remote.mjs), one env var registers any number of them (CUMCP_REMOTES). Verified on real hardware: the vision agent autonomously computed 7×3=21 on a remote, intranet, SSH-only Windows 11 box — twice. Full playbook: docs/remote-targets.md.

中文文档:docs/README.zh-CN.md

                    ZCode / Claude Code / Codex / OpenAI Agents / your agent
                                        │
                              ┌─────────┴─────────┐
                              │   Text or VLM     │   (any model, any size)
                              └─────────┬─────────┘
                                        │ MCP (stdio / Streamable HTTP)
                                        ▼
                     ┌──────────────────────────────────────┐
                     │      Computer-Use MCP Server         │
                     │  Orchestrator · Fusion Locator ·     │
                     │  Verifier · Recorder · DSL · Guard   │
                     └──────────────────┬───────────────────┘
                                        │
                              Platform Router (per-device sessions)
        ┌──────────┬──────────┬────────┴───┬────────────┬───────────┐
        ▼          ▼          ▼            ▼            ▼           ▼
     Windows     Linux      macOS      Android        iOS       Browser
   UIA+SendInput AT-SPI+    AX+     UIAutomator  XCUITest/   Playwright/
   PowerShell   xdotool   SystemEvents   +adb      WDA+simctl     CDP
        └──────────┴──────────┴────────────┴────────────┴───────────┘
                                        │  vision fallback (spec §4)
                                        ▼
                        UI-Venus-2-9B (remote GPU, OpenAI-compatible)

The product definition: give any AI agent — text-only or multimodal, small or large — unified, cross-platform, structure-first, vision-fallback computer-use capabilities. Not "an AI mouse for Windows", and not "UI-Venus wrapped in a few click APIs": UI-Venus acts as the cross-platform GUI expert (visual grounding, next-action decision, visual verification), while platform adapters execute reliably through native semantics whenever they exist.

Highlights

  • 17 MCP tools — targets, state, inspect, screenshot, locate, action, step, execute_task, verify, recorder (start/stop/to_script), run_script, run_ui_test, get/cancel_task. Full contract in docs/api.md.

  • Four agent modes (spec-level): delegate (autonomous loop for text-only agents), assist / direct (your agent stays in control; MCP provides primitives), auto (structured-first with per-step vision fallback).

  • Structure-first execution: click(elementRef) becomes UIA Invoke on Windows, AXPress on macOS, AT-SPI doAction on Linux, UIAutomator tap on Android, XCUITest tap on iOS, Playwright click in browsers. Raw coordinates are the last resort — and when vision produces a point, the fusion locator snaps it back onto a structured element before acting.

  • Cross-platform coordinate system: vision-normalized [0,1000] ↔ screenshot px ↔ logical points ↔ physical pixels, with DPI/Retina/density handled and unit-tested.

  • Honesty by construction: missing permissions, offline devices, Wayland restrictions, unsigned WDA — reported as permission_required / restricted / BLOCKED, never faked as success.

  • Anti-stagnation loop guard: same-action / same-element / same-screen (perceptual hash) detection with a recovery ladder (re-observe → alternate strategy → honest failure).

  • Security: app/device/action/domain allowlists + sensitive-action detection (删除/支付/转账/install…) parking tasks in WAITING_CONFIRMATION with confirm tokens.

  • Recorder → portable DSL: recordings become YAML scripts with semantic locators (never click(432,621); sleep(2)), replayable across platforms; a UI-test runtime reports PASS / FAIL / SKIP / BLOCKED.

  • Remote GPU architecture: the 9B vision model runs on a server (OpenAI-compatible endpoint); clients only send screenshots — phones and laptops need no VRAM.

  • Remote Windows control (spec §13): drive a Windows box over SSH from macOS/Linux with nothing installed on the target — a hidden session-1 queue agent runs the inbox PowerShell bridges; screenshots SCP back, SendInput goes forward. Verified on real hardware: the vision agent autonomously computed 7×3=21 in the remote Windows calculator (scripts/win-remote/remote-agent-e2e.mjs).

Related MCP server: screen-use

Quick start

# packaged (GitHub Release) — see docs/install/README.md for ZCode & clients:
npm install github:q1820926174-cpu/UI-Venus-MCP

# or from source:
git clone https://github.com/q1820926174-cpu/UI-Venus-MCP.git
cd UI-Venus-MCP
pnpm install
pnpm build

export VENUS_BASE_URL=http://<gpu-host>:8300/v1   # OpenAI-compatible
export VENUS_API_KEY=<your-key>
export VENUS_MODEL=UI-Venus-2-9B-W8A8

# stdio (ZCode / Claude Code / Codex)
node dist/index.js

# Streamable HTTP (remote agents / LAN)
node dist/index.js --http --port 8765

Smoke-test the vision endpoint (one calibration request):

pnpm smoke:venus

ZCode / Claude Code config

{
  "mcpServers": {
    "ui-venus-mcp": {
      "command": "node",
      "args": ["/absolute/path/to/UI-Venus-MCP/dist/index.js"],
      "env": {
        "VENUS_BASE_URL": "http://<gpu-host>:8300/v1",
        "VENUS_API_KEY": "<your-key>",
        "VENUS_MODEL": "UI-Venus-2-9B-W8A8"
      }
    }
  }
}

More configs (HTTP transport, Codex, OpenAI Agents): examples/mcp-config.md.

Usage patterns

Assist mode — a multimodal agent drives; the MCP executes and grounds:

computer_inspect  → screenshot + UI tree (+ optional vision description)
computer_locate   → "关闭按钮" → {element, point}   (structured first, vision fallback)
computer_action   → { type: "click", element: {...} }
computer_verify   → structured assertion, else vision verdict

Delegate mode — a text-only agent hands over the whole task:

{
  "tool": "computer_execute_task",
  "arguments": {
    "target": { "type": "device", "platform": "android", "deviceId": "emulator-5554" },
    "task": "打开设置,将Wi-Fi打开",
    "mode": "delegate",
    "maxSteps": 30
  }
}

Internally: Observe → Plan → Locate → Execute → Observe → Verify → Recover/Finish, with the state machine, security gates, and loop guard enforcing honest termination. Poll with computer_get_task, cancel with computer_cancel_task; sensitive actions return WAITING_CONFIRMATION + confirmToken.

Record → script → replay:

computer_record_start → (drive the app via computer_action) → computer_record_to_script
name: disable-auto-update
steps:
  - locate: { role: button, name: 设置 }
    action: click
  - locate: { name: 自动更新 }
    action: toggle
    value: false
assert:
  - element: { name: 自动更新 }
    property: { checked: false }

Replay with computer_run_script, or run as a UI test (PASS/FAIL/SKIP/BLOCKED) with computer_run_ui_test. DSL reference: examples/dsl/toggle-autoupdate.yaml.

Platform support matrix

Capability

Windows

Linux

macOS

Android

iOS

Browser

Screenshot

✅ CopyFromScreen

✅ import/scrot/grim

✅ screencapture

✅ screencap

✅ simctl/WDA

✅ Playwright

Accessibility tree

✅ UIA

✅ AT-SPI (pyatspi)

✅ AX (System Events)

✅ UIAutomator dump

⚠️ WDA/idb

✅ aria snapshot

Semantic actions

✅ UIA patterns

✅ doAction/setText

✅ AXPress/AXValue

✅ dump+tap center

⚠️ WDA elements

✅ role/text/testid locators

Global input

✅ SendInput

✅ xdotool / wtype

✅ CGEvent

✅ adb input

⚠️ WDA/idb

✅ keyboard/mouse

App control

✅

✅

✅ open/AppleScript

✅ monkey/am

✅ simctl/WDA

n/a

Unicode typing

✅ KEYEVENTF_UNICODE

✅ xdotool type

✅ CGEvent unicode

⚠️ ADBKeyboard IME

✅ WDA

✅

Vision fallback

✅

✅

✅

✅

✅

✅

⚠️ = capability depends on optional tooling/signing — the server reports exactly what is missing (capabilities.notes), per the honesty rules. Platform-specific setup: docs/install/ — macOS · Windows · Linux · Android · iOS.

The UI-Venus provider (verified endpoint contract)

The default provider speaks the OpenAI chat/completions format with chat_template_kwargs: {enable_thinking: false} and grounds natural-language elements to [0,1000]-normalized points. Calibrated against the W8A8 build (2026-09-28): button at pixel (1300,740) on 1920×1080 → [676, 680]; the official grounding prompt and the live verification live in ui-venus-service/README.md and tests/e2e/venus-live.test.ts.

Swapping providers (UI-TARS, Qwen-GUI, any VLM): implement ComputerVisionProvider (src/providers/types.ts) and register it in the registry — nothing else knows which model is in use.

Development

pnpm test                  # full hermetic suite (unit + integration + browser E2E)
pnpm test:e2e:macos        # real macOS E2E (needs permissions) — RUN_MACOS_E2E=1
pnpm test:e2e:venus        # live grounding E2E against your endpoint — RUN_VENUS_LIVE=1
pnpm typecheck && pnpm build

QA status per platform (what is really verified vs mocked): docs/qa-report.md. Architecture deep-dive: docs/architecture.md.

Repository layout

src/
├── mcp/            17 tools, server assembly
├── orchestrator/   task loop, fusion locator, verifier, loop guard, security
├── providers/      ComputerVisionProvider + UI-Venus implementation + mock
├── platforms/      adapter interface, router, macos/linux/windows/android/ios/browser
├── scripting/      YAML DSL + runner      ├── recorder/   semantic recording
├── coordinate/     space transforms      ├── screenshot/ pipeline (ROI/hash/JPEG)
├── core/           types, actions, state machine, errors
ui-venus-service/   endpoint contract & serving reference
tests/              unit · integration (mock+browser+mcp) · e2e (opt-in real)

License

MIT — see LICENSE.

Available Tools

17 tools
computer_actionExecute one actionB

Execute a single unified action (spec §17): click/double_click/right_click/long_press/move/drag/swipe/scroll/type/set_value/clear/press/hotkey/select/toggle/invoke/launch_app/terminate_app/focus/back/home/wait. Provide element (from computer_locate) or point. Records to the active recorder if recording.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesUnified action
targetNoUnified device target (spec §14). Defaults to this machine with auto-detected platform.
confirmTokenNoConfirmation token from a WAITING_CONFIRMATION result

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the entire behavioral burden. It discloses one useful trait ('Records to the active recorder if recording') but omits that several actions are destructive or irreversible (terminate_app, launch_app, clear), that finish/fail terminate the task, and any mention of the confirmation flow implied by the confirmToken parameter. These are significant undisclosed behaviors for a mutation-heavy tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then the action vocabulary, then the sourcing rule and recorder behavior. The long action list is dense but defensible since it mirrors the schema discriminators; only the missing finish/fail entries make it slightly untidy rather than bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a unified 23-variant action tool with no annotations and no output schema, the description covers the basics: what it does, what to pass, and recorder side effects. It stops short of the confirmation/approval flow, irreversible-action warnings, and terminal actions, leaving real gaps given the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the action union/point/element shapes are fully structured, so baseline is 3. The description adds marginal value by explaining that element comes from computer_locate and can be substituted with a point, but it does not explain the target object, confirmToken usage, or the omitted finish/fail action variants.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Execute a single unified action') and enumerates the supported actions, which is far more informative than the bare title. The word 'single' implicitly separates it from multi-step siblings like computer_execute_task, but no sibling is named explicitly, so differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The line 'Provide element (from computer_locate) or point' gives workflow guidance by pointing at the locating tool, which is genuinely useful. However, it never says when to use this tool versus computer_step, computer_execute_task, or computer_run_script, and gives no exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computer_cancel_taskCancel taskB

Request cancellation of a running task; also cancels its pending confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses two behaviors: that cancellation is a *request* (implying asynchronous, non-guaranteed termination) and that it cascades to a pending confirmation. It omits whether the task is force-terminated, what happens if it already finished, and any permission requirements, so it is partially but not fully transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the primary action front-loaded and the side effect appended. No filler, no repetition of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter mutation tool with no annotations and no output schema, the description covers the core action and one cascade effect but leaves open what is returned, whether cancellation is acknowledged or confirmed, and how taskId is sourced. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and there is one required parameter, taskId, which the description never explains (format, where to obtain it, e.g. from computer_get_task or computer_list_targets). The word "task" gives implicit context but no real added meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Request cancellation of a running task") and adds a non-obvious scope note that it also cancels the task's pending confirmation. It does not explicitly contrast itself with siblings like computer_execute_task or computer_get_task, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase "running task" weakly implies applicability, but there is no explicit when-to-use guidance, no mention of prerequisites, and no routing to alternatives for stopping or inspecting tasks. The agent must infer the trigger condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computer_execute_taskExecute a full GUI taskA

Delegate a natural-language task to the autonomous GUI agent loop (observe → decide → locate → execute → verify → recover). mode=delegate|auto. Returns immediately with a taskId unless wait=true. Sensitive actions park in WAITING_CONFIRMATION with a confirmToken — re-call with the token to approve.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoauto
taskYesNatural language task, e.g. 打开设置,将Wi-Fi打开
waitNoWait for completion instead of returning a taskId
targetNoUnified device target (spec §14). Defaults to this machine with auto-detected platform.
languageNo
maxStepsNo
confirmTokenNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose two valuable traits: asynchronous return with a taskId unless wait=true, and the WAITING_CONFIRMATION/confirmToken approval flow. It omits failure/error behavior, timeout handling, what maxSteps bounds, and what the taskId enables afterwards, so the safety profile of this powerful autonomous tool is only partially covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly packed sentences, front-loaded with the core purpose and no filler. It is dense rather than bloated, though the parenthetical pipeline enumeration and mode clause make it slightly more compressed than ideally scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex autonomous-loop tool with no output schema and no annotations, the description covers the pipeline, the async return contract, and the confirmation flow well. Gaps remain around the assist/direct modes, the target object, and post-launch usage of the taskId (e.g., computer_get_task), which an agent driving a rich sibling set would benefit from.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 43%, so the description must compensate and it does for wait and confirmToken and partially for mode (only delegate|auto mentioned though the enum has four values). The task, target (nested object), language, and maxSteps parameters get no added semantic detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Delegate a natural-language task to the autonomous GUI agent loop" gives a specific verb (delegate) and resource (GUI agent loop), and enumerates the internal pipeline phases. It clearly distinguishes a high-level delegated task from the low-level siblings, but never names alternatives like computer_action or computer_step explicitly to route the choice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the delegation use case and mentions mode=delegate|auto plus the wait semantics, giving some context. However, it never states when to prefer this tool over computer_action, computer_run_script, or computer_run_ui_test, and it does not explain the assist/direct modes in the enum either.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computer_get_stateGet computer stateA

Cheap structural state: screen info, frontmost app, windows, capabilities. No screenshot.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNoUnified device target (spec §14). Defaults to this machine with auto-detected platform.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses the cost profile ('Cheap') and that no image data is returned, but says nothing about side effects, permissions, or whether the target machine must be reachable. The read-only nature is only implied by 'get'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single terse sentence fragment, front-loaded with the cost characteristic and the returned fields, with zero filler. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description usefully enumerates the return contents, covering the main gap. It omits any mention of the target parameter's behavior or error conditions, but for a read-only state probe that is a minor omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the nested target object is self-documented in the schema, so baseline 3 applies. The description adds no syntax, default, or target-selection guidance beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific operation and enumerates exactly what state is returned (screen info, frontmost app, windows, capabilities), so the agent knows the payload without opening a schema. The explicit 'No screenshot' clause cleanly distinguishes it from the computer_screenshot sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Cheap' and 'No screenshot' imply the tradeoff against pixel-capturing siblings, but the description never states when to prefer this over computer_inspect or computer_screenshot, nor does it name an alternative. Usage is inferable rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computer_get_targetGet target detailsB

Resolve one target and report its capabilities, permission state and platform details.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNoUnified device target (spec §14). Defaults to this machine with auto-detected platform.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It states what is reported (capabilities, permission state, platform details) and implies a read-only resolve operation, but it does not disclose permissions required, error behavior, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with zero waste, front-loading the core action and the reported outputs. It is appropriately sized for the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool with a fully described schema and no output schema, the description covers the purpose and return content adequately. However, given no annotations, it leaves gaps around usage context, safety guarantees, and failure modes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the nested target object is already fully documented in the schema. The description adds no syntax or format details beyond what the schema provides, so a baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb 'resolve' and a specific resource 'one target', and lists the reported details. It is clear what the tool does, but it does not explicitly distinguish itself from siblings like computer_list_targets or computer_get_state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The phrase 'one target' implies a single-target lookup, but there are no explicit when/when-not conditions or named sibling alternates.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computer_get_taskGet task statusB

Fetch a task record: state, steps, outcome, pending confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskIdYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Fetch' implies a safe read and the listed fields give useful insight into the response shape (compensating for the absent output schema), but there is no disclosure of behavior on missing/expired tasks, permissions, or whether state is a live poll or cached snapshot.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the return contents front-loaded after the verb. No wasted words, though it is a fragment and could carry one more clause of routing or input guidance at little cost.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description partially compensates by naming the record fields, which is valuable. However, it leaves the required taskId unexplained and gives no usage context, so it is only minimally adequate for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required parameter taskId is undocumented in both the schema and the description. The description never mentions the identifier, its expected format, or how to obtain it, leaving the only input fully unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (fetch) and resource (task record) and enumerates the record's contents (state, steps, outcome, pending confirmation). It is clear what the tool does, though it does not explicitly differentiate itself from siblings like computer_execute_task or computer_cancel_task.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites (e.g., where a valid taskId comes from), and no mention of alternatives such as computer_get_state or computer_step. The agent must infer all routing context from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computer_inspectInspect screen (vision)B

Full observation: screenshot + UI tree, optionally with a vision-model description focused on a topic.

ParametersJSON Schema
NameRequiredDescriptionDefault
focusNoOptional focus instruction for the vision model
targetNoUnified device target (spec §14). Defaults to this machine with auto-detected platform.
includeTreeNo
maxTreeDepthNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry behavioral disclosure, and it does state what the tool returns (screenshot, UI tree, optional vision description) and that vision is opt-in. It omits whether the vision-model call adds latency/cost, whether the operation has side effects (it appears read-only but this is never stated), and permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that packs the key information with no filler. It is dense to the point of being terse — the optional/conditional nature of the vision description is the only nuance conveyed — but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-annotation, no-output-schema tool with a nested target object and two undocumented params, the description covers the return composition adequately. It remains incomplete on side-effect/permission posture and on the tree-depth and includeTree controls an agent would need to call it precisely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% and the nested target object is documented in-schema. The description usefully explains the 'focus' parameter as a vision-model topic instruction, but says nothing about includeTree or maxTreeDepth, leaving two parameters undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Inspect screen') and enumerates the composite output: 'screenshot + UI tree, optionally with a vision-model description'. This distinguishes it from the bare computer_screenshot sibling by conveying it is a fuller, multi-source observation. It stops short of explicitly contrasting itself with computer_get_state or computer_locate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Full observation' implies this is the superset call to reach for when a complete picture is needed, and the optional focus scopes it to a topic. However, no when-to-use vs when-not guidance is given and no alternative sibling (e.g., computer_screenshot for a quick capture, computer_get_state for non-visual state) is named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computer_list_targetsList computer targetsA

Enumerate every device/platform this MCP server can currently serve (local desktops, android devices, ios simulators, browser), with availability and honest capability notes.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations at all, the description carries the full burden and does useful work: it discloses that results are runtime-dependent ('currently serve'), that availability is reported, and that capability notes are 'honest' (i.e., caveats may be included). It does not mention permissions or auth requirements, which is the main remaining gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with the scope front-loaded and no filler. Every clause (what is listed, which platforms, availability, capability notes) carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter discovery tool with no output schema, the description adequately characterizes what comes back (targets with availability and capability notes). A brief note on output shape or that it is safe/read-only would make it fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4; there is nothing for the description to disambiguate. It correctly adds no parameter discussion, which is appropriate here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (enumerate) and resource (every device/platform the server can serve) and enumerates the categories (local desktops, android, ios simulators, browser). This clearly separates it from the singular sibling computer_get_target, which retrieves one target's details.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'can currently serve' plus 'availability' implies this is the discovery entry point to call before acting on a specific target, but the description never states when to prefer it over computer_get_target or computer_get_state, nor any prerequisite ordering.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computer_locateLocate element (fusion)A

Find an element by natural language or structured descriptor. Structured (UIA/AX/AT-SPI/DOM) first, UI-Venus vision grounding as fallback; vision points are snapped back to structured elements when possible. Coordinates in results are screenshot pixel space; use with computer_action.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNoUnified device target (spec §14). Defaults to this machine with auto-detected platform.
descriptorNo
instructionYesNatural language element description, e.g. "关闭按钮" or "the login button"
preferStructuredNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it does substantial work: it discloses the resolution strategy (structured first, vision fallback), the snapping behavior, and the coordinate space of results. It omits whether the call is read-only/side-effect-free and any auth or performance caveats, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, zero filler, and the core identification plus the most decision-relevant behavior (fallback order) is front-loaded. Every sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description steps in to explain the return coordinate space, which is the key thing an agent needs to pass results to computer_action. It is nearly complete, missing only target-selection guidance for the nested object and explicit read-only confirmation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%, so the description must compensate, and it only partially does: it explains the instruction/descriptor duality and the result coordinate space, but says nothing about the target block (platform, deviceId, cdpEndpoint) or preferStructured, which are the exact fields left undocumented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Find an element') and clarifies the two input modes (natural language vs structured descriptor). It also hints at its relationship to computer_action, though it does not explicitly differentiate itself from siblings like computer_inspect or computer_get_state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: structured lookup is attempted first with vision grounding as fallback, and it names computer_action as the downstream consumer of the result. No explicit 'when not to use this' or named alternative for locating-only scenarios, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computer_record_startStart recordingB

Start recording executed actions on a target session (semantic evidence: elements, hashes, apps).

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNoUnified device target (spec §14). Defaults to this machine with auto-detected platform.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden. It usefully discloses what is captured (elements, hashes, apps), which is real behavioral context, but says nothing about side effects, persistence of the recording, resource cost, or whether starting while a recording is already active is an error.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with the key action first and a compact parenthetical detail. No wasted words, though it is arguably under-specified rather than elegantly terse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema and no annotations, the definition covers the basic purpose but omits the return value (e.g. a recording handle needed by computer_record_stop) and any lifecycle context, leaving gaps an agent would have to guess at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single nested target object is fully documented in the schema itself. The phrase 'target session' loosely maps to the target parameter but adds no syntax, defaults, or format meaning beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb+resource ('Start recording executed actions') and scopes it to a target session, with a parenthetical clarifying what evidence is captured. It is distinguishable from siblings like computer_screenshot or computer_inspect, but it never references its obvious counterpart computer_record_stop.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance and no conditions or prerequisites. It does not say that a recording session must later be closed with computer_record_stop, nor when recording is preferable to computer_run_script or computer_run_ui_test.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computer_record_stopStop recordingC

Stop recording and return the recorded entries.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNoUnified device target (spec §14). Defaults to this machine with auto-detected platform.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It hints at a return value ("recorded entries") but says nothing about error behavior when no recording is in progress, whether the recording session is cleared, whether the target must match the one passed to record_start, or any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the action front-loaded and no wasted words. It is tight, though its brevity is partly a result of under-specification rather than deliberate economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-param stop tool with no output schema and no annotations, the definition is minimally adequate: the action and a rough sense of the return are conveyed. It omits the recording lifecycle context (matching target, no-active-recording behavior) that an agent would need to call it reliably.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single nested target object is already documented in the schema as the unified device target, so the description adds nothing beyond it. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Stop recording") and adds the outcome ("return the recorded entries"), so the agent knows what it does without opening the schema. It doesn't explicitly name or distinguish itself from siblings like computer_record_to_script, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no mention of pairing with computer_record_start, and no statement about what happens if no recording is active. The agent must infer the lifecycle from the sibling names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computer_record_to_scriptRecording → automation scriptC

Convert the current/last recording into a portable YAML automation script with semantic locators (never raw coordinates).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNorecorded-flow
stopNo
targetNoUnified device target (spec §14). Defaults to this machine with auto-detected platform.

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose that the generated script uses semantic locators rather than raw coordinates and is portable YAML, which is useful output context. However, it omits side effects, permissions, whether it stops or modifies the recording, and what the tool returns or writes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no redundant text. It efficiently conveys the core action, output format, and a distinguishing quality of the generated script.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a nested target object, no output schema, no annotations, and low parameter description coverage, the description is too thin. It explains the output artifact but not the inputs, side effects, or workflow position needed to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has only 33% description coverage across three parameters, including a nested target object, yet the description gives no information about name, stop, or target fields. It does not compensate for the low schema coverage, so parameter meaning remains largely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific transformation: converting the current/last recording into a portable YAML automation script. It names the output format and a key output property (semantic locators, never raw coordinates), so the agent can identify the tool's purpose. It does not explicitly distinguish itself from siblings such as computer_record_stop or computer_run_script.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'current/last recording' implies this is used after a recording exists, but there is no explicit when-to-use guidance, no prerequisites, and no alternatives named. It does not say whether to stop recording first, which sibling to use instead, or in what workflow context this conversion should be invoked.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computer_run_scriptRun automation scriptC

Execute a portable YAML automation DSL script (semantic locators, cross-platform) against a target.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoPath to a .yaml script file
yamlNoInline YAML script
targetNoUnified device target (spec §14). Defaults to this machine with auto-detected platform.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden, and it discloses very little: nothing about side effects on the target device, permissions, blocking/timeout behavior, or what happens on script failure. 'Portable', 'semantic locators', and 'cross-platform' describe the DSL flavor but not execution behavior an agent needs to invoke this safely.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler and no repetition of the title. It is arguably too terse for a tool with a nested target object and two mutually alternative script inputs, but it wastes nothing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a fairly complex tool (nested target object with six target types, two alternative script inputs, no output schema, no annotations), the description omits how the script is supplied, what a run produces, and any failure/timeout semantics. An agent would need to open the schema and guess at the rest.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the nested target object is fully documented in the schema, so the baseline of 3 applies. The description's 'against a target' adds no syntax or meaning beyond the schema, and it never clarifies the path-vs-yaml alternative for supplying the script.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description pairs a specific verb (Execute) with a specific resource (a portable YAML automation DSL script) and its scope (against a target), so an agent understands this runs whole scripts rather than single actions. It does not, however, differentiate itself from near-siblings like computer_run_ui_test, computer_execute_task, or computer_action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to reach for this tool versus computer_action (single step), computer_execute_task, or computer_run_ui_test, nor any prerequisite or exclusion. The only hint is the phrase 'against a target', which merely restates the schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computer_run_ui_testRun UI test suiteB

Run one or more YAML UI test cases; returns PASS/FAIL/SKIP/BLOCKED per case with step evidence. BLOCKED = environmental restriction (permissions/offline/unsupported), reported honestly.

ParametersJSON Schema
NameRequiredDescriptionDefault
casesYes
targetNoUnified device target (spec §14). Defaults to this machine with auto-detected platform.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does add real value by defining the result vocabulary and explaining that BLOCKED means an environmental restriction reported honestly, which is genuine behavioral context. However it says nothing about side effects on the target, whether tests mutate state, required permissions, or execution timing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences that lead with what the tool does and follow with return semantics. Every clause earns its place and the BLOCKED definition is front-loaded where an agent will read it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description rightly explains return statuses and does so adequately. But for a tool with a nested target object and no annotations, it omits target selection, environmental requirements, and whether execution has side effects, leaving meaningful gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%; the nested target object carries its own partial description. The description only implies the cases array by saying 'one or more YAML UI test cases' and adds nothing about target, browser, platform, or device selection, so it does not compensate for the uncovered half.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: run YAML UI test cases, returning per-case PASS/FAIL/SKIP/BLOCKED with step evidence. This is distinguishable from the sibling computer_run_script, though the description never explicitly contrasts the two.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to choose this over computer_run_script, computer_execute_task, or other siblings, and no prerequisites or preconditions stated. The agent must infer usage from the resource name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computer_screenshotTake screenshotC

Capture the screen/window/display. Returns metadata + image content; optionally saves to a file.

ParametersJSON Schema
NameRequiredDescriptionDefault
regionNo
targetNoUnified device target (spec §14). Defaults to this machine with auto-detected platform.
savePathNo
windowIdNo
displayIdNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries the full behavioral burden. It usefully states that it returns metadata plus image content and can optionally save to a file, but it omits permissions, read-only status, overwrite behavior, and target requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The one-sentence description is front-loaded and free of waste. It is concise, though arguably too sparse for a tool with complex nested parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, no annotations, and a nested target object, the description is incomplete. It does not explain return metadata, target selection, region capture, or the save-path behavior in enough detail for reliable invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20%, and the description does not explain the five parameters. It vaguely implies saving to a file but never covers region, target, windowId, or displayId.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Capture') and resource ('screen/window/display'), so the basic action is clear. However, it does not distinguish this tool from sibling computer_* tools or clarify which capture target is intended.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus alternatives such as computer_inspect or computer_locate. The description only says what it does, leaving selection to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computer_stepOne autonomous stepB

Observe + let the vision GUI expert decide the next action + execute it. Returns the decision and execution result. Useful for building your own loop in the calling agent.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalYes
dryRunNoDecide but do not execute
targetNoUnified device target (spec §14). Defaults to this machine with auto-detected platform.
languageNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the observe-decide-execute cycle plus what is returned ('the decision and execution result'). It does not say that the executed action is a real GUI mutation that may be irreversible, nor does it mention permission, session, or target-state requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core cycle and then the return value and intended usage. No filler, though the third sentence is a soft nudge rather than hard information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with a nested target object, no annotations, and no output schema, the description covers the return shape ('decision and execution result') but omits safety/permission context and says nothing about the target or goal parameters. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50% (goal and language are undocumented, and the nested target object carries most of the documentation burden), so the description needs to compensate and does not. It adds no meaning about goal semantics, language, or how target selection affects behavior; the only hint at dryRun is the word 'execute' in the summary sentence.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete verb chain ('Observe + let the vision GUI expert decide the next action + execute it'), which pins down the resource and scope as a single autonomous GUI step. The title 'One autonomous step' reinforces this and implicitly contrasts with whole-task siblings like computer_execute_task, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Useful for building your own loop in the calling agent' implies the usage context — the caller wants step-level control rather than a full autonomous run. However, it never states when NOT to use it or names the alternative (e.g. computer_execute_task or computer_action) that handles those cases, leaving the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

computer_verifyVerify stateB

Verify a goal against the current screen: structured UI-tree heuristics first, UI-Venus vision verification as fallback. Supports structured assertions (checked/exists/value/textVisible).

ParametersJSON Schema
NameRequiredDescriptionDefault
goalNo
targetNoUnified device target (spec §14). Defaults to this machine with auto-detected platform.
assertionNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses genuine internal behavior (structured UI-tree heuristics first, vision fallback) which is valuable. However, it does not state what happens on a failed verification (throw vs return), whether the vision fallback has cost/latency implications, or any return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact clauses with no filler, and the verification strategy is front-loaded ahead of the assertion syntax. Efficient, though the second clause mixes strategy detail with assertion trivia.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a nested target object and a nested assertion object, no output schema, and no annotations, the description is thin. It explains the verification approach but leaves the target parameter, the goal semantics, and the verification return/outcome entirely undocumented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is low (33%), so the description should compensate, but it only enumerates some assertion properties (checked/exists/value/textVisible) and notably omits valueEquals/valueContains that exist in the schema. The goal and target parameters receive no explanation at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (verify) and resource (a goal against the current screen), and separates itself from a pure read by naming the verification mechanism (UI-tree heuristics then vision fallback). It stops short of naming which sibling (e.g., computer_get_state, computer_inspect) to use instead, so it is clear but not sibling-differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the reader infers this is for checking a goal/assertion after an action. There is no explicit when-to-use vs when-not guidance and no mention of the alternative verification paths offered by siblings like computer_inspect or computer_get_state.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 17 tool updatesv0.2.1
    • First observedcomputer_action
    • First observedcomputer_cancel_task
    • First observedcomputer_execute_task
    • First observedcomputer_get_state
    • First observedcomputer_get_target
    • First observedcomputer_get_task
    • First observedcomputer_inspect
    • First observedcomputer_list_targets
    • First observedcomputer_locate
    • First observedcomputer_record_start
    • First observedcomputer_record_stop
    • First observedcomputer_record_to_script
    • First observedcomputer_run_script
    • First observedcomputer_run_ui_test
    • First observedcomputer_screenshot
    • First observedcomputer_step
    • First observedcomputer_verify

TDQS

A3.5/5.0

Scored across 17 tools

Disambiguation4/5

Most tools have clearly distinct scopes (single action vs. one autonomous step vs. full task delegation; different observation depths; recording vs. script execution). Some boundaries among computer_action, computer_step, computer_execute_task, and computer_run_script could blur, but descriptions and parameter differences largely disambiguate them.

Naming Consistency5/5

All tool names use consistent computer_ snake_case with clear verb/noun or record_* patterns. There is no mixed casing or conflicting verb style.

Tool Count4/5

17 tools is slightly above the typical 3-15 sweet spot, but the breadth of GUI automation capabilities (targets, observation, action, tasks, recording, scripting, testing) justifies most tools. A few tools could potentially be consolidated but are reasonable for this domain.

Completeness5/5

The surface covers target discovery, observation, element location, action execution, autonomous task delegation, verification, cancellation, recording/scripting, and UI testing. There are no obvious dead ends for the stated GUI automation purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI coding agents to automate Windows desktop applications through semantic UI Automation instead of brittle coordinate clicks, with tools for discovering windows, finding controls by stable identifiers, and verifying actions.
    39 PyPI
    2
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to see, locate UI elements, and operate any Windows desktop app through natural language, using accessibility-tree matching with optional vision-model fallback, plus an autonomous visual loop with introspection and meta-learning.
    3
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI agents to automate real desktop applications across Windows, Linux, and macOS using incremental screen perception, accessibility trees, OCR, and window management, dramatically reducing token usage compared to screenshot-per-step approaches.
    61 PyPI
    3
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to operate local desktops and Chromium browsers through MCP tools, unifying accessibility trees, physical input, screenshots, DOM/ARIA, visual grounding, and result verification.
    1,112 npm
    MIT