Skip to main content
Glama
WARNING

Luda is deprecated. Use LCU (Linux Computer Use) instead. Install LCU.

Luda

Let your coding agent use a Linux desktop.

Luda gives agents tools to see applications, read their controls, click, type, and check what changed. Use it to fill forms, edit documents, work with files, or test a graphical application.

It runs on the Linux machine that owns the desktop. No hosted service or model API is required by Luda.

How it works

Luda has two parts:

  • Tools connect your agent to the desktop through MCP, a protocol supported by many coding agents.

  • A skill teaches the agent how to choose controls, enter text, verify results, and recover when something changes.

The agent can read accessible controls directly or work from screenshots. Actions report whether their result was verified, merely sent, or uncertain. Clicking “Save,” for example, is not itself proof that a file was saved.

Related MCP server: MacWright

Install: tools and skill together

Run these commands on the Linux machine whose desktop the agent will control. You need an existing Linux X11 desktop, Python 3.12+, and Git. Ubuntu 24.04 with XFCE is the tested starting point. Wayland and Xwayland are not supported.

git clone --branch v0.3.4 --depth 1 https://github.com/0xpolarzero/luda.git
cd luda
sudo bash scripts/install.sh --user "$(id -un)"

The installer installs the runtime and Linux dependencies, shows detected and supported agents, and lets you select one or several. It then registers the computer-use tools and installs their skill for your account. Restart/reconnect your agent and ask: “Use Luda to inspect my desktop.” The skill is discoverable by the agent; the client decides when to load it.

Already know your agent? Use --agent codex --yes, for example:

sudo bash scripts/install.sh --user "$(id -un)" --agent codex --yes

Use bash scripts/install.sh --list-agents for all supported identifiers. Other agents can use a portable tools-and-skill export. Installation delegates to pinned versions of Vercel skills and add-mcp. It preserves unrelated settings, updates Luda’s own skill and MCP entry, and does not install or authenticate the agent itself. Required installer tooling is supplied automatically; no separate Node installation is needed.

Codex connected to a VM through its built-in SSH connection? Run installation inside the VM, selecting the Linux account used by that connection. Codex's backend there loads the skill and tool registration. Running an ordinary ssh command from a local agent does not automatically load the VM's configuration.

Building a VM or machine image? Run the same installer as root with an explicit account and all supported agents:

bash scripts/install.sh --user YOUR_ACCOUNT --agent all --yes

The account must already exist; use --user root explicitly when root is the intended agent account. Supported agents do not need to be installed yet, and existing agent settings can already be present. Use repeated --agent NAME options to select a subset instead. No running desktop or agent credentials are needed during the build. Start the desktop before using the tools. See the image recipe and first-boot checks.

You can also ask your agent:

Read Luda's README and installation guide. Install the released version on the Linux desktop machine, register its MCP tools and complete skill for my agent account, preserve unrelated configuration, and verify readiness. If this is an image build, configure everything without requiring a running desktop.

Installation, upgrades and custom agents · Agent connection details · Release downloads

What can it do?

  • Work with applications: find and activate windows, inspect controls, operate menus, select items, and manage windows or workspaces.

  • Enter and check text: Unicode, multiple lines, selections, clipboard paste, and readback where the application supports it.

  • Use the screen: screenshots, clicks, drags, scrolling, and keyboard shortcuts. Optional OCR, image matching, and recording provide additional ways to observe.

  • Handle interruptions: wait for changes, cancel work, pause agent input, and recover owned input after a disconnect.

Watch the agent work with as little disruption as possible. Luda shows a distinct agent cursor where supported and tries to leave your mouse, keyboard, and foreground windows undisturbed. The cursor is drawn into the desktop image, so ordinary desktop viewers can display it.

Luda automatically chooses background actions or independent input where supported. When compatibility requires it, Luda uses ordinary foreground mouse/keyboard control, which can move your pointer, redirect keyboard input, and bring windows forward. Application behavior can also change focus. The agent uses the same tools with no mode selection. See input routing and evidence.

The optional browser provider opens a temporary Chromium session for ordinary web fields. Existing browser profiles are not attached automatically.

Optional editor add-on

Editor Bridge adds exact rich-text verification for applications built with ProseMirror. It is a separate installation with its own tools and skill, and the application's developer must also register the adapter. It is not included or enabled by installing core Luda. Most desktop tasks do not need it.

Package it into an environment

You can preinstall Luda in a workstation, container with a graphical session, or VM image. Your integration provisions the desktop and accounts, then runs Luda's installer to configure the chosen agents.

Use the environment packaging guide to make tools and skills available across folders for each agent account. There is no universal installation directory that every agent automatically discovers.

Learn more

I want to…

Read

Install or attach to a graphical session

Installation

Connect an agent and install its skill

Agent integrations

Understand the tools and their parameters

Tool reference

Read the agent's operating instructions

Core skill

Check supported backends and known limits

Backend support · Validation

Build packages or contribute

Development and releases

Application accessibility varies, and human input can race with an agent. Luda reports these limits rather than treating every dispatched action as success.

MIT license.

Available Tools

36 tools
desktop_activateB

Reveal a window from desktop_windows and establish input focus using the automatically selected route. Observe again afterward.

ParametersJSON Schema
NameRequiredDescriptionDefault
window_idYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It clearly conveys that the tool mutates UI state by revealing a window and moving input focus, and it warns that the route is automatically selected. However, it does not explain potential side effects, failure conditions, or permissions, which is a meaningful gap for a focus-changing tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The core action is front-loaded, and the follow-up observation instruction adds workflow value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description provides a basic workflow and names the source of the window, which is good for a single-parameter tool. But with no output schema and no annotations, it omits expected return behavior, what 'automatically selected route' means for reliability, and what to do if activation fails. It is adequate but leaves clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions pulling a window 'from desktop_windows', which hints that window_id identifies a window from that list, but it never explicitly maps window_id to the window being activated or clarifies the expected format. The parameter meaning is left mostly to inference from the schema title.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action: reveal a window and establish input focus. It names desktop_windows as the source of the window, which differentiates it from sibling tools like desktop_observe or desktop_focus_element, though the phrase 'automatically selected route' is somewhat opaque.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a workflow: get a window from desktop_windows, activate it, then observe again afterward. This gives helpful context about when to call the tool, but it does not explicitly state when to choose desktop_activate over alternatives such as desktop_focus_element or desktop_recover_input.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_applicationsA
Read-only

Find installed desktop applications by name, description or ID. Returns application_id and file/URI support; works while input is paused. Use an exact returned ID with desktop_launch.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true, so the description correctly avoids repeating that. It adds value by noting it works while input is paused and describing the return value (application_id and file/URI support), which are not in annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The purpose is front-loaded, and the usage hint is succinct. Every sentence contributes useful information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description appropriately mentions it returns application_id and file/URI support. It also provides the paused-input context, which is useful. It lacks details on limit behavior, but for a simple search tool, it is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains that query can search by name, description or ID, giving meaning to the query parameter. However, it does not explain the 'limit' parameter, leaving a gap in understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a verb 'find' and a resource 'installed desktop applications', with search fields (name, description, ID). It also notes it returns application_id and file/URI support, and hints at use with desktop_launch, which helps differentiate it from other siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use an exact returned ID with desktop_launch', giving a clear usage path. It also mentions it works while input is paused, providing a contextual condition. However, it does not specify when not to use it or compare with alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_chooseA

Choose an observed list/radio/combo option or a visible table cell and verify selection. A table cell selects its whole row. Successful results verify selection only, including already-selected results. If selection alone was requested, stop. Otherwise verify the requested application effect; some controls apply on selection. If the effect did not occur, inspect for an advertised activation action or Apply/Open control, use it, then verify the outcome. Default makes the choice exclusive; extend preserves other list or table-row selections. Scroll offscreen rows into view and inspect again; reacquire after sorting/filtering. Open collapsed options and inspect first. range_end_id selects an inclusive range of at most 50 visible list items or table rows, using two endpoints from the same inspection; reversed endpoints are allowed. Every intermediate item must be inspected in unchanged order, and table endpoints use the same column. extend adds the range; otherwise it replaces the selection. Duplicate range labels, unsupported providers and unloaded gaps are refused. List replacement verifies one clear then each addition; an already exact set is unchanged. Limits: 50 range items and 500 selected items. Selection step receipts are historical verification, never instructions to retry a remainder.

ParametersJSON Schema
NameRequiredDescriptionDefault
extendNo
element_idYes
range_end_idNo

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so excellently. It discloses exclusive vs. extend behavior, whole-row selection for table cells, already-selected results, activation fallback, offscreen scrolling, reacquisition after sorting/filtering, range endpoint rules, hard limits, and receipt semantics. This is model transparency beyond what the schema could convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but front-loaded with the core purpose and organized in a logical sequence from verification, to activation fallback, to range semantics, to limits. Given the tool's complexity, most sentences earn their place; a few clauses are dense and slightly redundant, so it is not optimally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no annotations, no output schema, and minimal schema documentation, so the description must provide nearly all context. It covers verification semantics, application-effect handling, range behavior, error/refusal cases, limits, and the meaning of receipts. An agent has enough information to invoke the tool correctly and interpret its results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate for all parameters. It thoroughly explains extend, range_end_id, default exclusivity, range inversion, and the 50-item cap. The required element_id is implied as the observed list/radio/combo/table-cell reference, but it is never explicitly tied to an inspection-provided identifier, leaving a small semantic gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Choose an observed list/radio/combo option or a visible table cell and verify selection.' This clearly communicates the tool's scope. However, it does not explicitly distinguish itself from the sibling desktop_select or other selection-related tools, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives extensive procedural guidance once the tool is chosen, such as verifying effects and handling Apply/Open controls, but it never states when to choose this tool instead of desktop_select, desktop_invoke, or desktop_set_value. The intended use case is implied by the opening line but alternatives and exclusions are not mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_clickA

Click screenshot-image coordinates in the target window or its observed menus, automatically choosing independent or foreground pointer input after target validation. Rejects expired snapshots, changed window layout/identity and covered targets. Snapshot validity does not prove unchanged application content; observe again after content transitions before selecting a control.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
countNo
buttonNoleft
window_idYes
snapshot_idYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and discloses multiple important behaviors: automatic choice between pointer input methods, rejection of expired snapshots, changed layout/identity, and covered targets, plus the warning that snapshot validity doesn't prove content unchanged. This materially reduces surprise. It stops short of describing error responses or side effects, so a 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler; the core action is front-loaded and the caveats are dense but relevant. It could be slightly more organized (e.g., separating validation rules from usage guidance), but it is economical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a click tool with no output schema and no annotations, the description covers target validation and failure conditions but omits what the tool returns on success or failure, and doesn't explain how to obtain valid window/snapshot IDs. This is a moderate gap: an agent can invoke it correctly but may not interpret results properly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does explain that x/y are screenshot-image coordinates and that snapshot_id relates to an observed snapshot/window. However, it doesn't clarify how window_id/snapshot_id are obtained, the meaning of count, or button beyond the schema's enum. The added semantic is meaningful but incomplete for 6 parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action (Click) on a resource (screenshot-image coordinates in the target window or its observed menus), which clearly distinguishes it from scroll/drag/hover siblings. It doesn't explicitly name a sibling alternative, but the screenshot-coordinate mechanism is a strong differentiator, so 4 is appropriate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: you click when you have a validated target and screenshot coordinates from a snapshot. It cautions about expiring snapshots and content changes, but it never names alternatives (e.g., desktop_choose, desktop_invoke) or states when not to use this tool. The guidance is contextual but not comparative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_controlA

Pause/resume cooperating agent input across servers on this display. Pause interrupts at the next checkpoint; already-delivered input is not undone. Observation remains available. This does not stop arbitrary external input programs.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNostatus

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and delivers: checkpoint-based interruption, no undoing of already-delivered input, continued observation, and the limitation regarding external programs. This is well beyond the bare action enum and gives the agent a precise mental model of the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, each earning its place: purpose, interrupt mechanics, side effect of observation, and the boundary of what it does not affect. Information is front-loaded and there is no fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with no output schema and no annotations, the description covers purpose, side effects, boundaries, and behavioral semantics. The only missing context is the meaning or return value of the 'status' action, which is a minor omission given the default value and the enum.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clearly explains the semantics behind the 'pause' and 'resume' actions, which maps to two of the three enum values. However, it never mentions the 'status' action or what it returns; the enum and default are self-explanatory enough that this is a minor gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Pause/resume cooperating agent input across servers on this display.' It actively distinguishes itself from siblings by clarifying it only affects cooperating agent input and explicitly saying it does not stop arbitrary external programs, so the agent can tell it apart from tools like desktop_recover_input or desktop_observe.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear usage context: pausing interrupts at the next checkpoint, observation remains available, and it is not for stopping external input programs. It includes an explicit when-not but does not name sibling alternatives directly, so an agent has to infer which other tool to use instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_doctorA
Read-only

Check actual display access, desktop session, dependencies and accessibility availability. Reports driver version and content identities for tool declarations and the server-bundled skill; the latter does not identify the skill loaded by your agent.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint already marks the tool as safe, and the description adds meaningful behavioral context: it verifies actual access, reports driver version and content identities, and discloses a limitation about the server-bundled skill not matching the agent's loaded skill. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no filler. The first sentence immediately states the tool's core purpose, and the second adds necessary output and limitation details. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only diagnostic tool, the description covers what is checked, what is reported, and a key limitation. It does not specify the exact return structure, but the high-level reporting details are enough for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and schema coverage is complete, so there is no parameter ambiguity. The description clarifies what the tool inspects and reports, which is sufficient for a parameterless call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific diagnostic action ('Check actual display access, desktop session, dependencies and accessibility availability') with a concrete resource scope. It is clear, but it does not explicitly differentiate itself from sibling tools like desktop_status or desktop_report.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use as a diagnostic health-check tool is implied by the listed checks, but the description does not state when to prefer it over alternative desktop tools or provide exclusion conditions. The caveat about not identifying the loaded skill gives some usage-relevant context, but no explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_dragA

Drag between two observed points inside the same window, automatically choosing independent or foreground pointer input after validation; always attempts release of its own button. Use desktop_drag_to for another destination window.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
end_xYes
end_yYes
buttonNoleft
window_idYes
snapshot_idYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It goes beyond a generic 'drag' by explaining that input mode is chosen automatically after validation and that the tool always attempts to release its own button, which is highly relevant safety-relevant behavior. It does not fully cover failure modes or side effects, but the disclosed behaviors are specific and meaningful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, with the core action and key behavioral caveat front-loaded and the sibling alternative mentioned last. Every clause earns its place and the structure aids quick comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of annotations and output schema, the description covers the most important operational aspects: the drag's scope, the validation/input-mode behavior, the release guarantee, and the sibling alternative. The main omissions are explicit coordinate semantics and potential failure conditions, but the description is strong enough for an agent to invoke the tool with reasonable confidence.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It helps map 'two observed points' to the coordinate parameters, 'same window' to window_id, and 'observed points' to snapshot_id, but it does not explain the button parameter, coordinate units, or the exact role of the snapshot. This partial semantic addition justifies a middle score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Drag'), a resource scope ('between two observed points inside the same window'), and explicit behavioral details like input-mode selection and self-release. It additionally names the sibling 'desktop_drag_to' with the distinguishing condition 'for another destination window', so the agent can clearly tell them apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to use this tool for drags within the same window and points to 'desktop_drag_to' as the alternative for a different destination window. This gives clear when-to-use and when-not-to-use guidance without requiring the agent to infer it from schemas.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_drag_toA

Drag from a source into a second observed window, automatically activating the source after validating both endpoints. Coordinates refer to one screenshot. Verify transfer in the applications; dispatch does not prove a drop was accepted.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
end_xYes
end_yYes
buttonNoleft
snapshot_idYes
source_window_idYes
target_window_idYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations at all, the description carries the full burden and meets it: it discloses automatic source activation, pre-drag endpoint validation, that coordinates share one screenshot, and the key caveat that dispatch does not prove the drop was accepted. These go beyond the bare action and give the agent realistic expectations about side effects and result trustworthiness.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the action, the coordinate constraint, and the verification caveat. The most important information is front-loaded and there is no filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameter-heavy automation action with no output schema or annotations, the description covers the core elements an agent needs to call it correctly and interpret results cautiously. It omits only refinements such as expected return/error behavior or how it compares to desktop_drag, but the essential behavioral contract is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it only partially does. It clarifies that 'Coordinates refer to one screenshot', which gives meaning to x/y/end_x/end_y and connects them to snapshot_id, and 'both endpoints' maps to source_window_id/target_window_id. However, it leaves button and the exact coordinate frame/origin undefined, relying on self-explanatory names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with the specific action and object: 'Drag from a source into a second observed window', which clearly identifies what desktop_drag_to does and differentiates it from desktop_drag and other input tools. It adds a distinctive trait, 'automatically activating the source after validating both endpoints', so an agent can distinguish this tool from sibling drag/hover/click actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the intended context explicitly: dragging from a source into a second observed window, which tells the agent when to invoke it. It does not name sibling alternatives or exclude same-window drags, so it stops short of the explicit when/when-not guidance that would earn a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_focus_elementA

Request element focus for agent input; owned browser fields refuse active/unknown composition before focus. Native application callbacks can still change human focus. Inspect to confirm focused state.

ParametersJSON Schema
NameRequiredDescriptionDefault
element_idYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry behavioral disclosure, and it does: focus is only 'requested', not guaranteed, native callbacks can override it, and verification via inspect is required. The phrase about 'active/unknown composition' adds context but is somewhat opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with the primary purpose front-loaded makes this compact and mostly efficient. The middle sentence is terse but jargon-heavy, which slightly hurts readability without adding enough explanation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter focus operation with no output schema, the description covers the purpose, a key reliability caveat, and a verification step. The main gap is element_id semantics, but the tool is simple enough that an agent can likely use it correctly with inspection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one required parameter, element_id, with no description coverage. The tool description never explains where element_id comes from or how it should be formatted, so the description does not compensate for the missing parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening phrase 'Request element focus for agent input' names a concrete verb and resource, and adds the useful qualifier that this is for agent input. It is not explicitly contrasted with sibling tools like desktop_activate or desktop_hover, but the focus-specific purpose is clear enough to distinguish it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use it: before agent input, especially for owned browser fields that may refuse composition before focus. It also tells the agent to follow up with inspect to confirm the focused state, though it does not name explicit alternatives or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_hoverA

Move the pointer to a recent observed point without clicking, automatically activating the target after validation; observe tooltips/submenus afterward.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
window_idYes
snapshot_idYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the safety burden and largely meets it: it says the action is 'without clicking', that activation only happens 'after validation', and that the follow-up is to observe tooltips/submenus. It leaves validation-failure behavior and pointer-state details unspecified, but the main side effects are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence carries the action, the constraint, the side effect, and the recommended next step without filler. Front-loading 'Move the pointer' immediately tells the agent what operation this tool performs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple hover action, the description covers the core behavior and follow-up ('observe tooltips/submenus afterward'). However, with no output schema and no annotations, it omits return/error semantics and does not state what happens if the snapshot point is stale or invalid.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. 'Recent observed point' does tie x/y to snapshot_id and indicates the coordinates must come from a prior observation, but it does not explain coordinate origin/units, window_id's role, or validation constraints. The compensation is partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('move the pointer'), a specific resource ('recent observed point'), and a distinctive mechanism ('without clicking, automatically activating the target after validation'). It is clearly distinguishable from sibling click/drag/focus tools, and the reference to 'observe tooltips/submenus afterward' pins down the hover use case.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context: this is the tool to use for hover-only interaction that reveals tooltips/submenus, and it explicitly distinguishes from clicking. It does not name an alternative sibling, which keeps it a step below an explicit when-to-use/when-not-to-use statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_inspectA
Read-only

Inspect a window or find controls by name/role substring and required states. Returns bounded tree, parent IDs, supported actions and 60-second element IDs. Owned browser text_fields have distinct provider-bound IDs for ordinary HTML fields; role="entry" filters for fields. Field metadata describes line breaks and write scope. When extra owned pages/windows or frames make the owned provider unavailable, native nodes remain independently inspected; owned_browser reports unavailable/code and text_fields is empty. Cached owned fields still refuse unsupported scope; no mutation fallback. Empty matches and unavailable accessibility are distinct.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
roleNo
limitNo
statesNo
max_depthNo
window_idYes

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes far beyond the readOnlyHint annotation, disclosing return contents, bounded tree/parent IDs/supported actions, 60-second element ID validity, provider-bound text field IDs, native fallback behavior, no mutation fallback, and the distinction between empty matches and inaccessible accessibility. This is rich behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first sentence is front-loaded and clear, but the remainder is a dense, unstructured paragraph of conditional clauses about providers, caching, and edge cases. The content is informative but would benefit from bullet structure or clearer separation of concerns.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite missing explicit parameter semantics, the description covers many important edge cases: provider unavailability, cached field behavior, role filtering for fields, and empty-versus-unavailable results. For a complex tool with no output schemaaine, it gives a solid high-level contract; the main gap is clarity around limit/max_depth/window_id usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains role filtering and mentions required states, but does not clarify the meaning or allowed values of limit, max_depth, window_id, or states. Several parameters remain under-specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair: 'Inspect a window or find controls by name/role substring and required states.' This clearly states what the tool does and distinguishes it from action-oriented siblings like desktop_click, desktop_type, or desktop_control.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives useful contextual hints, such as role="entry" filtering for fields and behavior when the owned browser provider is unavailable, but it never explicitly says when to choose desktop_inspect over related tools like desktop_observe or desktop_control, nor when not to use it. Usage guidance is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_invokeA

Invoke the sole action returned by inspect, or supply its exact action name. Native actions address the control without activating its window; application callbacks may bring windows forward. Multiple actions require an explicit choice; no click/press naming guess is needed for a single-action button. Completion means dispatch, not verified application outcome.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNo
element_idYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden and does a good job: it discloses that native actions do not activate the window, application callbacks may bring windows forward, and completion means dispatch rather than verified outcome. This is valuable behavioral context beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is reasonably concise and front-loaded with the primary purpose. Each sentence adds information, though the third sentence somewhat restates the first regarding action naming, so it could be slightly tighter. Overall, it avoids fluff and stays readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema and no annotations, the description covers the most important operational details: how to select the action, window-activation behavior, and the meaning of completion. It lacks an explicit note about what the function returns, but the dispatch-vs-outcome statement partially addresses that concern.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It clearly explains the 'action' parameter semantics: it can be the sole action from inspect, an exact action name, or omitted when there is a single action. However, the required 'element_id' parameter is not addressed at all, leaving one of two parameters underspecified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb and resource: 'Invoke the sole action returned by inspect, or supply its exact action name.' It also distinguishes the tool from click/press-based siblings by noting that no naming guess is needed for a single-action button, making the tool's unique role explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete guidance on when to use this tool: after inspect returns an action, or when you know the exact action name. It also notes that multiple actions require an explicit choice, which helps the agent decide when additional information is needed. It does not explicitly name alternative tools, but the click/press contrast implies the boundary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_launchA

Launch an installed application by its desktop_applications ID, optionally opening absolute existing paths or URIs. Returns dispatched with process-bound window candidates observed for wait_timeout seconds (default 1, range 0–3; 0 skips observation). Candidates do not prove document readiness; singleton association is never guessed. Inspect candidates or desktop_windows before acting; never blindly retry an uncertain launch.

ParametersJSON Schema
NameRequiredDescriptionDefault
wait_timeoutNo
files_or_urisNo
application_idYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it delivers: it discloses that the tool returns 'dispatched with process-bound window candidates observed for wait_timeout seconds', that candidates do not prove document readiness, that singleton association is never guessed, and that retrying uncertain launches is unsafe. This is rich behavioral context beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action, then the return behavior, then the safety caveat. Every sentence earns its place; no filler or repetition of schema fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a launch tool with 3 params, no output schema, and no annotations, the description is complete: it covers the input semantics, the return value, the timeout behavior, and the follow-up action the agent should take. Nothing critical is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains application_id (the desktop_applications ID), files_or_uris (absolute existing paths or URIs), and wait_timeout (default 1, range 0–3, 0 skips observation). It does not explicitly restate the parameter names, but the semantics are clearly conveyed. Minor gap: it doesn't mention that files_or_uris is optional/nullable, though the schema already shows that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Launch'), a specific resource ('an installed application by its desktop_applications ID'), and the optional payload ('absolute existing paths or URIs'). It clearly distinguishes this from sibling tools like desktop_activate or desktop_invoke by focusing on launching an application by ID rather than interacting with an already-running window.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage guidance: use the desktop_applications ID, optionally pass paths/URIs, and set wait_timeout to observe window candidates. It also tells the agent when not to act: 'never blindly retry an uncertain launch' and 'Inspect candidates or desktop_windows before acting.' This is strong when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_match_imageA
Read-only

Find historical visual candidates from a selected screenshot crop in another retained screenshot. template_bounds requires x, y, width and height in source returned-image pixels; x/y are nonnegative and width/height are 8–512. Both IDs must be retained and fresh, from the same server and image scale; only the target must still have its captured layout. Historical source crops may come from a window that moved. Results use target returned-image pixels. Threshold is finite 0–1, limit 1–100; scores are uncalibrated correlation, never semantic identity or click permission. Returns non-overlapping candidates, preserving distinct duplicates; flat templates are refused. No new capture or input. Optional system OpenCV required.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
thresholdNo
snapshot_idYes
template_boundsYes
template_snapshot_idYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses important behavioral traits beyond the readOnlyHint annotation: it states that scores are uncalibrated correlation, never semantic identity or click permission, that flat templates are refused, and that no new capture or input occurs. It also notes the optional system OpenCV requirement. The readOnlyHint annotation is consistent with the description's statement that no new capture or input occurs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-organized, front-loading the core purpose before diving into parameter constraints and behavioral notes. Every sentence adds value, though the density of technical constraints could be slightly overwhelming. It is appropriately sized for the complexity of the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the essential context: prerequisites (retained and fresh IDs, same server and image scale), constraints (template_bounds ranges, threshold/limit ranges), behavioral caveats (uncalibrated scores, flat templates refused), and dependencies (optional OpenCV). While there is no output schema, the description adequately explains what results will look like in terms of non-overlapping candidates and distinct duplicates.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema description coverage is 0%, the description provides substantial parameter semantics: it explains that template_bounds requires x, y, width, and height in source returned-image pixels, with specific constraints on ranges. It also clarifies that results use target returned-image pixels and that threshold is finite 0–1 with limit 1–100. This compensates well for the lack of schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: finding historical visual candidates from a selected screenshot crop in another retained screenshot. It uses specific verbs and resources (find, screenshot crop, retained screenshot) and distinguishes itself from siblings by focusing on historical matching rather than live observation or interaction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool: when both IDs are retained and fresh, from the same server and image scale, and only the target must still have its captured layout. It also notes that historical source crops may come from a window that moved, providing clear context for when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_observeB
Read-only

Return screenshot plus window layout and a 15-second snapshot ID. Pointer coordinates and window/popup image_bounds use returned-image pixels; bounds/frame_bounds remain native X11 root pixels. image_bounds is null when no integer screenshot pixel maps into the client. Use integer image points.

ParametersJSON Schema
NameRequiredDescriptionDefault
max_widthNo

TDQS

B3.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already establishes safety, and the description adds meaningful behavioral detail: coordinate systems differ between image pixels and native X11 root pixels, image_bounds can be null, snapshot IDs expire after 15 seconds, and integer image points should be used. This goes beyond what the annotation provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is tight and information-dense, with the main return value front-loaded. Each sentence adds distinct technical detail, though the coordinate-system explanations are packed into a single sentence and may be hard to parse. Still, there is no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description carries the burden of explaining return-ish behavior and does cover screenshot, window layout, snapshot ID, coordinate systems, and null bounds. However, it omits the effect of max_width and does not describe the structure or shape of the returned window layout or snapshot ID, leaving gaps for an agent invoking this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description never mentions max_width, which is the only parameter. Although the parameter name and default are somewhat self-explanatory, the description does not compensate for the schema gap by explaining how max_width affects the returned screenshot or when an agent should override the default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it returns a screenshot, window layout, and a 15-second snapshot ID. This clearly identifies the tool as an observation/capture function and is distinct enough from siblings like desktop_status or desktop_ocr, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus the many sibling tools. It explains coordinate conventions and pixel mapping, which is useful after the tool is selected, but it does not help an agent decide between desktop_observe, desktop_status, desktop_report, or desktop_inspect.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_ocrA
Read-only

Read uncertain local OCR word candidates from this exact retained screenshot, never a new capture. Returns image-pixel boxes and uncalibrated engine scores, not exact text or action permission. Snapshot expires after 15 seconds or cache eviction; changed layout is refused. Optional Tesseract and the selected language must be installed. No input is sent.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
languageNoeng
snapshot_idYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the readOnly annotation by disclosing that no input is sent, results are local candidates with uncalibrated scores, the snapshot expires after 15 seconds or eviction, and changed layouts cause refusal. These are meaningful behavioral details an agent must know before calling the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences, each earning its place: purpose, return semantics, validity constraints, prerequisites/privacy. The most decision-relevant information is front-loaded, and there is no repetition of schema or annotation data.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with no output schema, the description fully covers what is returned, what is not returned, when the call can fail, prerequisites, and privacy behavior. The only minor omission is limit semantics, but that is a small parameter detail and does not undermine safe invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does clarify snapshot_id as the exact retained screenshot and language as the selected installed language, but it never explains the limit parameter or its default behavior. The compensation is partial at best.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Read uncertain local OCR word candidates from this exact retained screenshot.' It actively distinguishes itself from any capture-based or exact-text tool by saying 'never a new capture' and 'not exact text or action permission,' making it easy to tell apart from siblings like desktop_read_text or desktop_observe.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear conditions: use on the retained snapshot, not a new capture; the snapshot must still be valid and unchanged; Tesseract and the selected language must be installed. It implies exclusions by saying this is not exact text or action permission, though it does not explicitly name alternative sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_open_browserA

Open a fresh owned Chromium with explicit temporary_session lifetime. Browser/profile and unsaved content are deleted on server disconnect, backend close or reconnect. Requires optional browser dependencies and configured executable; never downloads automatically or attaches existing profiles. Inspect text_fields for ordinary HTML fields and explicit password-only secret entry. Ordinary protected-field operations, contenteditable editors and frames are unsupported.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
lifetimeYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure, and it does so thoroughly. It discloses deletion triggers, dependency requirements, the absence of auto-downloading/profile attachment, and unsupported UI interaction modes. This is exactly the kind of behavioral detail an agent needs to avoid incorrect assumptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: core behavior first, then lifecycle/deletion semantics, then prerequisites and limitations, then interaction constraints. It is front-loaded with the most important decision-relevant information and contains no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a browser-opening tool with no annotations and no output schema, the description covers prerequisites, lifecycle, and operational limitations well. It does not describe the return value or explicit failure behavior, but those are secondary for this kind of command. Overall, an agent has enough context to invoke it correctly and predict its effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the bare schema. It adds meaningful semantics to the lifetime parameter by explaining that the temporary session is deleted on disconnect, backend close, or reconnect. The url parameter remains simple and predictable from the tool's purpose and schema title, so the lack of extra explanation is not a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: "Open a fresh owned Chromium" and immediately specifies the lifetime mode. It also distinguishes the tool from generic launchers by detailing ephemerality and the fact that it never attaches existing profiles. This is clearly differentiated from the sibling automation tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: use this for a temporary, disposable browser session with explicit lifetime semantics, and understand that dependencies must already be configured. It does not name a sibling alternative, but no sibling tool appears to be a competing browser opener, so the omission is acceptable. The guidance to inspect text_fields and the list of unsupported operations also helps the agent know what interactions to expect after opening the browser.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_pasteA

Paste through the shared CLIPBOARD when semantic typing is unavailable, automatically choosing independent or foreground input for the target. Chooses common app shortcut from window class, with optional override. Destination is unverified; inspect dialogs/read back. Terminals can execute pasted newlines.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
shortcutNo
window_idYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses important behavioral traits: it automatically chooses input mode, selects a shortcut from window class, allows override, and warns that the destination is unverified and that terminals can execute pasted newlines. This goes beyond the schema and provides safety-relevant context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose, then adds behavioral caveats. Every sentence adds information, though the phrasing is slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with no output schema and no annotations, the description covers the key behavioral risks (unverified destination, terminal newline execution) and the shortcut override. It doesn't explain return values, but no output schema exists and the tool's purpose is clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden. It explains the 'shortcut' parameter's role (optional override) and the overall paste behavior, but doesn't detail 'window_id' or 'text' semantics beyond what the schema names. The description adds some meaning but not full parameter-level detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Paste') and resource ('shared CLIPBOARD'), and distinguishes it from semantic typing by noting it's used 'when semantic typing is unavailable.' It doesn't explicitly name a sibling alternative, but the context makes the tool's role clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use this tool ('when semantic typing is unavailable') and notes that terminals can execute pasted newlines, which is a useful caution. It doesn't explicitly list alternatives or exclusions, but the usage context is reasonably clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_press_keysA

Send a deliberate chord, e.g. ctrl+s, ctrl+plus, ctrl+minus, Return, Tab, Escape or Down. Punctuation uses X11 names (plus, equal, bracketleft, slash); implicit Shift follows the current layout. count is 1–20 complete press/release repetitions, default 1. Revalidates target identity, focus and input state between repetitions; stops on the first failure and never retries. Automatically chooses independent input where supported or foreground control otherwise. Refuses held input on the selected devices and preserves their keyboard mapping. Foreground fallback can change human focus. Unavailable symbols return UNSUPPORTED_KEYMAP; text belongs in desktop_type. When a final companion receipt is available, progress reports fully dispatched, possibly partial and not-started repetitions. Dispatched count is not application completion.

ParametersJSON Schema
NameRequiredDescriptionDefault
chordYes
countNo
window_idYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly: it discloses revalidation between repetitions, stop-on-first-failure with no retry, automatic input-device selection, refusal of held input, preservation of keyboard mapping, possible human-focus changes, the UNSUPPORTED_KEYMAP error, and the distinction between dispatched and application-completed actions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence adds distinct information, from examples and key naming rules to failure behavior and result semantics. It is front-loaded with the core purpose before moving into edge cases.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Considering there is no output schema and no annotations, the description covers purpose, parameter constraints, failure modes, side effects, and even progress-report semantics. It is slightly opaque about the exact return/receipt structure and about what window_id should contain, but overall it is unusually complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds strong semantics for chord (X11 key names, implicit Shift) and count (1-20 repetitions, default 1), which is essential because schema description coverage is 0%. However, the required window_id parameter is never explained, leaving a meaningful gap for a required argument.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Send a deliberate chord', followed by concrete examples like ctrl+s, Return, Tab and Down. It also distinguishes itself from keyboard text input by stating 'text belongs in desktop_type', so an agent can tell it apart from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly routes text input away from this tool with 'text belongs in desktop_type' and warns about focus-changing foreground fallback. It could be more explicit about when to prefer this over other siblings like desktop_paste or desktop_type_secret, but the chord-versus-text distinction is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_read_textA
Read-only

Read accessible text and representation metadata, preserving whitespace. limit counts Unicode code points (default 16000, maximum 1000000), not bytes. Opaque embedded objects are not exact logical plain text: check plain_text_verification_supported. Normalization reads the bounded full field; a smaller limit does not enable streaming. Protected fields are refused. Generic native readback reports composition known=false, active=null; pending preedit is not checked.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
element_idYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses substantial behavioral details: preservation of whitespace, code-point-based limits, the non-exactness of opaque embedded objects, normalization reading the full bounded field, refusal of protected fields, and specific native readback behavior (composition known=false, active=null, pending preedit not checked). This far exceeds what annotations provide and gives the agent critical edge-case knowledge.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and technical, but every sentence adds necessary nuance. It is front-loaded with the primary action and then layers on important caveats. While not as brief as a one-liner, the length is justified by the complexity of the behavior, and the structure flows logically from purpose to edge cases.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and the absence of an output schema, the description covers a wide range of scenarios: limit semantics, streaming behavior, normalization, opaque objects, protected fields, and native readback details. This is comprehensive enough for an agent to understand what to expect and avoid common pitfalls when invoking the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain parameters. It thoroughly explains 'limit' (default 16000, max 1000000, counts code points, not bytes) and its behavioral implications. 'element_id' is not elaborated, but its meaning is obvious from the name and required status. The description compensates well for the lack of schema documentation on the key parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb-resource pairing: 'Read accessible text and representation metadata, preserving whitespace.' This is specific and distinguishes the tool from OCR and other text-reading tools by focusing on accessibility metadata and whitespace preservation. It is not a tautology and immediately tells the agent what the tool accomplishes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides several usage constraints (e.g., 'limit counts Unicode code points', 'a smaller limit does not enable streaming', 'check plain_text_verification_supported') but does not explicitly state when to use this tool over siblings like desktop_ocr or desktop_inspect. The guidance is implicit rather than explicit, with no mention of alternatives or exclusions, so it falls short of a clear decision rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_reconnectA

Reconnect this MCP connection to a running XFCE session owned by this account after a desktop restart. Omit PID only when exactly one session exists. Validates display and bus before replacing the backend; failed validation preserves it. Returns BUSY during other operations, never restarts apps or replays input. All prior window, element and screenshot IDs expire; observe again. The selected display's pause state remains in force. Temporary owned browsers/profiles are closed and unsaved content is lost; browser_cleanup=unconfirmed reports cleanup that could not be proved.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_pidNo

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully carries the burden, and it delivers: it discloses validation and preservation of the backend on failure, the BUSY status, the non-destructive nature (no app restarts or input replay), expiry of all prior IDs, pause state persistence, and the closure of temporary browsers/profiles with unsaved content loss. This is exceptionally transparent about side effects and safety.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence carries essential information. It is front-loaded with purpose and usage, then systematically covers validation, safety, ID expiry, pause state, and browser cleanup. No filler or redundancy; each clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one optional parameter, no output schema, and no annotations, the description covers all critical aspects: when to call, how to handle the PID, validation behavior, concurrency, side effects, ID invalidation, pause state, and data loss. An agent has everything needed to decide and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides no description for session_pid (0% coverage), so the description must compensate. It does so by explaining when to omit the PID ('only when exactly one session exists') and implying its purpose. It does not elaborate on the PID's format or how to obtain it, but the name is self-explanatory and the condition is useful. This is adequate compensation for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('reconnect'), a precise resource ('MCP connection to a running XFCE session'), and a clear trigger ('after a desktop restart'). It also differentiates itself by explicitly noting it never restarts apps or replays input, setting it apart from many sibling interaction tools. This is a clear, unambiguous purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit conditions for use: 'after a desktop restart' and the PID omission rule ('Omit PID only when exactly one session exists'). It also cautions about BUSY during other operations. However, it does not mention alternative tools or when NOT to use this tool, so it lacks explicit exclusions, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_recordingA

Explicit temporary screen recording: start, status, stop or delete by ticket. Start records only this X11 display, no audio, at 10fps and at most 1280×720 for 1–60 seconds. Start is not completed-file verification; stop/status return a path only after decoding verifies completion. Files are private, bounded and deleted on backend close/reconnect/server death; explicitly copy elsewhere before closure to save durably. Stop/delete remain usable while paused. Optional local ffmpeg/ffprobe required.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes
max_secondsNo
recording_idNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so comprehensively. It discloses technical limits (X11 only, no audio, 10fps, max 1280x720, 1-60s), verification semantics (stop/status return path only after decoding), file lifecycle (deleted on backend close/reconnect/server death), and dependency (optional ffmpeg/ffprobe). This goes beyond basic mutation details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense yet efficient, with each sentence adding critical information. It opens with the core purpose, then layers constraints, verification behavior, lifecycle, and dependencies without redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (four actions, stateful recording, file persistence) and no output schema, the description covers all essential aspects: what it does, constraints, verification, persistence requirements, and external dependencies. It even mentions return paths for stop/status, providing a complete operational picture.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It mentions 'start, status, stop or delete' mapping to the action enum, 'by ticket' implying recording_id, and '1-60 seconds' hinting at max_seconds. However, it does not explicitly name the parameters or define their exact roles, leaving some inference needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool is for explicit temporary screen recording and lists the four actions (start, status, stop, delete) by ticket. This specific verb+resource distinguishes it from sibling tools like desktop_observe or desktop_inspect, which are likely for real-time viewing rather than recording.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context about the recording constraints (temporary, private, bounded) but does not explicitly state when to use this tool versus alternatives like desktop_observe. It implies usage for recording sessions but leaves the decision to the agent without direct exclusions or comparisons.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_recover_inputA

Retry cleanup of this server's interrupted supervised input, without replaying keys/clicks or resuming a paused desktop. Uses each operation's original session; a replaced X server is left untouched. Unproven cleanup stays blocked. Returns pending_count and recovery proofs; observe again before acting. Returns BUSY if another operation is still running.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full responsibility for behavioral disclosure, and it excels. It details non-replay of input, non-resumption of paused desktop, use of original session, untouched replaced X server, blocked unproven cleanup, return values, and BUSY error condition. This is comprehensive and leaves little to inference.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, front-loaded with the primary purpose, and every sentence adds distinct useful information: exclusion of replay/resume, session behavior, cleanup blocking, return contents, and BUSY condition. No filler or redundancy is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of a zero-parameter tool with no output schema, the description is remarkably complete. It explains the action, behavioral constraints, return values (pending_count, recovery proofs), a required follow-up ('observe again'), and an error condition (BUSY). An agent has enough context to call and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema confirms an empty object, so the description cannot add parameter-level meaning. The baseline for zero parameters is 4, and the description appropriately focuses on behavior and return semantics rather than parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Retry cleanup of this server's interrupted supervised input') on a specific resource, and explicitly contrasts it with what it does not do ('without replaying keys/clicks or resuming a paused desktop'). This distinguishes it from sibling tools like desktop_reconnect and desktop_play while making the tool's core role unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly conveys when this tool is appropriate: after an interrupted supervised input, when cleanup is needed but replaying or resuming is not desired. It also warns to 'observe again before acting', which gives operational guidance. However, it does not explicitly name alternative tools or state when not to use it, so it stops short of full exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_reportA
Read-only

Return a sanitized bug-report JSON: fixed environment/dependency versions, projected health and up to 32 recent operation IDs/methods/effects/timings, fixed error codes and semantic verbs from this MCP process. Excludes desktop content, paths, exceptions and action arguments. No files, uploads or replay. Supply synthetic repro steps separately; explicitly save/delete the returned report if needed.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description reveals sanitization behavior, explicit exclusions (desktop content, paths, exceptions, action arguments), the 32-result cap, and non-actions (no files, uploads, replay). This gives the agent accurate expectations about side effects and output constraints without contradicting the read-only annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded and the description is compact given the amount of detail. It is somewhat dense and run-on, but each clause provides decision-relevant information and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining the return value, and it does so by enumerating fields, exclusions, and limits. It stops short of specifying the exact JSON structure, and 'projected health' remains somewhat vague, but an agent can call the tool and interpret the result correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already declares zero properties with 100% coverage, so there are no parameters for the description to explain. The description compensates by characterizing the report contents, which is the only relevant semantic information for this parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and deliverable: 'Return a sanitized bug-report JSON.' It enumerates the contents (environment/dependency versions, projected health, up to 32 recent operation IDs/methods/effects/timings, fixed error codes and semantic verbs) and states this is scoped to the MCP process, making the tool's purpose clear and distinct from likely sibling operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear how-to-use context: supply synthetic repro steps separately and explicitly save/delete the returned report if needed, and it tells the agent what the tool will not do (no files, uploads, replay). It does not explicitly name alternative sibling tools or define a when-to-use vs. when-not-to-use rule, but the instructions are actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_scrollA

Scroll 1–20 wheel ticks at a point in the observed target, automatically choosing independent or foreground pointer input after validation. Read resulting state to confirm.

ParametersJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
ticksNo
directionYes
window_idYes
snapshot_idYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it discloses non-obvious behavior: it automatically chooses independent vs. foreground pointer input and validates before acting. The instruction to read resulting state also gives the agent a useful post-condition expectation. It does not describe every possible failure or side effect, but for a scrolling action it communicates the key behavioral surprise.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences carry the essential action, range, validation behavior, and a follow-up instruction. There is no filler, and the core scroll operation is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has six parameters, no output schema, and no annotations, so more detail would be justified. The description gives a usable mental model, but open questions remain around validation failure, coordinate space, and what 'read resulting state' means concretely. Despite this, it is neither misleading nor severely gapped for a simple scroll action.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds meaning to 'ticks' by bounding it to 1–20 and ties x/y to 'a point in the observed target,' which helps explain window_id/snapshot_id usage. However, it does not clarify the coordinate frame, direction semantics, or the relationship between the snapshot and window beyond the phrase 'observed target.'

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action—scroll—and specifies the operand (1–20 wheel ticks at a point in the observed target), making it immediately distinct from sibling pointer tools like desktop_click, desktop_hover, or desktop_drag. The wording is not a tautology and conveys both the operation and its scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly situates the tool's use case: scroll inside an observed target after validation, then read resulting state to confirm. It does not explicitly name alternatives or state when not to use it, so it lacks the exclusions needed for full guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_selectA

Select a text range using Unicode code-point offsets, or place the caret when equal; verify the result. Owned-browser HTML fields focus automatically after range validation. Native edits inside graphemes can be refused.

ParametersJSON Schema
NameRequiredDescriptionDefault
element_idYes
end_offsetYes
start_offsetYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses that the tool verifies the result, that owned-browser HTML fields focus automatically after range validation, and that native edits inside graphemes can be refused. These are useful traits, but 'verify the result' is ambiguous (does the tool verify or should the agent verify?), and the description does not state whether selection is non-destructive or what failure modes exist beyond grapheme refusal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three short sentences with no filler. The primary purpose is front-loaded, and each additional sentence adds a distinct behavioral note. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The core operation and offset semantics are covered, but there are gaps: the meaning of 'verify the result' is unclear, no return value or error behavior is described, and 'owned-browser HTML fields' is jargon that may confuse an agent. Given no output schema and no annotations, these omissions leave the agent to infer important runtime details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must clarify the parameters. It explains that offsets are Unicode code-point offsets, which is essential, and that when start_offset equals end_offset the tool places the caret. This adds meaning well beyond the raw schema. It does not explicitly describe element_id, but that identifier is self-explanatory in context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'Select a text range' using offsets, with the caret behavior when offsets are equal. It is distinct from siblings like desktop_set_value or desktop_focus_element because it addresses selection rather than value assignment or focus. However, it does not explicitly name a sibling or contrast itself, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: call this when you need to select a text range or place a caret in an element. There is no explicit when-to-use or when-not-to-use guidance, nor any mention of alternatives such as desktop_set_value or desktop_type. The note about owned-browser HTML fields provides a contextual condition but not a decision rule for selecting this tool over others.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_set_checkedA

Set a checkable control to the requested state; avoid a blind toggle when it already matches.

ParametersJSON Schema
NameRequiredDescriptionDefault
checkedYes
element_idYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are absent, so the description carries the full burden and it does add a behavioral guarantee: it will not blindly toggle when the desired state is already set. It stops short of describing failures, visibility requirements, or side effects, but the core state-setting behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence front-loads the main action and adds the toggle-avoidance caveat in a second clause. No filler or redundant restatement of the name or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter setter this is mostly complete, covering target element, desired state, and idempotence. It lacks guidance on error conditions, visibility/availability prerequisites, and what the tool returns, but no output schema makes return-value detail optional.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has no parameter descriptions, but 'requested state' maps checked to a target boolean value and 'checkable control' implies the element_id refers to a control capable of being checked. It does not explain element_id resolution, but the two parameters are simple enough that the added context is meaningful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('set') and resource ('a checkable control') and explicitly frames the action as bringing the control to a requested state. This differentiates it from sibling setters like desktop_set_expanded or desktop_set_value.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies this tool is for checkable controls and warns against blind toggling when the control already matches, which is a condition for calling it. It does not explicitly name alternatives or state when to prefer desktop_set_expanded or desktop_set_value.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_set_expandedB

Expand or collapse a supported control and verify state. Reinspect newly exposed children.

ParametersJSON Schema
NameRequiredDescriptionDefault
expandedYes
element_idYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It adds meaningful context beyond the bare action by stating it 'verify state' and 'Reinspect newly exposed children,' which suggests the tool confirms the outcome and updates the element tree. It does not cover failure handling, permissions, or side effects, so there are still gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The primary action is front-loaded, and the secondary behavior (reinspection) is a necessary caveat. Every clause adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool, the description covers the core action and even adds verification and reinspection context. However, it does not define what qualifies as a 'supported control,' what the return value is, or whether it waits for state changes; absent annotations and an output schema, these gaps matter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not explain element_id or expanded. The boolean expanded can be loosely inferred from 'Expand or collapse,' but there is no explicit mapping between the parameters and their roles, which is weak given the description must compensate for an empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Expand or collapse') and resource ('a supported control'), and adds a distinct second action ('verify state'). It clearly indicates what the tool does and is distinguishable from sibling tools like desktop_set_value or desktop_set_checked, though it does not explicitly name a sibling it is not.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for expanding or collapsing controls that support this behavior, which is a clear usage context. However, it does not explicitly state when not to use it, mention prerequisites, or point to alternative sibling tools, leaving some selection reasoning to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_set_valueA

Set a numeric control within its inspected range and verify its accessibility numeric value. Displayed formatting and application commit may differ; inspect/read both, then explicitly commit only when intended.

ParametersJSON Schema
NameRequiredDescriptionDefault
valueYes
element_idYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It does this well by warning that displayed formatting and application commit may differ, requiring verification and explicit commit only when intended. This is meaningful behavioral context beyond a generic 'set' operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two dense sentences with no filler. It front-loads the core action and then delivers an essential caveat about verification and commit behavior, earning its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter setter with no output schema, the description covers the main prerequisite (inspect/read both) and the key risk (commit/display mismatch). It is slightly sparse on return/outcome details, but generally complete for invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate by explaining parameters. It vaguely implies the value is numeric and should fall within an inspected range, but it does not explain how to obtain element_id or specify value constraints beyond that. The two parameters remain under-documented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and target: 'Set a numeric control...' and adds verification of the accessibility value. This differentiates it from sibling setters like desktop_set_checked and desktop_set_expanded by limiting scope to numeric controls, although it does not explicitly name any alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage context: use it to set numeric controls within an inspected range and verify the accessible value. It does not explicitly state when not to use this tool or name alternatives, leaving selection guidance somewhat implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_statusA
Read-only

Return recent operation outcomes after timeout/cancellation. This operation history excludes input text and screenshots.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint annotation already establishes that this is a read-only operation. The description adds useful behavioral context beyond that by disclosing that the operation history excludes input text and screenshots, which is important privacy-relevant information. It could go further by describing what fields or limits the history includes, but for a simple read-only tool this is solid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no filler. The primary behavior is stated first, and the exclusion detail is a single additional clause that earns its place. This is appropriately minimal for a zero-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, zero-parameter, read-only tool with no output schema, the description explains what it returns, when to use it, and what it deliberately excludes. It does not describe the exact shape of the operation outcomes or any history limits, but those are relatively minor gaps given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters and schema coverage is trivially 100%, so the input schema carries no burden. The description adds context about when results are relevant ('after timeout/cancellation'), which is appropriate for a no-argument tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Return') and resource ('recent operation outcomes'), making the tool's function clear. It also adds context with 'after timeout/cancellation'. It does not explicitly name or contrast with siblings like desktop_report or desktop_doctor, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'after timeout/cancellation' provides a clear context for when this tool should be used. However, it does not mention alternatives or give negative guidance about when not to use it, which would have made it fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_typeA

Type into an editable element and verify exact readback. Native supported edits address the control in the background; native input chooses independent or foreground control automatically. Insert replaces the selection; replace changes the entire field. Preserves Unicode, LF and tabs without submitting. Owned HTML fields require complete grapheme boundaries. Exact readback does not confirm saving or application commit; inspect before an explicit commit or suggestion selection.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoinsert
textYes
element_idYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses important behaviors: verification of exact readback, preservation of Unicode/LF/tabs without submitting, mode semantics (insert vs. replace), and the caveat that readback does not confirm saving/commit. This is substantial and honest, though it omits error handling or permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the primary purpose, followed by mode semantics and caveats. It is dense but every sentence adds value—no fluff. The structure flows logically from core action to edge cases, making it easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (modes, edge cases, commit caveat), the description covers key aspects: how it interacts with controls, what it preserves, and its limitations. It does not address error scenarios or prerequisites (e.g., element focus), but given no output schema and only three parameters, it is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain parameters. It explicitly defines the 'mode' parameter ('Insert replaces the selection; replace changes the entire field') and indirectly explains 'text' via grapheme boundaries. However, 'element_id' is not described beyond being an identifier, and no parameter-level details are given for text formatting or length limits.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Type into an editable element and verify exact readback.' It specifies the verb (type) and resource (editable element), and adds a verification aspect. It distinguishes from siblings like desktop_paste or desktop_set_value by emphasizing native input and readback verification, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides contextual conditions (e.g., 'Owned HTML fields require complete grapheme boundaries') but does not explicitly state when to prefer this tool over alternatives such as desktop_type_secret or desktop_set_value. It lacks exclusions or alternative routing, leaving the agent to infer usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_type_secretA

Replace an observed protected field. Never reads back or echoes the value, uses no clipboard, and reports dispatched only. Requires native protected EditableText or an owned password input with known inactive composition. Owned input rejects LF/CR and declared maxlength overflow; submission is separate. Applications control their own masking, which can change during input.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
element_idYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does so exceptionally well. It discloses that the tool never reads back or echoes the value, uses no clipboard, reports dispatched only, rejects LF/CR and maxlength overflow, and notes that masking is application-controlled. This is rich, security-relevant behavior far beyond a generic 'types text.'

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four dense sentences with no filler. The core action is front-loaded in the first sentence, and each subsequent sentence adds a distinct, necessary constraint or behavior. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of annotations and output schema, the description covers the critical behavioral and prerequisite context well: target requirements, input constraints, and response reporting. It does not fully clarify jargon like 'owned password input' or explain how to obtain the element_id, but these are relatively minor gaps for a two-parameter tool with such a detailed description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does by explaining constraints relevant to the text value (LF/CR rejection, maxlength behavior, no echo) and by clarifying the target element requirements (protected EditableText or owned password input). It does not explicitly map element_id and text to their meanings, but the parameter names plus these constraints make the semantics reasonably clear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Replace an observed protected field.' It clearly differentiates itself from sibling tools like desktop_type, desktop_paste, and desktop_set_value by emphasizing protected-field usage, no clipboard, and no read-back. This gives an agent a precise idea of what the tool does without inspecting the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when the tool is appropriate: it requires a native protected EditableText or an owned password input with known inactive composition. However, it does not explicitly name alternatives or state when-not-to-use, such as pointing to desktop_type for normal fields. This is clear context with no explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_waitA
Read-only

Wait up to timeout (0–10 seconds) for an observed condition. window_present/window_absent/window_active require only window_id, never a title in text; discover title matches with desktop_windows(query=...). text_equals/text_contains require only element_id and text. element_present/element_absent require window_id plus at least one name/role substring or states filter; they search native accessibility nodes, and absence requires complete coverage. pixels_stable requires window_id and optionally stable_for (0.1–10 seconds, at most timeout); it samples the client rectangle, not general application idleness. Omit parameters belonging to other condition types.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
roleNo
textNo
statesNo
timeoutNo
conditionYes
window_idNo
element_idNo
stable_forNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the description is expected to add context. It does: it explains that element search targets 'native accessibility nodes,' that element_absent 'requires complete coverage,' and that pixels_stable 'samples the client rectangle, not general application idleness.' These nuances help agents use the right parameters. It omits timeout-failure behavior (error vs. false return) which would be useful, but the read-only annotation covers the safety profile, so this is a mild gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description takes one dense paragraph, front-loading the core purpose before systematically enumerating each condition type and its parameter rules. Every sentence adds operational value—no filler or repetition. The structure flows from general to specific, making it easy to scan despite covering eight distinct modes. It is appropriately concise for the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 9 parameters and 8 condition types, the description covers the essential usage constraints and behavior nuances, including the crucial note that pixels_stable samples the client rectangle and the discovery pointer to desktop_windows. It does not describe what the tool returns on success or timeout (e.g., whether it throws), and it assumes element_id/window_id are obtained elsewhere without detailing how. These are minor omissions given the absence of an output schema, but a complete description would address them.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% (no descriptions in the input schema), so the description carries the full burden. It does so admirably: it maps each condition to its required parameters (e.g., 'element_present/element_absent require window_id plus at least one name/role substring or states filter'), gives value ranges ('timeout (0–10 seconds),' 'stable_for (0.1–10 seconds)'), and clarifies relationships. This exceeds what a bare schema offers, compensating fully for the missing parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific purpose: 'Wait up to timeout for an observed condition.' It enumerates the eight condition types and their parameter requirements, clearly distinguishing this wait/observe tool from its action-oriented siblings like desktop_click or desktop_type. The verb 'wait' plus the resource 'observed condition' is specific, and the coverage of each condition makes the scope unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit per-condition parameter guidance: 'window_present/window_absent/window_active require only window_id,' 'text_equals/text_contains require only element_id and text,' and so on. It also advises 'Omit parameters belonging to other condition types' and points to desktop_windows(query=...) for discovering titles, an alternative for that sub-task. However, it does not explicitly contrast with other observation tools like desktop_observe, so agents must infer when this wait-with-timeout tool is more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_windowA

Manage one window. move uses frame x/y; resize uses client width/height; workspace requires its index. fullscreen requests WM fullscreen; restore exits fullscreen/maximization/minimization; raise changes stacking without activation. Other actions take no extra parameters. Close reports an owned blocking dialog without confirming it. Geometry actions return fresh same-generation client/frame bounds and WM state; move/resize distinguish matched from nonmatching requests without upgrading dispatched effects. Restore compares captured pre-maximize bounds when observed geometry, state and size hints remain consistent; otherwise comparison is unknown. No saved geometry is forced, and unobserved external changes cannot be ruled out.

ParametersJSON Schema
NameRequiredDescriptionDefault
xNo
yNo
widthNo
actionYes
heightNo
window_idYes
workspaceNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It discloses several important behaviors: close reports a blocking dialog without confirming, geometry actions return fresh bounds and WM state, move/resize distinguish matched vs. nonmatching requests, and restore compares captured pre-maximize bounds. It also acknowledges limitations (no saved geometry forced, unobserved changes cannot be ruled out). This is thorough but could be clearer about return format and error conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-organized: it lists action-specific parameter requirements first, then describes behaviors. It is front-loaded with the key scoping and parameter notes. Each sentence provides new information, though the latter sentences about geometry comparisons are complex and could be simplified. There is no wasted wording, but the density might hinder readability for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (9 actions, 7 parameters, no output schema or annotations), the description covers action-specific parameters and key behavioral nuances such as blocking dialogs and geometry reporting. However, it does not mention return values in detail (e.g., what 'fresh bounds' look like), potential failure modes, or permissions required. These are minor gaps given the lack of output schema, but for a robust agent, more detail on response format would be helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the meaning of x/y (frame), width/height (client), workspace (requires index), and says other actions take no extra parameters. It does not explain the exact units or data types beyond schema, but the semantic mapping is helpful. The description adds value by linking parameters to actions, which the schema alone does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool manages a single window and enumerates each action with its specific semantics (e.g., move uses frame x/y, resize uses client width/height). It distinguishes from siblings like desktop_windows (plural) by emphasizing singular scope. The verb 'Manage' plus resource 'window' is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly outlines which parameters are needed for which actions (move/resize/workspace) and notes that other actions take no extra parameters. It does not explicitly state when not to use this tool or compare with siblings like desktop_activate or desktop_invoke, but the action enum and parameter guidance give clear usage context. A small gap exists in not mentioning alternatives for actions like close vs. desktop_control.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_windowsA
Read-only

List window identities, titles, focus and client bounds, optionally filtering title/class. Returns counts, unavailable rows and next_offset when paginated. Each call is a fresh enumeration.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNo
offsetNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the read-only nature is structured. The description adds value by disclosing that each call is a fresh enumeration and that results include counts, unavailable rows, and next_offset when paginated. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler: purpose is front-loaded, then return characteristics, then behavioral note. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with no output schema, the description covers enough: what is returned, optional filter, pagination signals, and fresh enumeration. It could be more explicit on limit/offset mechanics, but is largely complete for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden. It adds meaning by explaining that the query filters by title/class and that limit/offset relate to pagination via next_offset. However, it does not explicitly define each parameter's semantics or the default behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') with a precise resource ('window identities, titles, focus and client bounds') and states the optional filtering by title/class. This clearly distinguishes it from singular desktop_window by indicating a multi-window enumeration scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: it lists window details, supports filtering, and returns pagination info. However, it does not explicitly say when to use this tool versus siblings like desktop_window or desktop_status, leaving the choice to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

desktop_workspacesA

List workspaces, or switch to an existing index and verify the active workspace. Every dispatched switch expires this server’s screenshots, even when requesting the current workspace; observe again before using coordinates. External workspace changes between observations are not continuously tracked.

ParametersJSON Schema
NameRequiredDescriptionDefault
workspaceNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description discloses important side effects: every dispatched switch expires the server's screenshots, even when requesting the current workspace, and external workspace changes are not continuously tracked. These are non-obvious traits an agent must know to avoid coordinate-related errors.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences front-load the primary purpose and then deliver critical behavioral warnings. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema, the description covers purpose, parameter semantics, and critical behavioral caveats. It lacks explicit return-value details and error behavior for invalid indices, but the core usage is sufficiently clear for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only an optional integer or null 'workspace' with 0% description coverage. The description adds meaning by contrasting 'list workspaces' (null) with 'switch to an existing index' (integer), and warns that even switching to the current index has side effects. This compensates for the schema's lack of semantic detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states specific operations: listing workspaces and switching to an existing index, with verification of the active workspace. It identifies the resource (workspaces) and differentiates from sibling tools like desktop_windows by focusing on workspace-level management.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for enumerating or switching workspaces but does not explicitly compare to alternatives or state when to use this tool over siblings such as desktop_windows or desktop_status. The warning about screenshot expiration provides context for sequencing with observation tools, but no exclusions or alternative routing are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 38 tool updatesv0.3.4
    • Changeddesktop_activate1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
    • Addeddesktop_applications
    • Addeddesktop_choose
    • Changeddesktop_click1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
    • Addeddesktop_control
    • Changeddesktop_doctor1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
    • Changeddesktop_drag1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
    • Addeddesktop_drag_to
    • Removeddesktop_enter_text
    • Changeddesktop_focus_element1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
    • Addeddesktop_hover
    • Changeddesktop_inspect5 fields changed
      • addedInput schema / additionalProperties
        Added value: +false
      • addedInput schema / properties / max_depth
        Added value: +{
        +  "default": 30,
        +  "title": "Max Depth",
        +  "type": "integer"
        +}
      • addedInput schema / properties / name
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Name"
        +}
      • addedInput schema / properties / role
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Role"
        +}
      • addedInput schema / properties / states
        Added value: +{
        +  "anyOf": [
        +    {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "States"
        +}
    • Changeddesktop_invoke5 fields changed
      • addedInput schema / additionalProperties
        Added value: +false
      • addedInput schema / properties / action / anyOf
        Added value: +[
        +  {
        +    "type": "string"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • addedInput schema / properties / action / default
        Added value: +null
      • removedInput schema / properties / action / type
        Removed value: -"string"
      • changedInput schema / required
        Previous value: -[
        -  "element_id",
        -  "action"
        -]New value: +[
        +  "element_id"
        +]
    • Addeddesktop_launch
    • Addeddesktop_match_image
    • Changeddesktop_observe1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
    • Addeddesktop_ocr
    • Addeddesktop_open_browser
    • Addeddesktop_paste
    • Changeddesktop_press_keys2 fields changed
      • addedInput schema / additionalProperties
        Added value: +false
      • addedInput schema / properties / count
        Added value: +{
        +  "default": 1,
        +  "title": "Count",
        +  "type": "integer"
        +}
    • Changeddesktop_read_text1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
    • Addeddesktop_reconnect
    • Addeddesktop_recording
    • Addeddesktop_recover_input
    • Addeddesktop_report
    • Changeddesktop_scroll1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
    • Addeddesktop_select
    • Addeddesktop_set_checked
    • Addeddesktop_set_expanded
    • Removeddesktop_set_text
    • Addeddesktop_set_value
    • Addeddesktop_status
    • Addeddesktop_type
    • Addeddesktop_type_secret
    • Addeddesktop_wait
    • Addeddesktop_window
    • Changeddesktop_windows4 fields changed
      • addedInput schema / additionalProperties
        Added value: +false
      • addedInput schema / properties / limit
        Added value: +{
        +  "default": 50,
        +  "title": "Limit",
        +  "type": "integer"
        +}
      • addedInput schema / properties / offset
        Added value: +{
        +  "default": 0,
        +  "title": "Offset",
        +  "type": "integer"
        +}
      • addedInput schema / properties / query
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Query"
        +}
    • Addeddesktop_workspaces
  2. 14 tool updatesv0.1.0
    • First observeddesktop_activate
    • First observeddesktop_click
    • First observeddesktop_doctor
    • First observeddesktop_drag
    • First observeddesktop_enter_text
    • First observeddesktop_focus_element
    • First observeddesktop_inspect
    • First observeddesktop_invoke
    • First observeddesktop_observe
    • First observeddesktop_press_keys
    • First observeddesktop_read_text
    • First observeddesktop_scroll
    • First observeddesktop_set_text
    • First observeddesktop_windows

TDQS

A3.7/5.0

Scored across 36 tools

Disambiguation4/5

Most tools target clearly distinct operations, and the detailed descriptions resolve boundary cases such as desktop_select vs desktop_choose or desktop_inspect vs desktop_read_text. A few diagnostic/status tools (desktop_status, desktop_report, desktop_doctor) and singular/plural window tools could still be confused at first glance, but not to the point of causing frequent misselection.

Naming Consistency4/5

All tool names share the desktop_ prefix and use lowercase snake_case, which provides a strong overall pattern. Most names are verb-led (desktop_click, desktop_set_value, desktop_open_browser), though a few noun-only names like desktop_windows, desktop_applications, and desktop_status are minor deviations from a strict verb_noun convention.

Tool Count2/5

36 tools is well above the 25-tool threshold and feels heavy even for a desktop automation server. Each tool is individually meaningful, but the set could be consolidated or grouped more tightly without losing essential functionality.

Completeness5/5

The tool surface covers the full desktop automation lifecycle: discovery, observation, inspection, input, window management, verification, and diagnostics/recovery. There are no obvious dead ends; every interaction has a corresponding observation or inspection tool, and failure/health handling is well supported.

Maintenance

ActivityNo data
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    A desktop automation MCP server that enables AI agents to interact with Linux environments through screenshots, window inspection, and input simulation. It provides tools for mouse control, keyboard input, and screen capture using xdotool and XDG Desktop Portals.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server for reliable native macOS desktop control from AI agents, providing 72 tools for screenshots, mouse, keyboard, scroll, clipboard, window management, and more.
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    An MCP server that gives a model eyes and hands on a Linux Wayland desktop, enabling screenshot capture, mouse/keyboard control, OCR, and icon detection via OmniParser.
    1
    -
  • A
    license
    Not graded
    quality
    A
    maintenance
    MCP server for controlling Linux desktops over Wayland, enabling AI agents to perform mouse, keyboard, window, and screenshot operations on Fedora KDE Plasma.
    AGPL 3.0