Skip to main content
Glama

desktop_drag

Perform a drag between two points within a window, auto-selecting pointer input and ensuring button release for reliable desktop automation.

Instructions

Drag between two observed points inside the same window, automatically choosing independent or foreground pointer input after validation; always attempts release of its own button. Use desktop_drag_to for another destination window.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
xYes
yYes
end_xYes
end_yYes
buttonNoleft
window_idYes
snapshot_idYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.3.4
    • addedInput schema / additionalProperties
      Added value: +false
  2. First observedv0.1.0

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It goes beyond a generic 'drag' by explaining that input mode is chosen automatically after validation and that the tool always attempts to release its own button, which is highly relevant safety-relevant behavior. It does not fully cover failure modes or side effects, but the disclosed behaviors are specific and meaningful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, with the core action and key behavioral caveat front-loaded and the sibling alternative mentioned last. Every clause earns its place and the structure aids quick comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the absence of annotations and output schema, the description covers the most important operational aspects: the drag's scope, the validation/input-mode behavior, the release guarantee, and the sibling alternative. The main omissions are explicit coordinate semantics and potential failure conditions, but the description is strong enough for an agent to invoke the tool with reasonable confidence.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It helps map 'two observed points' to the coordinate parameters, 'same window' to window_id, and 'observed points' to snapshot_id, but it does not explain the button parameter, coordinate units, or the exact role of the snapshot. This partial semantic addition justifies a middle score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Drag'), a resource scope ('between two observed points inside the same window'), and explicit behavioral details like input-mode selection and self-release. It additionally names the sibling 'desktop_drag_to' with the distinguishing condition 'for another destination window', so the agent can clearly tell them apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to use this tool for drags within the same window and points to 'desktop_drag_to' as the alternative for a different destination window. This gives clear when-to-use and when-not-to-use guidance without requiring the agent to infer it from schemas.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.