Skip to main content
Glama

glass_drag

Drag the pointer between two window-relative coordinates while holding a button and optional modifiers, with a configurable duration to simulate mouse drags for GUI testing.

Instructions

Drag one pointer from (x1,y1) to (x2,y2), holding the button and modifiers throughout. Motion spans duration_ms. Either endpoint outside the window is refused. Use glass_gesture for multi-touch. For 2+ known steps, use glass_do.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
x1YesPress x in window-relative pixels.
x2YesRelease x in window-relative pixels.
y1YesPress y in window-relative pixels.
y2YesRelease y in window-relative pixels.
buttonNoButton held for the drag: "left" (default), "right", or "middle".
modifiersNoHeld modifiers, e.g. ["ctrl", "shift"].
duration_msNoMotion duration in ms (default 200). Faster motion gives the app fewer sampled frames.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed6 schema fields changedv1.8.0
    • changedInput schema / properties / duration_ms / description
      Previous value: -"Span the drag's motion over this many milliseconds so a frame-based GUI\nsamples the path across multiple frames (and registers the drag even while\nit repaints). Default 200. Lower = faster but coarser."New value: +"Motion duration in ms (default 200). Faster motion gives the app fewer sampled frames."
    • changedInput schema / properties / modifiers / description
      Previous value: -"Modifier keys to hold during the action, e.g. [\"ctrl\"] or [\"ctrl\",\"shift\"] for multi/range-select."New value: +"Held modifiers, e.g. [\"ctrl\", \"shift\"]."
    • changedInput schema / properties / x1 / description
      Previous value: -"Press-point x, window-relative — 0 is the window's left edge, not the screen's."New value: +"Press x in window-relative pixels."
    • changedInput schema / properties / x2 / description
      Previous value: -"Release-point x, window-relative."New value: +"Release x in window-relative pixels."
    • changedInput schema / properties / y1 / description
      Previous value: -"Press-point y, window-relative — 0 is the window's top edge, not the screen's."New value: +"Press y in window-relative pixels."
    • changedInput schema / properties / y2 / description
      Previous value: -"Release-point y, window-relative."New value: +"Release y in window-relative pixels."
  2. Changed6 schema fields changedv1.2.0
    • addedInput schema / properties / button / description
      Added value: +"Button held for the drag: \"left\" (default), \"right\", or \"middle\"."
    • addedInput schema / properties / x1 / description
      Added value: +"Press-point x, window-relative — 0 is the window's left edge, not the screen's."
    • addedInput schema / properties / x2 / description
      Added value: +"Release-point x, window-relative."
    • addedInput schema / properties / y1 / description
      Added value: +"Press-point y, window-relative — 0 is the window's top edge, not the screen's."
    • addedInput schema / properties / y2 / description
      Added value: +"Release-point y, window-relative."
    • removedInput schema / title
      Removed value: -"DragArgs"
  3. First observedv1.0.1

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (readOnlyHint=false, destructiveHint=false, openWorldHint=false), so the description must carry the behavioral burden. It does well by disclosing the holding of the button/modifiers, the duration control, and the endpoint refusal condition. It also notes in the parameter description that faster motion gives the app fewer sampled frames, which is valuable. It does not disclose side effects (e.g., whether this triggers navigation or state changes) but for a drag tool the key behaviors are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is succinct - three sentences that cover the core action, a key constraint, and routing to alternatives. It is front-loaded with the primary action and immediately provides differentiation. Every sentence earns its place, with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no output schema and the complexity is moderate, the description covers the core action, constraints, and sibling routing. It does not specify the return value or error behavior, but that is often implicit. The guidance on duration effects and endpoint refusal is sufficient for an agent to call it correctly. Missing a brief note on the default button and modifier behavior is negligible because the schema covers those defaults.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the schema already documents all parameters clearly. The description adds minimal extra semantic value: it explains the holding behavior and the effect of duration on sampling, but the parameter details are already comprehensive. Thus a baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (drag one pointer from (x1,y1) to (x2,y2)) and the resource (a pointer in a window), with specific details about holding the button and modifiers. It explicitly differentiates from siblings by naming glass_gesture for multi-touch and glass_do for 2+ known steps, which are the primary alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: for multi-touch use glass_gesture, for 2+ known steps use glass_do. It also mentions the constraint that either endpoint outside the window is refused, which sets expectations for when the tool cannot be used. This is thorough and specific.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.