Skip to main content
Glama

longpress

Hold a UI element to open a context menu or trigger a press-and-hold gesture. Target by snapshot ref, selector, or coordinates, with adjustable duration.

Instructions

Hold a UI target by snapshot ref, selector, or coordinates to open a context menu or perform another hold gesture. Set durationMs when the default hold duration is unsuitable.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
cwdNoWorking directory for command execution.
rawNoUse raw snapshot data during selector resolution.
udidNoiOS device UDID selector.
debugNoEnable debug diagnostics.
depthNoSnapshot traversal depth.
runIdNoLease run identifier.
scopeNoSnapshot scope selector used before resolution.
deviceNoDevice name selector.
serialNoAndroid device or Vega VVD serial selector.
settleNoAfter the action, wait for the UI to go quiet and return the settled diff vs the pre-action tree in the same response. Best-effort; never fails the action.
targetYesUI target. This is separate from deviceTarget, which selects the device form.
tenantNoRemote tenant identifier.
leaseIdNoExisting lease identifier.
sessionNoAgent-device session name.
noRecordNoDo not record this action.
platformNoPlatform selector used to resolve a device.
stateDirNoAgent-device state directory.
timeoutMsNoSettle: wait deadline in milliseconds (default 10000).
durationMsNoLong press duration in milliseconds.
includeCostNoInclude per-command agent-cost (cost.wallClockMs, …) in structuredContent. Defaults to off; the default response shape is unchanged.
deviceTargetNoDevice target form. Maps to the CLI --target flag.
daemonBaseUrlNoRemote daemon base URL.
responseLevelNoResponse verbosity: token-cheap digest / default (today) / full. Defaults to default; the default response shape is unchanged.
settleQuietMsNoSettle: quiet window in milliseconds (default 500).
daemonAuthTokenNoRemote daemon auth token.
iosXctestEnvDirNoWritable directory for iOS XCTest runner env overlays.
mcpOutputFormatNoMCP text content format. Defaults to optimized agent-friendly text; use json for JSON text. Structured content is always returned separately.
iosXctestrunFileNoExternally built iOS XCTest runner .xctestrun artifact path.
iosSimulatorDeviceSetNoiOS simulator device-set path used for device resolution.
androidDeviceAllowlistNoAndroid serial allowlist used for device resolution.
iosXctestDerivedDataPathNoDerived data path for external iOS XCTest runner execution.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
xNoResolved interaction x coordinate when available.
yNoResolved interaction y coordinate when available.
refNoSnapshot ref without the @ prefix when the target was an @ref.
costNo
hintNo
settleNo
gestureNo
messageNo
warningNo
evidenceNo
refLabelNo
selectorNoSelector expression when the target was a selector.
durationMsNo
resolutionNoPre-action disclosure of how the acting path resolved its target. Absent when resolutionDisclosure is inapplicable for the path.
targetKindYesResolved interaction target kind.
selectorChainNo
referenceWidthNoReference frame width for visualizing the interaction point.
targetHittableNo
referenceHeightNoReference frame height for visualizing the interaction point.
maestroFallbackReasonNo
maestroNonHittableCoordinateFallbackUsedNoWhether the direct iOS Maestro coordinate fallback was actually used.
maestroNonHittableCoordinateFallbackAllowedNoWhether the direct iOS Maestro coordinate fallback was allowed for this selector.
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries full responsibility for behavioral transparency. It reveals the action (hold) and its purpose, but doesn't disclose side effects, whether the action mutates UI state, how long the default hold is, or outcome/error behavior. This is a minimal disclosure for an interactive gesture tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, immediately front-loaded with the action and target. No redundant words; every sentence contributes meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 31 parameters, but the schema fully documents them and an output schema exists. The description is extremely brief; while adequate for selection and basic invocation, it lacks contextual information like when to prefer longpress over similar gestures, or any platform-specific nuances. Given the complexity, slightly more guidance would help.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so a baseline of 3 applies. The description adds value by summarizing the three target variants (ref, selector, coordinates) that are encoded in the target oneOf schema, and it clarifies when durationMs should be used ('when the default hold duration is unsuitable'). This exceeds bare schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool performs a long-press/hold gesture on a UI target, specifying target resolution methods (snapshot ref, selector, coordinates) and the purpose (open context menu or other hold gesture). The verb 'Hold' and resource 'UI target' are specific, and the description distinguishes from short-press actions like click or tap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clear context is provided: use when a hold/long-press gesture is needed, such as opening a context menu. The description offers parameter guidance ('Set durationMs when the default hold duration is unsuitable'), but doesn't explicitly mention alternatives like press or gesture or exclusion scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Install Server

Other Tools

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/callstack/agent-device'

If you have feedback or need assistance with the MCP directory API, please join our Discord server