Skip to main content
Glama

Robot Actions — Remote Device Control

flow_replay_start

Kick off a flow replay in the background. Returns a replayId immediately — pass it to flow_replay_status to poll progress. Use this instead of flow_recording_replay when you want to monitor live or do other work while the replay runs. Validation failures are auto-skipped (MCP has no interactive input).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
platformNoDevice platform — auto-detected from recording if omitted
targetUdidYesUDID of the device to replay on
recordingIdYesFlow recording ID to replay (from flow_recording_list)
resetAppDataNoWipe the recorded app's data (Android `pm clear`) BEFORE replay so it starts from a clean first-run state. Use this for recordings of enrollment / first-run / logged-out flows (e.g. create-passcode) that will NOT reproduce against an already-enrolled or logged-in app. Destructive — erases the app's local data on the target device. Android only; default false.
validateElementsNoUse recorded element locators to find targets before acting (default: true)
visualCheckEnabledNoOpt in to AI Visual Review: capture per-step baselines and run the end-of-replay visual-analysis phase. Slower; default false.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It clearly discloses that it runs in the background, returns immediately, auto-skips validation failures, and implicitly signals a poll-able async workflow. The destructive nature of the resetAppData parameter is well documented within the param description, and the disabled-interactive-input limitation surfaces a genuine behavior an agent could otherwise be surprised by. It doesn't detail return format or error behavior, but given zero annotations the description is strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences that earn their place: the action and return contract, the alternative-tool disambiguation, and the validation-failure behavior. No filler, no restating of the schema, and the most decision-relevant info (you get an ID back, poll separately) is front-loaded in the first sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool returns a replayId and is part of a multi-tool workflow (flow_replay_start → flow_replay_status/abort/step). The description adequately ties into that workflow by naming the poll tool and explaining the async contract. With no output schema, the description carries the burden of explaining what the caller receives, which it does. It could describe the replayId's other downstream consumers (flow_replay_abort, flow_replay_step) and what the other 3 params do behaviorally, but the core workflow is sufficiently specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds value beyond the schema by explaining the orchestration contract (replayId → flow_replay_status poll loop) and gives semantic guidance in resetAppData about when to enable it (enrollment/first-run/logged-out flows) and the destructive consequence. The description enriches the two required params' purpose by framing them as the minimal invocation while the optional params clarify tradeoffs. The main upstream param (recordingId) could reference where recordings come from, but flow_recording_list is already named in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses a specific verb ('Kick off a flow replay in the background') with a clear resource and key behavior (returns replayId immediately). It explicitly distinguishes from the sibling tool flow_recording_replay by noting it runs in background vs. monitoring live, and it clearly names what to pass the result to (flow_replay_status). This is a model example of a purpose that differentiates from alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use this tool ('Use this instead of flow_recording_replay when you want to monitor live or do other work while the replay runs'), naming the exact alternative and the disambiguating condition. Also discloses the behavioral constraint that validation failures are auto-skipped due to MCP lacking interactive input, informing when the tool should NOT be expected to pause for confirmation. This fully satisfies the when/when-not/alternatives requirement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

B3.1/5.0
Disambiguation2/5

The set contains near-identical duplicate families: web_* and playwright_* expose ~15 pairs of the same desktop-grid-browser operations (web_get_text/playwright_get_text, web_reload/playwright_reload), and screenshot/log/network/mock capabilities each have 5-8 entry points (device_screenshot vs android_mjpeg_screenshot vs ios_screenshot vs ios_fast_screenshot vs web_screenshot vs webpage_screenshot vs session_screenshot). Many individual descriptions carefully draw boundaries (devtools vs traffic, HID vs session), but an agent cannot reliably distinguish web_* from playwright_*, and ios_screenshot/ios_fast_screenshot/ios_mjpeg_screenshot blur together.

Naming Consistency2/5

The prefix scheme is broken: Android functionality is split arbitrarily between android_* and device_* (device_screenshot vs android_mjpeg_screenshot), the desktop browser gets two parallel prefixes (web_* and playwright_*), and verbs vary across equivalents (device_navigate_url vs web_navigate vs ios_safari_navigate). session_* uses bare verbs (session_url, session_back), and the same concept gets different names (ios_clipboard_get_hid vs ios_get_pasteboard; device_screen vs ios_orientation).

Tool Count1/5

333 tools is an extreme count by any measure — far beyond the 50+ threshold — and much of the bulk is duplicative (the web_*/playwright_* pairs alone double ~15 slots) or out-of-scope for a device-control server (TestRail, Jira, AzDO, agent memory, secret variables, feedback). Even granting that remote device control + test automation is a broad domain, this surface will devastate agent context budgets and is impossible to navigate coherently.

Completeness4/5

The core device-control and test-automation domain is remarkably thorough: Android and iOS each have full interaction, app-lifecycle, file, network/proxy, performance, crash, accessibility, recording, and replay coverage, with CRUD lifecycles for flows, suites, app uploads, TestRail cases, and visual-review baselines. Minor gaps exist at the margins — Jira/AzDO lack update/transition/comment operations, and iOS cannot open/close tabs — but the central workflows have no dead ends.

Resources