Skip to main content
Glama

flow-execute

Run a saved YAML flow end to end to replay a recorded path, re-run a QA regression, or verify a known journey still passes. Returns a per-step report; the first failure stops the run and remaining steps are skipped.

Instructions

Run a saved YAML flow end to end. Use when asked to replay a recorded path, re-run a QA regression, or check that a known journey still passes; for a one-off interaction use the gesture tools instead, and to author a flow use flow-start-recording. Pass exactly one flow source: name (under project_root) or flow_path. Returns a per-step report: the first failure stops the run and the rest report as skipped.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
nameNoName of a saved flow to run from `.argent/flows` (e.g. "settings-explore"). Omit when flow_path is set.
deviceNoDevice id to run against (iOS UDID, Android/Vega serial, Chromium id) — the id list-devices reports. Auto-detected when omitted, but only when exactly one booted device matches (optionally narrowed by `platform`); with several booted the run fails and lists them, so pass this explicitly whenever more than one device is up.
platformNoRestrict auto-detection to this platform when several devices are booted. `ios` selects local simulators only — pass `ios-remote` to select a remote one. `chromium` does more than filter: with no `device` it SELECTS the self-boot branch for an e2e flow - the runner boots an Electron instance from the `launch` step's chromium value and tears it down after the run (a single-key `launch: { chromium: … }` map selects it on its own, without this parameter). When it selects that branch it never falls back to device auto-detection (a fragment, or an e2e launch map with no `chromium` key, still does), and the launch value must be a real Electron app path on the tool-server host: a bare-string `launch:` - what the recorder writes - holds an installed-app bundle id, so passing `chromium` for one fails the whole run with `Electron boot: path does not exist`. Edit the launch to `{ chromium: <app path> }` first.
flow_fileNoPath to the flow .yaml as readable by the tool-server. Internal — the argent client derives it from project_root and name automatically; leave unset.
flow_pathNoOmit when name is set. Absolute path to a co-located flow .yaml on the client and tool server's shared filesystem. This must be supplied through the file-input boundary. For remote execution, pass name + project_root instead.
project_rootYesAbsolute path to the calling agent's project root — the cwd it is working in. With name, the saved flow is read from `.argent/flows/<name>.yaml` under this root; with flow_path, the flow, its run: siblings, its script: paths and baselines all resolve beside the YAML instead, so pass the agent's cwd. A script still RUNS in this root whichever source was used.
updateBaselinesNoWrite/refresh screenshot baselines for `snapshot` steps instead of diffing against them.
prerequisiteAcknowledgedNoSet to true to confirm the execution prerequisite has been met. Required (LLM path) when a fragment defines an executionPrerequisite.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changedv0.25.1
    • changedInput schema / properties / platform / description
      Previous value: -"Restrict auto-detection to this platform when several devices are booted. `chromium` does more than filter: with no `device` it SELECTS the self-boot branch for an e2e flow - the runner boots an Electron instance from the `launch` step's chromium value and tears it down after the run (a single-key `launch: { chromium: … }` map selects it on its own, without this parameter). When it selects that branch it never falls back to device auto-detection (a fragment, or an e2e launch map with no `chromium` key, still does), and the launch value must be a real Electron app path on the tool-server host: a bare-string `launch:` - what the recorder writes - holds an installed-app bundle id, so passing `chromium` for one fails the whole run with `Electron boot: path does not exist`. Edit the launch to `{ chromium: <app path> }` first."New value: +"Restrict auto-detection to this platform when several devices are booted. `ios` selects local simulators only — pass `ios-remote` to select a remote one. `chromium` does more than filter: with no `device` it SELECTS the self-boot branch for an e2e flow - the runner boots an Electron instance from the `launch` step's chromium value and tears it down after the run (a single-key `launch: { chromium: … }` map selects it on its own, without this parameter). When it selects that branch it never falls back to device auto-detection (a fragment, or an e2e launch map with no `chromium` key, still does), and the launch value must be a real Electron app path on the tool-server host: a bare-string `launch:` - what the recorder writes - holds an installed-app bundle id, so passing `chromium` for one fails the whole run with `Electron boot: path does not exist`. Edit the launch to `{ chromium: <app path> }` first."
    • changedInput schema / properties / platform / enum
      Previous value: -[
      -  "ios",
      -  "android",
      -  "chromium",
      -  "vega"
      -]New value: +[
      +  "ios",
      +  "android",
      +  "chromium",
      +  "vega",
      +  "ios-remote"
      +]
  2. Changed1 schema field changedv0.25.0
    • changedInput schema / properties / project_root / description
      Previous value: -"Absolute path to the calling agent's project root — the cwd it is working in. With name, the saved flow is read from `.argent/flows/<name>.yaml` under this root; with flow_path, the flow, its run: siblings, and baselines all resolve beside the YAML instead, so pass the agent's cwd."New value: +"Absolute path to the calling agent's project root — the cwd it is working in. With name, the saved flow is read from `.argent/flows/<name>.yaml` under this root; with flow_path, the flow, its run: siblings, its script: paths and baselines all resolve beside the YAML instead, so pass the agent's cwd. A script still RUNS in this root whichever source was used."
  3. Changed2 schema fields changedv0.22.0
    • removedInput schema / oneOf
      Removed value: -[
      -  {
      -    "required": [
      -      "name"
      -    ]
      -  },
      -  {
      -    "required": [
      -      "flow_path"
      -    ]
      -  }
      -]
    • changedInput schema / properties / flow_path / description
      Previous value: -"Absolute path to a co-located flow .yaml on the client and tool server's shared filesystem. This must be supplied through the file-input boundary. For remote execution, pass name + project_root instead."New value: +"Omit when name is set. Absolute path to a co-located flow .yaml on the client and tool server's shared filesystem. This must be supplied through the file-input boundary. For remote execution, pass name + project_root instead."
  4. Changed2 schema fields changedv0.20.0
    • changedInput schema / properties / device / description
      Previous value: -"Device id to run against (iOS UDID, Android/Vega serial, Chromium id). Auto-detected when omitted."New value: +"Device id to run against (iOS UDID, Android/Vega serial, Chromium id) — the id list-devices reports. Auto-detected when omitted, but only when exactly one booted device matches (optionally narrowed by `platform`); with several booted the run fails and lists them, so pass this explicitly whenever more than one device is up."
    • changedInput schema / properties / platform / description
      Previous value: -"Restrict auto-detection to this platform when several devices are booted."New value: +"Restrict auto-detection to this platform when several devices are booted. `chromium` does more than filter: with no `device` it SELECTS the self-boot branch for an e2e flow - the runner boots an Electron instance from the `launch` step's chromium value and tears it down after the run (a single-key `launch: { chromium: … }` map selects it on its own, without this parameter). When it selects that branch it never falls back to device auto-detection (a fragment, or an e2e launch map with no `chromium` key, still does), and the launch value must be a real Electron app path on the tool-server host: a bare-string `launch:` - what the recorder writes - holds an installed-app bundle id, so passing `chromium` for one fails the whole run with `Electron boot: path does not exist`. Edit the launch to `{ chromium: <app path> }` first."
  5. Changed5 schema fields changedv0.19.0
    • addedInput schema / oneOf
      Added value: +[
      +  {
      +    "required": [
      +      "name"
      +    ]
      +  },
      +  {
      +    "required": [
      +      "flow_path"
      +    ]
      +  }
      +]
    • addedInput schema / properties / flow_path
      Added value: +{
      +  "description": "Absolute path to a co-located flow .yaml on the client and tool server's shared filesystem. This must be supplied through the file-input boundary. For remote execution, pass name + project_root instead.",
      +  "type": "string"
      +}
    • changedInput schema / properties / name / description
      Previous value: -"Name of the flow to run (e.g. \"settings-explore\")"New value: +"Name of a saved flow to run from `.argent/flows` (e.g. \"settings-explore\"). Omit when flow_path is set."
    • changedInput schema / properties / project_root / description
      Previous value: -"Absolute path to the project root directory that contains `.argent/flows/<name>.yaml`."New value: +"Absolute path to the calling agent's project root — the cwd it is working in. With name, the saved flow is read from `.argent/flows/<name>.yaml` under this root; with flow_path, the flow, its run: siblings, and baselines all resolve beside the YAML instead, so pass the agent's cwd."
    • changedInput schema / required
      Previous value: -[
      -  "name",
      -  "project_root"
      -]New value: +[
      +  "project_root"
      +]
  6. Changed4 schema fields changedv0.16.0
    • addedInput schema / properties / device
      Added value: +{
      +  "description": "Device id to run against (iOS UDID, Android/Vega serial, Chromium id). Auto-detected when omitted.",
      +  "type": "string"
      +}
    • addedInput schema / properties / platform
      Added value: +{
      +  "description": "Restrict auto-detection to this platform when several devices are booted.",
      +  "enum": [
      +    "ios",
      +    "android",
      +    "chromium",
      +    "vega"
      +  ],
      +  "type": "string"
      +}
    • changedInput schema / properties / prerequisiteAcknowledged / description
      Previous value: -"Set to true to confirm the execution prerequisite has been met. Required when the flow defines an executionPrerequisite."New value: +"Set to true to confirm the execution prerequisite has been met. Required (LLM path) when a fragment defines an executionPrerequisite."
    • addedInput schema / properties / updateBaselines
      Added value: +{
      +  "description": "Write/refresh screenshot baselines for `snapshot` steps instead of diffing against them.",
      +  "type": "boolean"
      +}
  7. First observedv0.15.0

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds meaningful runtime behavior: 'Returns a per-step report: the first failure stops the run and the rest report as skipped.' This goes beyond the schema and helps the agent predict execution semantics, though it does not warn about side effects of actually running flows on a device.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose. Every sentence earns its place: when to use, which alternatives to prefer, the one-source rule, and return behavior. There is no redundant repetition of parameter details already covered in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex 8-parameter schema with 100% coverage and no output schema, the description supplies the missing high-level context: return format, stop-on-failure semantics, and source selection rules. The combination of a concise main description and richly documented parameters leaves an agent with enough information to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The main description adds value by enforcing the mutual-exclusivity rule for name and flow_path, and by clarifying where 'name' is resolved (under project_root). Since the parameter-level schema is already rich, this extra semantic guidance is useful but not extensive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific verb and resource: 'Run a saved YAML flow end to end.' It clearly differentiates from sibling tools by naming gesture tools for one-off interactions and flow-start-recording for authoring a flow. The use cases (replay recorded path, re-run QA regression, check a known journey) further clarify intent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool and when not to: use it for replaying recorded paths or QA regression, use gesture tools for one-off interactions, and use flow-start-recording to author a flow. It also provides the key routing rule 'Pass exactly one flow source: name or flow_path,' giving actionable guidance beyond the schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.