Skip to main content
Glama

maestro_test

Run Maestro flows on a booted simulator, stream the run live in the Viewer, and return pass/fail per flow with JUnit reports.

Instructions

Run the app's Maestro flows (.maestro/) on the booted simulator through tools/maestro-live: it starts maestro mcp, opens the Maestro Viewer (http://localhost:7777, or the next free port) in the browser so the founder watches the run live, and runs each flow with the MCP run tool. Returns pass/fail per flow (JUnit also in build/maestro/report.xml). tags: comma-separated include filter, e.g. "smoke". The app must already be installed (build_for_sim + simctl install). Writes .appfactory/verify/maestro.json: the features gate and testflight_ship require every smoke flow green. There is no headless option here: the Viewer is mandatory on the Mac.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
tagsNo
deviceNo
app_dirYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.1

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden and does so richly: it discloses that it spawns `maestro mcp`, forces a browser Viewer on localhost:7777 (next free port) with no headless option, writes .appfactory/verify/maestro.json, produces JUnit at build/maestro/report.xml, and returns pass/fail per flow. Side effects, environment constraints, and outputs are all made explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and Viewer behavior are front-loaded, and the sentences are dense with useful facts rather than filler. It is a long single paragraph, though, and details like the exact port number sit alongside higher-value prerequisites without structural separation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-step, side-effecting test runner the description covers prerequisites, mandatory Viewer, artifact paths, and gating consequences; an output schema exists so return values need not be restated (it summarizes them anyway). The notable gap is the undocumented `device` parameter and ambiguity about what `app_dir` should point to.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains `tags` well ('comma-separated include filter, e.g. "smoke"'), but says nothing about `app_dir` (only inferable) and nothing about `device`, leaving two of three parameters semantically undefined.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Run the app's Maestro flows on the booted simulator') plus the exact mechanism (tools/maestro-live, MCP `run` tool), which clearly separates it from generic build/test siblings like build_test. It never names a sibling to route away from, so it falls just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a concrete precondition ('The app must already be installed (build_for_sim + simctl install)') and explains downstream context (the features gate and testflight_ship require every `smoke` flow green), so an agent knows when this step is required. It does not explicitly state when another tool should be used instead, so no exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools