Skip to main content
Glama

review_workflow

Read-onlyIdempotent

Sanity-check a planned workflow before execution. Validates structure, catches delete-then-use, missing params, and approval gaps to prevent production failures.

Instructions

[READ] Sanity-check a planned workflow before execution.

Performs structural validation only — does NOT call into other skills. Catches the common authoring errors before they hit production:

  • Delete-then-use: a step deletes resource X, a later step references X

  • Missing required params: a step has empty params or placeholder values

  • Cross-skill order issues: surfacing the cross-skill dispatch sequence

  • Risk profile: count of destructive / write / read-only / unclassified steps

  • Approval coverage: is every destructive OR unclassifiable step gated behind a preceding require_approval?

Each step is placed in a tier from the tool's entry in get_skill_catalog first, then from name patterns. A step that matches neither is reported as ungated_unclassified rather than assumed safe: pilot dispatches nothing itself and cannot inspect a sibling skill's annotations, so "this tool is unknown to me" is the honest finding, and it needs the same gate a known destructive step does. The remedy is in the message — add the tool to SKILL_CATALOG with the risk its own skill declares, or add a gate.

Medium-risk writes (create / scale / enable) are classified and counted but not gated: the family gates destructive work, and several built-in templates deliberately create in staging before asking for approval.

Returns: Dict with keys: - verdict: "approved" if no structural issues, otherwise "needs_revision" - findings: list of {severity, kind, message, step_index}. Kinds ungated_destructive, ungated_unclassified and destructive_in_parallel_group are what run_workflow refuses on; force=True overrides that only for a built-in template, never for a custom workflow's missing approval gate. - summary: counts — total/destructive/write/read_only/approval_gates, parallel_groups, est_duration_min, plus classified_steps and unclassified_steps so an "approved" verdict can be told apart from a workflow this review could not read.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
workflow_idYesThe workflow ID returned by ``plan_workflow``.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changedv1.11.1
    • addedInput schema / additionalProperties
      Added value: +false
    • addedInput schema / properties / workflow_id / description
      Added value: +"The workflow ID returned by ``plan_workflow``."
  2. First observedv1.5.22

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses that this is a pure read/review operation that cannot inspect sibling skill annotations, explains why unclassified steps are reported as 'ungated_unclassified' rather than assumed safe, details the classification order (get_skill_catalog first, then name patterns), and clarifies force=True only applies to built-in templates. This is substantial behavioral nuance the annotations do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose and scoping, and uses clear bullet lists and a return-key breakdown. It is long, and there is some redundancy between 'does NOT call into other skills' and the later 'pilot dispatches nothing itself' passage, but the detail is mostly justified given the absence of an output schema and the nuanced validation policy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully documents the return structure (verdict, findings, summary), the finding kinds that cause run_workflow to refuse, the force=True caveat, and the classification rules. An agent has all necessary information to invoke the tool correctly and interpret its results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already fully describes workflow_id as 'The workflow ID returned by plan_workflow' with 100% coverage. The description adds only contextual framing like 'planned workflow' and 'before execution' but no additional parameter-specific semantics, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Sanity-check a planned workflow before execution.' It clearly differentiates from siblings by stating 'structural validation only' and 'does NOT call into other skills', so it cannot be confused with run_workflow, plan_workflow, or design_workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to use it 'before execution' and 'before they hit production', and states when-not: 'structural validation only' and 'does NOT call into other skills'. It also references run_workflow's refusal behavior, force=True semantics, and the gating relationship with require_approval, giving clear contextual guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.