Skip to main content
Glama

rush_test_heal

Run tests repeatedly to detect flaky race conditions and propose corrective fixes.

Instructions

Diagnose flaky test race conditions and suggest fixes

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
runsNo
seedNo
targetYes
dry_runNo
allow_slowNo
allow_buildNo
allow_artifact_writeNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed8 schema fields changedv0.2.2
    • addedInput schema / properties / allow_artifact_write
      Added value: +{
      +  "default": false,
      +  "title": "Allow Artifact Write",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / allow_build
      Added value: +{
      +  "default": false,
      +  "title": "Allow Build",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / allow_slow
      Added value: +{
      +  "default": false,
      +  "title": "Allow Slow",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / dry_run
      Added value: +{
      +  "default": true,
      +  "title": "Dry Run",
      +  "type": "boolean"
      +}
    • changedInput schema / properties / runs / default
      Previous value: -5New value: +20
    • addedInput schema / properties / seed
      Added value: +{
      +  "default": 0,
      +  "title": "Seed",
      +  "type": "integer"
      +}
    • changedInput schema / title
      Previous value: -"mcp_rush_test_healArguments"New value: +"rush_test_healArguments"
    • changedOutput schema / title
      Previous value: -"mcp_rush_test_healOutput"New value: +"rush_test_healOutput"
  2. First observedv0.3.0

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It indicates the tool diagnoses and suggests fixes, but does not state whether it actually runs tests, whether it can modify files, or what side effects may occur. Parameters like dry_run and allow_artifact_write hint at behavior, but the description itself leaves these unknown.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler, and the core action is front-loaded. It earns points for being concise and clear, though it does not use the available space to add necessary context. It is appropriately sized for a simple title but the tool itself is more complex.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 7 parameters, no annotations, and only a one-line description, which is insufficient for an agent to understand how to invoke it correctly. It lacks guidance on parameter semantics, expected behavior, and how it relates to overlapping siblings like rush_flaky or rush_fix. The presence of an output schema does not make up for the missing operational context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description provides no information about any of the 7 parameters, including the required 'target' or the meaningful 'dry_run', 'seed', and 'runs'. The tool name and description do not compensate for the complete lack of parameter explanation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Diagnose') and resource ('flaky test race conditions'), and mentions the outcome ('suggest fixes'). This distinguishes its core purpose from generic test runners like rush_test or linters like rush_lint. However, it does not explicitly differentiate itself from closely related siblings like rush_flaky or rush_fix, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when flaky test race conditions are encountered, but offers no explicit guidance on when to prefer this tool over alternatives such as rush_flaky or rush_fix. There is no mention of when not to use it, prerequisites, or context that would help an agent decide between the many sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.