rush_prompt_eval
Evaluate recorded prompt runs against golden task criteria to identify failures and generate a pass/fail summary with findings.
Instructions
Evaluate recorded prompt runs against golden task criteria at . Returns {status, findings[], summary}.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| options | Yes | ||
| records | No | ||
| allow_slow | No | ||
| allow_build | No | ||
| golden_path | No | ||
| allow_browser | No | ||
| allow_network | No | ||
| allow_download | No | ||
| allow_cache_write | No | ||
| max_cost_threshold | No | ||
| pass_rate_threshold | No | ||
| allow_artifact_write | No | ||
| max_tokens_threshold | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| raw | No | ||
| tool | No | ||
| engine | No | ||
| status | No | ||
| metrics | No | ||
| summary | No | ||
| findings | No | ||
| metadata | No | ||
| artifacts | No | ||
| duration_ms | No | ||
| review_kind | No | ||
| engine_version | No | ||
| review_provider | No |