martin_eval
Evaluate AI coding agent runs by grading task completion, verifier health, diff discipline, risk, and reviewability. Identify issues to prevent runaway loops, bad code, and token waste.
Instructions
Grade a Martin run for task completion, verifier health, diff discipline, risk, and reviewability.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | ||
| latest | No | ||
| loopId | No | ||
| runsDir | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||