run_eval
Tests an agent by simulating scripted scenarios, scoring close speed, confirmation accuracy, and urgency handling, then returns a /100 PASS/WARN/FAIL report.
Instructions
The agent tests itself. Simulates N scripted scenarios via the LLM (no audio cost), scores behavior (close speed, no price, spelled confirmation, urgency handling), returns a /100 report with PASS/WARN/FAIL.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| market | Yes | ||
| vertical | Yes | ||
| assistantId | Yes |