Run evaluations
run_evaluationsGrades up to 200 finished agent sessions using the agent's evaluation setup; results are queued for asynchronous polling until completed or failed.
Instructions
Grades up to 200 of the agent's finished sessions with its own evaluation configuration. Grading is asynchronous: each accepted session gets a result with status: queued; poll it with GET /agents/{agent_id}/evaluations/{evaluation_id} until it is completed or failed. A new result replaces the previous result for that session.
Sessions are skipped, not rejected, when they are unfinished, incognito, or not owned by this agent (ineligible), or already have a queued or running result (in_flight, with the existing result_id). The caller is charged one credit per queued session. Set dry_run: true to see the cost and skips without queuing anything.
Requires edit access on the agent and a plan with evaluations enabled. Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| dry_run | No | Report cost and skipped sessions without queuing. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| agent_id | Yes | ID of the agent that owns the sessions. | |
| session_ids | No | Sessions to grade. Duplicates are rejected. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. |