Run evaluation on sessions
run_organization_evaluationGrades up to 200 existing sessions with an organization evaluation, queues results asynchronously, and skips ineligible or in-flight sessions. Use dry_run to preview cost without queuing.
Instructions
Grades up to 200 existing sessions with this evaluation. Grading is asynchronous: each accepted session gets a result with status: queued; poll it with GET /evaluations/{evaluation_id}/results/{result_id} until it is completed or failed.
Sessions are skipped, not rejected, when they are not completed sessions of an agent the evaluation covers (ineligible) or already have a queued or running result for this evaluation (in_flight, with the existing result_id). The caller is charged one credit per queued session. Set dry_run: true to see the cost and skips without queuing anything.
Explicit confirmation is required for this exact account operation. Runs can spend credits or trigger downstream actions; never resubmit unknown outcomes automatically.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Named private Gumloop account; selects private credentials and user/team identity. | |
| confirm | No | Set true only when the user asked for exactly this action. | |
| dry_run | No | Report cost and skipped sessions without queuing. | |
| payload | No | Complete JSON request body instead of body flags. Preserves current endpoint fields and values. | |
| session_ids | No | Sessions to grade. Duplicates are rejected. | |
| payload_file | No | Regular local JSON body file, at most 5 MB. Cannot be mixed with body flags or payload. | |
| evaluation_id | Yes | ID of the organization evaluation. |