try_evaluator
Test a draft Evaluation Function against a Session to judge replayed runs, returning metrics and errors without persisting data.
Instructions
Tries Evaluator source on one Session and returns what it found. Persists nothing.
Use it to check a draft before showing it as ready, and to judge a replayed Session against
its source. The source is an Evaluation Function as the evaluation skill describes; to try a
saved Evaluator, read a version's python_code with get_evaluator and pass it. A try that
judges with a model costs a model call per judged turn on the customer's provider.
outcome ERRORED means the code broke: read error_detail and fix it. metrics are the
values per address with their rationales. trace_summary says which Session Tools the
evaluator called. If status is RUNNING the wait ran out: poll get_evaluation_try.
:param pipeline_name: Name of the pipeline the Session belongs to.
:param python_code: The Evaluation Function's source.
:param session_id: The Session to judge: a search_session_id, or a single-turn run's query_id.
:param wait_seconds: How long to wait for the result before handing back the id to poll.
:returns: The try report, or an error message.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | ||
| python_code | Yes | ||
| wait_seconds | No | ||
| pipeline_name | Yes |