create_experiment
Set up a new experiment configuration to benchmark AI operator performance against chosen metrics and benchmarks. Requires authorized access.
Instructions
Create an experiment configuration. REQUIRES AUTHORIZATION.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Experiment name (e.g., 'Q3 Claude vs ChatGPT operator comparison') | |
| configuration | No | Pilot configuration object (JSON) — see list_pilot_options for available metrics, eval families, and benchmark classes |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| tool | Yes | ||
| error | Yes | ||
| message | Yes |