Auto-optimize prompt (score-driven)
auto_optimize_promptOptimizes prompts with metric-driven test cases: scores the original, tests rewrites across train/dev, and returns the improved prompt plus a candidate scoreboard.
Instructions
DSPy-style, metric-driven optimization: builds a test set (given or auto-generated with per-case criteria), splits it into train/dev, scores the original prompt, then for several rounds proposes candidates — reflect (analyse failing outputs → targeted edits, GEPA-like), template (rewrite with a different optimization strategy, MIPRO-like), fewshot (insert the best passing outputs as Examples, BootstrapFewShot-like) — evaluates them on train, confirms finalists on dev, and keeps a candidate only if the dev score improves. Returns the best prompt, a scoreboard of every candidate, per-case before/after, cost, and a saved report. Budget-guarded: the search is scaled down to fit maxCalls and stops early on no improvement or time limit. Typical light run with 6 cases: ~40-60 API calls.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| goal | No | What a good response looks like — used for grading and reflection, e.g. 'short interview-style answers, always asks for missing info' | |
| seed | No | ||
| model | No | Model the prompt will run on (default EVAL_MODEL) | |
| assert | No | Assertions applied to every case (e.g. is-json, max-length) | |
| budget | No | light: 2 rounds × 2 candidates · medium: 3×3 · heavy: 5×4 | light |
| prompt | Yes | The prompt to optimize | |
| rounds | No | ||
| rubric | No | llm-rubric applied to every case | |
| maxCalls | No | Hard cap on API calls (default UPO_MAX_CALLS) | |
| numCases | No | Cases to generate when none are given | |
| testCases | No | ||
| testsFile | No | YAML/JSON/CSV tests file | |
| maxMinutes | No | ||
| promptMode | No | system | |
| strategies | No | ||
| temperature | No | ||
| candidatesPerRound | No |