Evaluate / A-B test prompts
evaluate_prompt_previewCompare and validate several prompt variants across selected models against shared test cases, applying assertions and graders while reporting pass rates, costs, and winners.
Instructions
promptfoo-style evaluation. Runs prompt variants (A, B, C…) × models × test cases × repeats against the real model, applies assertions and model graders, and reports pass rate, weighted score, per-metric averages, latency, tokens, cost, repeat consistency, an optional pairwise judge, and a winner. Writes an HTML + JSON report. Assertions (prefix not- to negate; weight/metric/transform/threshold supported): equals, contains, icontains, contains-any/all, icontains-any/all, starts-with, regex, is-json/contains-json (+JSON schema), javascript, levenshtein, latency, cost, min/max-length, is-refusal, llm-rubric, g-eval, factuality, answer-relevance, prompt-compliance, assert-set. Tests can come from testCases, a YAML/JSON/CSV testsFile, or a promptfoo-style configPath (prompts/providers/tests/defaultTest; array vars expand as a matrix); if none are given, cases are auto-generated. Responses are cached on disk so re-runs are free. The call count is estimated first and the run is refused if it exceeds maxCalls. mock=true → static lint + request preview, 0 calls.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| mock | No | ||
| labels | No | Names for the variants, in order | |
| models | No | Model ids to compare on the same endpoint, e.g. ['deepseek-flash','deepseek-v4-pro'] | |
| prompt | No | Variant A (e.g. the original). Optional when configPath provides prompts | |
| repeat | No | Run each case N times to measure consistency (default 1) | |
| rubric | No | llm-rubric applied to every case | |
| promptB | No | Variant B (e.g. the optimized prompt) | |
| maxCalls | No | Refuse to run if the estimated API calls exceed this (default UPO_MAX_CALLS) | |
| pairwise | No | With exactly 2 variants: LLM judge picks the better output per case (position-bias controlled) | |
| useCache | No | ||
| testCases | No | ||
| testsFile | No | Path to tests (.yaml/.json/.csv with __expected columns), relative to the project folder or absolute | |
| configPath | No | Path to a promptfoo-style config (.yaml/.json) | |
| promptMode | No | system: prompt = system message, case input = user message. user: prompt is the user message with {{vars}}. Default system (user when configPath is used) | |
| saveReport | No | ||
| temperature | No | ||
| extraPrompts | No | Variants C, D… | |
| defaultAssert | No | Assertions applied to every case | |
| maxOutputChars | No | ||
| autoGenerateCases | No | Generated when no tests are given (each with its own grading criteria) |