run_comparison
Runs test prompts against up to four models, scores each response with LLM judges and rule checks, then returns grades, cost, latency, and a verdict.
Instructions
Run a set of test-case prompts against a set of models, scoring each response.
Each test case may include an optional "rubric" (scored by an LLM judge) and/or
"checks" (rule-based checks). Returns per-model grades, cost/latency stats, and an
overall verdict. Rate-limited to 3 calls per 8 hours.
Models are "<catalog id>" (OpenRouter) or "<catalog id>@bedrock" / "<catalog id>@vertex" /
"<catalog id>@foundry"; pass matching creds ({"openrouter"?, "bedrock"?, "vertex"?, "foundry"?})
or a bare OpenRouter api_key.
judge_backend picks which backend runs the judge and policy gate.
At most 4 models. priority (balanced|quality|fastest|cheapest) ranks the results;
repeats (1-3) re-sends each prompt for timing accuracy. Returns suggestions (same
provider and backend only), advice, ranking, and best_for_priority, plus a per-run
`cost` total and, per cell, an `evaluation` (answered/quality/instruction_following/
completeness/helpfulness/safety scores, strengths, weaknesses, reasoning, overall).
If the operator has set server-side keys for a backend (see https://github.com/thejaredchapman/evalforge-lite/blob/main/docs/hosting-and-server-keys.md), those are used for it automatically and creds for it are not needed.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| creds | No | ||
| models | Yes | ||
| api_key | No | ||
| repeats | No | ||
| priority | No | balanced | |
| test_cases | Yes | ||
| judge_backend | No | openrouter |