Skip to main content
Glama
yanlong-iao

Ultimate Prompt Optimizer

by yanlong-iao

Evaluate / A-B test prompts

evaluate_prompt_preview

Compare and validate several prompt variants across selected models against shared test cases, applying assertions and graders while reporting pass rates, costs, and winners.

Instructions

promptfoo-style evaluation. Runs prompt variants (A, B, C…) × models × test cases × repeats against the real model, applies assertions and model graders, and reports pass rate, weighted score, per-metric averages, latency, tokens, cost, repeat consistency, an optional pairwise judge, and a winner. Writes an HTML + JSON report. Assertions (prefix not- to negate; weight/metric/transform/threshold supported): equals, contains, icontains, contains-any/all, icontains-any/all, starts-with, regex, is-json/contains-json (+JSON schema), javascript, levenshtein, latency, cost, min/max-length, is-refusal, llm-rubric, g-eval, factuality, answer-relevance, prompt-compliance, assert-set. Tests can come from testCases, a YAML/JSON/CSV testsFile, or a promptfoo-style configPath (prompts/providers/tests/defaultTest; array vars expand as a matrix); if none are given, cases are auto-generated. Responses are cached on disk so re-runs are free. The call count is estimated first and the run is refused if it exceeds maxCalls. mock=true → static lint + request preview, 0 calls.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
mockNo
labelsNoNames for the variants, in order
modelsNoModel ids to compare on the same endpoint, e.g. ['deepseek-flash','deepseek-v4-pro']
promptNoVariant A (e.g. the original). Optional when configPath provides prompts
repeatNoRun each case N times to measure consistency (default 1)
rubricNollm-rubric applied to every case
promptBNoVariant B (e.g. the optimized prompt)
maxCallsNoRefuse to run if the estimated API calls exceed this (default UPO_MAX_CALLS)
pairwiseNoWith exactly 2 variants: LLM judge picks the better output per case (position-bias controlled)
useCacheNo
testCasesNo
testsFileNoPath to tests (.yaml/.json/.csv with __expected columns), relative to the project folder or absolute
configPathNoPath to a promptfoo-style config (.yaml/.json)
promptModeNosystem: prompt = system message, case input = user message. user: prompt is the user message with {{vars}}. Default system (user when configPath is used)
saveReportNo
temperatureNo
extraPromptsNoVariants C, D…
defaultAssertNoAssertions applied to every case
maxOutputCharsNo
autoGenerateCasesNoGenerated when no tests are given (each with its own grading criteria)

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.3.0

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses disk response caching (free re-runs), call-count estimation with a refusal guard, mock mode performing static lint plus request preview at 0 calls, report file output, and a position-bias-controlled pairwise judge. It omits permission/auth requirements and what happens to existing reports on overwrite, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and behavior, but the second half is a dense run-on that re-lists the full assertion enum already present in the schema, duplicating structured data and inflating length. The report-metric sentence is also a long comma chain.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 20-parameter tool with no annotations and no output schema, the description covers the key operational facts an agent needs: inputs/fallbacks for test cases, caching, call budgeting, mock preview, and the reported outputs. Minor gaps remain around report paths and endpoint/permission assumptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 70% across 20 parameters, and the description adds real meaning beyond it: the assertion type catalogue with 'not-' negation and weight/metric/transform/threshold semantics, configPath expansion behavior, promptMode defaults, and autoGenerateCases behavior. This supplements rather than merely restates the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource set ('runs prompt variants × models × test cases × repeats against the real model, applies assertions and model graders') and enumerates the exact artifacts it produces. It is clearly a batch evaluation/benchmarking tool, readily separable from siblings like analyze_prompt or optimize_prompt_via_api.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete operating conditions: tests come from testCases, testsFile, or configPath, and auto-generate when none given; mock=true yields a 0-call preview; the run is refused if estimates exceed maxCalls. It does not explicitly contrast when to choose this over sibling tools such as analyze_prompt, leaving that inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.