Skip to main content
Glama
yanlong-iao

Ultimate Prompt Optimizer

by yanlong-iao

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
LLM_MODELNoMain model name for every tool.gpt-4o-mini
UPO_CACHENoDisk cache for identical requests.true
EVAL_MODELNoModel under test in evaluations.same as LLM_MODEL
PO_BACKENDNoOptional backend type for self-hosted prompt-optimizer.auto
JUDGE_MODELNoGrader and pairwise judge model (use a stronger model to grade a cheaper one).EVAL_MODEL
LLM_API_KEYYesAPI key for the main model. Required for cloud providers; optional for localhost endpoints (e.g., Ollama).
PO_BASE_URLNoOptional self-hosted prompt-optimizer base URL.empty
EVAL_API_KEYNoAPI key for the model under test in evaluations.same as LLM_API_KEY
LLM_BASE_URLNoMain model base URL for every tool.https://api.openai.com/v1
UPO_DATA_DIRNoWhere history/, reports/, .cache/ live.project folder
EVAL_BASE_URLNoBase URL for the model under test in evaluations.same as LLM_BASE_URL
UPO_MAX_CALLSNoHard cap on API calls per tool run.80
UPO_CONCURRENCYNoParallel requests during evaluation.4
UPO_TEMPLATES_DIRNoCustom templates directory.custom-templates
PO_ACCESS_PASSWORDNoOptional access password for self-hosted prompt-optimizer.empty
LLM_PRICE_INPUT_PER_MNoUSD per 1M input tokens, only for cost estimates.
LLM_PRICE_OUTPUT_PER_MNoUSD per 1M output tokens, only for cost estimates.

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": true
}

Tools

Functions exposed to the LLM to take actions

NameDescription
suggest_promptsA

START HERE for any new topic. Given a topic/goal in plain words, returns 2-4 genuinely different ready-to-use prompts (concise / LangGPT structured / interactive coach / strict output format), each with when-to-use, expected effect and a LIVE sample reply produced by the real model on the same first message, so the user can pick by seeing the effect. Present the returned menu to the user and let them choose by number; then offer optimize_prompt_via_api / auto_optimize_prompt / evaluate_prompt_preview on the chosen one. Cost ≈ 1 design call + 1 LangGPT call + 1 preview call per candidate (≈ 6 calls for 4 candidates); previews can be turned off.

generate_langgpt_structureA

Expand a brief idea into a strict LangGPT Markdown system prompt. Default style 'v2' follows the current LangGPT spec (Role, Profile, Background, Goal with Outcome/Done Criteria/Non-Goals, Skill-N subsections, Rules, Workflow, OutputFormat, optional Commands/Reminder/Examples, Initialization); 'classic' uses Goals/Constraints/Skills lists. With an LLM configured the content is domain-specific (the model fills a validated JSON spec; Markdown is rendered deterministically, so structure is guaranteed). Checks that every points to an existing section. 1 API call (0 with strategy=template).

optimize_prompt_via_apiA

Iteratively rewrite a prompt with the built-in template library (see list_templates): round 1 optimizes with a strategy template (default: langgpt-strict for system prompts, user-basic for user prompts); rounds 2..N apply your requirements with the iterate template, which sees both the original draft and the latest version. Every round is checked for LangGPT structure, dangling and lost {{placeholders}}, which are repaired automatically. Cost: 1 API call per round (+1 repair if needed). Uses a prompt-optimizer Docker deployment instead only if PO_BASE_URL is configured and PO_BACKEND allows it. For score-driven optimization with test cases, use auto_optimize_prompt.

analyze_promptA

Design review of a prompt WITHOUT running it: 0-100 scores on goal clarity, instruction completeness, structural executability, ambiguity control and robustness, plus strengths, issues, improvements and an exact patch plan (oldText → newText). Optionally apply the patches locally (no extra call) and/or rewrite the prompt from the analysis (+1 call). Also reports LangGPT structure and static lint. Cost: 1 API call (cached if unchanged).

evaluate_prompt_previewA

promptfoo-style evaluation. Runs prompt variants (A, B, C…) × models × test cases × repeats against the real model, applies assertions and model graders, and reports pass rate, weighted score, per-metric averages, latency, tokens, cost, repeat consistency, an optional pairwise judge, and a winner. Writes an HTML + JSON report. Assertions (prefix not- to negate; weight/metric/transform/threshold supported): equals, contains, icontains, contains-any/all, icontains-any/all, starts-with, regex, is-json/contains-json (+JSON schema), javascript, levenshtein, latency, cost, min/max-length, is-refusal, llm-rubric, g-eval, factuality, answer-relevance, prompt-compliance, assert-set. Tests can come from testCases, a YAML/JSON/CSV testsFile, or a promptfoo-style configPath (prompts/providers/tests/defaultTest; array vars expand as a matrix); if none are given, cases are auto-generated. Responses are cached on disk so re-runs are free. The call count is estimated first and the run is refused if it exceeds maxCalls. mock=true → static lint + request preview, 0 calls.

auto_optimize_promptA

DSPy-style, metric-driven optimization: builds a test set (given or auto-generated with per-case criteria), splits it into train/dev, scores the original prompt, then for several rounds proposes candidates — reflect (analyse failing outputs → targeted edits, GEPA-like), template (rewrite with a different optimization strategy, MIPRO-like), fewshot (insert the best passing outputs as Examples, BootstrapFewShot-like) — evaluates them on train, confirms finalists on dev, and keeps a candidate only if the dev score improves. Returns the best prompt, a scoreboard of every candidate, per-case before/after, cost, and a saved report. Budget-guarded: the search is scaled down to fit maxCalls and stops early on no improvement or time limit. Typical light run with 6 cases: ~40-60 API calls.

list_templatesA

List built-in and custom optimization templates (ids usable as template / iterateTemplate). Custom templates are Markdown files with frontmatter in the custom-templates folder and override built-ins with the same id. 0 API calls.

prompt_historyA

Browse saved results of generate/optimize/analyze/auto-optimize runs (every version is stored locally). action=list shows recent runs; action=get returns the full record with all versions. 0 API calls.

optimizer_connection_statusA

Diagnose the setup: which auth scheme the prompt-optimizer deployment uses and whether login succeeds, whether its /mcp backend is reachable, which upstream templates exist, and whether the local/eval LLMs are configured. Call this first when optimize_prompt_via_api fails.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A4/5.0

Scored across 9 tools

Disambiguation4/5

Most tools have distinct purposes and descriptions give explicit 'when to use' cues (e.g. suggest_prompts is 'START HERE', auto_optimize for metric-driven). However, generate_langgpt_structure vs optimize_prompt_via_api vs suggest_prompts all produce/rewrite prompts, and optimize_prompt_via_api vs auto_optimize_prompt both 'optimize', so an agent could occasionally misselect among these clusters.

Naming Consistency4/5

Almost all tools use snake_case with verb_noun or verb_phrase patterns (suggest_prompts, optimize_prompt_via_api, evaluate_prompt_preview, auto_optimize_prompt). Minor deviations: prompt_history and optimizer_connection_status are noun-only, and prefixes alternate between 'prompt_' and 'optimizer_', but the scheme is broadly readable and consistent.

Tool Count5/5

Nine tools is well within the ideal 3-15 range and each covers a distinct facet of the prompt-optimization workflow (generation, optimization, analysis, evaluation, templates, history, diagnostics). The set feels deliberately scoped with no obviously redundant tools.

Completeness4/5

Covers the full prompt lifecycle: generate suggestions, build LangGPT structure, iterative optimize, metric-driven auto-optimize, static analysis, live evaluation, template listing, and history browsing plus connection diagnostics. Minor gaps: no explicit save/delete/export of named prompts or history cleanup, but core workflows are fully supported.

Maintenance

ActivityMaintained
ResponsivenessNo issues