Ultimate Prompt Optimizer
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| LLM_MODEL | No | Main model name for every tool. | gpt-4o-mini |
| UPO_CACHE | No | Disk cache for identical requests. | true |
| EVAL_MODEL | No | Model under test in evaluations. | same as LLM_MODEL |
| PO_BACKEND | No | Optional backend type for self-hosted prompt-optimizer. | auto |
| JUDGE_MODEL | No | Grader and pairwise judge model (use a stronger model to grade a cheaper one). | EVAL_MODEL |
| LLM_API_KEY | Yes | API key for the main model. Required for cloud providers; optional for localhost endpoints (e.g., Ollama). | |
| PO_BASE_URL | No | Optional self-hosted prompt-optimizer base URL. | empty |
| EVAL_API_KEY | No | API key for the model under test in evaluations. | same as LLM_API_KEY |
| LLM_BASE_URL | No | Main model base URL for every tool. | https://api.openai.com/v1 |
| UPO_DATA_DIR | No | Where history/, reports/, .cache/ live. | project folder |
| EVAL_BASE_URL | No | Base URL for the model under test in evaluations. | same as LLM_BASE_URL |
| UPO_MAX_CALLS | No | Hard cap on API calls per tool run. | 80 |
| UPO_CONCURRENCY | No | Parallel requests during evaluation. | 4 |
| UPO_TEMPLATES_DIR | No | Custom templates directory. | custom-templates |
| PO_ACCESS_PASSWORD | No | Optional access password for self-hosted prompt-optimizer. | empty |
| LLM_PRICE_INPUT_PER_M | No | USD per 1M input tokens, only for cost estimates. | |
| LLM_PRICE_OUTPUT_PER_M | No | USD per 1M output tokens, only for cost estimates. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| suggest_promptsA | START HERE for any new topic. Given a topic/goal in plain words, returns 2-4 genuinely different ready-to-use prompts (concise / LangGPT structured / interactive coach / strict output format), each with when-to-use, expected effect and a LIVE sample reply produced by the real model on the same first message, so the user can pick by seeing the effect. Present the returned menu to the user and let them choose by number; then offer optimize_prompt_via_api / auto_optimize_prompt / evaluate_prompt_preview on the chosen one. Cost ≈ 1 design call + 1 LangGPT call + 1 preview call per candidate (≈ 6 calls for 4 candidates); previews can be turned off. |
| generate_langgpt_structureA | Expand a brief idea into a strict LangGPT Markdown system prompt. Default style 'v2' follows the current LangGPT spec (Role, Profile, Background, Goal with Outcome/Done Criteria/Non-Goals, Skill-N subsections, Rules, Workflow, OutputFormat, optional Commands/Reminder/Examples, Initialization); 'classic' uses Goals/Constraints/Skills lists. With an LLM configured the content is domain-specific (the model fills a validated JSON spec; Markdown is rendered deterministically, so structure is guaranteed). Checks that every points to an existing section. 1 API call (0 with strategy=template). |
| optimize_prompt_via_apiA | Iteratively rewrite a prompt with the built-in template library (see list_templates): round 1 optimizes with a strategy template (default: langgpt-strict for system prompts, user-basic for user prompts); rounds 2..N apply your requirements with the iterate template, which sees both the original draft and the latest version. Every round is checked for LangGPT structure, dangling and lost {{placeholders}}, which are repaired automatically. Cost: 1 API call per round (+1 repair if needed). Uses a prompt-optimizer Docker deployment instead only if PO_BASE_URL is configured and PO_BACKEND allows it. For score-driven optimization with test cases, use auto_optimize_prompt. |
| analyze_promptA | Design review of a prompt WITHOUT running it: 0-100 scores on goal clarity, instruction completeness, structural executability, ambiguity control and robustness, plus strengths, issues, improvements and an exact patch plan (oldText → newText). Optionally apply the patches locally (no extra call) and/or rewrite the prompt from the analysis (+1 call). Also reports LangGPT structure and static lint. Cost: 1 API call (cached if unchanged). |
| evaluate_prompt_previewA | promptfoo-style evaluation. Runs prompt variants (A, B, C…) × models × test cases × repeats against the real model, applies assertions and model graders, and reports pass rate, weighted score, per-metric averages, latency, tokens, cost, repeat consistency, an optional pairwise judge, and a winner. Writes an HTML + JSON report. Assertions (prefix not- to negate; weight/metric/transform/threshold supported): equals, contains, icontains, contains-any/all, icontains-any/all, starts-with, regex, is-json/contains-json (+JSON schema), javascript, levenshtein, latency, cost, min/max-length, is-refusal, llm-rubric, g-eval, factuality, answer-relevance, prompt-compliance, assert-set. Tests can come from testCases, a YAML/JSON/CSV testsFile, or a promptfoo-style configPath (prompts/providers/tests/defaultTest; array vars expand as a matrix); if none are given, cases are auto-generated. Responses are cached on disk so re-runs are free. The call count is estimated first and the run is refused if it exceeds maxCalls. mock=true → static lint + request preview, 0 calls. |
| auto_optimize_promptA | DSPy-style, metric-driven optimization: builds a test set (given or auto-generated with per-case criteria), splits it into train/dev, scores the original prompt, then for several rounds proposes candidates — reflect (analyse failing outputs → targeted edits, GEPA-like), template (rewrite with a different optimization strategy, MIPRO-like), fewshot (insert the best passing outputs as Examples, BootstrapFewShot-like) — evaluates them on train, confirms finalists on dev, and keeps a candidate only if the dev score improves. Returns the best prompt, a scoreboard of every candidate, per-case before/after, cost, and a saved report. Budget-guarded: the search is scaled down to fit maxCalls and stops early on no improvement or time limit. Typical light run with 6 cases: ~40-60 API calls. |
| list_templatesA | List built-in and custom optimization templates (ids usable as |
| prompt_historyA | Browse saved results of generate/optimize/analyze/auto-optimize runs (every version is stored locally). action=list shows recent runs; action=get returns the full record with all versions. 0 API calls. |
| optimizer_connection_statusA | Diagnose the setup: which auth scheme the prompt-optimizer deployment uses and whether login succeeds, whether its /mcp backend is reachable, which upstream templates exist, and whether the local/eval LLMs are configured. Call this first when optimize_prompt_via_api fails. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 9 tools
Most tools have distinct purposes and descriptions give explicit 'when to use' cues (e.g. suggest_prompts is 'START HERE', auto_optimize for metric-driven). However, generate_langgpt_structure vs optimize_prompt_via_api vs suggest_prompts all produce/rewrite prompts, and optimize_prompt_via_api vs auto_optimize_prompt both 'optimize', so an agent could occasionally misselect among these clusters.
Almost all tools use snake_case with verb_noun or verb_phrase patterns (suggest_prompts, optimize_prompt_via_api, evaluate_prompt_preview, auto_optimize_prompt). Minor deviations: prompt_history and optimizer_connection_status are noun-only, and prefixes alternate between 'prompt_' and 'optimizer_', but the scheme is broadly readable and consistent.
Nine tools is well within the ideal 3-15 range and each covers a distinct facet of the prompt-optimization workflow (generation, optimization, analysis, evaluation, templates, history, diagnostics). The set feels deliberately scoped with no obviously redundant tools.
Covers the full prompt lifecycle: generate suggestions, build LangGPT structure, iterative optimize, metric-driven auto-optimize, static analysis, live evaluation, template listing, and history browsing plus connection diagnostics. Minor gaps: no explicit save/delete/export of named prompts or history cleanup, but core workflows are fully supported.