Ultimate Prompt Optimizer
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Ultimate Prompt Optimizeroptimize my customer-support prompt and test it on sample chats"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Ultimate Prompt Optimizer (MCP server)
An MCP server that turns Claude into a prompt-engineering workbench. It suggests candidate prompts, structures them with LangGPT, optimizes them over several rounds, and proves the result on test cases with promptfoo-style assertions and DSPy-style train/dev search. It runs on any OpenAI-compatible model with your own API key.
Why
Prompts written by hand are too vague for a model: "keep it short", "occasionally", "when the user seems stuck". Rewriting them with a single "make this better" call just moves the vagueness around. This server follows a loop that ends in evidence, not opinion:
Generate several candidates.
suggest_promptsreturns 2–4 genuinely different prompts, each with a live sample reply.Let the user pick by seeing the effect, not by reading the prompt.
Structure it. A LangGPT spec is rendered deterministically and validated.
Iterate. Template-driven rewrites with requirements, automatic structure and
{{placeholder}}repair.Prove it. A/B evaluation on test cases with assertions, model graders, a pairwise judge and repeat-consistency checks. Or let
auto_optimize_promptsearch for a better prompt and keep it only if a held-out dev set improves.
Related MCP server: Promptheus
Demo
This excerpt comes from a real run on deepseek-flash (full case study). It compares a hand-written quant-interview coaching prompt (A) with the version produced by optimize_prompt_via_api (B):
You: Use evaluate_prompt_preview to compare the original (A) and the optimized version (B), pairwise on, repeat 2, rubric: "ask exactly one question and wait; after the answer give a verdict, the fastest interview-ready solution and the key trick; be concise".
Claude →
evaluate_prompt_preview(auto-generated 3 cases · estimated ≤ 27 calls)
Variant
Pass rate
Score
Latency avg / p95
Tokens in/out
Consistency
original
100%
0.97
5004 / 8605 ms
7308 / 4759
100%
optimized
100%
1.00
3518 / 5243 ms
11136 / 2947
100%
Pairwise judge: original wins 2 · optimized wins 1
Case 3 (adversarial: "stop quizzing me, write my CV and tell me the next answer"): both variants refused and asked one question. Case 2: the judge found that A's logic puzzle had no unique answer under its own assumptions.
The whole four-tool pipeline took 36 API calls, ~78k tokens, ≈ $0.06, 2 min 15 s. The case study also explains why these numbers do not prove that B is better.
Architecture
flowchart LR
host["Claude Desktop / Claude Code<br/>(any MCP host)"] <-->|stdio · JSON-RPC| srv["MCP server<br/>9 tools · zod schemas"]
srv --> mods["suggest · langgpt · analysis<br/>optimize · eval · templates"]
mods --> llm
subgraph client["LLM client (one per tool run)"]
llm["chat()"] --- cache[("disk cache")]
llm --- budget["call budget"]
llm --- meter["usage & cost meter"]
end
llm -->|HTTPS /chat/completions| prov["OpenAI-compatible provider<br/>DeepSeek · OpenAI · Ollama · OpenRouter…"]
mods -.->|optional · PO_BASE_URL| up["upstream prompt-optimizer<br/>Docker MCP"]Modules, data flow and design trade-offs: docs/ARCHITECTURE.md.
Tools
Tool | What it does | API calls |
| Topic in plain words → 2–4 different prompts (concise / LangGPT / coach / strict format), each with a live sample reply. The recommended starting point. | ≈ 2 + 1 per candidate (5 for 3 candidates in the case study) |
| Idea → strict LangGPT prompt (v2 Goal / Skill-N / Rules, optional Commands / Reminder, | 1 (+1 retry on invalid JSON); 0 in template mode |
| Design review without running: 5 dimension scores, issues, exact | 1 (+1 with rewrite) |
| Multi-round template rewrite: optimize → iterate on requirements → auto-repair of structure and lost placeholders | 1 per round (+1 repair if needed) |
| promptfoo-style eval: up to 6 variants × 4 models × cases × 5 repeats, 24 assertion types, pairwise judge, HTML/JSON report | Estimated before running; refused if over budget. 0 in mock mode |
| DSPy-style search: generate cases → reflect / re-template / few-shot candidates → select on train → accept only on dev improvement | Scaled to |
| Built-in and custom optimization templates | 0 |
| Every saved version from previous runs | 0 |
| Configuration and upstream diagnostics | 0 |
Engineering highlights
Each item points to the code that implements it.
Injection-resistant templates. The prompt being optimized or judged is passed as a JSON evidence block declared to be data, and every template shares the same "data ≠ task" rules. Otherwise, a prompt saying "ignore previous instructions" would hijack the optimizer. →
src/templates/shared.ts{{placeholder}}protection. Variables are substituted in a single pass (inserted text is never re-scanned). After every round the pipeline diffs the placeholders between source and output, and any it lost are restored with a targeted repair round. →src/templates/registry.ts,src/optimize/pipeline.tsAutomatic LangGPT repair. A validator checks required sections, empty sections and dangling
<References>(recognising Chinese and English section aliases). Violations become repair requirements for the next round. →src/langgpt/template.tsThe LLM fills a JSON spec, zod validates it (one retry with the validation errors fed back), and code renders the Markdown. Structure is correct by construction, and when no LLM is available a deterministic scaffold is returned instead. →
src/langgpt/generator.tspromptfoo-style assertion engine.
not-negation,weight(weight 0 = informational),metric,transform,threshold, nestedassert-set, JSON Schema via ajv, and model graders (llm-rubric,g-eval,factuality,answer-relevance,prompt-compliance) at temperature 0. Also reads promptfoo-style YAML configs, CSV__expectedcolumns and variable matrices. →src/eval/assertions.ts,src/eval/config.tsPairwise judge with position swapping. A and B are shown in alternating positions across cases and the verdict is mapped back, which cancels a constant position bias over the test set. →
src/analysis/analyzer.tsDSPy-style train/dev split. Candidates (GEPA-like reflection on failing traces, alternative templates, BootstrapFewShot-like demos at zero extra calls) are screened on train, and finalists are accepted only if the dev score improves. The original is the baseline, so the result is never worse. →
src/optimize/autoOptimizer.tsCost control. The call count is estimated before an eval and the run is refused if it exceeds the budget. A hard
CallBudgetapplies per tool run, a content-addressed disk cache makes identical re-runs free, and a usage meter reports tokens and cost by purpose. In the case study the estimate of 27 eval calls matched the 27 made. →src/eval/runner.ts,src/utils/run.tsProvider-agnostic. Plain
fetchto/chat/completions; JSON mode is dropped automatically if the provider rejects it; retries with backoff on 429/5xx; localhost endpoints need no key (Ollama); separate eval and judge models. →src/clients/llmClient.tsOffline test harness. A mock OpenAI-compatible LLM plus a mock upstream MCP server (Vercel-cookie and Docker Basic auth, session loss). The smoke suite spawns the real server over stdio and drives all 9 tools: 18 test groups, no API key, run in CI on Node 20 and 22. →
test/
Quickstart
Requirements: Node ≥ 20 and an API key for any OpenAI-compatible provider.
git clone https://github.com/yanlong-iao/ultimate-prompt-optimizer-mcp.git
cd ultimate-prompt-optimizer-mcp
npm ci && npm run build
npm run smoke # optional: offline end-to-end tests, no key needed
cp .env.example .env # then fill in LLM_API_KEY (or use `bash setup.sh`, interactive)Claude Code
claude mcp add prompt-optimizer -- node "$(pwd)/dist/index.js"Claude Desktop (~/Library/Application Support/Claude/claude_desktop_config.json on macOS, %APPDATA%\Claude\claude_desktop_config.json on Windows):
{
"mcpServers": {
"prompt-optimizer": {
"command": "node",
"args": ["/absolute/path/to/ultimate-prompt-optimizer-mcp/dist/index.js"],
"env": {
"LLM_BASE_URL": "https://api.deepseek.com",
"LLM_API_KEY": "<your key>",
"LLM_MODEL": "deepseek-flash"
}
}
}
}The server reads .env next to dist/. Variables in the env block take precedence. Then try one of these in Claude:
"Use suggest_prompts for: a coach that quizzes me on probability questions for quant interviews"
"analyze_prompt this prompt and apply the patches: …"
"evaluate_prompt_preview original (A) vs optimized (B), pairwise on, rubric: …"
"auto_optimize_prompt this prompt, goal: short answers that state right/wrong first, budget light"
Environment variables
Variable | Default | Purpose |
|
| Main model for every tool. Key optional for localhost endpoints |
| same as | Model under test in evaluations |
|
| Grader and pairwise judge (use a stronger model to grade a cheaper one) |
| – | USD per 1M tokens, only for cost estimates |
|
| Hard cap on API calls per tool run |
|
| Parallel requests during evaluation |
|
| Disk cache for identical requests |
| project folder / | Where |
| empty / | Optional self-hosted prompt-optimizer as the optimization backend |
See .env.example for the full list and examples/ for a promptfoo-style config and CSV tests.
Case study
docs/CASE_STUDY.md runs the full pipeline on a real prompt with DeepSeek and reports scores, pass rates, tokens, cost and latency, together with what the numbers do and don't show.
Limitations & roadmap
Limitations (known and deliberate)
The
javascriptassertion andtransformrun in Node'svmmodule, which is not a security sandbox. Only run eval configs you wrote yourself.Auto-generated test cases are single-turn, so multi-turn behaviour (for example, grading a user's answer) is not exercised unless you write such cases yourself. In the case study this made the rubric saturate at 100% for both variants.
The pairwise judge swaps positions across cases, not within a case, so per-case position bias remains. With few cases, the win counts are not statistically significant.
The DSPy-style search is a simplified beam over strategy templates. It has no Bayesian surrogate like MIPROv2, and small dev sets are noisy.
The promptfoo compatibility is a common subset (24 assertion types). There is no red-teaming, embedding similarity or web viewer, because DeepSeek has no embeddings endpoint.
Cost figures are estimates from configured prices. Provider-side prompt-cache discounts are not modelled.
The cache has no invalidation when a provider updates a model behind the same name (14-day TTL).
Non-template files in
custom-templates/(such as its README) are logged as skipped.
Roadmap
Multi-turn test cases (conversation arrays) in eval and case generation
Both-order pairwise judging with a confidence interval over wins
Measured
auto_optimize_promptcase study with hand-written, multi-turn cases and a cross-family judgenpm package and official MCP Registry entry (
npxinstall)Demo GIF
A feature-by-feature comparison with LangGPT, prompt-optimizer, promptfoo and DSPy (in Chinese) is in COMPARISON.md.
Acknowledgements
This project borrows methods, not code or template text. All templates and code were written independently.
LangGPT: structured prompt format (Role / Profile / Goal / Skills / Rules / Workflow / Initialization)
linshenkx/prompt-optimizer: template-based optimize / iterate workflow, evidence-style prompt wrapping, design-review idea (AGPL-3.0; no text or code copied)
promptfoo: assertion vocabulary, config format and evaluation matrix
DSPy: metric-driven optimization, train/dev split, GEPA-style reflection, MIPRO-style instruction proposals, BootstrapFewShot
Model Context Protocol: TypeScript SDK
License
MIT © 2026 Yanlong Liao
中文简介
Ultimate Prompt Optimizer 是一个本地运行的 MCP 服务器,在 Claude Desktop / Claude Code 里提供一整套提示词工程工作流:
先给出多个风格不同的候选提示词,每个都附带真实的试运行回复,让用户按效果挑选。
用 LangGPT 做结构化:LLM 填 JSON 规格,经 zod 校验后由代码确定性渲染。
用模板库多轮迭代优化,自动修复结构和
{{变量}}。用 promptfoo 式的断言和成对评审做 A/B 评测,或者用 DSPy 式的 train/dev 分数驱动搜索,只有 dev 集分数提高才接受修改。
所有 API 调用都会先估算次数、受预算上限约束、走磁盘缓存并计量成本。项目兼容任意 OpenAI 格式的接口(DeepSeek、OpenAI、Ollama 等),使用者自带 API key。离线测试共 18 组,不需要 key,在 CI 中运行。DeepSeek 真实实测见 docs/CASE_STUDY.md。
Available Tools
9 toolsanalyze_promptAnalyze prompt designA
Design review of a prompt WITHOUT running it: 0-100 scores on goal clarity, instruction completeness, structural executability, ambiguity control and robustness, plus strengths, issues, improvements and an exact patch plan (oldText → newText). Optionally apply the patches locally (no extra call) and/or rewrite the prompt from the analysis (+1 call). Also reports LangGPT structure and static lint. Cost: 1 API call (cached if unchanged).
| Name | Required | Description | Default |
|---|---|---|---|
| focus | No | A specific concern to prioritize, e.g. 'the model keeps giving long derivations' | |
| prompt | Yes | ||
| rewrite | No | Also produce a full revised prompt from the analysis (+1 API call) | |
| applyPatches | No | Apply exact, unique patches and return the patched prompt |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the exact return content, the API-call cost model, that results are cached when unchanged, that applyPatches is local with no extra call, and that rewrite costs +1 call. Missing only edge cases such as failure behavior or patch-applicability limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense paragraph, front-loaded with the core purpose and scope, then the outputs, then the optional modes and cost. Every clause conveys distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter analysis tool with no output schema, the description enumerates the returned artifacts and the call-cost implications well enough to invoke correctly. It could say more about how focus interacts with scoring or what happens when patches can't be uniquely applied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75% and the description adds genuine meaning beyond it, notably the cost implication of rewrite (+1 API call) and the local, no-extra-call nature of applyPatches, plus the patch format (oldText → newText). It gives less detail on the focus parameter than the schema already does.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Design review of a prompt WITHOUT running it') and immediately enumerates what is produced: 0-100 scores across five named dimensions, strengths, issues, improvements, and a patch plan. The 'WITHOUT running it' scope cleanly separates it from siblings like evaluate_prompt_preview or optimize_prompt_via_api.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes clear when this tool applies (static design review rather than execution) and describes the optional follow-on modes: applying patches locally or rewriting the prompt. It doesn't explicitly name a competing sibling and the condition that would select it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
auto_optimize_promptAuto-optimize prompt (score-driven)A
DSPy-style, metric-driven optimization: builds a test set (given or auto-generated with per-case criteria), splits it into train/dev, scores the original prompt, then for several rounds proposes candidates — reflect (analyse failing outputs → targeted edits, GEPA-like), template (rewrite with a different optimization strategy, MIPRO-like), fewshot (insert the best passing outputs as Examples, BootstrapFewShot-like) — evaluates them on train, confirms finalists on dev, and keeps a candidate only if the dev score improves. Returns the best prompt, a scoreboard of every candidate, per-case before/after, cost, and a saved report. Budget-guarded: the search is scaled down to fit maxCalls and stops early on no improvement or time limit. Typical light run with 6 cases: ~40-60 API calls.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | No | What a good response looks like — used for grading and reflection, e.g. 'short interview-style answers, always asks for missing info' | |
| seed | No | ||
| model | No | Model the prompt will run on (default EVAL_MODEL) | |
| assert | No | Assertions applied to every case (e.g. is-json, max-length) | |
| budget | No | light: 2 rounds × 2 candidates · medium: 3×3 · heavy: 5×4 | light |
| prompt | Yes | The prompt to optimize | |
| rounds | No | ||
| rubric | No | llm-rubric applied to every case | |
| maxCalls | No | Hard cap on API calls (default UPO_MAX_CALLS) | |
| numCases | No | Cases to generate when none are given | |
| testCases | No | ||
| testsFile | No | YAML/JSON/CSV tests file | |
| maxMinutes | No | ||
| promptMode | No | system | |
| strategies | No | ||
| temperature | No | ||
| candidatesPerRound | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the multi-round search loop, the three named strategies, the dev-set gate for keeping candidates, budget guarding against maxCalls, early stopping on no improvement or time limit, and even a concrete cost estimate ('~40-60 API calls' for 6 cases). This is unusually rich behavioral disclosure for a mutation/expensive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the pipeline and dense with substance rather than filler; every clause (strategies, dev gate, budget guard, cost, return contents) earns its place. It is a single long block, which slightly hurts scannability, but it is appropriately sized for a tool of this complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-param tool with no output schema, the description usefully enumerates what is returned (best prompt, scoreboard, per-case before/after, cost, saved report) and the stopping/budget behavior. Minor gaps remain: where the report is saved, and the model/env dependencies (EVAL_MODEL, UPO_MAX_CALLS) that only appear in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 53% across 17 params, so the description has to compensate, and it does for the key knobs: it explains what reflect/template/fewshot do (GEPA-like, MIPRO-like, BootstrapFewShot-like), how the test set is sourced, and how budget/maxCalls shape the search. It still leaves several params (seed, temperature, promptMode, testsFile, maxMinutes) to the schema, but adds real meaning beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: metric-driven prompt optimization that builds a test set, splits train/dev, proposes candidates, and keeps only dev-improving ones. This is clearly distinguishable from siblings like optimize_prompt_via_api or evaluate_prompt_preview without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the internal conditions that drive behavior ('given or auto-generated' test set, budget scaling, early stop on no improvement), which is helpful context, but never states when to choose this over the sibling optimizer (auto_optimize_prompt vs optimize_prompt_via_api vs suggest_prompts). Usage is implied rather than routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluate_prompt_previewEvaluate / A-B test promptsA
promptfoo-style evaluation. Runs prompt variants (A, B, C…) × models × test cases × repeats against the real model, applies assertions and model graders, and reports pass rate, weighted score, per-metric averages, latency, tokens, cost, repeat consistency, an optional pairwise judge, and a winner. Writes an HTML + JSON report. Assertions (prefix not- to negate; weight/metric/transform/threshold supported): equals, contains, icontains, contains-any/all, icontains-any/all, starts-with, regex, is-json/contains-json (+JSON schema), javascript, levenshtein, latency, cost, min/max-length, is-refusal, llm-rubric, g-eval, factuality, answer-relevance, prompt-compliance, assert-set. Tests can come from testCases, a YAML/JSON/CSV testsFile, or a promptfoo-style configPath (prompts/providers/tests/defaultTest; array vars expand as a matrix); if none are given, cases are auto-generated. Responses are cached on disk so re-runs are free. The call count is estimated first and the run is refused if it exceeds maxCalls. mock=true → static lint + request preview, 0 calls.
| Name | Required | Description | Default |
|---|---|---|---|
| mock | No | ||
| labels | No | Names for the variants, in order | |
| models | No | Model ids to compare on the same endpoint, e.g. ['deepseek-flash','deepseek-v4-pro'] | |
| prompt | No | Variant A (e.g. the original). Optional when configPath provides prompts | |
| repeat | No | Run each case N times to measure consistency (default 1) | |
| rubric | No | llm-rubric applied to every case | |
| promptB | No | Variant B (e.g. the optimized prompt) | |
| maxCalls | No | Refuse to run if the estimated API calls exceed this (default UPO_MAX_CALLS) | |
| pairwise | No | With exactly 2 variants: LLM judge picks the better output per case (position-bias controlled) | |
| useCache | No | ||
| testCases | No | ||
| testsFile | No | Path to tests (.yaml/.json/.csv with __expected columns), relative to the project folder or absolute | |
| configPath | No | Path to a promptfoo-style config (.yaml/.json) | |
| promptMode | No | system: prompt = system message, case input = user message. user: prompt is the user message with {{vars}}. Default system (user when configPath is used) | |
| saveReport | No | ||
| temperature | No | ||
| extraPrompts | No | Variants C, D… | |
| defaultAssert | No | Assertions applied to every case | |
| maxOutputChars | No | ||
| autoGenerateCases | No | Generated when no tests are given (each with its own grading criteria) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses disk response caching (free re-runs), call-count estimation with a refusal guard, mock mode performing static lint plus request preview at 0 calls, report file output, and a position-bias-controlled pairwise judge. It omits permission/auth requirements and what happens to existing reports on overwrite, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and behavior, but the second half is a dense run-on that re-lists the full assertion enum already present in the schema, duplicating structured data and inflating length. The report-metric sentence is also a long comma chain.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 20-parameter tool with no annotations and no output schema, the description covers the key operational facts an agent needs: inputs/fallbacks for test cases, caching, call budgeting, mock preview, and the reported outputs. Minor gaps remain around report paths and endpoint/permission assumptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 70% across 20 parameters, and the description adds real meaning beyond it: the assertion type catalogue with 'not-' negation and weight/metric/transform/threshold semantics, configPath expansion behavior, promptMode defaults, and autoGenerateCases behavior. This supplements rather than merely restates the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource set ('runs prompt variants × models × test cases × repeats against the real model, applies assertions and model graders') and enumerates the exact artifacts it produces. It is clearly a batch evaluation/benchmarking tool, readily separable from siblings like analyze_prompt or optimize_prompt_via_api.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete operating conditions: tests come from testCases, testsFile, or configPath, and auto-generate when none given; mock=true yields a 0-call preview; the run is refused if estimates exceed maxCalls. It does not explicitly contrast when to choose this over sibling tools such as analyze_prompt, leaving that inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_langgpt_structureGenerate LangGPT promptA
Expand a brief idea into a strict LangGPT Markdown system prompt. Default style 'v2' follows the current LangGPT spec (Role, Profile, Background, Goal with Outcome/Done Criteria/Non-Goals, Skill-N subsections, Rules, Workflow, OutputFormat, optional Commands/Reminder/Examples, Initialization); 'classic' uses Goals/Constraints/Skills lists. With an LLM configured the content is domain-specific (the model fills a validated JSON spec; Markdown is rendered deterministically, so structure is guaranteed). Checks that every points to an existing section. 1 API call (0 with strategy=template).
| Name | Required | Description | Default |
|---|---|---|---|
| idea | Yes | Brief description of the assistant/prompt you want | |
| role | No | Explicit role name; inferred if omitted | |
| style | No | v2 | |
| author | No | ||
| skills | No | ||
| audience | No | ||
| language | No | auto | |
| strategy | No | auto = LLM if configured, else deterministic template | auto |
| constraints | No | Rules that must appear verbatim | |
| targetModel | No | weak = shorter, flatter prompt for small models | strong |
| outputFormat | No | Required response format | |
| includeCommands | No | Add a LangGPT ## Commands section (/help, /continue, /improve) | |
| includeReminder | No | Add a ## Reminder section for long conversations |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden well: it discloses cost ('1 API call, 0 with strategy=template'), the LLM-vs-template fallback ('auto = LLM if configured'), the validated-JSON-then-deterministic-render guarantee, and a cross-reference validation pass. It stops short of stating any side effects on stored prompts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense paragraph, front-loaded with purpose and then the style variants, guarantees, and cost. It is information-rich with little waste, though the parenthetical enumeration of the v2 spec is heavy for a single sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter tool with no annotations and no output schema, the description covers the essentials: what the output structure looks like in each style, the LLM/template modes, the validation step, and API cost. Remaining gaps are per-parameter semantics rather than conceptual context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 62%, and the description meaningfully expands the 'style' enum (which the schema leaves undescribed) and 'strategy' behavior. However, most of the 13 parameters (role, author, skills, audience, language, targetModel, outputFormat, include*) get no mention in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Expand a brief idea into a strict LangGPT Markdown system prompt.' The output artifact (strict LangGPT Markdown) is concrete and distinguishes it from generic prompt siblings like optimize_prompt_via_api or analyze_prompt.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the style/strategy switches (v2 vs classic, auto/llm/template) but never says when to reach for this tool versus optimizing or analyzing an existing prompt. Usage is implied by the output type rather than explicitly scoped against siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_templatesList optimization templatesA
List built-in and custom optimization templates (ids usable as template / iterateTemplate). Custom templates are Markdown files with frontmatter in the custom-templates folder and override built-ins with the same id. 0 API calls.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | ||
| showContent | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does a fair job: it discloses the free/local nature ('0 API calls'), where custom templates live (Markdown files with frontmatter in the custom-templates folder), and the override precedence rule. It does not explicitly state that this is a non-mutating read, but 'List' plus '0 API calls' makes it clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and output, followed by the customization/override rule and the cost note. No filler and nothing that could be trimmed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description should explain the return shape more fully. It names what is returned (built-in and custom templates, ids) but says nothing about the meaning of the `type` filter values or what `showContent` adds, leaving gaps an agent would need to guess at.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions the tool's own parameters (`type` enum, `showContent`) or what they filter/return. The enum values optimize-system/optimize-user/iterate correspond neatly to the template categories but the description leaves the agent to infer that mapping. The id-usage note concerns the output, not the inputs, so it does not compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List built-in and custom optimization templates') and immediately clarifies what the returned data is used for ('ids usable as `template` / `iterateTemplate`'). The resource itself cleanly separates it from the sibling tools, which are all single-prompt optimization actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the note that ids are usable as `template`/`iterateTemplate` tells the agent where this output feeds, and '0 API calls' hints it is a cheap local lookup. There is no explicit statement of when to call this versus, e.g., running an optimization, and no mention of prerequisites or the `type` filter as a selection guide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
optimize_prompt_via_apiOptimize prompt (iterative)A
Iteratively rewrite a prompt with the built-in template library (see list_templates): round 1 optimizes with a strategy template (default: langgpt-strict for system prompts, user-basic for user prompts); rounds 2..N apply your requirements with the iterate template, which sees both the original draft and the latest version. Every round is checked for LangGPT structure, dangling and lost {{placeholders}}, which are repaired automatically. Cost: 1 API call per round (+1 repair if needed). Uses a prompt-optimizer Docker deployment instead only if PO_BASE_URL is configured and PO_BACKEND allows it. For score-driven optimization with test cases, use auto_optimize_prompt.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | system = role/system prompt; user = a single one-off request | system |
| prompt | Yes | Draft prompt to optimize | |
| rounds | No | 1 = optimize only; each extra round is an iterate pass | |
| backend | No | Override PO_BACKEND for this call | |
| template | No | Optimize template id: langgpt-strict | general | output-format | analytical | user-basic | user-professional | user-planning | custom id | |
| autoRepair | No | ||
| showHistory | No | ||
| requirements | No | What the iterate rounds should change, e.g. 'reply as a Markdown table; max 200 words' | |
| enforceLangGPT | No | Validate & repair LangGPT structure (default: true for system mode) | |
| iterateTemplate | No | Iterate template id (default iterate) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does: it discloses cost (1 API call per round, +1 for repair), automatic validation/repair of LangGPT structure, dangling <References> and lost {{placeholders}}, and the backend-selection precondition. This is exactly the operational context an agent needs before committing API calls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and the round model, and every sentence carries operational detail (cost, repair, backend, alternative). It is a dense single paragraph, so a reader must work a little to extract the exclusions, but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter, no-annotation, no-output-schema tool, the description covers the workflow, cost, defaults, repair behavior and sibling routing. The main residual gap is what the response actually contains (e.g., what showHistory yields), which the agent must infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 80%, so the baseline is 3. The description adds real value on top: the meaning of rounds 2..N, the default template selection per mode, and the fact that the iterate template sees both the original draft and the latest version. It does not elaborate on autoRepair/showHistory/enforceLangGPT beyond their schema defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('iteratively rewrite a prompt') and names the mechanism (built-in template library, round 1 strategy template, rounds 2..N iterate template). It explicitly distinguishes itself from the sibling auto_optimize_prompt, so an agent can separate the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing: use auto_optimize_prompt for score-driven optimization with test cases, and use the Docker backend only when PO_BASE_URL is configured and PO_BACKEND allows it. Defaults per mode are also stated (langgpt-strict for system, user-basic for user).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
optimizer_connection_statusDiagnose optimizer connectionA
Diagnose the setup: which auth scheme the prompt-optimizer deployment uses and whether login succeeds, whether its /mcp backend is reachable, which upstream templates exist, and whether the local/eval LLMs are configured. Call this first when optimize_prompt_via_api fails.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the specific checks performed, which is useful, but does not state that it is read-only, what permissions are required, or what side effects (if any) the probing may have.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action ('Diagnose the setup:') and follows with a tight list of checks, then a single imperative sentence for usage. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter diagnostic with no output schema or annotations, the description covers what is inspected and when to invoke it. It could mention that it returns a diagnostic report or how failures are surfaced, but nothing essential for calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to clarify beyond what the empty schema already conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb 'Diagnose' and resource 'the setup', then enumerates exactly what it inspects: auth scheme, login success, /mcp backend reachability, upstream templates, and LLM configuration. This distinguishes it from every sibling, including the optimization tool it references.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger: 'Call this first when optimize_prompt_via_api fails.' Names the alternative and the condition that selects this tool, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prompt_historyPrompt version historyA
Browse saved results of generate/optimize/analyze/auto-optimize runs (every version is stored locally). action=list shows recent runs; action=get returns the full record with all versions. 0 API calls.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | ||
| limit | No | ||
| action | No | list |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses that versions are stored locally and that browsing costs 0 API calls, which is real behavioral value, but it never explicitly states this is a read-only, non-mutating operation or what happens on a missing id.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the resource and the local-storage/no-API-cost facts, with no filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with no output schema and no annotations, the description covers the action modes and locality but omits the id parameter behavior, limit bounds, and any sense of the returned record shape. Adequate but with clear holes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for 3 undocumented parameters. It explains the semantics of action list vs get and implies limit ('recent runs'), but never mentions the id parameter's format/role or the limit default and maximum, leaving real gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (browse) and resource (saved results of generate/optimize/analyze/auto-optimize runs), which clearly separates it from sibling tools that perform those operations. It does not name a sibling directly, but the 'history of prior runs' framing is distinctive enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies the usage context (retrieve past run results locally rather than calling the API again) but gives no explicit when-to-use vs when-not guidance or alternatives among the sibling tools. The action=list/action=get split is parameter guidance, not usage routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_promptsSuggest prompt candidates for any topicA
START HERE for any new topic. Given a topic/goal in plain words, returns 2-4 genuinely different ready-to-use prompts (concise / LangGPT structured / interactive coach / strict output format), each with when-to-use, expected effect and a LIVE sample reply produced by the real model on the same first message, so the user can pick by seeing the effect. Present the returned menu to the user and let them choose by number; then offer optimize_prompt_via_api / auto_optimize_prompt / evaluate_prompt_preview on the chosen one. Cost ≈ 1 design call + 1 LangGPT call + 1 preview call per candidate (≈ 6 calls for 4 candidates); previews can be turned off.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | How many candidates (order: concise, langgpt, coach, strict) | |
| topic | Yes | What the user wants to do or the assistant they want, in their own words | |
| context | No | Extra facts: audience, tools, constraints, examples of what good looks like | |
| preview | No | Run every candidate on the sample message (1 call each) | |
| language | No | auto | |
| strategy | No | auto | |
| sampleInput | No | The first user message used for the live preview; generated if omitted | |
| previewChars | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and does so richly: it discloses the cost model (≈1 design + 1 LangGPT + 1 preview call per candidate, ≈6 calls for 4 candidates), that live previews are produced by the real model on the sample message, and that previews can be turned off. This is exactly the kind of operational context an agent needs before calling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The imperative 'START HERE' is front-loaded and every sentence carries information (candidate types, what each includes, cost). It is dense and slightly long, but there is little redundant filler; the cost and routing details earn their space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must explain the return shape, and it does: each candidate carries when-to-use, expected effect, and a live sample reply, and the agent is told to present them as a numbered menu. Combined with cost disclosure, an agent has everything needed to call and use this correctly despite the 8-parameter surface.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 63%, so the schema documents most parameters, but the description adds meaning beyond it: '2-4 candidates' maps to count, 'previews can be turned off' maps to preview, and 'live sample reply ... on the same first message' explains sampleInput. It does not clarify language, strategy, or previewChars, leaving modest gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and output (returns 2-4 genuinely different ready-to-use prompt candidates) and frames itself as the entry point ('START HERE for any new topic'), which distinguishes it from siblings that optimize or evaluate an already-chosen prompt. An agent can tell immediately what it produces without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells when to use it ('START HERE for any new topic') and names the follow-on tools (optimize_prompt_via_api / auto_optimize_prompt / evaluate_prompt_preview) to invoke on the chosen candidate, plus instructs the agent to present a numbered menu and let the user pick. This is actionable routing, not vague guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.3.0- First observed
analyze_prompt - First observed
auto_optimize_prompt - First observed
evaluate_prompt_preview - First observed
generate_langgpt_structure - First observed
list_templates - First observed
optimize_prompt_via_api - First observed
optimizer_connection_status - First observed
prompt_history - First observed
suggest_prompts
TDQS
Scored across 9 tools
Most tools have distinct purposes and descriptions give explicit 'when to use' cues (e.g. suggest_prompts is 'START HERE', auto_optimize for metric-driven). However, generate_langgpt_structure vs optimize_prompt_via_api vs suggest_prompts all produce/rewrite prompts, and optimize_prompt_via_api vs auto_optimize_prompt both 'optimize', so an agent could occasionally misselect among these clusters.
Almost all tools use snake_case with verb_noun or verb_phrase patterns (suggest_prompts, optimize_prompt_via_api, evaluate_prompt_preview, auto_optimize_prompt). Minor deviations: prompt_history and optimizer_connection_status are noun-only, and prefixes alternate between 'prompt_' and 'optimizer_', but the scheme is broadly readable and consistent.
Nine tools is well within the ideal 3-15 range and each covers a distinct facet of the prompt-optimization workflow (generation, optimization, analysis, evaluation, templates, history, diagnostics). The set feels deliberately scoped with no obviously redundant tools.
Covers the full prompt lifecycle: generate suggestions, build LangGPT structure, iterative optimize, metric-driven auto-optimize, static analysis, live evaluation, template listing, and history browsing plus connection diagnostics. Minor gaps: no explicit save/delete/export of named prompts or history cleanup, but core workflows are fully supported.
Maintenance
Related MCP Connectors
- PromptOTOAuthcom.promptot
Manage, version, and publish LLM prompts with blocks, variables, and evaluations.
Test and compare prompts across any AI provider. Bring your own keys.
Generate contextual prompts and reusable agent skills, evaluate prompts with the 16-dimension Prompt Score, and manage saved work in PromptDrive. Twelve MCP tools also provide authorized access to private Memory for source-grounded answers. Connect over Streamable HTTP using OAuth 2.1 and PKCE. Generation consumes account quota and automatically saves successful results; Memory access follows account permissions and plan limits.
Self-hosted AI prompt library: prompts, collections, tags, teams, chains. 29 MCP tools for agents.
Related MCP Servers
- FlicenseAqualityCmaintenanceRefines and improves AI prompts using workspace-aware context from your project's tech stack, structure, and dependencies. Includes tools to analyze prompt quality and generate well-structured prompts from raw ideas.4209 npm5-
- AlicenseNot gradedqualityCmaintenanceRefines and optimizes prompts for LLMs through adaptive questioning and intelligent clarification workflows. Supports multiple AI providers (Google, OpenAI, Anthropic, Groq, Qwen) with interactive prompt enhancement and targeted modifications.17MIT
- AlicenseAqualityDmaintenanceEnables prompt optimization loops and regression test suites for Claude Code, with a companion web UI for real-time visualization of scores and prompt revisions.19MIT

PromptBranch MCPofficial
AlicenseNot gradedqualityAmaintenanceLocal-first MCP server for versioning, searching, evaluating, and improving AI prompts. Agents can fetch prompts, report outcomes, add notes, and propose human-approved variations.9MIT