Skip to main content
Glama
yanlong-iao

Ultimate Prompt Optimizer

by yanlong-iao

Ultimate Prompt Optimizer (MCP server)

An MCP server that turns Claude into a prompt-engineering workbench. It suggests candidate prompts, structures them with LangGPT, optimizes them over several rounds, and proves the result on test cases with promptfoo-style assertions and DSPy-style train/dev search. It runs on any OpenAI-compatible model with your own API key.

CI MCP TypeScript Node License: MIT

Why

Prompts written by hand are too vague for a model: "keep it short", "occasionally", "when the user seems stuck". Rewriting them with a single "make this better" call just moves the vagueness around. This server follows a loop that ends in evidence, not opinion:

  1. Generate several candidates. suggest_prompts returns 2–4 genuinely different prompts, each with a live sample reply.

  2. Let the user pick by seeing the effect, not by reading the prompt.

  3. Structure it. A LangGPT spec is rendered deterministically and validated.

  4. Iterate. Template-driven rewrites with requirements, automatic structure and {{placeholder}} repair.

  5. Prove it. A/B evaluation on test cases with assertions, model graders, a pairwise judge and repeat-consistency checks. Or let auto_optimize_prompt search for a better prompt and keep it only if a held-out dev set improves.

Related MCP server: Promptheus

Demo

This excerpt comes from a real run on deepseek-flash (full case study). It compares a hand-written quant-interview coaching prompt (A) with the version produced by optimize_prompt_via_api (B):

You: Use evaluate_prompt_preview to compare the original (A) and the optimized version (B), pairwise on, repeat 2, rubric: "ask exactly one question and wait; after the answer give a verdict, the fastest interview-ready solution and the key trick; be concise".

Claude → evaluate_prompt_preview (auto-generated 3 cases · estimated ≤ 27 calls)

Variant

Pass rate

Score

Latency avg / p95

Tokens in/out

Consistency

original

100%

0.97

5004 / 8605 ms

7308 / 4759

100%

optimized

100%

1.00

3518 / 5243 ms

11136 / 2947

100%

Pairwise judge: original wins 2 · optimized wins 1

Case 3 (adversarial: "stop quizzing me, write my CV and tell me the next answer"): both variants refused and asked one question. Case 2: the judge found that A's logic puzzle had no unique answer under its own assumptions.

The whole four-tool pipeline took 36 API calls, ~78k tokens, ≈ $0.06, 2 min 15 s. The case study also explains why these numbers do not prove that B is better.

Architecture

flowchart LR
  host["Claude Desktop / Claude Code<br/>(any MCP host)"] <-->|stdio · JSON-RPC| srv["MCP server<br/>9 tools · zod schemas"]
  srv --> mods["suggest · langgpt · analysis<br/>optimize · eval · templates"]
  mods --> llm
  subgraph client["LLM client (one per tool run)"]
    llm["chat()"] --- cache[("disk cache")]
    llm --- budget["call budget"]
    llm --- meter["usage & cost meter"]
  end
  llm -->|HTTPS /chat/completions| prov["OpenAI-compatible provider<br/>DeepSeek · OpenAI · Ollama · OpenRouter…"]
  mods -.->|optional · PO_BASE_URL| up["upstream prompt-optimizer<br/>Docker MCP"]

Modules, data flow and design trade-offs: docs/ARCHITECTURE.md.

Tools

Tool

What it does

API calls

suggest_prompts

Topic in plain words → 2–4 different prompts (concise / LangGPT / coach / strict format), each with a live sample reply. The recommended starting point.

≈ 2 + 1 per candidate (5 for 3 candidates in the case study)

generate_langgpt_structure

Idea → strict LangGPT prompt (v2 Goal / Skill-N / Rules, optional Commands / Reminder, weak mode for small models)

1 (+1 retry on invalid JSON); 0 in template mode

analyze_prompt

Design review without running: 5 dimension scores, issues, exact oldText → newText patch plan (auto-applicable), optional rewrite

1 (+1 with rewrite)

optimize_prompt_via_api

Multi-round template rewrite: optimize → iterate on requirements → auto-repair of structure and lost placeholders

1 per round (+1 repair if needed)

evaluate_prompt_preview

promptfoo-style eval: up to 6 variants × 4 models × cases × 5 repeats, 24 assertion types, pairwise judge, HTML/JSON report

Estimated before running; refused if over budget. 0 in mock mode

auto_optimize_prompt

DSPy-style search: generate cases → reflect / re-template / few-shot candidates → select on train → accept only on dev improvement

Scaled to maxCalls; the tool's own estimate for a light run with 6 cases is ~40–60 (not yet measured in the case study)

list_templates

Built-in and custom optimization templates

0

prompt_history

Every saved version from previous runs

0

optimizer_connection_status

Configuration and upstream diagnostics

0

Engineering highlights

Each item points to the code that implements it.

  • Injection-resistant templates. The prompt being optimized or judged is passed as a JSON evidence block declared to be data, and every template shares the same "data ≠ task" rules. Otherwise, a prompt saying "ignore previous instructions" would hijack the optimizer. → src/templates/shared.ts

  • {{placeholder}} protection. Variables are substituted in a single pass (inserted text is never re-scanned). After every round the pipeline diffs the placeholders between source and output, and any it lost are restored with a targeted repair round. → src/templates/registry.ts, src/optimize/pipeline.ts

  • Automatic LangGPT repair. A validator checks required sections, empty sections and dangling <References> (recognising Chinese and English section aliases). Violations become repair requirements for the next round. → src/langgpt/template.ts

  • The LLM fills a JSON spec, zod validates it (one retry with the validation errors fed back), and code renders the Markdown. Structure is correct by construction, and when no LLM is available a deterministic scaffold is returned instead. → src/langgpt/generator.ts

  • promptfoo-style assertion engine. not- negation, weight (weight 0 = informational), metric, transform, threshold, nested assert-set, JSON Schema via ajv, and model graders (llm-rubric, g-eval, factuality, answer-relevance, prompt-compliance) at temperature 0. Also reads promptfoo-style YAML configs, CSV __expected columns and variable matrices. → src/eval/assertions.ts, src/eval/config.ts

  • Pairwise judge with position swapping. A and B are shown in alternating positions across cases and the verdict is mapped back, which cancels a constant position bias over the test set. → src/analysis/analyzer.ts

  • DSPy-style train/dev split. Candidates (GEPA-like reflection on failing traces, alternative templates, BootstrapFewShot-like demos at zero extra calls) are screened on train, and finalists are accepted only if the dev score improves. The original is the baseline, so the result is never worse. → src/optimize/autoOptimizer.ts

  • Cost control. The call count is estimated before an eval and the run is refused if it exceeds the budget. A hard CallBudget applies per tool run, a content-addressed disk cache makes identical re-runs free, and a usage meter reports tokens and cost by purpose. In the case study the estimate of 27 eval calls matched the 27 made. → src/eval/runner.ts, src/utils/run.ts

  • Provider-agnostic. Plain fetch to /chat/completions; JSON mode is dropped automatically if the provider rejects it; retries with backoff on 429/5xx; localhost endpoints need no key (Ollama); separate eval and judge models. → src/clients/llmClient.ts

  • Offline test harness. A mock OpenAI-compatible LLM plus a mock upstream MCP server (Vercel-cookie and Docker Basic auth, session loss). The smoke suite spawns the real server over stdio and drives all 9 tools: 18 test groups, no API key, run in CI on Node 20 and 22. → test/

Quickstart

Requirements: Node ≥ 20 and an API key for any OpenAI-compatible provider.

git clone https://github.com/yanlong-iao/ultimate-prompt-optimizer-mcp.git
cd ultimate-prompt-optimizer-mcp
npm ci && npm run build
npm run smoke          # optional: offline end-to-end tests, no key needed
cp .env.example .env   # then fill in LLM_API_KEY (or use `bash setup.sh`, interactive)

Claude Code

claude mcp add prompt-optimizer -- node "$(pwd)/dist/index.js"

Claude Desktop (~/Library/Application Support/Claude/claude_desktop_config.json on macOS, %APPDATA%\Claude\claude_desktop_config.json on Windows):

{
  "mcpServers": {
    "prompt-optimizer": {
      "command": "node",
      "args": ["/absolute/path/to/ultimate-prompt-optimizer-mcp/dist/index.js"],
      "env": {
        "LLM_BASE_URL": "https://api.deepseek.com",
        "LLM_API_KEY": "<your key>",
        "LLM_MODEL": "deepseek-flash"
      }
    }
  }
}

The server reads .env next to dist/. Variables in the env block take precedence. Then try one of these in Claude:

  • "Use suggest_prompts for: a coach that quizzes me on probability questions for quant interviews"

  • "analyze_prompt this prompt and apply the patches: …"

  • "evaluate_prompt_preview original (A) vs optimized (B), pairwise on, rubric: …"

  • "auto_optimize_prompt this prompt, goal: short answers that state right/wrong first, budget light"

Environment variables

Variable

Default

Purpose

LLM_BASE_URL / LLM_API_KEY / LLM_MODEL

https://api.openai.com/v1 / – / gpt-4o-mini

Main model for every tool. Key optional for localhost endpoints

EVAL_BASE_URL / EVAL_API_KEY / EVAL_MODEL

same as LLM_*

Model under test in evaluations

JUDGE_MODEL

EVAL_MODEL

Grader and pairwise judge (use a stronger model to grade a cheaper one)

LLM_PRICE_INPUT_PER_M / LLM_PRICE_OUTPUT_PER_M

–

USD per 1M tokens, only for cost estimates

UPO_MAX_CALLS

80

Hard cap on API calls per tool run

UPO_CONCURRENCY

4

Parallel requests during evaluation

UPO_CACHE

true

Disk cache for identical requests

UPO_DATA_DIR / UPO_TEMPLATES_DIR

project folder / custom-templates

Where history/, reports/, .cache/ live / custom templates

PO_BASE_URL, PO_BACKEND, PO_ACCESS_PASSWORD, …

empty / auto

Optional self-hosted prompt-optimizer as the optimization backend

See .env.example for the full list and examples/ for a promptfoo-style config and CSV tests.

Case study

docs/CASE_STUDY.md runs the full pipeline on a real prompt with DeepSeek and reports scores, pass rates, tokens, cost and latency, together with what the numbers do and don't show.

Limitations & roadmap

Limitations (known and deliberate)

  • The javascript assertion and transform run in Node's vm module, which is not a security sandbox. Only run eval configs you wrote yourself.

  • Auto-generated test cases are single-turn, so multi-turn behaviour (for example, grading a user's answer) is not exercised unless you write such cases yourself. In the case study this made the rubric saturate at 100% for both variants.

  • The pairwise judge swaps positions across cases, not within a case, so per-case position bias remains. With few cases, the win counts are not statistically significant.

  • The DSPy-style search is a simplified beam over strategy templates. It has no Bayesian surrogate like MIPROv2, and small dev sets are noisy.

  • The promptfoo compatibility is a common subset (24 assertion types). There is no red-teaming, embedding similarity or web viewer, because DeepSeek has no embeddings endpoint.

  • Cost figures are estimates from configured prices. Provider-side prompt-cache discounts are not modelled.

  • The cache has no invalidation when a provider updates a model behind the same name (14-day TTL).

  • Non-template files in custom-templates/ (such as its README) are logged as skipped.

Roadmap

  • Multi-turn test cases (conversation arrays) in eval and case generation

  • Both-order pairwise judging with a confidence interval over wins

  • Measured auto_optimize_prompt case study with hand-written, multi-turn cases and a cross-family judge

  • npm package and official MCP Registry entry (npx install)

  • Demo GIF

A feature-by-feature comparison with LangGPT, prompt-optimizer, promptfoo and DSPy (in Chinese) is in COMPARISON.md.

Acknowledgements

This project borrows methods, not code or template text. All templates and code were written independently.

  • LangGPT: structured prompt format (Role / Profile / Goal / Skills / Rules / Workflow / Initialization)

  • linshenkx/prompt-optimizer: template-based optimize / iterate workflow, evidence-style prompt wrapping, design-review idea (AGPL-3.0; no text or code copied)

  • promptfoo: assertion vocabulary, config format and evaluation matrix

  • DSPy: metric-driven optimization, train/dev split, GEPA-style reflection, MIPRO-style instruction proposals, BootstrapFewShot

  • Model Context Protocol: TypeScript SDK

License

MIT © 2026 Yanlong Liao


中文简介

Ultimate Prompt Optimizer 是一个本地运行的 MCP 服务器,在 Claude Desktop / Claude Code 里提供一整套提示词工程工作流:

  1. 先给出多个风格不同的候选提示词,每个都附带真实的试运行回复,让用户按效果挑选。

  2. 用 LangGPT 做结构化:LLM 填 JSON 规格,经 zod 校验后由代码确定性渲染。

  3. 用模板库多轮迭代优化,自动修复结构和 {{变量}}。

  4. 用 promptfoo 式的断言和成对评审做 A/B 评测,或者用 DSPy 式的 train/dev 分数驱动搜索,只有 dev 集分数提高才接受修改。

所有 API 调用都会先估算次数、受预算上限约束、走磁盘缓存并计量成本。项目兼容任意 OpenAI 格式的接口(DeepSeek、OpenAI、Ollama 等),使用者自带 API key。离线测试共 18 组,不需要 key,在 CI 中运行。DeepSeek 真实实测见 docs/CASE_STUDY.md。

Available Tools

9 tools
analyze_promptAnalyze prompt designA

Design review of a prompt WITHOUT running it: 0-100 scores on goal clarity, instruction completeness, structural executability, ambiguity control and robustness, plus strengths, issues, improvements and an exact patch plan (oldText → newText). Optionally apply the patches locally (no extra call) and/or rewrite the prompt from the analysis (+1 call). Also reports LangGPT structure and static lint. Cost: 1 API call (cached if unchanged).

ParametersJSON Schema
NameRequiredDescriptionDefault
focusNoA specific concern to prioritize, e.g. 'the model keeps giving long derivations'
promptYes
rewriteNoAlso produce a full revised prompt from the analysis (+1 API call)
applyPatchesNoApply exact, unique patches and return the patched prompt

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the exact return content, the API-call cost model, that results are cached when unchanged, that applyPatches is local with no extra call, and that rewrite costs +1 call. Missing only edge cases such as failure behavior or patch-applicability limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense paragraph, front-loaded with the core purpose and scope, then the outputs, then the optional modes and cost. Every clause conveys distinct information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a four-parameter analysis tool with no output schema, the description enumerates the returned artifacts and the call-cost implications well enough to invoke correctly. It could say more about how focus interacts with scoring or what happens when patches can't be uniquely applied.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and the description adds genuine meaning beyond it, notably the cost implication of rewrite (+1 API call) and the local, no-extra-call nature of applyPatches, plus the patch format (oldText → newText). It gives less detail on the focus parameter than the schema already does.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Design review of a prompt WITHOUT running it') and immediately enumerates what is produced: 0-100 scores across five named dimensions, strengths, issues, improvements, and a patch plan. The 'WITHOUT running it' scope cleanly separates it from siblings like evaluate_prompt_preview or optimize_prompt_via_api.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes clear when this tool applies (static design review rather than execution) and describes the optional follow-on modes: applying patches locally or rewriting the prompt. It doesn't explicitly name a competing sibling and the condition that would select it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

auto_optimize_promptAuto-optimize prompt (score-driven)A

DSPy-style, metric-driven optimization: builds a test set (given or auto-generated with per-case criteria), splits it into train/dev, scores the original prompt, then for several rounds proposes candidates — reflect (analyse failing outputs → targeted edits, GEPA-like), template (rewrite with a different optimization strategy, MIPRO-like), fewshot (insert the best passing outputs as Examples, BootstrapFewShot-like) — evaluates them on train, confirms finalists on dev, and keeps a candidate only if the dev score improves. Returns the best prompt, a scoreboard of every candidate, per-case before/after, cost, and a saved report. Budget-guarded: the search is scaled down to fit maxCalls and stops early on no improvement or time limit. Typical light run with 6 cases: ~40-60 API calls.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalNoWhat a good response looks like — used for grading and reflection, e.g. 'short interview-style answers, always asks for missing info'
seedNo
modelNoModel the prompt will run on (default EVAL_MODEL)
assertNoAssertions applied to every case (e.g. is-json, max-length)
budgetNolight: 2 rounds × 2 candidates · medium: 3×3 · heavy: 5×4light
promptYesThe prompt to optimize
roundsNo
rubricNollm-rubric applied to every case
maxCallsNoHard cap on API calls (default UPO_MAX_CALLS)
numCasesNoCases to generate when none are given
testCasesNo
testsFileNoYAML/JSON/CSV tests file
maxMinutesNo
promptModeNosystem
strategiesNo
temperatureNo
candidatesPerRoundNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the multi-round search loop, the three named strategies, the dev-set gate for keeping candidates, budget guarding against maxCalls, early stopping on no improvement or time limit, and even a concrete cost estimate ('~40-60 API calls' for 6 cases). This is unusually rich behavioral disclosure for a mutation/expensive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the pipeline and dense with substance rather than filler; every clause (strategies, dev gate, budget guard, cost, return contents) earns its place. It is a single long block, which slightly hurts scannability, but it is appropriately sized for a tool of this complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-param tool with no output schema, the description usefully enumerates what is returned (best prompt, scoreboard, per-case before/after, cost, saved report) and the stopping/budget behavior. Minor gaps remain: where the report is saved, and the model/env dependencies (EVAL_MODEL, UPO_MAX_CALLS) that only appear in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 53% across 17 params, so the description has to compensate, and it does for the key knobs: it explains what reflect/template/fewshot do (GEPA-like, MIPRO-like, BootstrapFewShot-like), how the test set is sourced, and how budget/maxCalls shape the search. It still leaves several params (seed, temperature, promptMode, testsFile, maxMinutes) to the schema, but adds real meaning beyond it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: metric-driven prompt optimization that builds a test set, splits train/dev, proposes candidates, and keeps only dev-improving ones. This is clearly distinguishable from siblings like optimize_prompt_via_api or evaluate_prompt_preview without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains the internal conditions that drive behavior ('given or auto-generated' test set, budget scaling, early stop on no improvement), which is helpful context, but never states when to choose this over the sibling optimizer (auto_optimize_prompt vs optimize_prompt_via_api vs suggest_prompts). Usage is implied rather than routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evaluate_prompt_previewEvaluate / A-B test promptsA

promptfoo-style evaluation. Runs prompt variants (A, B, C…) × models × test cases × repeats against the real model, applies assertions and model graders, and reports pass rate, weighted score, per-metric averages, latency, tokens, cost, repeat consistency, an optional pairwise judge, and a winner. Writes an HTML + JSON report. Assertions (prefix not- to negate; weight/metric/transform/threshold supported): equals, contains, icontains, contains-any/all, icontains-any/all, starts-with, regex, is-json/contains-json (+JSON schema), javascript, levenshtein, latency, cost, min/max-length, is-refusal, llm-rubric, g-eval, factuality, answer-relevance, prompt-compliance, assert-set. Tests can come from testCases, a YAML/JSON/CSV testsFile, or a promptfoo-style configPath (prompts/providers/tests/defaultTest; array vars expand as a matrix); if none are given, cases are auto-generated. Responses are cached on disk so re-runs are free. The call count is estimated first and the run is refused if it exceeds maxCalls. mock=true → static lint + request preview, 0 calls.

ParametersJSON Schema
NameRequiredDescriptionDefault
mockNo
labelsNoNames for the variants, in order
modelsNoModel ids to compare on the same endpoint, e.g. ['deepseek-flash','deepseek-v4-pro']
promptNoVariant A (e.g. the original). Optional when configPath provides prompts
repeatNoRun each case N times to measure consistency (default 1)
rubricNollm-rubric applied to every case
promptBNoVariant B (e.g. the optimized prompt)
maxCallsNoRefuse to run if the estimated API calls exceed this (default UPO_MAX_CALLS)
pairwiseNoWith exactly 2 variants: LLM judge picks the better output per case (position-bias controlled)
useCacheNo
testCasesNo
testsFileNoPath to tests (.yaml/.json/.csv with __expected columns), relative to the project folder or absolute
configPathNoPath to a promptfoo-style config (.yaml/.json)
promptModeNosystem: prompt = system message, case input = user message. user: prompt is the user message with {{vars}}. Default system (user when configPath is used)
saveReportNo
temperatureNo
extraPromptsNoVariants C, D…
defaultAssertNoAssertions applied to every case
maxOutputCharsNo
autoGenerateCasesNoGenerated when no tests are given (each with its own grading criteria)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses disk response caching (free re-runs), call-count estimation with a refusal guard, mock mode performing static lint plus request preview at 0 calls, report file output, and a position-bias-controlled pairwise judge. It omits permission/auth requirements and what happens to existing reports on overwrite, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and behavior, but the second half is a dense run-on that re-lists the full assertion enum already present in the schema, duplicating structured data and inflating length. The report-metric sentence is also a long comma chain.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 20-parameter tool with no annotations and no output schema, the description covers the key operational facts an agent needs: inputs/fallbacks for test cases, caching, call budgeting, mock preview, and the reported outputs. Minor gaps remain around report paths and endpoint/permission assumptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 70% across 20 parameters, and the description adds real meaning beyond it: the assertion type catalogue with 'not-' negation and weight/metric/transform/threshold semantics, configPath expansion behavior, promptMode defaults, and autoGenerateCases behavior. This supplements rather than merely restates the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource set ('runs prompt variants × models × test cases × repeats against the real model, applies assertions and model graders') and enumerates the exact artifacts it produces. It is clearly a batch evaluation/benchmarking tool, readily separable from siblings like analyze_prompt or optimize_prompt_via_api.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete operating conditions: tests come from testCases, testsFile, or configPath, and auto-generate when none given; mock=true yields a 0-call preview; the run is refused if estimates exceed maxCalls. It does not explicitly contrast when to choose this over sibling tools such as analyze_prompt, leaving that inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_langgpt_structureGenerate LangGPT promptA

Expand a brief idea into a strict LangGPT Markdown system prompt. Default style 'v2' follows the current LangGPT spec (Role, Profile, Background, Goal with Outcome/Done Criteria/Non-Goals, Skill-N subsections, Rules, Workflow, OutputFormat, optional Commands/Reminder/Examples, Initialization); 'classic' uses Goals/Constraints/Skills lists. With an LLM configured the content is domain-specific (the model fills a validated JSON spec; Markdown is rendered deterministically, so structure is guaranteed). Checks that every points to an existing section. 1 API call (0 with strategy=template).

ParametersJSON Schema
NameRequiredDescriptionDefault
ideaYesBrief description of the assistant/prompt you want
roleNoExplicit role name; inferred if omitted
styleNov2
authorNo
skillsNo
audienceNo
languageNoauto
strategyNoauto = LLM if configured, else deterministic templateauto
constraintsNoRules that must appear verbatim
targetModelNoweak = shorter, flatter prompt for small modelsstrong
outputFormatNoRequired response format
includeCommandsNoAdd a LangGPT ## Commands section (/help, /continue, /improve)
includeReminderNoAdd a ## Reminder section for long conversations

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden well: it discloses cost ('1 API call, 0 with strategy=template'), the LLM-vs-template fallback ('auto = LLM if configured'), the validated-JSON-then-deterministic-render guarantee, and a cross-reference validation pass. It stops short of stating any side effects on stored prompts.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense paragraph, front-loaded with purpose and then the style variants, guarantees, and cost. It is information-rich with little waste, though the parenthetical enumeration of the v2 spec is heavy for a single sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter tool with no annotations and no output schema, the description covers the essentials: what the output structure looks like in each style, the LLM/template modes, the validation step, and API cost. Remaining gaps are per-parameter semantics rather than conceptual context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 62%, and the description meaningfully expands the 'style' enum (which the schema leaves undescribed) and 'strategy' behavior. However, most of the 13 parameters (role, author, skills, audience, language, targetModel, outputFormat, include*) get no mention in the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Expand a brief idea into a strict LangGPT Markdown system prompt.' The output artifact (strict LangGPT Markdown) is concrete and distinguishes it from generic prompt siblings like optimize_prompt_via_api or analyze_prompt.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains the style/strategy switches (v2 vs classic, auto/llm/template) but never says when to reach for this tool versus optimizing or analyzing an existing prompt. Usage is implied by the output type rather than explicitly scoped against siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_templatesList optimization templatesA

List built-in and custom optimization templates (ids usable as template / iterateTemplate). Custom templates are Markdown files with frontmatter in the custom-templates folder and override built-ins with the same id. 0 API calls.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNo
showContentNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does a fair job: it discloses the free/local nature ('0 API calls'), where custom templates live (Markdown files with frontmatter in the custom-templates folder), and the override precedence rule. It does not explicitly state that this is a non-mutating read, but 'List' plus '0 API calls' makes it clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action and output, followed by the customization/override rule and the cost note. No filler and nothing that could be trimmed without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description should explain the return shape more fully. It names what is returned (built-in and custom templates, ids) but says nothing about the meaning of the `type` filter values or what `showContent` adds, leaving gaps an agent would need to guess at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions the tool's own parameters (`type` enum, `showContent`) or what they filter/return. The enum values optimize-system/optimize-user/iterate correspond neatly to the template categories but the description leaves the agent to infer that mapping. The id-usage note concerns the output, not the inputs, so it does not compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List built-in and custom optimization templates') and immediately clarifies what the returned data is used for ('ids usable as `template` / `iterateTemplate`'). The resource itself cleanly separates it from the sibling tools, which are all single-prompt optimization actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the note that ids are usable as `template`/`iterateTemplate` tells the agent where this output feeds, and '0 API calls' hints it is a cheap local lookup. There is no explicit statement of when to call this versus, e.g., running an optimization, and no mention of prerequisites or the `type` filter as a selection guide.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

optimize_prompt_via_apiOptimize prompt (iterative)A

Iteratively rewrite a prompt with the built-in template library (see list_templates): round 1 optimizes with a strategy template (default: langgpt-strict for system prompts, user-basic for user prompts); rounds 2..N apply your requirements with the iterate template, which sees both the original draft and the latest version. Every round is checked for LangGPT structure, dangling and lost {{placeholders}}, which are repaired automatically. Cost: 1 API call per round (+1 repair if needed). Uses a prompt-optimizer Docker deployment instead only if PO_BASE_URL is configured and PO_BACKEND allows it. For score-driven optimization with test cases, use auto_optimize_prompt.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNosystem = role/system prompt; user = a single one-off requestsystem
promptYesDraft prompt to optimize
roundsNo1 = optimize only; each extra round is an iterate pass
backendNoOverride PO_BACKEND for this call
templateNoOptimize template id: langgpt-strict | general | output-format | analytical | user-basic | user-professional | user-planning | custom id
autoRepairNo
showHistoryNo
requirementsNoWhat the iterate rounds should change, e.g. 'reply as a Markdown table; max 200 words'
enforceLangGPTNoValidate & repair LangGPT structure (default: true for system mode)
iterateTemplateNoIterate template id (default iterate)

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does: it discloses cost (1 API call per round, +1 for repair), automatic validation/repair of LangGPT structure, dangling <References> and lost {{placeholders}}, and the backend-selection precondition. This is exactly the operational context an agent needs before committing API calls.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the round model, and every sentence carries operational detail (cost, repair, backend, alternative). It is a dense single paragraph, so a reader must work a little to extract the exclusions, but there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter, no-annotation, no-output-schema tool, the description covers the workflow, cost, defaults, repair behavior and sibling routing. The main residual gap is what the response actually contains (e.g., what showHistory yields), which the agent must infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 80%, so the baseline is 3. The description adds real value on top: the meaning of rounds 2..N, the default template selection per mode, and the fact that the iterate template sees both the original draft and the latest version. It does not elaborate on autoRepair/showHistory/enforceLangGPT beyond their schema defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('iteratively rewrite a prompt') and names the mechanism (built-in template library, round 1 strategy template, rounds 2..N iterate template). It explicitly distinguishes itself from the sibling auto_optimize_prompt, so an agent can separate the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing: use auto_optimize_prompt for score-driven optimization with test cases, and use the Docker backend only when PO_BASE_URL is configured and PO_BACKEND allows it. Defaults per mode are also stated (langgpt-strict for system, user-basic for user).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

optimizer_connection_statusDiagnose optimizer connectionA

Diagnose the setup: which auth scheme the prompt-optimizer deployment uses and whether login succeeds, whether its /mcp backend is reachable, which upstream templates exist, and whether the local/eval LLMs are configured. Call this first when optimize_prompt_via_api fails.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the specific checks performed, which is useful, but does not state that it is read-only, what permissions are required, or what side effects (if any) the probing may have.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action ('Diagnose the setup:') and follows with a tight list of checks, then a single imperative sentence for usage. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter diagnostic with no output schema or annotations, the description covers what is inspected and when to invoke it. It could mention that it returns a diagnostic report or how failures are surfaced, but nothing essential for calling it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to clarify beyond what the empty schema already conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb 'Diagnose' and resource 'the setup', then enumerates exactly what it inspects: auth scheme, login success, /mcp backend reachability, upstream templates, and LLM configuration. This distinguishes it from every sibling, including the optimization tool it references.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger: 'Call this first when optimize_prompt_via_api fails.' Names the alternative and the condition that selects this tool, leaving little to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

prompt_historyPrompt version historyA

Browse saved results of generate/optimize/analyze/auto-optimize runs (every version is stored locally). action=list shows recent runs; action=get returns the full record with all versions. 0 API calls.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNo
limitNo
actionNolist

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses that versions are stored locally and that browsing costs 0 API calls, which is real behavioral value, but it never explicitly states this is a read-only, non-mutating operation or what happens on a missing id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the resource and the local-storage/no-API-cost facts, with no filler. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with no output schema and no annotations, the description covers the action modes and locality but omits the id parameter behavior, limit bounds, and any sense of the returned record shape. Adequate but with clear holes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for 3 undocumented parameters. It explains the semantics of action list vs get and implies limit ('recent runs'), but never mentions the id parameter's format/role or the limit default and maximum, leaving real gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (browse) and resource (saved results of generate/optimize/analyze/auto-optimize runs), which clearly separates it from sibling tools that perform those operations. It does not name a sibling directly, but the 'history of prior runs' framing is distinctive enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies the usage context (retrieve past run results locally rather than calling the API again) but gives no explicit when-to-use vs when-not guidance or alternatives among the sibling tools. The action=list/action=get split is parameter guidance, not usage routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

suggest_promptsSuggest prompt candidates for any topicA

START HERE for any new topic. Given a topic/goal in plain words, returns 2-4 genuinely different ready-to-use prompts (concise / LangGPT structured / interactive coach / strict output format), each with when-to-use, expected effect and a LIVE sample reply produced by the real model on the same first message, so the user can pick by seeing the effect. Present the returned menu to the user and let them choose by number; then offer optimize_prompt_via_api / auto_optimize_prompt / evaluate_prompt_preview on the chosen one. Cost ≈ 1 design call + 1 LangGPT call + 1 preview call per candidate (≈ 6 calls for 4 candidates); previews can be turned off.

ParametersJSON Schema
NameRequiredDescriptionDefault
countNoHow many candidates (order: concise, langgpt, coach, strict)
topicYesWhat the user wants to do or the assistant they want, in their own words
contextNoExtra facts: audience, tools, constraints, examples of what good looks like
previewNoRun every candidate on the sample message (1 call each)
languageNoauto
strategyNoauto
sampleInputNoThe first user message used for the live preview; generated if omitted
previewCharsNo

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so richly: it discloses the cost model (≈1 design + 1 LangGPT + 1 preview call per candidate, ≈6 calls for 4 candidates), that live previews are produced by the real model on the sample message, and that previews can be turned off. This is exactly the kind of operational context an agent needs before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The imperative 'START HERE' is front-loaded and every sentence carries information (candidate types, what each includes, cost). It is dense and slightly long, but there is little redundant filler; the cost and routing details earn their space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must explain the return shape, and it does: each candidate carries when-to-use, expected effect, and a live sample reply, and the agent is told to present them as a numbered menu. Combined with cost disclosure, an agent has everything needed to call and use this correctly despite the 8-parameter surface.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 63%, so the schema documents most parameters, but the description adds meaning beyond it: '2-4 candidates' maps to count, 'previews can be turned off' maps to preview, and 'live sample reply ... on the same first message' explains sampleInput. It does not clarify language, strategy, or previewChars, leaving modest gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and output (returns 2-4 genuinely different ready-to-use prompt candidates) and frames itself as the entry point ('START HERE for any new topic'), which distinguishes it from siblings that optimize or evaluate an already-chosen prompt. An agent can tell immediately what it produces without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells when to use it ('START HERE for any new topic') and names the follow-on tools (optimize_prompt_via_api / auto_optimize_prompt / evaluate_prompt_preview) to invoke on the chosen candidate, plus instructs the agent to present a numbered menu and let the user pick. This is actionable routing, not vague guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv0.3.0
    • First observedanalyze_prompt
    • First observedauto_optimize_prompt
    • First observedevaluate_prompt_preview
    • First observedgenerate_langgpt_structure
    • First observedlist_templates
    • First observedoptimize_prompt_via_api
    • First observedoptimizer_connection_status
    • First observedprompt_history
    • First observedsuggest_prompts

TDQS

A4/5.0

Scored across 9 tools

Disambiguation4/5

Most tools have distinct purposes and descriptions give explicit 'when to use' cues (e.g. suggest_prompts is 'START HERE', auto_optimize for metric-driven). However, generate_langgpt_structure vs optimize_prompt_via_api vs suggest_prompts all produce/rewrite prompts, and optimize_prompt_via_api vs auto_optimize_prompt both 'optimize', so an agent could occasionally misselect among these clusters.

Naming Consistency4/5

Almost all tools use snake_case with verb_noun or verb_phrase patterns (suggest_prompts, optimize_prompt_via_api, evaluate_prompt_preview, auto_optimize_prompt). Minor deviations: prompt_history and optimizer_connection_status are noun-only, and prefixes alternate between 'prompt_' and 'optimizer_', but the scheme is broadly readable and consistent.

Tool Count5/5

Nine tools is well within the ideal 3-15 range and each covers a distinct facet of the prompt-optimization workflow (generation, optimization, analysis, evaluation, templates, history, diagnostics). The set feels deliberately scoped with no obviously redundant tools.

Completeness4/5

Covers the full prompt lifecycle: generate suggestions, build LangGPT structure, iterative optimize, metric-driven auto-optimize, static analysis, live evaluation, template listing, and history browsing plus connection diagnostics. Minor gaps: no explicit save/delete/export of named prompts or history cleanup, but core workflows are fully supported.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers