Prompt Optimizer MCP
Prompt Optimizer MCP is a local-first MCP server for improving, inspecting, adapting, evaluating, comparing, and assembling AI agent prompts and context without requiring an API key.
Optimize vague or substantial prompts into bounded execution contracts while preserving intent.
Inspect prompts for intent, ambiguities, missing requirements, conflicts, and context issues without rewriting.
Adapt already-good prompts to specific target agent conventions like Codex, Claude Code, Cursor, and OpenCode.
Evaluate prompt quality using transparent heuristic diagnostics across clarity, specificity, goal preservation, and more.
Compare two prompts by heuristic dimensions and task context without declaring an arbitrary winner.
Rank, deduplicate, and budget supplied context while retaining critical items.
Build an execution instruction from an explicit goal, requirements, and constraints.
Explain caller-supplied optimization results and their pass decisions for auditing.
Suggest a proportional local optimization profile: minimal, balanced, or thorough.
Run locally by default with no network or API key; non-local modes may send redacted data to a configured provider.
Operate read-only and non-destructive with no server-side history retention.
Supports optional OpenAI-compatible provider APIs for remote model calls or a separate evaluator during prompt optimization, requiring trusted server configuration and explicit selection.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Prompt Optimizer MCPimprove this instruction, then execute it: fix auth"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Install an MCP server + portable Agent Skill. Your existing agent gains the ability to turn vague, complex or context-heavy requests into focused instructions, then execute them. No API key required.
Quick start · Install guides · Examples · API · Troubleshooting
Why Prompt Optimizer?
Small gaps in a request can become large implementation assumptions. Prompt Optimizer helps the agent clarify relevant requirements, preserve constraints and verify its work without turning every task into a lengthy specification.
Local-first and deterministic: useful optimization without network access or an external model.
Selective and proportional: complex tasks get structure; trivial requests stay out of the optimizer.
Repository-aware and scope-aware: reuse project conventions, avoid unnecessary abstractions and preserve the requested features.
Portable and inspectable: MCP, an instruction-only Agent Skill and a TypeScript SDK share one engine. Context ranking, compression and pass explanations are built in.
Related MCP server: deep-thinking-engine
How it works
Prompt Optimizer improves instructions. Your AI agent executes the work.
The skill decides whether optimization would help, reads only necessary repository context, calls optimize_prompt, then uses the validated result as an execution contract. The host implements and verifies that contract while respecting higher-priority instructions and explicit user requirements.
Supported AI agents
Host | MCP | Agent Skill | Evidence |
OpenAI Codex | Yes | Yes | Integration supported; MCP protocol tested |
Claude Code | Yes | Yes | Integration supported; MCP protocol tested |
Google Antigravity | Yes | Yes | CLI live tested, user-reported v0.1.x; IDE untested |
Cursor | Yes | Yes | Integration supported; MCP protocol tested |
OpenCode | Yes | Yes | Integration supported; MCP protocol tested |
Generic MCP hosts | Yes | If supported by host | Protocol tested; portable skill available |
Protocol tests use the official MCP client; they do not substitute for live testing each host. Agent names identify integrations and imply no endorsement. See the compatibility matrix and evidence.
Quick start
With Node.js 22+, check the package and local server health:
npx -y prompt-optimizer-mcp-engine@0.1.1 --version
npx -y prompt-optimizer-mcp-engine@0.1.1 doctorInstall MCP
Codex:
codex mcp add prompt-optimizer -- npx -y prompt-optimizer-mcp-engine@0.1.1Claude Code, project scope:
claude mcp add --transport stdio --scope project prompt-optimizer -- npx -y prompt-optimizer-mcp-engine@0.1.1For Antigravity, Cursor, OpenCode and other clients, use the host guides. Stdio and Streamable HTTP are supported.
Install Agent Skill
The skill teaches the host when and how to optimize. Install the package globally to access its portable installer; from your target project, in a POSIX shell:
npm install --global prompt-optimizer-mcp-engine@0.1.1
PROMPTOPT_ROOT="$(npm root -g)/prompt-optimizer-mcp-engine"
node "$PROMPTOPT_ROOT/scripts/install-skill.mjs" --agent codex --scope project --project "$PWD"Replace codex with claude-code, antigravity, cursor, opencode or generic as appropriate. Antigravity CLI global installs use --agent antigravity-cli --scope user. The installer refuses to overwrite an existing skill: review and refresh a previous copy when upgrading. Updating npm alone does not update a copied skill.
Restart the host. For a quick discovery test, ask: “Use Prompt Optimizer's inspect_prompt tool to analyze ‘fix auth’ without executing it.” For explicit optimization: “Use the prompt-optimizer skill to optimize this task, then execute it: …”
Install guides
Codex · Claude Code · Google Antigravity · Cursor · OpenCode · Generic MCP / VS Code / Gemini · Local build · Portable Agent Skill
Example: request → contract → execution
Before, from the Antigravity CLI field test:
Buat dashboard admin untuk platform AI yang menampilkan penggunaan token,
status model, biaya, dan riwayat request.
Gunakan Next.js dan sesuaikan dengan project yang sudah ada.In the user's v0.1.1 fresh-session replay, the skill recognized this multi-feature, repository-aware task and invoked optimize_prompt before implementation.
Optimized guidance, abbreviated and representative rather than a verbatim trace:
Build the requested Next.js admin dashboard within the existing project. Cover token usage, model status, costs and request history. Reuse existing components and conventions. Keep changes focused; include responsive behavior, accessibility and relevant interaction states. Run the appropriate checks and report the result.
Execution: the same host agent implements that scope. It should not invent a prompt studio, model registry, webhook system or other major features. Necessary repository-specific details may be added to carry out the task.
Selective optimization
Request | Expected skill decision |
“Buat dashboard AI modern pakai Next.js.” | Optimize before implementation |
“Fix authentication ini dan rapikan arsitekturnya.” | Optimize and surface scope/constraints |
“Ubah warna tombol utama jadi merah.” | Skip; make the direct edit |
“Apa arti dependency?” / “jalankan npm run build” | Skip; answer or execute directly |
“Optimize this prompt…” | Optimize explicitly |
A successful result stays the execution contract, not a starting point for a broader specification. Normally use one primary optimization pass; do not recursively optimize its output or run all nine tools on every task. Automatic discovery remains host-controlled. Activation troubleshooting.
Tools
MCP tool | Purpose |
| Improve a request with proportional execution guidance |
| Analyze intent, gaps, conflicts and context without rewriting |
| Apply documented target-agent conventions |
| Return clearly labeled heuristic quality diagnostics |
| Compare strengths and weaknesses by dimension |
| Rank, deduplicate and budget relevant context |
| Assemble explicit task information into instructions |
| Explain supplied optimization records |
| Recommend a proportional local profile |
Also available: five read-only resources and five optional task templates. Full API.
Architecture
The independent TypeScript core powers the MCP server and SDK. The portable skill contains workflow guidance, not a duplicate optimization engine. Tools work without skills; skills offer a lightweight fallback without MCP.
import { PromptOptimizer } from 'prompt-optimizer-mcp-engine';
const result = await new PromptOptimizer().optimize({
prompt: 'buat dashboard ai keren pake nextjs',
targetAgent: 'codex',
profile: 'balanced',
mode: 'local',
});
console.log(result.optimizedPrompt);Architecture · Optional OpenAI-compatible, Anthropic and Gemini providers
Security and privacy
Local mode makes no remote model requests and needs no API key. No telemetry is collected. Remote AI requires explicit configuration and mode selection. Context is untrusted data; optimizer advice cannot override higher-priority instructions or host approvals.
Secret redaction is defense-in-depth, not a guarantee. Classification, conflict checks, token estimates and quality diagnostics are heuristic. They do not prove semantic fidelity or safe execution. Security model · Report a vulnerability
Development
pnpm install --frozen-lockfile
pnpm build
pnpm check
pnpm examples
pnpm run doctorVerification · Release process · Reference projects and licenses
Roadmap
Broaden live host coverage, benchmark multilingual intent preservation and improve context provenance. New capabilities must preserve selective activation and bounded scope.
Contributing
Bug reports, reproducible host traces and focused contributions are welcome. See CONTRIBUTING.md and the code of conduct.
License
MIT. Built for the open agent ecosystem.
Available Tools
9 toolsadapt_promptBRead-onlyIdempotent
Add documented host guidance to an already good instruction. Use when changing agent environments; not for expanding task scope.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | ||
| audience | No | auto | |
| targetAgent | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| ir | Yes | |
| changes | Yes | |
| context | Yes | |
| metadata | Yes | |
| confidence | Yes | |
| assumptions | Yes | |
| diagnostics | Yes | |
| skippedPasses | Yes | |
| tokenEstimate | Yes | |
| detectedIntent | Yes | |
| selectedPasses | Yes | |
| optimizedPrompt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds the useful boundary that scope is not expanded, but says nothing about what the tool actually rewrites or how the original prompt is preserved.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with zero filler; the core action leads and the exclusion follows. Nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but the tool transforms a user-supplied prompt using two enum parameters that are undocumented in both the schema and the description. That leaves an agent unable to judge how to pick targetAgent or audience.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description mentions none of the three parameters (prompt, targetAgent, audience). With two enum-constrained params and no schema prose, the description needed to compensate for choosing a target agent/audience and does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a verb (add) and object (documented host guidance to an existing instruction), which is more than a tautology, but 'host guidance' is undefined jargon and no sibling is named. It only partially differentiates from the cluster of prompt tools (optimize_prompt, build_agent_instruction, etc.).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger ('use when changing agent environments') and an explicit exclusion ('not for expanding task scope'). No alternative sibling is named, but the when/when-not framing is unusually clear for this set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_agent_instructionARead-onlyIdempotent
Assemble an execution instruction from an explicit goal, requirements, and constraints. Use for multi-part tasks; does not execute them.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | Yes | ||
| context | No | ||
| constraints | No | ||
| targetAgent | No | auto | |
| requirements | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ir | Yes | |
| changes | Yes | |
| context | Yes | |
| metadata | Yes | |
| confidence | Yes | |
| assumptions | Yes | |
| diagnostics | Yes | |
| skippedPasses | Yes | |
| tokenEstimate | Yes | |
| detectedIntent | Yes | |
| selectedPasses | Yes | |
| optimizedPrompt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world, so the safety profile is fully covered. The description's 'does not execute them' adds a modest clarification of the boundary but is largely implied by readOnlyHint and gives no detail on determinism, targetAgent behavior, or output shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences: the purpose leads and the usage note follows. No filler, no repetition, and the key scope constraint is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. However, with 5 parameters, a 0%-covered schema, a nested context array with a kind enum, and the targetAgent enum, the description omits enough (especially targetAgent and the constraints-vs-requirements distinction) to be only adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It names goal, requirements, and constraints (3 of 5), which is meaningful partial compensation, but leaves context and the significant targetAgent enum (8 target formats, default 'auto') entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (assemble) and resource (execution instruction) and names the three input sources (goal, requirements, constraints). This clearly differentiates it from prompt-manipulation siblings like optimize_prompt or inspect_prompt, though it doesn't explicitly name any sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear usage cue ('Use for multi-part tasks') and a boundary ('does not execute them'). It does not name alternatives among the sibling prompt tools or state when NOT to use it versus them, but the context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_promptsARead-onlyIdempotent
Compare two instructions by heuristic dimensions and task context. Use to inspect trade-offs; does not declare an arbitrary winner.
| Name | Required | Description | Default |
|---|---|---|---|
| promptA | Yes | ||
| promptB | Yes | ||
| taskContext | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| kind | Yes | |
| promptA | Yes | |
| promptB | Yes | |
| conclusion | Yes | |
| dimensions | Yes | |
| taskContextUsed | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly, idempotent, non-destructive, closed-world. The description adds a meaningful behavioral trait beyond them: output is descriptive trade-off analysis rather than a ranked winner. With annotations carrying the safety profile, this extra contextual disclosure earns a 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two terse sentences, front-loaded with the verb+resource and immediately followed by the differentiator. Zero waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need not be explained. Annotations cover safety. The remaining gap is that, with 0% parameter descriptions, the agent must guess parameter semantics; the description's phrase 'task context' helps but does not fully close it. Still complete enough for a comparison tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the schema names promptA/promptB/taskContext but explains nothing. The description compensates partially by implying the two prompts and 'task context' are what get compared, but it never states syntax, limits, or whether taskContext is optional-in-practice. Baseline 3 given partial compensation for a 0%-coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb 'compare' plus resource 'two instructions' and the mechanism 'by heuristic dimensions and task context'. The negation 'does not declare an arbitrary winner' distinguishes it from sibling evaluate_prompt, which presumably ranks or scores.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Contains one useful boundary marker: it produces comparisons, not verdicts. However, it never says when to reach for this versus evaluate_prompt, optimize_prompt, or inspect_prompt, so the agent must infer the selection context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compress_contextARead-onlyIdempotent
Rank, deduplicate, and budget supplied context while retaining critical items. Use for noisy context; do not expect semantic summarization or instruction promotion.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| context | Yes | ||
| maxTokens | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| items | Yes | |
| droppedIds | Yes | |
| diagnostics | Yes | |
| afterEstimated | Yes | |
| budgetExceeded | Yes | |
| beforeEstimated | Yes | |
| duplicatesRemoved | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), so the description's real contribution is the algorithmic boundary — it clarifies this is mechanical ranking/budgeting, not summarization or instruction rewriting. That is meaningful added context, though it says nothing about failure modes, truncation behavior, or how the budget interacts with critical items.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The positive capability is front-loaded and the negative expectation is second, which is exactly the right ordering for an agent scanning for fit.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description covers purpose, trigger, and limits. The gap is parameter-level detail for a 0%-coverage schema, especially the undocumented query parameter that determines ranking relevance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across three parameters, but the description maps loosely onto them: 'rank' implies the relevance-driving query, 'budget' implies maxTokens, and 'retaining critical items' implies the critical flag. It adds no format, default, or constraint detail, so the compensation is partial at best.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses three specific verbs (rank, deduplicate, budget) on a named resource (supplied context) and states the retention goal ('while retaining critical items'). This is far more precise than the prompt-oriented siblings, though it never explicitly names an alternative to differentiate itself from them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use for noisy context' gives an explicit triggering condition, and 'do not expect semantic summarization or instruction promotion' is a clear negative scoping statement that prevents misuse. It lacks routing guidance naming a sibling alternative, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluate_promptBRead-onlyIdempotent
Explain instruction quality using transparent heuristic diagnostics. Use for review, not as a scientific benchmark or proof of correctness.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | ||
| originalPrompt | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| kind | Yes | |
| metrics | Yes | |
| limitations | Yes | |
| estimatedTokens | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent and non-destructive, so the safety profile is covered. The description adds genuine context beyond that: the tool is heuristic and 'transparent', and its output is explicitly not a correctness proof, which is meaningful expectation-setting for an evaluation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the core purpose and followed by the limitation. Every sentence carries weight, though the opening could be more concrete about operating on the input prompt.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. The gap is the undocumented optional 'originalPrompt' input and the absence of any link to the other prompt tools, which for a nine-sibling family is a notable omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description mentions no parameters at all. In particular the optional 'originalPrompt' is left entirely unexplained, so an agent cannot tell whether supplying it changes behavior (e.g. diffing a revised prompt).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states it explains instruction quality via heuristic diagnostics, which separates it somewhat from optimize_prompt or compare_prompts. However, it never plainly says it evaluates the supplied prompt, and the verb 'Explain' is weaker than the tool name's 'evaluate', leaving the core action slightly vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use for review, not as a scientific benchmark or proof of correctness' gives a directional use case and a caveat, which is real guidance. It offers no explicit when-not conditions tied to sibling tools like compare_prompts or optimize_prompt, so the agent must infer which of the nine siblings to pick.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_optimizationARead-onlyIdempotent
Explain a supplied optimization result and its pass decisions. Use for audit; the server stores no history and does not authenticate caller-supplied results.
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| applied | Yes | |
| skipped | Yes | |
| summary | Yes | |
| provenance | Yes | |
| diagnostics | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world). The description adds genuinely new behavioral context: the server persists no history and does not authenticate caller-supplied results, which tells the agent the input is untrusted and non-persistent. It does not describe the output payload, but that is largely covered by the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the core action front-loaded and the caveats second. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, so return values need no explanation, and the description adequately covers the trust/persistence model. However, for a tool whose only input is a huge nested result object, the description gives no orientation about what must be supplied beyond 'optimization result,' leaving a real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for a single required parameter whose shape is a very large, deeply nested required IR object. 'Supplied optimization result' usefully implies the input is the payload returned by optimize_prompt, but the description adds nothing about the required members, so an agent still depends entirely on the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Explain a supplied optimization result and its pass decisions.' The scope (result plus pass decisions) is concrete enough to distinguish it from optimize_prompt or evaluate_prompt, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use for audit' gives a clear use context, and the clause about no stored history implies the result must be produced and handed over in the same session. It stops short of naming alternatives (e.g., vs. evaluate_prompt), so it is context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_promptARead-onlyIdempotent
Inspect intent, ambiguity, conflicts, and context without rewriting. Use to diagnose instructions; skip simple factual questions.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | ||
| context | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ir | Yes | |
| intent | Yes | |
| conflicts | Yes | |
| ambiguities | Yes | |
| contextIssues | Yes | |
| recommendedPasses | Yes | |
| missingRequirements | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered. The description's 'without rewriting' adds a small amount of behavioral context (input is not mutated) but is essentially a restatement of readOnlyHint, with no detail on depth of analysis or output shape beyond the separate output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, with the core capability front-loaded before the routing guidance. There is no filler and nothing that could be cut without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema and annotations relieve the description of explaining return values and safety, and the usage sentence is crisp. However, for a tool whose only real surface is its two parameters, the complete absence of parameter guidance leaves an agent guessing how to populate 'context' and what mark the 'critical' flag has.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both parameters, so the description carries the full burden of explaining them and it never mentions 'prompt' or 'context' at all. It leaves the context array's kind enum, the 'critical' flag, and the 128-item/64000-char limits entirely undocumented in prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (inspect) and enumerates exactly what is inspected: intent, ambiguity, conflicts, and context. The phrase 'without rewriting' implicitly separates it from the optimize/adapt siblings, though no sibling is named outright.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives both a positive trigger ('use to diagnose instructions') and an exclusion ('skip simple factual questions'), which is more than most definitions offer. It stops short of naming alternative tools such as evaluate_prompt or optimize_prompt when the diagnosis turns into scoring or rewriting.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
optimize_promptARead-only
Improve a vague or substantial task before execution. Preserves the request and adds bounded guidance. Skip trivial or exact-wording requests. Local by default; non-local mode may send redacted data to a configured provider.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | local | |
| prompt | Yes | ||
| context | No | ||
| explain | No | ||
| profile | No | balanced | |
| audience | No | auto | |
| maxTokens | No | ||
| targetAgent | No | auto | |
| disabledPasses | No | ||
| previousOptimization | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| ir | Yes | |
| changes | Yes | |
| context | Yes | |
| metadata | Yes | |
| confidence | Yes | |
| assumptions | Yes | |
| diagnostics | Yes | |
| skippedPasses | Yes | |
| tokenEstimate | Yes | |
| detectedIntent | Yes | |
| selectedPasses | Yes | |
| optimizedPrompt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so safety is covered. The description adds genuinely new behavioral context: the request is preserved, guidance is bounded, and non-local mode may send redacted data to a configured provider — an important privacy disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the primary action, then the preservation guarantee, then the skip condition and the privacy caveat. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, but a 10-parameter tool with nested objects and four enums needs far more than this description provides. Key controls (profile, audience, targetAgent, previousOptimization for re-optimization) are unaddressed, leaving an agent unable to use the tool's main levers correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 10 parameters, including nested context objects and enums like profile, audience, and targetAgent. The description only gestures at 'local vs non-local mode' and 'bounded guidance', leaving the semantics of mode, profile, audience, targetAgent, disabledPasses, and previousOptimization entirely undocumented. It does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (improve) and resource (a vague or substantial task/prompt) before execution, which is distinguishable from siblings like evaluate_prompt and inspect_prompt. It stops short of explicitly naming which sibling to use when the request is already clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-not rule ('Skip trivial or exact-wording requests') and a default posture ('Local by default'), which is more than most definitions provide. It does not name alternative siblings (e.g., evaluate_prompt, adapt_prompt) for adjacent cases, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_optimization_profileBRead-onlyIdempotent
Choose a proportional local optimization profile for a task. Use when complexity is unclear; does not select a paid model.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| mode | Yes | |
| reason | Yes | |
| profile | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, closed-world behavior, so the bar is lower. The description adds one meaningful boundary — it does not select a paid model — but says nothing about how a profile is chosen, what it contains, or how to use the result downstream.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, with the action front-loaded and the constraint second. No wasted words, though the terseness contributes to the underspecification elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and annotations cover the safety profile. The gaps are the undefined notion of a 'profile' and the undocumented prompt parameter for a tool whose whole purpose is profile selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the single 'prompt' parameter is undocumented anywhere. The description only implies via 'for a task' that the prompt is the task text, but gives no format, length, or content expectations to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Choose/suggest) and resource (a proportional local optimization profile) for a task. The qualifiers 'proportional' and 'local' plus the boundary 'does not select a paid model' partially distinguish it from optimize_prompt, though what a 'profile' concretely is remains undefined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives one explicit triggering condition, 'Use when complexity is unclear,' which is real guidance. However it names no sibling alternative (e.g., optimize_prompt, evaluate_prompt) and gives no when-not condition beyond the paid-model exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
adapt_prompt - First observed
build_agent_instruction - First observed
compare_prompts - First observed
compress_context - First observed
evaluate_prompt - First observed
explain_optimization - First observed
inspect_prompt - First observed
optimize_prompt - First observed
suggest_optimization_profile
TDQS
Scored across 9 tools
Most tools have clearly distinct purposes, and the descriptions actively draw boundaries (inspect vs evaluate vs optimize, skip trivial requests, not for scope expansion). However, the trio inspect_prompt/evaluate_prompt/optimize_prompt shares enough surface area that an agent could still hesitate over which applies to a borderline prompt.
Every tool follows a clean verb_noun snake_case pattern (inspect_prompt, optimize_prompt, compress_context, build_agent_instruction). The only deviation is the plural noun in compare_prompts, which is trivial and natural.
Nine tools is well within the ideal range and each maps to a distinct capability in the prompt-engineering lifecycle. No filler or redundant tools are apparent.
The surface covers diagnosis, optimization, adaptation, evaluation, comparison, context compression, instruction building, and audit explanation — a broad, coherent set. Minor gaps exist, such as no persistence/history or explicit application/validation step, but the server states these are intentional.
Maintenance
Related MCP Connectors
Turn PRDs and product ideas into structured specs so coding agents build your intent, not theirs.
Turns vague automation requests into tool stacks, prompts, QA checks, and human boundaries.
Pre-execution governance for AI agents. Deterministic PASS/FAIL/REVIEW verdicts, replayable proof.
JSON/YAML, regex, diff, JWT, SQL dialects — the keyless millisecond ops an agent needs mid-task.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceLocal-first memory, pipelines, learning, feedback, and safe code tools for AI coding agents.MIT
- AlicenseNot gradedqualityDmaintenanceEnables deep reasoning and cognitive enhancement through multi-agent debate, bias detection, and structured thinking, with privacy-first local execution.1MIT
- AlicenseCqualityCmaintenanceEnables local-first model routing and persistent tool library management, allowing automated workflow detection and conversion of stable patterns into reusable tools.40Mozilla Public 2.0
- AlicenseAqualityBmaintenanceEnables coding agents to compact conversation contexts verbatim, make fast decisions through choice, boolean, and rubric scoring, and enforce command safety guardrails.6MIT