PromptTuner MCP
Provides direct API integration for prompt refinement using Google's Gemini models (2.0 Flash, 1.5 Pro), enabling LLM-powered prompt optimization and analysis.
Provides direct API integration for prompt refinement using OpenAI models (GPT-4o, GPT-4o-mini, GPT-4 Turbo), enabling LLM-powered prompt optimization and analysis.
Provides session store support for distributed multi-instance deployments and caching of refined prompts to improve performance.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@PromptTuner MCPrefine this prompt to make it clearer and more effective"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
PromptTuner MCP
PromptTuner MCP is an MCP server that fixes and boosts prompts using OpenAI, Anthropic, or Google Gemini.
What it does
Validates and trims input prompts (enforces
MAX_PROMPT_LENGTH).Wraps the prompt as JSON inside sentinel markers (sanitizing markers, bidi control chars, and null bytes).
Calls the selected provider.
Normalizes LLM output (strips code fences / labels if present).
Returns human-readable text plus machine-friendly
structuredContent.
Related MCP server: PromptArchitect MCP
Features
Polish and refine a prompt for clarity and flow (
fix_prompt).Boost and enhance a prompt for clarity and effectiveness (
boost_prompt).Craft a reusable workflow prompt for complex tasks (
crafting_prompt).Simple structured outputs.
Retry logic with exponential backoff for transient provider failures.
Quick Start
PromptTuner runs over stdio only. The dev:http and start:http scripts are compatibility aliases (no HTTP transport yet).
Claude Desktop
Add to claude_desktop_config.json:
{
"mcpServers": {
"prompttuner": {
"command": "npx",
"args": ["-y", "@j0hanz/prompt-tuner-mcp-server@latest"],
"env": {
"LLM_PROVIDER": "openai",
"OPENAI_API_KEY": "sk-..."
}
}
}
}Replace the API key and provider with your preferred LLM. Only configure the key for the active provider.
Configuration
PromptTuner uses minimal configuration. Set the provider and API key, and you're ready to go.
Variable | Default | Description |
|
|
|
| - | Required for all tools when |
| - | Required for all tools when |
| - | Required for all tools when |
| - | Override the default model. |
|
| Enable debug logging. |
All tools are LLM-backed and require an API key for the selected provider.
Default Models
Provider | Default Model |
|
|
|
|
|
|
CLI Options
Flag | Description |
| Show help text. |
| Print version. |
| Enable/disable debug logging. |
|
|
| Override the default model. |
Tools
All tools accept plain text, Markdown, or XML prompts. Responses include content (human-readable) and structuredContent (machine-readable).
Inputs are strict: extra fields are rejected. For fix_prompt/boost_prompt, only the prompt field is accepted.
fix_prompt
Polish and refine a prompt for clarity and flow while preserving intent and structure.
Parameter | Type | Required | Notes |
| string | Yes | Trimmed, length-checked; extra fields rejected. |
Returns: ok, fixed.
boost_prompt
Refine and enhance a prompt for clarity and effectiveness.
Parameter | Type | Required | Notes |
| string | Yes | Trimmed, length-checked; extra fields rejected. |
Returns: ok, boosted.
crafting_prompt
Generate a structured, reusable workflow prompt for complex tasks based on a raw request and a few settings.
Parameter | Type | Required | Notes |
| string | Yes | Trimmed, length-checked; strict input. |
| string | No | Hard requirements to enforce (bullet list recommended). |
| string | No |
|
| string | No |
|
| string | No |
|
| string | No |
|
Returns: ok, prompt, settings.
Response Format
content: array of content blocks. First block is JSON forstructuredContent, second is a short human message (orError: ...).structuredContent: machine-parseable results.Errors return
structuredContent.ok=falseand anerrorobject withcode,message, optionalcontext(sanitized, up to 200 chars),details, andrecoveryHint.Error responses also include
isError: true.
Development
Prerequisites
Node.js >= 22.0.0
npm
Scripts
Command | Description |
| Compile TypeScript and set permissions. |
| Build on install (publishing helper). |
| Run from source in watch mode. |
| Alias of |
| TypeScript compiler in watch mode. |
| Run the compiled server from |
| Alias of |
| Run |
| Run |
| Run |
| Run ESLint. |
| Run Prettier. |
| TypeScript type checking. |
| Run MCP Inspector against |
| Alias of |
| Run jscpd duplication report. |
| Lint, type-check, and build before publish. |
Project Structure
src/
index.ts Entry point
cli.ts CLI parsing, logging bootstrap, shutdown handling
server.ts MCP server setup (stdio transport)
tools.ts Tool implementations
schemas.ts Zod input/output schemas
config.ts Configuration and constants
types.ts Shared types and error codes
lib/ Shared utilities (LLM, retry, telemetry, prompt utils)
tests/ node:test suites
dist/ Compiled output (generated)
docs/ Static assetsSecurity
API keys are supplied only via environment variables.
Inputs are validated with Zod and additional length checks.
Error context is included in debug mode (sanitized and truncated to 200 chars).
Google safety filters are always enabled.
Contributing
Pull requests are welcome. Please include a short summary, tests run, and note any configuration changes.
License
MIT License. See LICENSE for details.
Available Tools
3 toolsboost_promptBoost PromptBRead-only
Transform a prompt using prompt engineering best practices for maximum clarity and effectiveness.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Prompt to transform and optimize |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate readOnlyHint=true (safe operation), openWorldHint=true (handles diverse inputs), and idempotentHint=false (non-idempotent). The description adds value by specifying the transformation is for 'clarity and effectiveness' and involves 'prompt engineering best practices', which provides behavioral context beyond annotations. However, it doesn't detail aspects like rate limits, error handling, or output format, keeping the score moderate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core action ('Transform a prompt') and adds necessary context without waste. Every word contributes to understanding the tool's purpose, making it appropriately sized and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (transformation operation), annotations cover safety and input handling, but there's no output schema to explain return values. The description adequately states the purpose but lacks details on usage guidelines, behavioral nuances like transformation specifics, or how it differs from siblings, leaving gaps in completeness for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with the 'prompt' parameter well-documented as 'Prompt to transform and optimize'. The description adds marginal meaning by implying optimization for 'clarity and effectiveness', but it doesn't provide additional syntax, examples, or constraints beyond the schema. Baseline 3 is appropriate given high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Transform') and resource ('prompt'), and it adds context about 'prompt engineering best practices' and goals like 'maximum clarity and effectiveness'. However, it doesn't explicitly differentiate from sibling tools like 'crafting_prompt' or 'fix_prompt', which might have overlapping or distinct functions, preventing a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as 'crafting_prompt' or 'fix_prompt'. It implies usage for optimizing prompts but lacks explicit when/when-not scenarios, prerequisites, or comparisons to siblings, leaving the agent with minimal context for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
crafting_promptCrafting PromptBRead-only
Generate a structured, reusable workflow prompt for complex tasks based on a raw request and a few settings.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | Raw user request / task description to turn into a workflow prompt | |
| constraints | No | Optional: hard requirements to enforce (e.g., no breaking changes) | |
| mode | No | general | |
| approach | No | balanced | |
| tone | No | direct | |
| verbosity | No | normal |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate read-only and open-world behavior, which the description doesn't contradict, but it adds no behavioral context beyond that—no details on rate limits, authentication needs, or output characteristics. With annotations covering safety, a 3 reflects minimal added value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose without unnecessary details, though it could be slightly more structured by explicitly listing key parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity with 6 parameters, low schema coverage, no output schema, and annotations that only cover safety, the description is incomplete—it lacks details on parameter meanings, output format, and usage scenarios, making it inadequate for full understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is low at 33%, with only 'request' and 'constraints' described in the schema. The description mentions 'a few settings' but doesn't explain parameters like 'mode,' 'approach,' 'tone,' or 'verbosity,' failing to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Generate') and resource ('structured, reusable workflow prompt for complex tasks'), distinguishing it from sibling tools like 'boost_prompt' and 'fix_prompt' by focusing on creating prompts from raw requests rather than enhancing or repairing existing ones.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by mentioning 'based on a raw request and a few settings,' suggesting it's for turning user input into prompts, but it lacks explicit guidance on when to use this tool versus alternatives like 'boost_prompt' or 'fix_prompt,' or any context-specific exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fix_promptFix PromptBRead-only
Polish and refine a prompt for better clarity, readability, and flow.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Prompt to polish and refine |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate readOnlyHint=true (safe read operation), openWorldHint=true (broad applicability), and idempotentHint=false (non-idempotent). The description adds context by specifying the refinement goals (clarity, readability, flow), which goes beyond the annotations. However, it does not disclose other behavioral traits like potential side effects, rate limits, or detailed output expectations, keeping the score moderate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core action ('polish and refine') and purpose. It avoids redundancy and wastes no words, making it highly concise and well-structured for quick understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (single parameter, no output schema), the description is adequate but has gaps. It explains what the tool does but lacks details on when to use it versus siblings, output format, or error handling. With annotations covering safety and scope, it meets minimum viability but isn't fully comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, fully documenting the single parameter 'prompt'. The description does not add any parameter-specific semantics beyond what the schema provides (e.g., it doesn't explain format or constraints). With high schema coverage, the baseline score is 3, as the description doesn't compensate but doesn't need to.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('polish and refine') and the resource ('a prompt'), and it specifies the improvement goals ('better clarity, readability, and flow'). However, it does not explicitly distinguish this tool from its siblings (boost_prompt, crafting_prompt), which would be needed for a score of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus its siblings (boost_prompt, crafting_prompt) or any alternatives. It lacks explicit instructions on context, prerequisites, or exclusions, offering only a general purpose without usage differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
The three tools have overlapping purposes in prompt improvement, with 'boost_prompt' and 'fix_prompt' both focusing on clarity and effectiveness, which could cause confusion. However, 'crafting_prompt' is more distinct as it generates structured workflows, providing some differentiation.
All tool names follow a consistent verb_noun pattern with clear, descriptive verbs ('boost', 'crafting', 'fix') and the same noun ('prompt'), making them predictable and easy to understand.
With only 3 tools, the server feels thin for a domain like prompt tuning, which might involve more operations such as evaluating prompts, testing variations, or managing prompt libraries. This limited set could restrict agent capabilities.
The toolset is incomplete for prompt tuning, missing essential operations like evaluating prompt effectiveness, comparing different versions, or storing/retrieving prompts. This creates gaps that could lead to agent failures in comprehensive prompt management tasks.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
PQS scores any prompt before the model runs. 8 dimensions. 5 frameworks. Pre-flight, not post-hoc.
Turns rough requests into sharp Role/Task/Context/Format prompts. Thai and English.
The Wikipedia of AI prompts: search 900+ curated prompts by model, style and type, in 7 languages
Generate tailored quality criteria and scoring guides from your task descriptions. Refine objectiv…
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnhances and cleans raw prompts using AI to make them more clear, actionable, and effective. Provides quality assessment, suggestions, and supports both general and code-specific optimization modes.1MIT
- FlicenseAqualityCmaintenanceRefines and improves AI prompts using workspace-aware context from your project's tech stack, structure, and dependencies. Includes tools to analyze prompt quality and generate well-structured prompts from raw ideas.42095
- AlicenseNot gradedqualityCmaintenanceRefines and optimizes prompts for LLMs through adaptive questioning and intelligent clarification workflows. Supports multiple AI providers (Google, OpenAI, Anthropic, Groq, Qwen) with interactive prompt enhancement and targeted modifications.17MIT
- AlicenseBqualityDmaintenanceAutomatically analyzes and optimizes AI prompts by calculating clarity scores, detecting risks, asking clarifying questions, and adding domain-specific requirements to improve AI interaction quality.1MIT
Appeared in Searches
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/j0hanz/prompt-tuner-mcp-server'
If you have feedback or need assistance with the MCP directory API, please join our Discord server