Prompt Auto-Optimizer MCP
# Prompt Auto-Optimizer MCP
> **AI-Powered Prompt Evolution** - An MCP server that automatically optimizes your AI prompts using evolutionary algorithms.
[](https://www.typescriptlang.org/)
[](https://nodejs.org/)
## 🎯 Purpose
Automatically evolve and optimize AI prompts to improve performance, creativity, and reliability. Uses genetic algorithms to iteratively improve prompts based on real performance data.
## 🛠️ Installation
```bash
# Clone and install
git clone https://github.com/your-org/prompt-auto-optimizer-mcp.git
cd prompt-auto-optimizer-mcp
npm install
npm run build
# Start the MCP server
npm run mcp:start
```
## ⚙️ Configuration
Add to your Claude Code settings (`.claude/settings.json`):
```json
{
"mcp": {
"servers": {
"prompt-optimizer": {
"command": "npx",
"args": ["prompt-auto-optimizer-mcp"],
"cwd": "./path/to/prompt-auto-optimizer-mcp"
}
}
}
}
```
## 🔧 Available Tools
### Core Optimization Tools
#### `gepa_start_evolution`
Start optimizing a prompt using evolutionary algorithms.
```typescript
{
taskDescription: string; // What you want to optimize for
seedPrompt?: string; // Starting prompt (optional)
config?: {
populationSize?: number; // How many variants to test (default: 20)
generations?: number; // How many iterations (default: 10)
mutationRate?: number; // How much to change prompts (default: 0.15)
};
}
```
#### `gepa_evaluate_prompt`
Test how well a prompt performs on specific tasks.
```typescript
{
promptId: string; // Which prompt to test
taskIds: string[]; // What tasks to test it on
rolloutCount?: number; // How many times to test (default: 5)
}
```
#### `gepa_reflect`
Analyze why prompts fail and get improvement suggestions.
```typescript
{
trajectoryIds: string[]; // Which test runs to analyze
targetPromptId: string; // Which prompt needs improvement
analysisDepth?: 'shallow' | 'deep'; // How detailed (default: 'deep')
}
```
#### `gepa_get_pareto_frontier`
Get the best prompt candidates that balance multiple goals.
```typescript
{
minPerformance?: number; // Minimum quality threshold
limit?: number; // Max results to return (default: 10)
}
```
#### `gepa_select_optimal`
Choose the best prompt for your specific use case.
```typescript
{
taskContext?: string; // Describe your use case
performanceWeight?: number; // How much to prioritize accuracy (default: 0.7)
diversityWeight?: number; // How much to prioritize creativity (default: 0.3)
}
```
#### `gepa_record_trajectory`
Log the results of prompt executions for analysis.
```typescript
{
promptId: string; // Which prompt was used
taskId: string; // What task was performed
executionSteps: ExecutionStep[]; // What happened during execution
result: {
success: boolean; // Did it work?
score: number; // How well did it work?
};
}
```
### Backup & Recovery Tools
- `gepa_create_backup` - Save current optimization state
- `gepa_restore_backup` - Restore from a previous backup
- `gepa_list_backups` - Show available backups
- `gepa_recovery_status` - Check system health
- `gepa_integrity_check` - Verify data integrity
## 📝 Basic Usage
1. **Start Evolution**: Use `gepa_start_evolution` with your task description
2. **Record Results**: Use `gepa_record_trajectory` to log how prompts perform
3. **Analyze Failures**: Use `gepa_reflect` to understand what went wrong
4. **Get Best Prompts**: Use `gepa_select_optimal` to find the best candidates
## 🔧 Environment Variables
```bash
# Optional performance tuning
GEPA_MAX_CONCURRENT_PROCESSES=3 # Parallel execution limit
GEPA_DEFAULT_POPULATION_SIZE=20 # Default prompt variants
GEPA_DEFAULT_GENERATIONS=10 # Default iterations
```
---
**Built for better AI prompts** • [📚 Docs](./docs/) • [🐛 Issues](https://github.com/your-org/prompt-auto-optimizer-mcp/issues)
TDQS
Scored across 12 tools
Most tools have distinct purposes, but some overlap exists: gepa_evaluate_prompt and gepa_record_trajectory both relate to prompt evaluation, and gepa_select_optimal and gepa_get_pareto_frontier both involve selecting optimal candidates. The descriptions help clarify differences, but an agent might occasionally confuse these pairs.
All tool names follow a consistent 'gepa_verb_noun' pattern with snake_case, using clear verbs like create, evaluate, get, list, record, recover, restore, select, and start. This predictability makes it easy for agents to understand and use the toolset.
With 12 tools, the server is well-scoped for prompt auto-optimization, covering key areas like evaluation, evolution, backup, recovery, and selection. Each tool has a clear role, and the count is neither too sparse nor overwhelming for the domain.
The toolset provides comprehensive coverage for prompt optimization workflows, including evolution (start_evolution), evaluation (evaluate_prompt, record_trajectory), selection (select_optimal, get_pareto_frontier), reflection (reflect), and disaster recovery (backup, restore, integrity_check, recovery). A minor gap is the lack of tools for modifying or deleting components, but core operations are well-covered.