Skip to main content
Glama
sloth-wq

Prompt Auto-Optimizer MCP

by sloth-wq
README.md
# Prompt Auto-Optimizer MCP

> **AI-Powered Prompt Evolution** - An MCP server that automatically optimizes your AI prompts using evolutionary algorithms.

[![TypeScript](https://img.shields.io/badge/TypeScript-007ACC?style=for-the-badge&logo=typescript&logoColor=white)](https://www.typescriptlang.org/)
[![Node.js](https://img.shields.io/badge/Node.js-43853D?style=for-the-badge&logo=node.js&logoColor=white)](https://nodejs.org/)

## 🎯 Purpose

Automatically evolve and optimize AI prompts to improve performance, creativity, and reliability. Uses genetic algorithms to iteratively improve prompts based on real performance data.

## 🛠️ Installation

```bash
# Clone and install
git clone https://github.com/your-org/prompt-auto-optimizer-mcp.git
cd prompt-auto-optimizer-mcp
npm install
npm run build

# Start the MCP server
npm run mcp:start
```

## ⚙️ Configuration

Add to your Claude Code settings (`.claude/settings.json`):

```json
{
  "mcp": {
    "servers": {
      "prompt-optimizer": {
        "command": "npx",
        "args": ["prompt-auto-optimizer-mcp"],
        "cwd": "./path/to/prompt-auto-optimizer-mcp"
      }
    }
  }
}
```

## 🔧 Available Tools

### Core Optimization Tools

#### `gepa_start_evolution`

Start optimizing a prompt using evolutionary algorithms.

```typescript
{
  taskDescription: string;           // What you want to optimize for
  seedPrompt?: string;              // Starting prompt (optional)
  config?: {
    populationSize?: number;        // How many variants to test (default: 20)
    generations?: number;           // How many iterations (default: 10)
    mutationRate?: number;         // How much to change prompts (default: 0.15)
  };
}
```

#### `gepa_evaluate_prompt`

Test how well a prompt performs on specific tasks.

```typescript
{
  promptId: string;                 // Which prompt to test
  taskIds: string[];               // What tasks to test it on
  rolloutCount?: number;           // How many times to test (default: 5)
}
```

#### `gepa_reflect`

Analyze why prompts fail and get improvement suggestions.

```typescript
{
  trajectoryIds: string[];         // Which test runs to analyze
  targetPromptId: string;          // Which prompt needs improvement
  analysisDepth?: 'shallow' | 'deep'; // How detailed (default: 'deep')
}
```

#### `gepa_get_pareto_frontier`

Get the best prompt candidates that balance multiple goals.

```typescript
{
  minPerformance?: number;         // Minimum quality threshold
  limit?: number;                  // Max results to return (default: 10)
}
```

#### `gepa_select_optimal`

Choose the best prompt for your specific use case.

```typescript
{
  taskContext?: string;            // Describe your use case
  performanceWeight?: number;      // How much to prioritize accuracy (default: 0.7)
  diversityWeight?: number;        // How much to prioritize creativity (default: 0.3)
}
```

#### `gepa_record_trajectory`

Log the results of prompt executions for analysis.

```typescript
{
  promptId: string;                // Which prompt was used
  taskId: string;                  // What task was performed
  executionSteps: ExecutionStep[]; // What happened during execution
  result: {
    success: boolean;              // Did it work?
    score: number;                 // How well did it work?
  };
}
```

### Backup & Recovery Tools

- `gepa_create_backup` - Save current optimization state
- `gepa_restore_backup` - Restore from a previous backup
- `gepa_list_backups` - Show available backups
- `gepa_recovery_status` - Check system health
- `gepa_integrity_check` - Verify data integrity

## 📝 Basic Usage

1. **Start Evolution**: Use `gepa_start_evolution` with your task description
2. **Record Results**: Use `gepa_record_trajectory` to log how prompts perform
3. **Analyze Failures**: Use `gepa_reflect` to understand what went wrong
4. **Get Best Prompts**: Use `gepa_select_optimal` to find the best candidates

## 🔧 Environment Variables

```bash
# Optional performance tuning
GEPA_MAX_CONCURRENT_PROCESSES=3           # Parallel execution limit
GEPA_DEFAULT_POPULATION_SIZE=20            # Default prompt variants
GEPA_DEFAULT_GENERATIONS=10               # Default iterations
```

---

**Built for better AI prompts** • [📚 Docs](./docs/) • [🐛 Issues](https://github.com/your-org/prompt-auto-optimizer-mcp/issues)

TDQS

B3.2/5.0

Scored across 12 tools

Disambiguation4/5

Most tools have distinct purposes, but some overlap exists: gepa_evaluate_prompt and gepa_record_trajectory both relate to prompt evaluation, and gepa_select_optimal and gepa_get_pareto_frontier both involve selecting optimal candidates. The descriptions help clarify differences, but an agent might occasionally confuse these pairs.

Naming Consistency5/5

All tool names follow a consistent 'gepa_verb_noun' pattern with snake_case, using clear verbs like create, evaluate, get, list, record, recover, restore, select, and start. This predictability makes it easy for agents to understand and use the toolset.

Tool Count5/5

With 12 tools, the server is well-scoped for prompt auto-optimization, covering key areas like evaluation, evolution, backup, recovery, and selection. Each tool has a clear role, and the count is neither too sparse nor overwhelming for the domain.

Completeness4/5

The toolset provides comprehensive coverage for prompt optimization workflows, including evolution (start_evolution), evaluation (evaluate_prompt, record_trajectory), selection (select_optimal, get_pareto_frontier), reflection (reflect), and disaster recovery (backup, restore, integrity_check, recovery). A minor gap is the lack of tools for modifying or deleting components, but core operations are well-covered.

Maintenance

ActivityInactive
ResponsivenessNo issues