SIMPA
Integrates with Google Gemini API for prompt refinement and embeddings.
Integrates with locally running Ollama models (e.g., llama3.2, nomic-embed-text) for prompt refinement and embeddings.
Integrates with OpenAI API to use GPT models for prompt refinement and embeddings during the self-improvement process.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@SIMPArefine this prompt: 'explain quantum computing to a child'"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
SIMPA - Self-Improving Meta Prompt Agent
π Transform your AI agents with self-optimizing prompt intelligence
SIMPA is a Model Context Protocol (MCP) service that learns from every interaction to continuously improve prompt quality. It remembers what worked, refines what didn't, and automatically selects the best prompts for any situation.
π Why SIMPA?
Every agent you deploy faces the same challenge: getting the prompt right. SIMPA solves this by:
π Learning from feedback - Automatically improves based on execution scores
π Semantic search - Finds similar successful prompts using vector similarity
π§ Smart selection - Chooses between refinement and reuse based on proven performance
π MCP Native - Seamlessly integrates with any MCP-compatible agent controller
Related MCP server: Engram MCP
ποΈ Architecture
flowchart TB
subgraph Controller["Agent Controller"]
A[Agent Request]
end
subgraph SIMPA["SIMPA MCP Service"]
direction TB
R[Refiner] --> S[Selector]
S --> V[Vector Store]
S --> L[LLM Service]
L --> E[Embedding Service]
end
subgraph Storage["Knowledge Base"]
direction TB
P[(PostgreSQL + pgvector)]
H[Prompt History]
end
A -->|original_prompt| R
V -->|similar_prompts| S
S -->|refined_prompt| A
S -->|store & learn| P
P -->|usage_stats| S
H -->|feedback_loop| S
style Controller fill:#e1f5fe
style SIMPA fill:#fff3e0
style Storage fill:#e8f5e9π Prompt Lifecycle
SIMPA sits between the Agent Orchestrator and Implementation Agents, continuously learning from each interaction:
flowchart LR
AO[Agent Orchestrator] -->|prompt| SR[SIMPA Prompt<br/>Refinement]
SR -->|refined-prompt| IA[Implementation<br/>Agent]
IA -->|Actions, Results<br/>& Products| RA[Reviewing Agent]
RA -->|refined-prompt-score| SR2[SIMPA]
SR2 -->|learn & improve| SR
style AO fill:#e1f5fe,color:#000000
style SR fill:#fff3e0,color:#000000
style IA fill:#fce4ec,color:#000000
style RA fill:#f3e5f5,color:#000000
style SR2 fill:#fff3e0,color:#000000The Flow:
Agent Orchestrator β Sends raw
promptto SIMPASIMPA β Returns
refined-prompt(structured with ROLE, GOAL, REQUIREMENTS)Implementation Agent β Executes actions using refined prompt, produces results/products
Reviewing Agent β Evaluates outcomes, generates
refined-prompt-scoreSIMPA β Receives score, learns what works, improves future refinements
This closed feedback loop ensures prompts get better with every execution.
β¨ Features
Feature | Description |
π€ MCP Protocol | Native Model Context Protocol support for universal agent integration |
π Vector Search | pgvector-powered similarity search for prompt retrieval |
π Self-Improvement | Sigmoid-based probability for intelligent refinement vs reuse |
π― Multi-Provider | OpenAI, Anthropic, and Ollama support for embeddings and LLM |
π Observability | Structured logging with structlog and comprehensive metrics |
π‘οΈ Security | PII detection and input validation built-in |
π§ͺ Tested | 274 automated tests with 100% pass rate |
π Prerequisites
Before installing SIMPA, ensure you have the following:
Required
Component | Version | Purpose |
Python | 3.10+ | Runtime environment |
PostgreSQL | 14+ | Database with pgvector extension |
Docker | Latest | Required for running tests with TestContainers |
Git | Latest | Clone repository |
Note: PostgreSQL and Ollama are expected to be installed and running separately (not via Docker) for normal operation. Docker is only required for the automated test suite.
For Ollama (Local Models - Recommended)
Component | Purpose |
Ollama | Local LLM & embedding inference |
nomic-embed-text | Embedding model (pull via |
llama3.2 | LLM for prompt refinement (pull via |
For Cloud Providers (Optional)
π Security Best Practice: Provider API keys (OpenAI, Anthropic, Google, Azure) should be kept in your user home directory at
~/.envrather than in the project.envfile. This prevents accidental commits of sensitive credentials to version control.Create
~/.envwith your provider keys:# ~/.env - User-level secrets (not committed) OPENAI_API_KEY=sk-... ANTHROPIC_API_KEY=sk-ant-... GOOGLE_API_KEY=... AZURE_OPENAI_KEY=...SIMPA will automatically load keys from
~/.envif available.
OpenAI API Key - Get from platform.openai.com
Anthropic API Key - Get from console.anthropic.com
Azure OpenAI - Azure subscription with OpenAI service
Google Gemini - Get from Google AI Studio
System Requirements
Resource | Minimum | Recommended |
RAM | 4 GB | 8 GB+ |
Disk | 2 GB free | 10 GB+ |
CPU | 2 cores | 4 cores+ |
Note: For local Ollama models, CPU is sufficient but GPU acceleration significantly improves performance.
π Quick Start
Option 1: Docker Compose (Recommended for Development/Testing)
This option runs PostgreSQL and Ollama in Docker containers for easy development and testing:
# Clone and setup
git clone https://github.com/yourusername/simpa-mcp.git
cd simpa-mcp
cp .env.example .env
# Start all services (PostgreSQL + Ollama in Docker)
make dev-setup
# Download models (one-time)
make pull-models
# Run migrations
make migrate
# Run tests
make testFor Production Use: Install PostgreSQL and Ollama directly on your system instead of using Docker. See the Manual Setup section below.
Option 2: Manual Setup (Production/Existing Services)
Use this if you already have PostgreSQL and Ollama installed locally.
Prerequisites:
PostgreSQL 14+ with pgvector extension installed
Ollama running locally (with
nomic-embed-textandllama3.2pulled)
# Install dependencies
pip install -e ".[dev]"
# Configure environment
cp .env.example .env
# Edit .env to match your PostgreSQL and Ollama settings
# Run migrations
alembic upgrade head
# Start MCP server
python -m src.mainQuick PostgreSQL setup with Docker (if needed):
# Only if you don't have PostgreSQL installed locally
docker run -d --name simpa-db \
-e POSTGRES_USER=simpa \
-e POSTGRES_PASSWORD=simpa \
-e POSTGRES_DB=simpa \
-p 5432:5432 \
pgvector/pgvector:pg16π Adding SIMPA to Your MCP Configuration
SIMPA works with any MCP-compatible client (Cursor, Claude Desktop, Windsurf, etc.).
Step 1: Install SIMPA Server
Option A: Global Installation (Easiest)
# Clone the repository
git clone https://github.com/dsidlo/simpa-mcp.git
cd simpa-mcp
# Create virtual environment
python -m venv .venv
# Activate virtual environment
# On macOS/Linux:
source .venv/bin/activate
# On Windows:
# .venv\Scripts\activate
# Install in editable mode
pip install -e .
# Install MCP dependencies
pip install fastmcp asyncpg pgvector sqlalchemy
# Setup environment
cp .env.example .env
# Edit .env with your configuration (see Configuration section below)
# Run database migrations
alembic upgrade headOption B: Docker (Recommended for Production)
# Build the MCP server image
docker build --target production -t simpa-mcp:latest .
# Or use docker compose (includes PostgreSQL + pgvector)
docker-compose up -dStep 2: Configure Your MCP Client
Add SIMPA to your MCP client's configuration file:
Cursor (~/.cursor/mcp.json)
{
"mcpServers": {
"simpa-mcp": {
"command": "uv",
"args": [
"--directory",
"/absolute/path/to/simpa-mcp",
"run",
"--env",
"/absolute/path/to/simpa-mcp/.env",
"python",
"-m",
"src.main"
],
"env": {
"PYTHONPATH": "/absolute/path/to/simpa-mcp/src"
}
}
}
}Claude Desktop (~/Library/Application Support/Claude/claude_desktop_config.json)
{
"mcpServers": {
"simpa-mcp": {
"command": "/absolute/path/to/simpa-mcp/.venv/bin/python",
"args": [
"-m",
"src.main"
],
"env": {
"DATABASE_URL": "postgresql://simpa:simpa@localhost:5432/simpa",
"EMBEDDING_PROVIDER": "ollama",
"EMBEDDING_MODEL": "nomic-embed-text",
"OLLAMA_BASE_URL": "http://localhost:11434",
"LLM_MODEL": "ollama/llama3.2",
"PYTHONPATH": "/absolute/path/to/simpa-mcp/src"
}
}
}
}Generic MCP Configuration
{
"mcpServers": {
"simpa-mcp": {
"name": "SIMPA Prompt Refinement",
"description": "Self-improving prompt optimization service",
"command": "python",
"args": [
"-m",
"src.main",
"--mcp",
"stdio"
],
"workingDirectory": "/absolute/path/to/simpa-mcp",
"envFile": "/absolute/path/to/simpa-mcp/.env"
}
}
}Using uv (Recommended)
This configuration ensures the server runs from the source directory and uses uv for dependency management:
{
"mcpServers": {
"simpa-mcp": {
"command": "/bin/bash",
"args": [
"-c",
"cd /path/to/simpa-mcp && uv run python src/main.py --log-level debug --log-file /tmp/simpa-mcp.log"
]
}
}
}Note: Replace
/path/to/simpa-mcpwith your actual installation path. Usingbash -cwithcdensures the server runs from the project root wherepyproject.tomland.envare located.
Step 3: Install MCP Inspector (Optional, for Testing)
# Install MCP Inspector globally
npm install -g @anthropics/mcp-inspector
# Test your SIMPA server
mcp-inspector --server "uv --directory /path/to/simpa-mcp run python -m src.main"Step 4: Verify Installation
In your MCP client (Cursor/Claude Desktop), you should see:
β Available Tools:
refine_prompt,update_prompt_resultsβ Server Status: Connected
β Capabilities: Prompt refinement enabled
π οΈ Troubleshooting
"Command not found: uv"
Install uv first:
curl -LsSf https://astral.sh/uv/install.sh | sh"ModuleNotFoundError: No module named 'src'"
Ensure PYTHONPATH includes the src directory:
export PYTHONPATH="/absolute/path/to/simpa-mcp/src:$PYTHONPATH"Database Connection Errors
Verify PostgreSQL is running with pgvector:
# Check if pgvector extension is available
psql -d simpa -c "CREATE EXTENSION IF NOT EXISTS vector;"MCP Server Not Responding
Test manually:
cd /path/to/simpa-mcp
source .venv/bin/activate
python -m src.main --helpπ§ Configuration
SIMPA can be configured via environment variables and command-line arguments.
Environment Variables
How configuration works: SIMPA uses Pydantic Settings to automatically load environment variables from
.envfiles. When you set an environment variable, it automatically becomes available viasettings.VARIABLE_NAMEin the codeβno explicitos.getenv()calls needed. Environment variables are case-insensitive (EMBEDDING_MODELandembedding_modelwork the same).
β‘ Critical Parameters (Required)
These parameters must be configured to bring up the MCP service:
Variable | Description | Why Required |
| PostgreSQL connection URL | Stores prompt knowledge base |
| OpenAI API key | Required only if using OpenAI models. Other providers need their respective keys. |
All other parameters can be left undefined β they default to known, usable values suitable for most deployments.
Minimal Configuration Example
The simplest working .env file (using local Ollama models):
# Only REQUIRED parameter - everything else defaults automatically
DATABASE_URL=postgresql://user@localhost:5432/simpaFor OpenAI instead of Ollama, just add the API key:
# Required
DATABASE_URL=postgresql://user@localhost:5432/simpa
OPENAI_API_KEY=sk-your-key-here
# LLM_MODEL defaults to ollama/llama3.2, but you can override:
# LLM_MODEL=openai/gpt-4Optional Parameters (With Working Defaults)
All sections below have sensible defaults. You only need to change them if you have specific requirements:
Database Connection Details
Variable | Description | Default |
| PostgreSQL connection URL |
|
Embedding Service
Variable | Description | Default |
| Embedding provider ( |
|
| Embedding model name |
|
| Vector dimensions (768 for nomic-embed-text) |
|
| Ollama API base URL |
|
LLM Service
Variable | Description | Default |
| LLM model (LiteLLM format: |
|
| Sampling temperature (0.0 - 2.0) |
|
Supported Models (via LiteLLM):
ollama/llama3.2- Local Ollama modelsgpt-4,gpt-3.5-turbo- OpenAIclaude-3-opus-20240229,claude-3-sonnet-20240229- Anthropicgemini/gemini-pro,gemini/gemini-ultra- Googleazure/<deployment-name>- Azure OpenAI
API Keys (Only if using cloud LLM providers)
Only needed if you use cloud-based LLMs instead of local Ollama models. These are loaded automatically by LiteLLM based on model prefix:
Variable | Description |
| OpenAI API key |
| Anthropic API key |
| Google Gemini API key |
| Azure OpenAI API key |
| Azure OpenAI endpoint base URL |
| Cohere API key |
Embedding Cache
Variable | Description | Default |
| Enable LRU cache for embeddings |
|
| Maximum cache entries (100-10000) |
|
| Maximum text length to cache |
|
LLM Cache
Variable | Description | Default |
| Enable LLM response caching |
|
| Cache TTL in seconds (60-86400) |
|
| Maximum cache entries (100-100000) |
|
| Path to cache SQLite database |
|
Fast-Path Hash Match
Variable | Description | Default |
| Enable hash-based exact match lookup |
|
| Minimum score for hash reuse (1.0-5.0) |
|
Refinement Strategy
Variable | Description | Default |
| Cosine similarity threshold for bypass (0.9-1.0) |
|
| Minimum score for high-similarity bypass (1.0-5.0) |
|
| Sigmoid steepness parameter |
|
| Sigmoid midpoint (50% threshold) |
|
| Minimum refinement probability (0.0-1.0) |
|
Vector Search
Variable | Description | Default |
| Number of similar prompts to retrieve (1-50) |
|
| Minimum similarity score (0.0-1.0) |
|
BM25 Hybrid Search
Variable | Description | Default |
| Enable BM25 keyword search |
|
| BM25 term saturation parameter (0.1-3.0) |
|
| BM25 document length normalization (0.0-1.0) |
|
| Number of BM25 results (1-20) |
|
| Number of vector results in hybrid (1-20) |
|
| Enable hybrid search combining vector + BM25 |
|
| Enable LLM re-ranking of results |
|
| Number of candidates for re-ranking (2-20) |
|
MCP Server
Variable | Description | Default |
| Transport protocol ( |
|
| Server port for SSE transport (1024-65535) |
|
Logging
Variable | Description | Default |
| Logging level ( |
|
| Enable structured JSON logging |
|
Security
Variable | Description | Default |
| Maximum prompt text length (100-100000) |
|
| Enable basic PII detection |
|
Project Association
Variable | Description | Default |
| Require project_id for all refinements |
|
Diff Saliency
Variable | Description | Default |
| Enable diff saliency filtering |
|
| Minimum saliency score (0.0-1.0) |
|
| Maximum diffs per request (1-50) |
|
Complete Example .env File
# Database (Required)
DATABASE_URL=postgresql://simpa:simpa@localhost:5432/simpa
# Embedding Service
EMBEDDING_PROVIDER=ollama
EMBEDDING_MODEL=nomic-embed-text
EMBEDDING_DIMENSIONS=768
OLLAMA_BASE_URL=http://localhost:11434
# LLM Service
LLM_MODEL=ollama/llama3.2
LLM_TEMPERATURE=0.7
# API Keys (if using cloud providers)
# OPENAI_API_KEY=sk-...
# ANTHROPIC_API_KEY=sk-ant-...
# GEMINI_API_KEY=...
# MCP Server
MCP_TRANSPORT=stdio
MCP_PORT=8000
# Caching (optional, defaults are reasonable)
EMBEDDING_CACHE_ENABLED=true
LLM_CACHE_ENABLED=true
# Refinement behavior (optional)
SIMILARITY_BYPASS_THRESHOLD=0.95
SIGMOID_K=1.5
SIGMOID_MU=3.0
# Logging
LOG_LEVEL=INFO
JSON_LOGGING=trueCommand Line Options
SIMPA supports several command line flags for runtime configuration:
# Show all available options
python -m src.main --helpOption | Description | Default |
| Initialize the database schema and exit | - |
| MCP transport protocol |
|
| Logging level |
|
| Path to log file |
|
| Also log to console (stderr) β οΈ Not recommended for MCP stdio mode | - |
| Path to |
|
| Require | - |
Examples:
# Initialize database
python -m src.main --init-db
# Run with SSE transport on custom port (also set MCP_PORT in .env)
python -m src.main --transport sse
# Debug logging to custom file
python -m src.main --log-level debug --log-file /var/log/simpa.log
# Use custom env file
python -m src.main --env ~/my-project/.env
# Require project_id for all prompts
python -m src.main --project-id-required
# Combination of options
python -m src.main --env ./.env.local --log-level debug --transport sseEnvironment File (--env)
By default, SIMPA loads environment variables from your home directory at ~/.env. You can specify a custom .env file using the --env option:
# Use a custom env file
python -m src.main --env ~/my-project/.env
# Or use a project-specific .env
python -m src.main --env ./.env.localLoading order:
If
--envis specified and the file exists, it is loaded firstIf
--envis not specified,~/.envis loaded if it existsThe project
./.envin the current directory is loaded last (overrides previous values)
This allows you to keep sensitive credentials (API keys) in ~/.env while keeping project-specific settings in the project .env.
Project-Associated Prompt Development (--project-id-required)
Enable strict project association mode to enforce that all prompts must be linked to a project:
# Require project_id for all prompt refinements
python -m src.main --project-id-requiredWhen enabled, calling refine_prompt without a project_id returns a helpful response guiding the agent to:
List existing projects - View available projects to find a suitable match
Create a new project - Use
create_projectif no suitable project existsResubmit with project_id - Retry the refinement with the chosen project
Why use project association?
Cross-project learning: Prompts refined for one Python web project can benefit similar Flask/Django projects
Knowledge clustering: Projects with similar tech stacks (React+Node, Python+PostgreSQL) share prompt patterns
Relevance scoring: Prompt selection considers project context for better matches
Team organization: Different teams/projects have distinct prompt preferences and patterns
Example workflow:
# Start server with strict project mode
python -m src.main --project-id-required
# Agent workflow:
# 1. First call without project_id β returns list of existing projects
# 2. Agent picks or creates project β gets project_id
# 3. Resubmit with project_id β prompt is refined and associated with projectπ οΈ MCP Tools
refine_prompt
Intelligently refine prompts before agent execution.
# Request
{
"original_prompt": "Write a function to sort a list",
"agent_type": "developer",
"main_language": "python"
}
# Response
{
"refined_prompt": "Write a Python function that takes a list of integers...",
"prompt_key": "uuid-v4",
"action": "refine|new|reuse",
"confidence_score": 0.95,
"similar_prompts_found": 3
}update_prompt_results
Provide feedback to improve future prompts.
# Request
{
"prompt_key": "uuid-v4",
"action_score": 4.5,
"test_passed": true,
"files_modified": ["main.py"],
"lint_score": 0.95
}
# Response
{
"success": true,
"usage_count": 5,
"average_score": 4.25
}π Prompt Refinement Examples
SIMPA transforms vague user requests into structured, actionable specifications.
Example 1: Developer Agent
Original Prompt:
Build a REST API for managing tasks.Refined Prompt:
ROLE: Senior Backend Developer
GOAL: Build a REST API for managing tasks.
CONSTRAINTS: Your output will be only a descriptive overview of what the API will do.
REQUIREMENTS:
- Define all REST endpoints (GET, POST, PUT, DELETE) with their URLs and purposes
- Explicitly specify request/response JSON formats for each endpoint
- Include pagination, filtering, and sorting capabilities for task listing
- Describe authentication mechanism (JWT or API key based)
- Define error response formats and standard HTTP status codes
- Outline rate limiting strategy if applicable
OUTPUT: Structured API specification document.
SUCCESS: All endpoints are well-defined, formats are validated, and specification is ready to implement.
AUTONOMY: You may choose implementation technology and framework.
FALLBACK: If a task requirement is unclear, ask for clarification or document your assumptions.Example 2: Architect Agent
Original Prompt:
Design a caching strategy for the product catalog.Refined Prompt:
ROLE: Software Architect
GOAL: Design a caching strategy for the product catalog.
CONSTRAINTS: Output will be descriptive only; no actual implementation code.
REQUIREMENTS:
- Identify data access patterns (read-heavy, write-heavy, or mixed)
- Design cache-invalidation strategy (time-based, event-driven, or hybrid)
- Define cache key structure and naming conventions
- Specify cache-aside, write-through, or write-behind patterns
- Address cache consistency, eviction policies, and TTL configuration
- Consider multi-tier caching (in-memory + distributed)
- Document potential issues: stale data, thundering herd, cache penetration
OUTPUT: Architecture document detailing caching strategy.
SUCCESS: Design addresses scalability, performance, and data consistency.
AUTONOMY: Choose appropriate caching technology (Redis, Memcached, Caffeine).
FALLBACK: If data volume or access patterns are unclear, document assumptions.Example 3: Tester Agent
Original Prompt:
Write tests for the login module.Refined Prompt:
ROLE: QA Engineer
GOAL: Write tests for the login module.
CONSTRAINTS: Provide complete test scenarios and detailed expected outcomes, no code.
TEST CATEGORIES:
- Unit Tests: Each function/method tested in isolation with mocks
- Integration Tests: Service interactions (DB, external auth providers)
- Edge Cases: Empty input, SQL injection attempts, XSS payloads
- Boundary Tests: Max username length, password complexity rules
- Error Paths: Invalid credentials, account locked, service unavailable
- Load Tests: Concurrent login attempts
REQUIREMENTS:
- Use GIVEN/WHEN/THEN format for each scenario
- Define specific test data (usernames, passwords)
- Specify expected status codes and error messages
- Identify test dependencies and setup requirements
OUTPUT: Comprehensive test specification document.
SUCCESS: All happy paths and edge cases are covered.
AUTONOMY: Choose testing frameworks and assertion styles.
FALLBACK: If authentication flow is unclear, document assumptions.Example 4: Reviewer Agent (with scope constraints)
Original Prompt:
Review error handling in the codebase.Scope Context:
{
"target_dirs": ["src/services/", "src/utils/"],
"target_files": ["src/services/user_service.py"],
"scope": "error handling patterns only",
"focus": ["security", "logging", "user-experience"]
}Refined Prompt:
ROLE: Security-Focused Code Reviewer
GOAL: Review error handling patterns in the codebase.
CONSTRAINTS: Review only within assigned scope and files:
- target_dirs: src/services/, src/utils/
- target_files: src/services/user_service.py
- focus: security, logging, user-experience
- scope: error handling patterns only
CONTEXT: Production code review process
OUTPUT: Line-by-line comments and summary report
SUCCESS: Critical issues identified, recommendations actionable
AUTONOMY: Can use static analysis tools within scope
FALLBACK: Ask if scope unclear
Review Checklist:
- Security: Exception leaks sensitive data, proper sanitization
- Logging: Appropriate log levels, no PII exposure
- User Experience: Helpful error messages, graceful degradation
- Code Quality: Consistent patterns, avoid catch-all exceptions
- Documentation: Error scenarios documented, recovery paths clearNote: When scope context is provided (target_dirs, target_files, scope, focus), SIMPA injects these constraints into the refined prompt above the CONSTRAINTS section, limiting the agent's work to the specified boundaries.
π§ Self-Improvement Algorithm
SIMPA uses a sigmoid function to intelligently balance exploration (refinement) vs exploitation (reuse):
p_refine(S) = 1 / (1 + exp(k * (S - mu)))Where:
S= Average score (1.0 - 5.0)k= Steepness (default: 1.5)mu= Midpoint (default: 3.0)
Refinement Probability:
Score | Probability |
β 1.0 | ~95% π Refine heavily |
ββ 2.0 | ~82% π Likely refine |
βββ 3.0 | ~50% βοΈ Balance point |
ββββ 4.0 | ~18% β Start reusing |
βββββ 5.0 | ~5% β Reuse proven |
π Database Schema
refined_prompts - The Prompt Knowledge Base
Column | Type | Purpose |
| UUID | Primary key |
| UUID | Public identifier for MCP tools |
| TIMESTAMP | When prompt was first refined |
| TIMESTAMP | Last modification time |
| TIMESTAMP | Last time this prompt was executed |
| vector(768) | Semantic embedding for similarity search |
| VARCHAR(100) | Agent specialization (e.g., "developer") |
| VARCHAR(20) | Strategy used (default: "sigmoid") |
| VARCHAR(50) | Primary programming language |
| JSON | Additional languages used |
| VARCHAR(100) | Domain/topic classification |
| JSON | Array of descriptive tags |
| VARCHAR(64) | Hash for fast exact-match lookup |
| TEXT | Raw input prompt |
| TEXT | Optimized/expanded version |
| INTEGER | Version number for iterative refinements |
| UUID | Self-reference for refinement chains |
| UUID | FK to projects (optional context) |
| INTEGER | Total times used |
| FLOAT | Running average of action scores (1.0-5.0) |
| FLOAT | Bayesian-weighted score for ranking |
| JSON | Scope context (focus, target_dirs, etc.) |
| BOOLEAN | Soft delete flag |
projects - Project Context
Column | Type | Purpose |
| UUID | Primary key |
| VARCHAR(255) | Unique project name |
| TEXT | Project description |
| VARCHAR(50) | Primary language for this project |
| JSON | Other languages used |
| JSON | Frameworks/libraries (e.g., ["react", "django"]) |
| JSON | Directory structure hints (src_dirs, test_dirs, etc.) |
| TIMESTAMP | Project creation time |
| TIMESTAMP | Last update time |
| BOOLEAN | Soft delete flag |
prompt_history - Learning Data
Column | Type | Purpose |
| UUID | Primary key |
| UUID | FK to projects (optional context) |
| UUID | FK to refined_prompts |
| TIMESTAMP | When this record was created |
| UUID | Optional trace/request ID |
| VARCHAR(100) | Which agent executed this prompt |
| TIMESTAMP | Execution timestamp |
| FLOAT | Quality score for this execution (1.0-5.0) |
| BOOLEAN | Whether tests passed |
| FLOAT | Code quality score |
| BOOLEAN | Security check results |
| JSON | List of modified files |
| JSON | List of new files created |
| JSON | List of deleted files |
| JSON | Code diffs organized by language |
| INTEGER | Time taken to execute (milliseconds) |
| TEXT | Summary of agent output |
| JSON | Test/lint/validation details |
| JSON | Diff saliency analysis data |
Relationships
projects ||--o{ refined_prompts : "has many"
projects ||--o{ prompt_history : "has many"
refined_prompts ||--o{ prompt_history : "has many"
refined_prompts ||--o{ refined_prompts : "refinement chain"projects β refined_prompts: One-to-many (a project has multiple prompts)
projects β prompt_history: One-to-many (a project has multiple history entries)
refined_prompts β prompt_history: One-to-many (a prompt has multiple execution records)
refined_prompts β refined_prompts: Self-referential (refinement chains via
prior_refinement_id)
Indexes
Performance-optimized indexes on frequently queried columns:
Table | Column(s) | Purpose |
|
| Unique lookup by public key |
|
| Filter by agent specialization |
|
| Filter by language |
|
| Filter by domain/topic |
|
| Join with projects table |
|
| Vector similarity search (pgvector HNSW) |
|
| Unique project name lookup |
|
| Filter by language |
|
| Join with refined_prompts |
|
| Join with projects table |
π§ͺ Development
Running Tests
# All tests (requires Docker)
pytest
# Integration tests only
pytest tests/integration -v
# With coverage
pytest --cov=src --cov-report=htmlCurrent Status: 274 tests passing β
Database Migrations
# Create new migration after model changes
alembic revision --autogenerate -m "description"
# Apply migrations
alembic upgrade head
# Rollback
alembic downgrade -1π³ Docker
Note: Docker is primarily used for testing SIMPA in an isolated environment. It can also be used as an alternative to installing PostgreSQL directly on your machine during development.
For production deployments, you may prefer running SIMPA directly with your existing PostgreSQL instance rather than containerizing both services.
Quick Start with Docker Compose (Testing)
The easiest way to test SIMPA without installing PostgreSQL locally:
# Start PostgreSQL with pgvector in Docker
docker-compose up -d postgres
# Initialize the database
python -m src.main --init-db
# Run the MCP server
python -m src.mainThis uses the docker-compose.test.yml which only starts the PostgreSQL serviceβSAMPA runs natively on your machine using the containerized database.
Production Deployment
# Build optimized image
docker build --target production -t simpa-mcp:latest .
# Run with environment
docker run -d \
--name simpa-mcp \
-e DATABASE_URL=postgresql://... \
-e OPENAI_API_KEY=sk-... \
simpa-mcp:latestMulti-stage Targets
Target | Purpose | Size |
| Compile dependencies | Base |
| Live code mounting | ~2GB |
| Optimized runtime | ~700MB |
π Documentation
Document | Description |
System architecture, data flow, and component design | |
Comprehensive testing guide and test development | |
API Reference - MCP tool documentation | |
Architecture Decisions - ADRs and design patterns |
π What's Next?
Multi-agent prompt coordination
Prompt lineage tracking
A/B testing framework
Prompt security scanning
Custom embedding models
π€ Contributing
Contributions are welcome! Please:
Fork the repository
Create a feature branch
Make your changes
Add tests (we have 274 as examples!)
Submit a pull request
π License
MIT License - see LICENSE for details
Available Tools
8 toolsactivate_promptActivate PromptA
Activate a previously deactivated prompt.
Reactivates a prompt so it can be used in future refinement searches.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | Activate request with prompt key |
Output Schema
| Name | Required | Description |
|---|---|---|
| message | Yes | |
| success | Yes | |
| is_active | Yes | |
| prompt_key | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the core behavioral effect: reactivating a previously deactivated prompt makes it available for future refinement searches. It does not discuss idempotency, error behavior, or permissions, but for a simple state-change tool the main behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. The first sentence states the operation, and the second adds the purpose without redundancy that harms clarity. Key information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with a fully documented schema and an output schema, the description provides sufficient context: state prior, state after, and purpose. The only notable omission is guidance about the inverse sibling tool or already-active prompts, but this is a minor gap given the low complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents the request object and required prompt_key. The tool description adds no additional parameter-level meaning, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly identifies the verb (activate/reactivate), the resource (prompt), and the state transition (previously deactivated to active). It also explains why the tool exists β enabling the prompt for future refinement searches β which distinguishes it from sibling tools like refine_prompt and deactivate_prompt.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: after a prompt has been deactivated and needs to be usable again. It does not explicitly name deactivate_prompt as the inverse alternative or state exclusions, such as what happens if the prompt is already active, so it falls just short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_projectCreate ProjectA
Create a new project for organizing prompts.
Creates a project with language and dependency metadata to enable better prompt selection based on project context.
Project Structure & Scoping:
Projects can define their default structure via project_structure:
Default directories agents should focus on (e.g., ["src/", "tests/"])
Default exclusions (e.g., [".venv/", "node_modules/"])
Known entry points (e.g., ["src/main.py", "src/app.py"])
This helps downstream agents understand the project layout and scope their work appropriately without needing to explore the entire codebase.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | Project creation request |
Output Schema
| Name | Required | Description |
|---|---|---|
| success | No | |
| created_at | Yes | |
| project_id | Yes | |
| description | No | |
| project_name | Yes | |
| project_structure | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It usefully discloses how project_structure is interpreted (focus directories, exclusions, entry points) and its effect on downstream agents. However, it does not disclose idempotency, duplicate handling, validation, or any error/side-effect behavior beyond the obvious creation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then uses a scoped section for project_structure details. It is longer than strictly necessary, but the extra detail earns its place by clarifying a non-obvious parameter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a creation tool with a nested request schema and an output schema, the description covers the main purpose and the most complex parameter's behavior. It does not discuss duplicate-name behavior or edge cases, but the structured schema and output schema cover most remaining invocation details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is met. The description adds meaningful semantics by explaining that project_structure defines 'default directories agents should focus on', 'default exclusions', and 'known entry points', which goes beyond the schema's terse 'Project structure hints'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Create a new project') and resource ('project for organizing prompts'), and adds the metadata purpose ('language and dependency metadata'). It is clearly distinct from the sibling get/list/update tools, so an agent can infer what this tool is for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than explicit: the description says 'Create a new project' and explains downstream benefits, but it never states when to prefer this over get_project or list_projects, nor does it mention preconditions such as checking for an existing project or naming constraints.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deactivate_promptDeactivate PromptA
Deactivate a prompt so it won't be used in searches.
Soft-deletes a prompt by marking it as inactive. The prompt remains in the database but won't appear in search results or be used for finding similar prompts.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | Deactivate request with prompt key |
Output Schema
| Name | Required | Description |
|---|---|---|
| message | Yes | |
| success | Yes | |
| is_active | Yes | |
| prompt_key | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden, and it clearly states the key side effect: this is a soft-delete, not a removal from the database. It also specifies search-related consequences, though it does not mention reversibility or permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The main purpose is front-loaded and the overall length is appropriate. There is slight redundancy between 'won't be used in searches' and 'won't appear in search results', but the text remains tight and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with an output schema, the description conveys the essential action and side effects without needing to explain return values. The main omission is explicit guidance around reactivation, but that is not required to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the only parameter, prompt_key, is self-describing. The description adds no additional parameter semantics beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states the exact verb ('Deactivate'), resource ('prompt'), and consequence ('won't be used in searches'). The second sentence clarifies soft-delete semantics and persistence, which separates it semantically from sibling tools like activate_prompt and update_prompt_results.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The rationale is implied: use this when you want a prompt excluded from searches. However, it does not explicitly say when not to use it or mention activate_prompt as the inverse/alternative, so selection guidance is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_projectGet ProjectA
Retrieve project information by ID or name.
Look up a project by either its ID (UUID) or name.
Project Scoping for Agents:
Returns project_structure which defines default scoping for this project:
src_dirs: Recommended directories to focus ontest_dirs: Test directory locationsentry_points: Main entry point filesexclude: Paths to ignore
Use this structure when refining prompts to help agents understand the codebase layout and scope their work appropriately.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | Get project request with project_id or project_name |
Output Schema
| Name | Required | Description |
|---|---|---|
| project_id | Yes | |
| description | Yes | |
| project_name | Yes | |
| prompt_count | Yes | |
| main_language | Yes | |
| other_languages | Yes | |
| project_structure | Yes | |
| project_created_at | Yes | |
| project_updated_at | Yes | |
| library_dependencies | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It clearly states that the tool retrieves project information and returns a project_structure object, detailing its fields and intended use. It does not cover error behavior or side effects, but 'retrieve' and the absence of mutation language convey a read-only operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The structure is clean with a summary, lookup detail, and a bulleted scoping section. The first two sentences are slightly redundant ('by ID or name' appears twice), so it is not maximally concise, but the scoping guidance is compact and useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and has an output schema, so the description need not restate return values. It adequately explains the returned project_structure and its agent-facing purpose, though it omits guidance on when to prefer this tool over siblings and does not clarify behavior when both or neither identifiers are supplied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds the UUID nuance for project_id and implies that either project_id or project_name can be used for lookup, which enriches the schema's bare string/null types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb+resource ('Retrieve project information') and specifies lookup by ID or name. It clearly conveys a single-project lookup, but it does not explicitly contrast with sibling tools like list_projects or create_project, so differentiation is implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description indicates when the tool is usefulβwhen project information or the project_structure is neededβand instructs agents to use the returned structure when refining prompts. However, it offers no explicit alternatives, exclusions, or when-not-to-use guidance relative to sibling tools such as list_projects or health_check.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
health_checkHealth CheckA
Health check endpoint.
Returns the health status of the SIMPA MCP service.
Examples:
Request (no parameters needed):
json {}
Returns: Service health status with version and timestamp
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| status | Yes | |
| service | Yes | |
| version | Yes | |
| timestamp | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the output (health status with version and timestamp) and implies a read-only operation through 'Returns' and 'health check'. However, it does not explicitly state that it has no side effects or requires no special permissions, which would make the safety profile clearer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the purpose and includes a redundant example request that duplicates the empty schema. The example is unnecessary but not harmful; overall it is concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with an output schema, the description covers the purpose and return summary. It does not mention error conditions or authentication, but these are less critical for a health check. It is complete enough for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema fully documents the input. The description adds no parameter semantics but correctly shows an example with an empty request. Baseline for zero parameters is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action and resource: 'Returns the health status of the SIMPA MCP service'. It is immediately distinct from sibling tools which handle prompts and projects, so an agent can easily identify its purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context that this tool reports service health, implying it should be used to check service status. It does not explicitly mention alternatives, but no sibling provides health check functionality, so the context is sufficient. No exclusions are needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_projectsList ProjectsA
List all projects with optional filtering.
Retrieves a paginated list of projects, optionally filtered by programming language.
Project Scoping:
Each project may include project_structure metadata that defines:
src_dirs: Recommended source directories (e.g., ["src/", "lib/"])test_dirs: Test directories (e.g., ["tests/"])entry_points: Main entry points (e.g., ["src/main.py"])exclude: Paths to ignore (e.g., [".venv/", "pycache/"])
Use get_project to retrieve full structure details for a specific project,
then use this information when scoping agent work via refine_prompt.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | List projects request with optional filters |
Output Schema
| Name | Required | Description |
|---|---|---|
| limit | Yes | |
| offset | Yes | |
| projects | Yes | |
| total_count | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does disclose pagination, optional filtering, and the variable project_structure metadata, which is useful, but it omits ordering, error behavior, authentication needs, and an explicit read-only statement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a clear summary and uses bullet points effectively for project_structure metadata. The first two sentences are somewhat redundant, which prevents a perfect score, but the overall structure is well organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a paginated list tool with a rich output schema and well-documented parameters, the description covers pagination, filtering, and downstream use of project_structure metadata. It does not address ordering or empty-result behavior, but these are not critical for successful invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description confirms that main_language is the filter and implies pagination via limit/offset, but it adds no additional semantic meaning beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'List all projects' and 'Retrieves a paginated list of projects'. It distinguishes itself from get_project by pointing out that full structure details belong to get_project, clarifying the division of labor between siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly directs agents to use get_project when full structure details for a specific project are needed, and mentions refine_prompt for scoping work. It does not state explicit 'when not to use' conditions, but the routing context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
refine_promptRefine PromptA
Refine a prompt before sending to an agent.
Given an original prompt and context, either selects an existing refined prompt or creates a new one optimized for the agent type and language.
Scoping Options (Optional but Recommended):
To improve focus and reduce context overload, include in the context dict:
target_dirs: List of directories the agent should focus on (e.g., ["src/", "tests/"])target_files: Specific files to modify (e.g., ["src/main.py", "src/config.py"])exclude_paths: Paths to ignore (e.g., [".venv/", "node_modules/"])scope: High-level scope description (e.g., "backend API layer only")focus: Priority aspects (e.g., ["performance", "security", "error-handling"])
Note: A project_id is required. If not provided, the response will include a list of existing projects or instructions to create one. The agent should:
Call list_projects to see available projects, or
Call create_project to create a new one, then
Resubmit the refine_prompt request with the project_id
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | Refinement request with prompt details |
Output Schema
| Name | Required | Description |
|---|---|---|
| action | Yes | |
| source | Yes | |
| prompt_key | Yes | |
| usage_count | No | |
| average_score | No | |
| refined_prompt | Yes | |
| confidence_score | No | |
| similar_prompts_found | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well: it discloses that the tool either selects an existing refined prompt or creates a new one, explains what happens when project_id is absent, and documents how context keys like target_dirs and focus affect behavior. This is substantial behavioral context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a crisp one-line purpose and uses markdown headings, bullets, and numbered steps effectively. It is longer than average, but the added detail is functional rather than filler, especially given the nested request object and the project_id fallback workflow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a nested request object, no annotations, and an output schema, the description is remarkably complete. It covers the main use case, optional scoping keys, required workflow around project_id, and sibling-tool handoffs. Nothing essential for invoking the tool correctly appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds meaning by documenting the recommended shape of the context object and clarifying that project_id is important for the workflow. It explains the intended use of agent_type and language more concretely than the schema alone, though it does not discuss every nested field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Refine a prompt before sending to an agent.' It clearly states the tool selects an existing refined prompt or creates a new one, distinguishing it from sibling tools like activate_prompt, deactivate_prompt, and project-management tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool ('before sending to an agent') and provides recommended scoping options to improve focus. It also gives an explicit fallback workflow when project_id is missing, directing the agent to list_projects or create_project before resubmitting, though it does not explicitly contrast against non-project siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_prompt_resultsUpdate Prompt ResultsA
Update prompt performance metrics after agent execution.
Records the outcome of using a refined prompt and updates the prompt's statistics for future refinement decisions.
Scoping Feedback:
Use files_modified and files_added to record which files were actually
touched. This helps the system understand the effective scope of the prompt
and can be used to suggest narrower scopes for similar future tasks.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | Update request with prompt key and results |
Output Schema
| Name | Required | Description |
|---|---|---|
| success | Yes | |
| usage_count | Yes | |
| last_used_at | No | |
| average_score | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It transparently states that the tool updates prompt statistics and can influence future refinement decisions, and it adds useful context about scoping feedback. Yet it does not disclose details like whether scores are overwritten or accumulated, idempotency, or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, focused sections. The main purpose is front-loaded in the first sentence, and the bolded 'Scoping Feedback' section earns its place by providing actionable guidance. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is complex with a nested schema and many result fields, but the description covers when to call it and clarifies the purpose of two key file-tracking parameters. The output schema presumably covers return values, so the description doesn't need to. It could be more complete by explaining the meaning of action_score or how metrics are aggregated, but it is largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The top-level schema has 100% description coverage, giving a baseline of 3. The tool description goes beyond the schema by explaining the semantic purpose of files_modified and files_added β recording actual touched files and enabling narrower scope suggestions β which adds real value for the agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Update') and resource ('prompt performance metrics'), and clarifies that it records outcomes after agent execution. This distinguishes it from sibling tools like refine_prompt and activate_prompt, which focus on generating or toggling prompts rather than recording results.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: this tool should be used after agent execution to record the outcome of a refined prompt, and it explains why files_modified and files_added matter for scoping feedback. However, it does not explicitly name when to use a sibling tool instead, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
activate_prompt - First observed
create_project - First observed
deactivate_prompt - First observed
get_project - First observed
health_check - First observed
list_projects - First observed
refine_prompt - First observed
update_prompt_results
TDQS
Scored across 8 tools
Each tool has a distinct purpose: refine_prompt handles prompt optimization, update_prompt_results tracks metrics, health_check monitors service status, create/get/list_projects manage project metadata, and activate/deactivate_prompt control prompt lifecycle. No two tools overlap in function.
Most tools follow a consistent verb_noun pattern (refine_prompt, create_project, list_projects, activate_prompt, etc.). The exception is health_check, which is a noun phrase rather than a verb_noun, making it slightly inconsistent but still readable.
With 8 tools, the server is well-scoped for a prompt management system. Each tool covers a necessary functionβproject management, prompt refinement, result tracking, and lifecycle controlβwithout redundancy or bloat.
The core workflows are covered: project creation and lookup, prompt refinement and deactivation, result feedback, and health checks. Missing project update/delete operations and a dedicated get_prompt tool are minor gaps that agents can work around, but the surface is largely complete.
Maintenance
Related MCP Connectors
Persistent memory for AI agents to retain, retrieve, and recall conversation context through MCP.
Generate contextual prompts and reusable agent skills, evaluate prompts with the 16-dimension Prompt Score, and manage saved work in PromptDrive. Twelve MCP tools also provide authorized access to private Memory for source-grounded answers. Connect over Streamable HTTP using OAuth 2.1 and PKCE. Generation consumes account quota and automatically saves successful results; Memory access follows account permissions and plan limits.
An MCP memory server. One memory your agents share β across models, devices and apps.
Persistent AI memory shared across Claude, ChatGPT, coding agents, and compatible MCP clients.
Related MCP Servers
- AlicenseBqualityDmaintenanceAn MCP server that automatically optimizes AI prompts using evolutionary algorithms, helping improve prompt performance, creativity, and reliability through iterative testing and refinement.123MIT
- AlicenseNot gradedqualityDmaintenanceEngram MCP provides persistent, cross-session memory for AI agents by automatically encoding errors, decisions, and discoveries during development sessions. It enables local, intelligent recall and automated context management to help AI learn from experience and avoid recurring mistakes.6 npmBusiness Source 1.1
- AlicenseAqualityDmaintenanceAn MCP server that automatically enhances user prompts by applying advanced engineering techniques like chain-of-thought and few-shot reasoning based on identified intent. It optimizes technique selection through local learning and integrates directly into Claude sessions to improve output quality without additional API costs.6MIT
- AlicenseAqualityBmaintenanceMCP server for persistent, compounding memory that automatically captures corrections and insights across AI sessions, enabling agents to learn and improve over time.5371MIT