Skip to main content
Glama

SIMPA - Self-Improving Meta Prompt Agent

πŸš€ Transform your AI agents with self-optimizing prompt intelligence

SIMPA is a Model Context Protocol (MCP) service that learns from every interaction to continuously improve prompt quality. It remembers what worked, refines what didn't, and automatically selects the best prompts for any situation.

🌟 Why SIMPA?

Every agent you deploy faces the same challenge: getting the prompt right. SIMPA solves this by:

  • πŸ“Š Learning from feedback - Automatically improves based on execution scores

  • πŸ” Semantic search - Finds similar successful prompts using vector similarity

  • 🧠 Smart selection - Chooses between refinement and reuse based on proven performance

  • πŸ”— MCP Native - Seamlessly integrates with any MCP-compatible agent controller

Related MCP server: Engram MCP

πŸ—οΈ Architecture

flowchart TB
    subgraph Controller["Agent Controller"]
        A[Agent Request]
    end
    
    subgraph SIMPA["SIMPA MCP Service"]
        direction TB
        R[Refiner] --> S[Selector]
        S --> V[Vector Store]
        S --> L[LLM Service]
        L --> E[Embedding Service]
    end
    
    subgraph Storage["Knowledge Base"]
        direction TB
        P[(PostgreSQL + pgvector)]
        H[Prompt History]
    end
    
    A -->|original_prompt| R
    V -->|similar_prompts| S
    S -->|refined_prompt| A
    S -->|store & learn| P
    P -->|usage_stats| S
    H -->|feedback_loop| S
    
    style Controller fill:#e1f5fe
    style SIMPA fill:#fff3e0
    style Storage fill:#e8f5e9

πŸ”„ Prompt Lifecycle

SIMPA sits between the Agent Orchestrator and Implementation Agents, continuously learning from each interaction:

flowchart LR
    AO[Agent Orchestrator] -->|prompt| SR[SIMPA Prompt<br/>Refinement]
    SR -->|refined-prompt| IA[Implementation<br/>Agent]
    IA -->|Actions, Results<br/>& Products| RA[Reviewing Agent]
    RA -->|refined-prompt-score| SR2[SIMPA]
    SR2 -->|learn & improve| SR
    
    style AO fill:#e1f5fe,color:#000000
    style SR fill:#fff3e0,color:#000000
    style IA fill:#fce4ec,color:#000000
    style RA fill:#f3e5f5,color:#000000
    style SR2 fill:#fff3e0,color:#000000

The Flow:

  1. Agent Orchestrator β†’ Sends raw prompt to SIMPA

  2. SIMPA β†’ Returns refined-prompt (structured with ROLE, GOAL, REQUIREMENTS)

  3. Implementation Agent β†’ Executes actions using refined prompt, produces results/products

  4. Reviewing Agent β†’ Evaluates outcomes, generates refined-prompt-score

  5. SIMPA β†’ Receives score, learns what works, improves future refinements

This closed feedback loop ensures prompts get better with every execution.

✨ Features

Feature

Description

πŸ€– MCP Protocol

Native Model Context Protocol support for universal agent integration

πŸ”Ž Vector Search

pgvector-powered similarity search for prompt retrieval

πŸ“ˆ Self-Improvement

Sigmoid-based probability for intelligent refinement vs reuse

🎯 Multi-Provider

OpenAI, Anthropic, and Ollama support for embeddings and LLM

πŸ“Š Observability

Structured logging with structlog and comprehensive metrics

πŸ›‘οΈ Security

PII detection and input validation built-in

πŸ§ͺ Tested

274 automated tests with 100% pass rate

πŸ“‹ Prerequisites

Before installing SIMPA, ensure you have the following:

Required

Component

Version

Purpose

Python

3.10+

Runtime environment

PostgreSQL

14+

Database with pgvector extension

Docker

Latest

Required for running tests with TestContainers

Git

Latest

Clone repository

Note: PostgreSQL and Ollama are expected to be installed and running separately (not via Docker) for normal operation. Docker is only required for the automated test suite.

Component

Purpose

Ollama

Local LLM & embedding inference

nomic-embed-text

Embedding model (pull via ollama pull nomic-embed-text)

llama3.2

LLM for prompt refinement (pull via ollama pull llama3.2)

For Cloud Providers (Optional)

πŸ” Security Best Practice: Provider API keys (OpenAI, Anthropic, Google, Azure) should be kept in your user home directory at ~/.env rather than in the project .env file. This prevents accidental commits of sensitive credentials to version control.

Create ~/.env with your provider keys:

# ~/.env - User-level secrets (not committed)
OPENAI_API_KEY=sk-...
ANTHROPIC_API_KEY=sk-ant-...
GOOGLE_API_KEY=...
AZURE_OPENAI_KEY=...

SIMPA will automatically load keys from ~/.env if available.

System Requirements

Resource

Minimum

Recommended

RAM

4 GB

8 GB+

Disk

2 GB free

10 GB+

CPU

2 cores

4 cores+

Note: For local Ollama models, CPU is sufficient but GPU acceleration significantly improves performance.

πŸš€ Quick Start

This option runs PostgreSQL and Ollama in Docker containers for easy development and testing:

# Clone and setup
git clone https://github.com/yourusername/simpa-mcp.git
cd simpa-mcp
cp .env.example .env

# Start all services (PostgreSQL + Ollama in Docker)
make dev-setup

# Download models (one-time)
make pull-models

# Run migrations
make migrate

# Run tests
make test

For Production Use: Install PostgreSQL and Ollama directly on your system instead of using Docker. See the Manual Setup section below.

Option 2: Manual Setup (Production/Existing Services)

Use this if you already have PostgreSQL and Ollama installed locally.

Prerequisites:

  • PostgreSQL 14+ with pgvector extension installed

  • Ollama running locally (with nomic-embed-text and llama3.2 pulled)

# Install dependencies
pip install -e ".[dev]"

# Configure environment
cp .env.example .env
# Edit .env to match your PostgreSQL and Ollama settings

# Run migrations
alembic upgrade head

# Start MCP server
python -m src.main

Quick PostgreSQL setup with Docker (if needed):

# Only if you don't have PostgreSQL installed locally
docker run -d --name simpa-db \
  -e POSTGRES_USER=simpa \
  -e POSTGRES_PASSWORD=simpa \
  -e POSTGRES_DB=simpa \
  -p 5432:5432 \
  pgvector/pgvector:pg16

πŸ”Œ Adding SIMPA to Your MCP Configuration

SIMPA works with any MCP-compatible client (Cursor, Claude Desktop, Windsurf, etc.).

Step 1: Install SIMPA Server

Option A: Global Installation (Easiest)

# Clone the repository
git clone https://github.com/dsidlo/simpa-mcp.git
cd simpa-mcp

# Create virtual environment
python -m venv .venv

# Activate virtual environment
# On macOS/Linux:
source .venv/bin/activate
# On Windows:
# .venv\Scripts\activate

# Install in editable mode
pip install -e .

# Install MCP dependencies
pip install fastmcp asyncpg pgvector sqlalchemy

# Setup environment
cp .env.example .env
# Edit .env with your configuration (see Configuration section below)

# Run database migrations
alembic upgrade head
# Build the MCP server image
docker build --target production -t simpa-mcp:latest .

# Or use docker compose (includes PostgreSQL + pgvector)
docker-compose up -d

Step 2: Configure Your MCP Client

Add SIMPA to your MCP client's configuration file:

Cursor (~/.cursor/mcp.json)

{
  "mcpServers": {
    "simpa-mcp": {
      "command": "uv",
      "args": [
        "--directory",
        "/absolute/path/to/simpa-mcp",
        "run",
        "--env",
        "/absolute/path/to/simpa-mcp/.env",
        "python",
        "-m",
        "src.main"
      ],
      "env": {
        "PYTHONPATH": "/absolute/path/to/simpa-mcp/src"
      }
    }
  }
}

Claude Desktop (~/Library/Application Support/Claude/claude_desktop_config.json)

{
  "mcpServers": {
    "simpa-mcp": {
      "command": "/absolute/path/to/simpa-mcp/.venv/bin/python",
      "args": [
        "-m",
        "src.main"
      ],
      "env": {
        "DATABASE_URL": "postgresql://simpa:simpa@localhost:5432/simpa",
        "EMBEDDING_PROVIDER": "ollama",
        "EMBEDDING_MODEL": "nomic-embed-text",
        "OLLAMA_BASE_URL": "http://localhost:11434",
        "LLM_MODEL": "ollama/llama3.2",
        "PYTHONPATH": "/absolute/path/to/simpa-mcp/src"
      }
    }
  }
}

Generic MCP Configuration

{
  "mcpServers": {
    "simpa-mcp": {
      "name": "SIMPA Prompt Refinement",
      "description": "Self-improving prompt optimization service",
      "command": "python",
      "args": [
        "-m",
        "src.main",
        "--mcp",
        "stdio"
      ],
      "workingDirectory": "/absolute/path/to/simpa-mcp",
      "envFile": "/absolute/path/to/simpa-mcp/.env"
    }
  }
}

This configuration ensures the server runs from the source directory and uses uv for dependency management:

{
  "mcpServers": {
    "simpa-mcp": {
      "command": "/bin/bash",
      "args": [
        "-c",
        "cd /path/to/simpa-mcp && uv run python src/main.py --log-level debug --log-file /tmp/simpa-mcp.log"
      ]
    }
  }
}

Note: Replace /path/to/simpa-mcp with your actual installation path. Using bash -c with cd ensures the server runs from the project root where pyproject.toml and .env are located.

Step 3: Install MCP Inspector (Optional, for Testing)

# Install MCP Inspector globally
npm install -g @anthropics/mcp-inspector

# Test your SIMPA server
mcp-inspector --server "uv --directory /path/to/simpa-mcp run python -m src.main"

Step 4: Verify Installation

In your MCP client (Cursor/Claude Desktop), you should see:

  • βœ… Available Tools: refine_prompt, update_prompt_results

  • βœ… Server Status: Connected

  • βœ… Capabilities: Prompt refinement enabled

πŸ› οΈ Troubleshooting

"Command not found: uv"

Install uv first:

curl -LsSf https://astral.sh/uv/install.sh | sh

"ModuleNotFoundError: No module named 'src'"

Ensure PYTHONPATH includes the src directory:

export PYTHONPATH="/absolute/path/to/simpa-mcp/src:$PYTHONPATH"

Database Connection Errors

Verify PostgreSQL is running with pgvector:

# Check if pgvector extension is available
psql -d simpa -c "CREATE EXTENSION IF NOT EXISTS vector;"

MCP Server Not Responding

Test manually:

cd /path/to/simpa-mcp
source .venv/bin/activate
python -m src.main --help

πŸ”§ Configuration

SIMPA can be configured via environment variables and command-line arguments.

Environment Variables

How configuration works: SIMPA uses Pydantic Settings to automatically load environment variables from .env files. When you set an environment variable, it automatically becomes available via settings.VARIABLE_NAME in the codeβ€”no explicit os.getenv() calls needed. Environment variables are case-insensitive (EMBEDDING_MODEL and embedding_model work the same).

⚑ Critical Parameters (Required)

These parameters must be configured to bring up the MCP service:

Variable

Description

Why Required

DATABASE_URL

PostgreSQL connection URL

Stores prompt knowledge base

OPENAI_API_KEY

OpenAI API key

Required only if using OpenAI models. Other providers need their respective keys.

All other parameters can be left undefined β€” they default to known, usable values suitable for most deployments.

Minimal Configuration Example

The simplest working .env file (using local Ollama models):

# Only REQUIRED parameter - everything else defaults automatically
DATABASE_URL=postgresql://user@localhost:5432/simpa

For OpenAI instead of Ollama, just add the API key:

# Required
DATABASE_URL=postgresql://user@localhost:5432/simpa
OPENAI_API_KEY=sk-your-key-here
# LLM_MODEL defaults to ollama/llama3.2, but you can override:
# LLM_MODEL=openai/gpt-4

Optional Parameters (With Working Defaults)

All sections below have sensible defaults. You only need to change them if you have specific requirements:

Database Connection Details

Variable

Description

Default

DATABASE_URL

PostgreSQL connection URL

postgresql://dsidlo@localhost:5432/simpa

Embedding Service

Variable

Description

Default

EMBEDDING_PROVIDER

Embedding provider (ollama or openai)

ollama

EMBEDDING_MODEL

Embedding model name

nomic-embed-text

EMBEDDING_DIMENSIONS

Vector dimensions (768 for nomic-embed-text)

768

OLLAMA_BASE_URL

Ollama API base URL

http://localhost:11434

LLM Service

Variable

Description

Default

LLM_MODEL

LLM model (LiteLLM format: provider/model)

ollama/llama3.2

LLM_TEMPERATURE

Sampling temperature (0.0 - 2.0)

0.7

Supported Models (via LiteLLM):

  • ollama/llama3.2 - Local Ollama models

  • gpt-4, gpt-3.5-turbo - OpenAI

  • claude-3-opus-20240229, claude-3-sonnet-20240229 - Anthropic

  • gemini/gemini-pro, gemini/gemini-ultra - Google

  • azure/<deployment-name> - Azure OpenAI

API Keys (Only if using cloud LLM providers)

Only needed if you use cloud-based LLMs instead of local Ollama models. These are loaded automatically by LiteLLM based on model prefix:

Variable

Description

OPENAI_API_KEY

OpenAI API key

ANTHROPIC_API_KEY

Anthropic API key

GEMINI_API_KEY

Google Gemini API key

AZURE_API_KEY

Azure OpenAI API key

AZURE_API_BASE

Azure OpenAI endpoint base URL

COHERE_API_KEY

Cohere API key

Embedding Cache

Variable

Description

Default

EMBEDDING_CACHE_ENABLED

Enable LRU cache for embeddings

true

EMBEDDING_CACHE_MAX_SIZE

Maximum cache entries (100-10000)

1000

EMBEDDING_CACHE_MAX_TEXT_LENGTH

Maximum text length to cache

10000

LLM Cache

Variable

Description

Default

LLM_CACHE_ENABLED

Enable LLM response caching

true

LLM_CACHE_TTL_SECONDS

Cache TTL in seconds (60-86400)

3600

LLM_CACHE_MAX_ENTRIES

Maximum cache entries (100-100000)

10000

LLM_CACHE_DB_PATH

Path to cache SQLite database

./llm_cache.db

Fast-Path Hash Match

Variable

Description

Default

HASH_FAST_PATH_ENABLED

Enable hash-based exact match lookup

true

HASH_FAST_PATH_MIN_SCORE

Minimum score for hash reuse (1.0-5.0)

4.0

Refinement Strategy

Variable

Description

Default

SIMILARITY_BYPASS_THRESHOLD

Cosine similarity threshold for bypass (0.9-1.0)

0.95

SIMILARITY_BYPASS_MIN_SCORE

Minimum score for high-similarity bypass (1.0-5.0)

4.5

SIGMOID_K

Sigmoid steepness parameter

1.5

SIGMOID_MU

Sigmoid midpoint (50% threshold)

3.0

MIN_REFINEMENT_PROBABILITY

Minimum refinement probability (0.0-1.0)

0.05

Variable

Description

Default

VECTOR_SEARCH_LIMIT

Number of similar prompts to retrieve (1-50)

5

VECTOR_SIMILARITY_THRESHOLD

Minimum similarity score (0.0-1.0)

0.7

Variable

Description

Default

BM25_SEARCH_ENABLED

Enable BM25 keyword search

true

BM25_K1

BM25 term saturation parameter (0.1-3.0)

1.2

BM25_B

BM25 document length normalization (0.0-1.0)

0.75

BM25_LIMIT

Number of BM25 results (1-20)

5

BM25_VECTOR_LIMIT

Number of vector results in hybrid (1-20)

5

HYBRID_SEARCH_ENABLED

Enable hybrid search combining vector + BM25

true

LLM_RERANK_ENABLED

Enable LLM re-ranking of results

true

LLM_RERANK_CANDIDATES

Number of candidates for re-ranking (2-20)

10

MCP Server

Variable

Description

Default

MCP_TRANSPORT

Transport protocol (stdio or sse)

stdio

MCP_PORT

Server port for SSE transport (1024-65535)

8000

Logging

Variable

Description

Default

LOG_LEVEL

Logging level (TRACE, DEBUG, INFO, WARNING, ERROR, CRITICAL)

INFO

JSON_LOGGING

Enable structured JSON logging

true

Security

Variable

Description

Default

MAX_PROMPT_LENGTH

Maximum prompt text length (100-100000)

10000

ENABLE_PII_DETECTION

Enable basic PII detection

true

Project Association

Variable

Description

Default

REQUIRE_PROJECT_ID

Require project_id for all refinements

false

Diff Saliency

Variable

Description

Default

DIFF_SALIENCY_ENABLED

Enable diff saliency filtering

true

DIFF_SALIENCY_THRESHOLD

Minimum saliency score (0.0-1.0)

0.6

DIFF_MAX_STORED_PER_REQUEST

Maximum diffs per request (1-50)

10

Complete Example .env File

# Database (Required)
DATABASE_URL=postgresql://simpa:simpa@localhost:5432/simpa

# Embedding Service
EMBEDDING_PROVIDER=ollama
EMBEDDING_MODEL=nomic-embed-text
EMBEDDING_DIMENSIONS=768
OLLAMA_BASE_URL=http://localhost:11434

# LLM Service
LLM_MODEL=ollama/llama3.2
LLM_TEMPERATURE=0.7

# API Keys (if using cloud providers)
# OPENAI_API_KEY=sk-...
# ANTHROPIC_API_KEY=sk-ant-...
# GEMINI_API_KEY=...

# MCP Server
MCP_TRANSPORT=stdio
MCP_PORT=8000

# Caching (optional, defaults are reasonable)
EMBEDDING_CACHE_ENABLED=true
LLM_CACHE_ENABLED=true

# Refinement behavior (optional)
SIMILARITY_BYPASS_THRESHOLD=0.95
SIGMOID_K=1.5
SIGMOID_MU=3.0

# Logging
LOG_LEVEL=INFO
JSON_LOGGING=true

Command Line Options

SIMPA supports several command line flags for runtime configuration:

# Show all available options
python -m src.main --help

Option

Description

Default

--init-db

Initialize the database schema and exit

-

--transport {stdio,sse}

MCP transport protocol

stdio

--log-level {trace,debug,info,warn,error,fatal}

Logging level

info

--log-file PATH

Path to log file

/tmp/simpa-mcp.log

--log-console

Also log to console (stderr) ⚠️ Not recommended for MCP stdio mode

-

--env PATH

Path to .env file

~/.env

--project-id-required

Require project_id for all refinement requests

-

Examples:

# Initialize database
python -m src.main --init-db

# Run with SSE transport on custom port (also set MCP_PORT in .env)
python -m src.main --transport sse

# Debug logging to custom file
python -m src.main --log-level debug --log-file /var/log/simpa.log

# Use custom env file
python -m src.main --env ~/my-project/.env

# Require project_id for all prompts
python -m src.main --project-id-required

# Combination of options
python -m src.main --env ./.env.local --log-level debug --transport sse

Environment File (--env)

By default, SIMPA loads environment variables from your home directory at ~/.env. You can specify a custom .env file using the --env option:

# Use a custom env file
python -m src.main --env ~/my-project/.env

# Or use a project-specific .env
python -m src.main --env ./.env.local

Loading order:

  1. If --env is specified and the file exists, it is loaded first

  2. If --env is not specified, ~/.env is loaded if it exists

  3. The project ./.env in the current directory is loaded last (overrides previous values)

This allows you to keep sensitive credentials (API keys) in ~/.env while keeping project-specific settings in the project .env.

Project-Associated Prompt Development (--project-id-required)

Enable strict project association mode to enforce that all prompts must be linked to a project:

# Require project_id for all prompt refinements
python -m src.main --project-id-required

When enabled, calling refine_prompt without a project_id returns a helpful response guiding the agent to:

  1. List existing projects - View available projects to find a suitable match

  2. Create a new project - Use create_project if no suitable project exists

  3. Resubmit with project_id - Retry the refinement with the chosen project

Why use project association?

  • Cross-project learning: Prompts refined for one Python web project can benefit similar Flask/Django projects

  • Knowledge clustering: Projects with similar tech stacks (React+Node, Python+PostgreSQL) share prompt patterns

  • Relevance scoring: Prompt selection considers project context for better matches

  • Team organization: Different teams/projects have distinct prompt preferences and patterns

Example workflow:

# Start server with strict project mode
python -m src.main --project-id-required

# Agent workflow:
# 1. First call without project_id β†’ returns list of existing projects
# 2. Agent picks or creates project β†’ gets project_id
# 3. Resubmit with project_id β†’ prompt is refined and associated with project

πŸ› οΈ MCP Tools

refine_prompt

Intelligently refine prompts before agent execution.

# Request
{
  "original_prompt": "Write a function to sort a list",
  "agent_type": "developer",
  "main_language": "python"
}

# Response
{
  "refined_prompt": "Write a Python function that takes a list of integers...",
  "prompt_key": "uuid-v4",
  "action": "refine|new|reuse",
  "confidence_score": 0.95,
  "similar_prompts_found": 3
}

update_prompt_results

Provide feedback to improve future prompts.

# Request
{
  "prompt_key": "uuid-v4",
  "action_score": 4.5,
  "test_passed": true,
  "files_modified": ["main.py"],
  "lint_score": 0.95
}

# Response
{
  "success": true,
  "usage_count": 5,
  "average_score": 4.25
}

πŸ“ Prompt Refinement Examples

SIMPA transforms vague user requests into structured, actionable specifications.

Example 1: Developer Agent

Original Prompt:

Build a REST API for managing tasks.

Refined Prompt:

ROLE: Senior Backend Developer
GOAL: Build a REST API for managing tasks.
CONSTRAINTS: Your output will be only a descriptive overview of what the API will do.
REQUIREMENTS:
- Define all REST endpoints (GET, POST, PUT, DELETE) with their URLs and purposes
- Explicitly specify request/response JSON formats for each endpoint
- Include pagination, filtering, and sorting capabilities for task listing
- Describe authentication mechanism (JWT or API key based)
- Define error response formats and standard HTTP status codes
- Outline rate limiting strategy if applicable
OUTPUT: Structured API specification document.
SUCCESS: All endpoints are well-defined, formats are validated, and specification is ready to implement.
AUTONOMY: You may choose implementation technology and framework.
FALLBACK: If a task requirement is unclear, ask for clarification or document your assumptions.

Example 2: Architect Agent

Original Prompt:

Design a caching strategy for the product catalog.

Refined Prompt:

ROLE: Software Architect
GOAL: Design a caching strategy for the product catalog.
CONSTRAINTS: Output will be descriptive only; no actual implementation code.
REQUIREMENTS:
- Identify data access patterns (read-heavy, write-heavy, or mixed)
- Design cache-invalidation strategy (time-based, event-driven, or hybrid)
- Define cache key structure and naming conventions
- Specify cache-aside, write-through, or write-behind patterns
- Address cache consistency, eviction policies, and TTL configuration
- Consider multi-tier caching (in-memory + distributed)
- Document potential issues: stale data, thundering herd, cache penetration
OUTPUT: Architecture document detailing caching strategy.
SUCCESS: Design addresses scalability, performance, and data consistency.
AUTONOMY: Choose appropriate caching technology (Redis, Memcached, Caffeine).
FALLBACK: If data volume or access patterns are unclear, document assumptions.

Example 3: Tester Agent

Original Prompt:

Write tests for the login module.

Refined Prompt:

ROLE: QA Engineer
GOAL: Write tests for the login module.
CONSTRAINTS: Provide complete test scenarios and detailed expected outcomes, no code.
TEST CATEGORIES:
- Unit Tests: Each function/method tested in isolation with mocks
- Integration Tests: Service interactions (DB, external auth providers)
- Edge Cases: Empty input, SQL injection attempts, XSS payloads
- Boundary Tests: Max username length, password complexity rules
- Error Paths: Invalid credentials, account locked, service unavailable
- Load Tests: Concurrent login attempts
REQUIREMENTS:
- Use GIVEN/WHEN/THEN format for each scenario
- Define specific test data (usernames, passwords)
- Specify expected status codes and error messages
- Identify test dependencies and setup requirements
OUTPUT: Comprehensive test specification document.
SUCCESS: All happy paths and edge cases are covered.
AUTONOMY: Choose testing frameworks and assertion styles.
FALLBACK: If authentication flow is unclear, document assumptions.

Example 4: Reviewer Agent (with scope constraints)

Original Prompt:

Review error handling in the codebase.

Scope Context:

{
  "target_dirs": ["src/services/", "src/utils/"],
  "target_files": ["src/services/user_service.py"],
  "scope": "error handling patterns only",
  "focus": ["security", "logging", "user-experience"]
}

Refined Prompt:

ROLE: Security-Focused Code Reviewer
GOAL: Review error handling patterns in the codebase.
CONSTRAINTS: Review only within assigned scope and files:
  - target_dirs: src/services/, src/utils/
  - target_files: src/services/user_service.py
  - focus: security, logging, user-experience
  - scope: error handling patterns only
CONTEXT: Production code review process
OUTPUT: Line-by-line comments and summary report
SUCCESS: Critical issues identified, recommendations actionable
AUTONOMY: Can use static analysis tools within scope
FALLBACK: Ask if scope unclear

Review Checklist:
- Security: Exception leaks sensitive data, proper sanitization
- Logging: Appropriate log levels, no PII exposure
- User Experience: Helpful error messages, graceful degradation
- Code Quality: Consistent patterns, avoid catch-all exceptions
- Documentation: Error scenarios documented, recovery paths clear

Note: When scope context is provided (target_dirs, target_files, scope, focus), SIMPA injects these constraints into the refined prompt above the CONSTRAINTS section, limiting the agent's work to the specified boundaries.

🧠 Self-Improvement Algorithm

SIMPA uses a sigmoid function to intelligently balance exploration (refinement) vs exploitation (reuse):

p_refine(S) = 1 / (1 + exp(k * (S - mu)))

Where:

  • S = Average score (1.0 - 5.0)

  • k = Steepness (default: 1.5)

  • mu = Midpoint (default: 3.0)

Refinement Probability:

Score

Probability

⭐ 1.0

~95% πŸ”„ Refine heavily

⭐⭐ 2.0

~82% πŸ”„ Likely refine

⭐⭐⭐ 3.0

~50% βš–οΈ Balance point

⭐⭐⭐⭐ 4.0

~18% βœ… Start reusing

⭐⭐⭐⭐⭐ 5.0

~5% βœ… Reuse proven

πŸ“Š Database Schema

refined_prompts - The Prompt Knowledge Base

Column

Type

Purpose

id

UUID

Primary key

prompt_key

UUID

Public identifier for MCP tools

created_at

TIMESTAMP

When prompt was first refined

updated_at

TIMESTAMP

Last modification time

last_used_at

TIMESTAMP

Last time this prompt was executed

embedding

vector(768)

Semantic embedding for similarity search

agent_type

VARCHAR(100)

Agent specialization (e.g., "developer")

refinement_type

VARCHAR(20)

Strategy used (default: "sigmoid")

main_language

VARCHAR(50)

Primary programming language

other_languages

JSON

Additional languages used

domain

VARCHAR(100)

Domain/topic classification

tags

JSON

Array of descriptive tags

original_prompt_hash

VARCHAR(64)

Hash for fast exact-match lookup

original_prompt

TEXT

Raw input prompt

refined_prompt

TEXT

Optimized/expanded version

refinement_version

INTEGER

Version number for iterative refinements

prior_refinement_id

UUID

Self-reference for refinement chains

project_id

UUID

FK to projects (optional context)

usage_count

INTEGER

Total times used

average_score

FLOAT

Running average of action scores (1.0-5.0)

score_weighted

FLOAT

Bayesian-weighted score for ranking

context

JSON

Scope context (focus, target_dirs, etc.)

is_active

BOOLEAN

Soft delete flag

projects - Project Context

Column

Type

Purpose

id

UUID

Primary key

project_name

VARCHAR(255)

Unique project name

description

TEXT

Project description

main_language

VARCHAR(50)

Primary language for this project

other_languages

JSON

Other languages used

library_dependencies

JSON

Frameworks/libraries (e.g., ["react", "django"])

project_structure

JSON

Directory structure hints (src_dirs, test_dirs, etc.)

created_at

TIMESTAMP

Project creation time

updated_at

TIMESTAMP

Last update time

is_active

BOOLEAN

Soft delete flag

prompt_history - Learning Data

Column

Type

Purpose

id

UUID

Primary key

project_id

UUID

FK to projects (optional context)

prompt_id

UUID

FK to refined_prompts

created_at

TIMESTAMP

When this record was created

request_id

UUID

Optional trace/request ID

executed_by_agent

VARCHAR(100)

Which agent executed this prompt

executed_at

TIMESTAMP

Execution timestamp

action_score

FLOAT

Quality score for this execution (1.0-5.0)

test_passed

BOOLEAN

Whether tests passed

lint_score

FLOAT

Code quality score

security_scan_passed

BOOLEAN

Security check results

files_modified

JSON

List of modified files

files_added

JSON

List of new files created

files_deleted

JSON

List of deleted files

diffs

JSON

Code diffs organized by language

execution_duration_ms

INTEGER

Time taken to execute (milliseconds)

agent_output_summary

TEXT

Summary of agent output

validation_results

JSON

Test/lint/validation details

saliency_metadata

JSON

Diff saliency analysis data

Relationships

projects ||--o{ refined_prompts : "has many"
projects ||--o{ prompt_history : "has many"
refined_prompts ||--o{ prompt_history : "has many"
refined_prompts ||--o{ refined_prompts : "refinement chain"
  • projects β†’ refined_prompts: One-to-many (a project has multiple prompts)

  • projects β†’ prompt_history: One-to-many (a project has multiple history entries)

  • refined_prompts β†’ prompt_history: One-to-many (a prompt has multiple execution records)

  • refined_prompts β†’ refined_prompts: Self-referential (refinement chains via prior_refinement_id)

Indexes

Performance-optimized indexes on frequently queried columns:

Table

Column(s)

Purpose

refined_prompts

prompt_key

Unique lookup by public key

refined_prompts

agent_type

Filter by agent specialization

refined_prompts

main_language

Filter by language

refined_prompts

domain

Filter by domain/topic

refined_prompts

project_id

Join with projects table

refined_prompts

embedding

Vector similarity search (pgvector HNSW)

projects

project_name

Unique project name lookup

projects

main_language

Filter by language

prompt_history

prompt_id

Join with refined_prompts

prompt_history

project_id

Join with projects table

πŸ§ͺ Development

Running Tests

# All tests (requires Docker)
pytest

# Integration tests only
pytest tests/integration -v

# With coverage
pytest --cov=src --cov-report=html

Current Status: 274 tests passing βœ…

Database Migrations

# Create new migration after model changes
alembic revision --autogenerate -m "description"

# Apply migrations
alembic upgrade head

# Rollback
alembic downgrade -1

🐳 Docker

Note: Docker is primarily used for testing SIMPA in an isolated environment. It can also be used as an alternative to installing PostgreSQL directly on your machine during development.

For production deployments, you may prefer running SIMPA directly with your existing PostgreSQL instance rather than containerizing both services.

Quick Start with Docker Compose (Testing)

The easiest way to test SIMPA without installing PostgreSQL locally:

# Start PostgreSQL with pgvector in Docker
docker-compose up -d postgres

# Initialize the database
python -m src.main --init-db

# Run the MCP server
python -m src.main

This uses the docker-compose.test.yml which only starts the PostgreSQL serviceβ€”SAMPA runs natively on your machine using the containerized database.

Production Deployment

# Build optimized image
docker build --target production -t simpa-mcp:latest .

# Run with environment
docker run -d \
  --name simpa-mcp \
  -e DATABASE_URL=postgresql://... \
  -e OPENAI_API_KEY=sk-... \
  simpa-mcp:latest

Multi-stage Targets

Target

Purpose

Size

builder

Compile dependencies

Base

development

Live code mounting

~2GB

production

Optimized runtime

~700MB

πŸ“š Documentation

Document

Description

SIMPA Process Architecture

System architecture, data flow, and component design

Test Suite Development

Comprehensive testing guide and test development

API Reference - MCP tool documentation

Architecture Decisions - ADRs and design patterns

πŸ“ˆ What's Next?

  • Multi-agent prompt coordination

  • Prompt lineage tracking

  • A/B testing framework

  • Prompt security scanning

  • Custom embedding models

🀝 Contributing

Contributions are welcome! Please:

  1. Fork the repository

  2. Create a feature branch

  3. Make your changes

  4. Add tests (we have 274 as examples!)

  5. Submit a pull request

πŸ“„ License

MIT License - see LICENSE for details


Available Tools

8 tools
activate_promptActivate PromptA

Activate a previously deactivated prompt.

Reactivates a prompt so it can be used in future refinement searches.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesActivate request with prompt key

Output Schema

ParametersJSON Schema
NameRequiredDescription
messageYes
successYes
is_activeYes
prompt_keyYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the core behavioral effect: reactivating a previously deactivated prompt makes it available for future refinement searches. It does not discuss idempotency, error behavior, or permissions, but for a simple state-change tool the main behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler. The first sentence states the operation, and the second adds the purpose without redundancy that harms clarity. Key information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with a fully documented schema and an output schema, the description provides sufficient context: state prior, state after, and purpose. The only notable omission is guidance about the inverse sibling tool or already-active prompts, but this is a minor gap given the low complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents the request object and required prompt_key. The tool description adds no additional parameter-level meaning, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly identifies the verb (activate/reactivate), the resource (prompt), and the state transition (previously deactivated to active). It also explains why the tool exists β€” enabling the prompt for future refinement searches β€” which distinguishes it from sibling tools like refine_prompt and deactivate_prompt.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it: after a prompt has been deactivated and needs to be usable again. It does not explicitly name deactivate_prompt as the inverse alternative or state exclusions, such as what happens if the prompt is already active, so it falls just short of full guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_projectCreate ProjectA

Create a new project for organizing prompts.

Creates a project with language and dependency metadata to enable better prompt selection based on project context.

Project Structure & Scoping: Projects can define their default structure via project_structure:

  • Default directories agents should focus on (e.g., ["src/", "tests/"])

  • Default exclusions (e.g., [".venv/", "node_modules/"])

  • Known entry points (e.g., ["src/main.py", "src/app.py"])

This helps downstream agents understand the project layout and scope their work appropriately without needing to explore the entire codebase.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesProject creation request

Output Schema

ParametersJSON Schema
NameRequiredDescription
successNo
created_atYes
project_idYes
descriptionNo
project_nameYes
project_structureNo

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden. It usefully discloses how project_structure is interpreted (focus directories, exclusions, entry points) and its effect on downstream agents. However, it does not disclose idempotency, duplicate handling, validation, or any error/side-effect behavior beyond the obvious creation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and then uses a scoped section for project_structure details. It is longer than strictly necessary, but the extra detail earns its place by clarifying a non-obvious parameter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a creation tool with a nested request schema and an output schema, the description covers the main purpose and the most complex parameter's behavior. It does not discuss duplicate-name behavior or edge cases, but the structured schema and output schema cover most remaining invocation details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is met. The description adds meaningful semantics by explaining that project_structure defines 'default directories agents should focus on', 'default exclusions', and 'known entry points', which goes beyond the schema's terse 'Project structure hints'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Create a new project') and resource ('project for organizing prompts'), and adds the metadata purpose ('language and dependency metadata'). It is clearly distinct from the sibling get/list/update tools, so an agent can infer what this tool is for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than explicit: the description says 'Create a new project' and explains downstream benefits, but it never states when to prefer this over get_project or list_projects, nor does it mention preconditions such as checking for an existing project or naming constraints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

deactivate_promptDeactivate PromptA

Deactivate a prompt so it won't be used in searches.

Soft-deletes a prompt by marking it as inactive. The prompt remains in the database but won't appear in search results or be used for finding similar prompts.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesDeactivate request with prompt key

Output Schema

ParametersJSON Schema
NameRequiredDescription
messageYes
successYes
is_activeYes
prompt_keyYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden, and it clearly states the key side effect: this is a soft-delete, not a removal from the database. It also specifies search-related consequences, though it does not mention reversibility or permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main purpose is front-loaded and the overall length is appropriate. There is slight redundancy between 'won't be used in searches' and 'won't appear in search results', but the text remains tight and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with an output schema, the description conveys the essential action and side effects without needing to explain return values. The main omission is explicit guidance around reactivation, but that is not required to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the only parameter, prompt_key, is self-describing. The description adds no additional parameter semantics beyond what the schema already provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states the exact verb ('Deactivate'), resource ('prompt'), and consequence ('won't be used in searches'). The second sentence clarifies soft-delete semantics and persistence, which separates it semantically from sibling tools like activate_prompt and update_prompt_results.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The rationale is implied: use this when you want a prompt excluded from searches. However, it does not explicitly say when not to use it or mention activate_prompt as the inverse/alternative, so selection guidance is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_projectGet ProjectA

Retrieve project information by ID or name.

Look up a project by either its ID (UUID) or name.

Project Scoping for Agents: Returns project_structure which defines default scoping for this project:

  • src_dirs: Recommended directories to focus on

  • test_dirs: Test directory locations

  • entry_points: Main entry point files

  • exclude: Paths to ignore

Use this structure when refining prompts to help agents understand the codebase layout and scope their work appropriately.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesGet project request with project_id or project_name

Output Schema

ParametersJSON Schema
NameRequiredDescription
project_idYes
descriptionYes
project_nameYes
prompt_countYes
main_languageYes
other_languagesYes
project_structureYes
project_created_atYes
project_updated_atYes
library_dependenciesYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It clearly states that the tool retrieves project information and returns a project_structure object, detailing its fields and intended use. It does not cover error behavior or side effects, but 'retrieve' and the absence of mutation language convey a read-only operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The structure is clean with a summary, lookup detail, and a bulleted scoping section. The first two sentences are slightly redundant ('by ID or name' appears twice), so it is not maximally concise, but the scoping guidance is compact and useful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple and has an output schema, so the description need not restate return values. It adequately explains the returned project_structure and its agent-facing purpose, though it omits guidance on when to prefer this tool over siblings and does not clarify behavior when both or neither identifiers are supplied.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds the UUID nuance for project_id and implies that either project_id or project_name can be used for lookup, which enriches the schema's bare string/null types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb+resource ('Retrieve project information') and specifies lookup by ID or name. It clearly conveys a single-project lookup, but it does not explicitly contrast with sibling tools like list_projects or create_project, so differentiation is implicit rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description indicates when the tool is usefulβ€”when project information or the project_structure is neededβ€”and instructs agents to use the returned structure when refining prompts. However, it offers no explicit alternatives, exclusions, or when-not-to-use guidance relative to sibling tools such as list_projects or health_check.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

health_checkHealth CheckA

Health check endpoint.

Returns the health status of the SIMPA MCP service.

Examples: Request (no parameters needed): json {}

Returns: Service health status with version and timestamp

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
statusYes
serviceYes
versionYes
timestampYes

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It states the output (health status with version and timestamp) and implies a read-only operation through 'Returns' and 'health check'. However, it does not explicitly state that it has no side effects or requires no special permissions, which would make the safety profile clearer.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the purpose and includes a redundant example request that duplicates the empty schema. The example is unnecessary but not harmful; overall it is concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with an output schema, the description covers the purpose and return summary. It does not mention error conditions or authentication, but these are less critical for a health check. It is complete enough for an agent to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema fully documents the input. The description adds no parameter semantics but correctly shows an example with an empty request. Baseline for zero parameters is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action and resource: 'Returns the health status of the SIMPA MCP service'. It is immediately distinct from sibling tools which handle prompts and projects, so an agent can easily identify its purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context that this tool reports service health, implying it should be used to check service status. It does not explicitly mention alternatives, but no sibling provides health check functionality, so the context is sufficient. No exclusions are needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_projectsList ProjectsA

List all projects with optional filtering.

Retrieves a paginated list of projects, optionally filtered by programming language.

Project Scoping: Each project may include project_structure metadata that defines:

  • src_dirs: Recommended source directories (e.g., ["src/", "lib/"])

  • test_dirs: Test directories (e.g., ["tests/"])

  • entry_points: Main entry points (e.g., ["src/main.py"])

  • exclude: Paths to ignore (e.g., [".venv/", "pycache/"])

Use get_project to retrieve full structure details for a specific project, then use this information when scoping agent work via refine_prompt.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesList projects request with optional filters

Output Schema

ParametersJSON Schema
NameRequiredDescription
limitYes
offsetYes
projectsYes
total_countYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It does disclose pagination, optional filtering, and the variable project_structure metadata, which is useful, but it omits ordering, error behavior, authentication needs, and an explicit read-only statement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a clear summary and uses bullet points effectively for project_structure metadata. The first two sentences are somewhat redundant, which prevents a perfect score, but the overall structure is well organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a paginated list tool with a rich output schema and well-documented parameters, the description covers pagination, filtering, and downstream use of project_structure metadata. It does not address ordering or empty-result behavior, but these are not critical for successful invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description confirms that main_language is the filter and implies pagination via limit/offset, but it adds no additional semantic meaning beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'List all projects' and 'Retrieves a paginated list of projects'. It distinguishes itself from get_project by pointing out that full structure details belong to get_project, clarifying the division of labor between siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly directs agents to use get_project when full structure details for a specific project are needed, and mentions refine_prompt for scoping work. It does not state explicit 'when not to use' conditions, but the routing context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

refine_promptRefine PromptA

Refine a prompt before sending to an agent.

Given an original prompt and context, either selects an existing refined prompt or creates a new one optimized for the agent type and language.

Scoping Options (Optional but Recommended): To improve focus and reduce context overload, include in the context dict:

  • target_dirs: List of directories the agent should focus on (e.g., ["src/", "tests/"])

  • target_files: Specific files to modify (e.g., ["src/main.py", "src/config.py"])

  • exclude_paths: Paths to ignore (e.g., [".venv/", "node_modules/"])

  • scope: High-level scope description (e.g., "backend API layer only")

  • focus: Priority aspects (e.g., ["performance", "security", "error-handling"])

Note: A project_id is required. If not provided, the response will include a list of existing projects or instructions to create one. The agent should:

  1. Call list_projects to see available projects, or

  2. Call create_project to create a new one, then

  3. Resubmit the refine_prompt request with the project_id

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesRefinement request with prompt details

Output Schema

ParametersJSON Schema
NameRequiredDescription
actionYes
sourceYes
prompt_keyYes
usage_countNo
average_scoreNo
refined_promptYes
confidence_scoreNo
similar_prompts_foundNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does well: it discloses that the tool either selects an existing refined prompt or creates a new one, explains what happens when project_id is absent, and documents how context keys like target_dirs and focus affect behavior. This is substantial behavioral context beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a crisp one-line purpose and uses markdown headings, bullets, and numbered steps effectively. It is longer than average, but the added detail is functional rather than filler, especially given the nested request object and the project_id fallback workflow.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a nested request object, no annotations, and an output schema, the description is remarkably complete. It covers the main use case, optional scoping keys, required workflow around project_id, and sibling-tool handoffs. Nothing essential for invoking the tool correctly appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds meaning by documenting the recommended shape of the context object and clarifying that project_id is important for the workflow. It explains the intended use of agent_type and language more concretely than the schema alone, though it does not discuss every nested field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Refine a prompt before sending to an agent.' It clearly states the tool selects an existing refined prompt or creates a new one, distinguishing it from sibling tools like activate_prompt, deactivate_prompt, and project-management tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool ('before sending to an agent') and provides recommended scoping options to improve focus. It also gives an explicit fallback workflow when project_id is missing, directing the agent to list_projects or create_project before resubmitting, though it does not explicitly contrast against non-project siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_prompt_resultsUpdate Prompt ResultsA

Update prompt performance metrics after agent execution.

Records the outcome of using a refined prompt and updates the prompt's statistics for future refinement decisions.

Scoping Feedback: Use files_modified and files_added to record which files were actually touched. This helps the system understand the effective scope of the prompt and can be used to suggest narrower scopes for similar future tasks.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesUpdate request with prompt key and results

Output Schema

ParametersJSON Schema
NameRequiredDescription
successYes
usage_countYes
last_used_atNo
average_scoreYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It transparently states that the tool updates prompt statistics and can influence future refinement decisions, and it adds useful context about scoping feedback. Yet it does not disclose details like whether scores are overwritten or accumulated, idempotency, or permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, focused sections. The main purpose is front-loaded in the first sentence, and the bolded 'Scoping Feedback' section earns its place by providing actionable guidance. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is complex with a nested schema and many result fields, but the description covers when to call it and clarifies the purpose of two key file-tracking parameters. The output schema presumably covers return values, so the description doesn't need to. It could be more complete by explaining the meaning of action_score or how metrics are aggregated, but it is largely sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The top-level schema has 100% description coverage, giving a baseline of 3. The tool description goes beyond the schema by explaining the semantic purpose of files_modified and files_added – recording actual touched files and enabling narrower scope suggestions – which adds real value for the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Update') and resource ('prompt performance metrics'), and clarifies that it records outcomes after agent execution. This distinguishes it from sibling tools like refine_prompt and activate_prompt, which focus on generating or toggling prompts rather than recording results.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: this tool should be used after agent execution to record the outcome of a refined prompt, and it explains why files_modified and files_added matter for scoping feedback. However, it does not explicitly name when to use a sibling tool instead, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv0.1.0
    • First observedactivate_prompt
    • First observedcreate_project
    • First observeddeactivate_prompt
    • First observedget_project
    • First observedhealth_check
    • First observedlist_projects
    • First observedrefine_prompt
    • First observedupdate_prompt_results

TDQS

A4.1/5.0

Scored across 8 tools

Disambiguation5/5

Each tool has a distinct purpose: refine_prompt handles prompt optimization, update_prompt_results tracks metrics, health_check monitors service status, create/get/list_projects manage project metadata, and activate/deactivate_prompt control prompt lifecycle. No two tools overlap in function.

Naming Consistency4/5

Most tools follow a consistent verb_noun pattern (refine_prompt, create_project, list_projects, activate_prompt, etc.). The exception is health_check, which is a noun phrase rather than a verb_noun, making it slightly inconsistent but still readable.

Tool Count5/5

With 8 tools, the server is well-scoped for a prompt management system. Each tool covers a necessary functionβ€”project management, prompt refinement, result tracking, and lifecycle controlβ€”without redundancy or bloat.

Completeness4/5

The core workflows are covered: project creation and lookup, prompt refinement and deactivation, result feedback, and health checks. Missing project update/delete operations and a dedicated get_prompt tool are minor gaps that agents can work around, but the surface is largely complete.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Engram MCP provides persistent, cross-session memory for AI agents by automatically encoding errors, decisions, and discoveries during development sessions. It enables local, intelligent recall and automated context management to help AI learn from experience and avoid recurring mistakes.
    6 npm
    Business Source 1.1
  • A
    license
    A
    quality
    D
    maintenance
    An MCP server that automatically enhances user prompts by applying advanced engineering techniques like chain-of-thought and few-shot reasoning based on identified intent. It optimizes technique selection through local learning and integrates directly into Claude sessions to improve output quality without additional API costs.
    6
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    MCP server for persistent, compounding memory that automatically captures corrections and insights across AI sessions, enabling agents to learn and improve over time.
    5
    371
    MIT