Skip to main content
Glama
LumabyteCo

Clarifyprompt-MCP

by LumabyteCo

ClarifyPrompt MCP

npm version License: Apache-2.0 Node.js

A context-aware MCP prompt compiler that transforms vague prompts into platform-optimized prompts for 58+ AI platforms across 7 categories — grounded in your workspace signals (CLAUDE.md, AGENTS.md, .cursorrules, package.json), resolved intent, and the capabilities of the target model.

Send a raw prompt. ClarifyPrompt gathers the right context, resolves what you're actually trying to do, and returns a version specifically optimized for Midjourney, DALL-E, Sora, Runway, ElevenLabs, Claude, ChatGPT, Cursor, or any of the 58+ supported platforms — with the right syntax, parameters, structure, and grounding.

New in 1.2.0: Context Engine — automatic workspace signal gathering, intent resolution, target-model capability hints, and local JSONL tracing. See CHANGELOG.md.

How It Works

You write:    "a dragon flying over a castle at sunset"

ClarifyPrompt returns (for Midjourney):
  "a majestic dragon flying over a medieval castle at sunset
   --ar 16:9 --v 6.1 --style raw --q 2 --chaos 30 --s 700"

ClarifyPrompt returns (for DALL-E):
  "A majestic dragon flying over a castle at sunset. Size: 1024x1024"

Same prompt, different platform, completely different output. ClarifyPrompt knows what each platform expects — and in 1.2.0, it also knows what you're working on.

Related MCP server: Refine Prompt

What's in the box (1.2.0)

  • Context Engine — auto-gathers workspace rules (CLAUDE.md, AGENTS.md, .cursorrules, .clinerules, clarify.md), detects frameworks and languages from package.json and sibling manifests, tracks an active file excerpt, and maintains a per-session ring buffer of recent optimizations and their outcomes.

  • Unified PromptAnalyzer — one LLM call produces { category, intent, recommendedMode, confidence } together. 10 intents: production-code, brand-voice, stakeholder-comm, data-extract, creative-media, technical-spec, analysis, quick-draft, exploration, unknown. Intent beats surface keywords on ambiguity.

  • Target-model-aware prompt shaping — system prompt, maxTokens, and temperature adapt to the downstream LLM's context window and the resolved intent. Small local models get a compact prompt; Claude/GPT-4/Gemini get the full richness.

  • Grounding Context (single, priority-ordered) — user pinned instructions → project rules → active file → prior accepted examples → web search → workspace metadata → target-model hints → custom platform instructions → built-in syntax hints. No more parallel context silos.

  • Session retrieval (save_outcome) — the caller reports accepted | edited | rejected per optimization; similar accepted outputs in the same session get injected as few-shot examples into future similar prompts. Persistent memory lands in 1.3.

  • Local JSONL tracing — every optimization writes a structured trace line (now with shape, groundingSources, error fields) to $CLARIFYPROMPT_HOME/traces/YYYY-MM-DD.jsonl. Nothing is uploaded. Toggle via CLARIFYPROMPT_TRACE=off.

  • Unified $CLARIFYPROMPT_HOME — one env var for everything ClarifyPrompt writes. Legacy CLARIFYPROMPT_CONFIG_DIR / CLARIFYPROMPT_DATA_DIR still work (deprecation hint, silenceable).

  • 58+ platforms, 7 categories, custom platforms — the original core is unchanged and fully backward-compatible.

  • Any LLM, any provider. One code path works with any OpenAI-compatible API — Ollama (local + cloud), LM Studio, vLLM, OpenAI, Google Gemini, xAI Grok, Groq, Mistral, DeepSeek, Cohere, Perplexity, Together, Fireworks, OpenRouter — plus Anthropic Claude directly. Reasoning models (o1/o3/o4, deepseek-reasoner, gpt-oss, *-thinking) are auto-detected and given a larger token budget so they actually produce content. See 15+ pre-configured provider examples below.

  • Apache-2.0, forever. Open-source core, no relicensing.

Quick Start

With Claude Desktop

Add to your claude_desktop_config.json:

{
  "mcpServers": {
    "clarifyprompt": {
      "command": "npx",
      "args": ["-y", "clarifyprompt-mcp"],
      "env": {
        "LLM_API_URL": "http://localhost:11434/v1",
        "LLM_MODEL": "qwen2.5:7b"
      }
    }
  }
}

With Claude Code

claude mcp add clarifyprompt -- npx -y clarifyprompt-mcp

Set the environment variables in your shell before launching:

export LLM_API_URL=http://localhost:11434/v1
export LLM_MODEL=qwen2.5:7b

With Cursor

Add to your .cursor/mcp.json:

{
  "mcpServers": {
    "clarifyprompt": {
      "command": "npx",
      "args": ["-y", "clarifyprompt-mcp"],
      "env": {
        "LLM_API_URL": "http://localhost:11434/v1",
        "LLM_MODEL": "qwen2.5:7b"
      }
    }
  }
}

Supported Platforms (58+ built-in, unlimited custom)

Category

Platforms

Default

Image (10)

Midjourney, DALL-E 3, Stable Diffusion, Flux, Ideogram, Leonardo AI, Adobe Firefly, Grok Aurora, Google Imagen 3, Recraft

Midjourney

Video (11)

Sora, Runway Gen-3, Pika Labs, Kling AI, Luma, Minimax/Hailuo, Google Veo 2, Wan, HeyGen, Synthesia, CogVideoX

Runway

Chat (9)

Claude, ChatGPT, Gemini, Llama, DeepSeek, Qwen, Kimi, GLM, Minimax

Claude

Code (9)

Claude, ChatGPT, Cursor, GitHub Copilot, Windsurf, DeepSeek Coder, Qwen Coder, Codestral, Gemini

Claude

Document (8)

Claude, ChatGPT, Gemini, Jasper, Copy.ai, Notion AI, Grammarly, Writesonic

Claude

Voice (7)

ElevenLabs, OpenAI TTS, Fish Audio, Sesame, Google TTS, PlayHT, Kokoro

ElevenLabs

Music (4)

Suno AI, Udio, Stable Audio, MusicGen

Suno

Tools

optimize_prompt

The main tool. Optimizes a prompt for a specific AI platform.

{
  "prompt": "a cat sitting on a windowsill",
  "category": "image",
  "platform": "midjourney",
  "mode": "concise"
}

All parameters except prompt are optional. When category and platform are omitted, ClarifyPrompt auto-detects them from the prompt content.

Three calling modes:

Mode

Example

Zero-config

{ "prompt": "sunset over mountains" }

Category only

{ "prompt": "...", "category": "image" }

Fully explicit

{ "prompt": "...", "category": "image", "platform": "dall-e" }

Parameters:

Parameter

Required

Description

prompt

Yes

The prompt to optimize

category

No

chat, image, video, voice, music, code, document. Auto-detected when omitted.

platform

No

Platform ID (e.g. midjourney, dall-e, sora, claude). Uses category default when omitted.

mode

No

Output style: concise, detailed, structured, step-by-step, bullet-points, technical, simple. Default: detailed.

enrich_context

No

Set true to use web search for context enrichment. Default: false.

session_id

No

Stitches related optimizations together so session memory can bias subsequent calls. Auto-generated when omitted.

file_path

No

Active file path — infers language and shapes platform hints.

file_language

No

Explicit language override for the active file.

file_excerpt

No

Short excerpt (≤2 KB) of the active file to ground the rewrite.

cwd

No

Working directory to scan for CLAUDE.md / AGENTS.md / .cursorrules / package.json. Defaults to server cwd.

user_locale

No

Locale hint (e.g. en-US, ar-EG) to inform tone and language.

user_pinned_instructions

No

Pinned, always-applied user instructions (short core-memory block).

include_bundle

No

Include the resolved ContextBundle summary in the response. Default: false.

skip_intent_resolution

No

Skip the intent classifier LLM call (faster; loses intent signal). Default: false.

Response (1.2.0):

{
  "id": "opt_mo9vlg9i_foohjx",
  "sessionId": "sess_mo9vlfn3_abc123",
  "originalPrompt": "a dragon flying over a castle at sunset",
  "optimizedPrompt": "a majestic dragon flying over a medieval castle at sunset --ar 16:9 --v 6.1 --style raw --q 2 --s 700",
  "category": "image",
  "platform": "midjourney",
  "mode": "concise",
  "modeSource": "analyzer",
  "analysis": {
    "category": "image",
    "intent": "creative-media",
    "recommendedMode": "detailed",
    "confidence": "high",
    "source": "llm"
  },
  "grounding": {
    "sources": ["project-rules", "workspace-meta", "target-model", "platform-hints"],
    "acceptedExamplesUsed": 0
  },
  "shape": {
    "systemPromptBudget": "standard",
    "maxTokens": 2048,
    "temperature": 0.9
  },
  "metadata": {
    "model": "qwen2.5:14b-instruct-q4_K_M",
    "processingTimeMs": 3911,
    "strategy": "ImageStrategy"
  },
  "detection": { "autoDetected": true, "detectedCategory": "image", "detectedPlatform": "midjourney", "confidence": "high" },
  "intent": { "detected": "creative-media", "confidence": "high" }
}

The canonical classification field is analysis. The detection and intent fields are deprecated aliases kept for 1.x back-compat; they will be removed in 2.x.

modeSource tells you how the final mode was decided (user if you passed one, analyzer if intent-driven, default if neither).

grounding.sources lists which Grounding Context sections contributed, in priority order. grounding.acceptedExamplesUsed tells you how many few-shot examples the engine pulled from save_outcome history.

shape tells you how the system prompt was sized for your target model.

inspect_context (new in 1.2.0)

Preview the ContextBundle ClarifyPrompt would assemble for a given prompt — workspace rules, frameworks, target-model capabilities, resolved intent, and session history — without running the full optimization. Useful for debugging why an optimization turned out the way it did.

{
  "prompt": "Write an email to finance explaining the Q2 spend variance",
  "category": "document",
  "cwd": "/path/to/your/project"
}

Returns the full ContextBundle as JSON.

list_traces (new in 1.2.0)

Summary list of recent optimization traces captured by the local tracer (when CLARIFYPROMPT_TRACE=local, the default).

{ "day": "2026-04-22", "limit": 50 }

Returns trace IDs, inputs previews, resolved intents, target families, and latencies — never the full system prompt (use get_trace for that). Omit day to get the most recent day with data.

get_trace (new in 1.2.0)

Fetch the full trace for a single optimization by ID, including the exact system prompt, bundle summary, and output.

{ "id": "opt_xxx", "lookback_days": 7 }

save_outcome (new in 1.2.0)

Tell ClarifyPrompt whether a past optimization was accepted, edited, or rejected. Accepted outputs become few-shot examples for similar future prompts in the same session. In 1.3+ this will also feed the persistent memory layer. The IDE / agent / caller is expected to invoke this after the user acts on the optimization.

{
  "optimization_id": "opt_xxx",
  "session_id": "sess_yyy",
  "verdict": "accepted",
  "diff": "optional: the user's edited version or a patch"
}

list_categories

Lists all 7 categories with platform counts (built-in and custom) and defaults.

list_platforms

Lists available platforms for a given category, including custom registered platforms. Shows which is the default and whether custom instructions are configured.

list_modes

Lists all 7 output modes with descriptions.

register_platform

Register a new custom AI platform for prompt optimization.

{
  "id": "my-llm",
  "category": "chat",
  "label": "My Custom LLM",
  "description": "Internal fine-tuned model",
  "syntax_hints": ["JSON mode", "max 2000 tokens"],
  "instructions": "Always use structured output format",
  "instructions_file": "my-llm.md"
}

Parameter

Required

Description

id

Yes

Unique ID (lowercase, alphanumeric with hyphens)

category

Yes

Category this platform belongs to

label

Yes

Human-readable platform name

description

Yes

Short description

syntax_hints

No

Platform-specific syntax hints

instructions

No

Inline optimization instructions

instructions_file

No

Path to a .md file with detailed instructions

update_platform

Update a custom platform or add instruction overrides to a built-in platform.

For built-in platforms (e.g. Midjourney, Claude), you can add custom instructions and extra syntax hints without modifying the originals:

{
  "id": "midjourney",
  "category": "image",
  "instructions": "Always use --v 6.1, prefer --style raw",
  "syntax_hints_append": ["--no plants", "--tile for patterns"]
}

For custom platforms, all fields can be updated.

unregister_platform

Remove a custom platform or clear instruction overrides from a built-in platform.

{
  "id": "my-llm",
  "category": "chat"
}

For built-in platforms, use remove_override_only: true to clear your custom instructions without affecting the platform itself.

Custom Platforms & Instructions

ClarifyPrompt supports registering custom platforms and providing optimization instructions — similar to how .cursorrules or CLAUDE.md guide AI behavior.

How It Works

  1. Register a custom platform via register_platform

  2. Provide instructions inline or as a .md file

  3. Optimize prompts targeting your custom platform — instructions are injected into the optimization pipeline

Instruction Files

Instructions can be provided as markdown files stored at ~/.clarifyprompt/instructions/:

~/.clarifyprompt/
  config.json                    # custom platforms + overrides
  instructions/
    my-llm.md                   # instructions for custom platform
    midjourney-overrides.md     # extra instructions for built-in platform

Example instruction file (my-llm.md):

# My Custom LLM Instructions

## Response Format
- Always output valid JSON
- Include a "reasoning" field before the answer

## Constraints
- Max 2000 tokens
- Temperature should be set low (0.1-0.3) for factual queries

## Style
- Be concise and technical
- Avoid filler phrases

Override Built-in Platforms

You can add custom instructions to any of the 58 built-in platforms using update_platform. This lets you customize how prompts are optimized for platforms like Midjourney, Claude, or Sora without modifying the defaults.

Config Directory

The config directory defaults to ~/.clarifyprompt/ and can be changed via the CLARIFYPROMPT_CONFIG_DIR environment variable. Custom platforms and overrides persist across server restarts.

LLM Configuration

ClarifyPrompt uses an LLM to optimize prompts. It works with any OpenAI-compatible API and with the Anthropic API directly.

Environment Variables

Variable

Required

Description

LLM_API_URL

Yes

API endpoint URL

LLM_API_KEY

Depends

API key (not needed for local Ollama)

LLM_MODEL

Yes

Model name/ID

CLARIFYPROMPT_HOME

No

Canonical (1.2.0+) root for everything ClarifyPrompt writes — custom platforms, instruction .md files, traces, and (1.3+) memory + packs. Default: $XDG_DATA_HOME/clarifyprompt or ~/.clarifyprompt.

CLARIFYPROMPT_TRACE

No

off | local | otel. Default: local. Traces are strictly local JSONL; nothing is uploaded.

CLARIFYPROMPT_SUPPRESS_LEGACY_WARN

No

Set to 1 to silence the one-line deprecation hint when CLARIFYPROMPT_CONFIG_DIR / CLARIFYPROMPT_DATA_DIR are used.

CLARIFYPROMPT_CONFIG_DIR

No

Legacy alias for CLARIFYPROMPT_HOME. Still works; will be removed in 2.x.

CLARIFYPROMPT_DATA_DIR

No

Legacy alias for CLARIFYPROMPT_HOME. Still works; will be removed in 2.x.

Provider Examples

Ollama (local, free):

LLM_API_URL=http://localhost:11434/v1
LLM_MODEL=qwen2.5:7b

Ollama — cloud models via local passthrough (recommended):

If your local Ollama is signed in to Ollama Cloud, any :cloud model routes through it transparently — same URL, no separate API key. The capability table auto-detects reasoning / thinking variants (gpt-oss, kimi-k2-thinking, qwen3-thinking, deepseek-r1, etc.) and bumps maxTokens so they finish thinking and actually produce content.

LLM_API_URL=http://localhost:11434/v1
LLM_MODEL=gpt-oss:20b-cloud        # or kimi-k2.6:cloud, qwen3-next:80b-cloud, glm-4.6:cloud, etc.

Ollama — direct cloud endpoint (no local install):

LLM_API_URL=https://ollama.com/v1
LLM_API_KEY=your-ollama-cloud-key
LLM_MODEL=qwen2.5:7b

OpenAI:

LLM_API_URL=https://api.openai.com/v1
LLM_API_KEY=sk-...
LLM_MODEL=gpt-4o

Anthropic Claude:

LLM_API_URL=https://api.anthropic.com/v1
LLM_API_KEY=sk-ant-...
LLM_MODEL=claude-sonnet-4-20250514

Google Gemini:

LLM_API_URL=https://generativelanguage.googleapis.com/v1beta/openai
LLM_API_KEY=your-gemini-key
LLM_MODEL=gemini-2.0-flash

Groq:

LLM_API_URL=https://api.groq.com/openai/v1
LLM_API_KEY=gsk_...
LLM_MODEL=llama-3.3-70b-versatile

DeepSeek:

LLM_API_URL=https://api.deepseek.com/v1
LLM_API_KEY=your-deepseek-key
LLM_MODEL=deepseek-chat

OpenRouter (any model):

LLM_API_URL=https://openrouter.ai/api/v1
LLM_API_KEY=your-openrouter-key
LLM_MODEL=anthropic/claude-sonnet-4

See .env.example for the full list of 20+ supported providers including Together AI, Fireworks, Mistral, xAI, Cohere, Perplexity, LM Studio, vLLM, LocalAI, Jan, GPT4All, and more.

Web Search (Optional)

Enable context enrichment by setting enrich_context: true in your optimize_prompt call. ClarifyPrompt will search the web for relevant context before optimizing.

Supported search providers:

Provider

Variable

URL

Tavily (default)

SEARCH_API_KEY

tavily.com

Brave Search

SEARCH_API_KEY

brave.com/search/api

Serper

SEARCH_API_KEY

serper.dev

SerpAPI

SEARCH_API_KEY

serpapi.com

Exa

SEARCH_API_KEY

exa.ai

SearXNG (self-hosted)

—

github.com/searxng/searxng

SEARCH_PROVIDER=tavily
SEARCH_API_KEY=your-key

Before and After

Image (Midjourney)

Before: "a cat sitting on a windowsill"

After:  "a tabby cat sitting on a sunlit windowsill, warm golden hour
         lighting, shallow depth of field, dust particles in light beams,
         cozy interior background, shot on 35mm film, warm amber color
         palette --ar 16:9 --v 6.1 --style raw --q 2"

Video (Sora)

Before: "a timelapse of a city"

After:  "Cinematic timelapse of a sprawling metropolitan skyline
         transitioning from golden hour to blue hour to full night.
         Camera slowly dollies forward from an elevated vantage point.
         Light trails from traffic appear as the city illuminates.
         Clouds move rapidly overhead. Duration: 10s.
         Style: documentary cinematography, 4K."

Code (Claude)

Before: "write a function to validate emails"

After:  "Write a TypeScript function `validateEmail(input: string): boolean`
         that validates email addresses against RFC 5322. Handle edge cases:
         quoted local parts, IP address domains, internationalized domain
         names. Return boolean, no exceptions. Include JSDoc with examples
         of valid and invalid inputs. No external dependencies."

Music (Suno)

Before: "compose a chill lo-fi beat for studying"

After:  "Compose an instrumental chill lo-fi beat for studying.
         [Tempo: medium] [Genre: lo-fi] [Length: 2 minutes]"

Context Engine (1.2.0)

Every optimization runs through five integrated passes that flow one bundle of context end-to-end:

  1. Analysis — a single analyzePrompt() LLM call produces category, intent, and recommendedMode together so they can't disagree. Intent beats surface keywords when they conflict (e.g. "validate emails" → code not document).

  2. Mode reconciliation — explicit user mode wins; otherwise the analyzer's intent-derived recommendation applies; modeSource in the response tells you which.

  3. Prompt shaping — target-model capability signal drives systemPromptBudget (compact for small local models, rich for 100K+ ctx models), maxTokens, temperature (intent-aware), and whether examples are included.

  4. Intent overlay — a short overlay per intent (production-code: demand error handling + tests; data-extract: demand strict schema; brand-voice: lead with tone; etc.) folded into the strategy's system prompt.

  5. Grounding Context — a single priority-ordered block that merges user pinned instructions → project rules → active file → session few-shot examples → web search → workspace metadata → target-model hints → custom platform instructions → built-in syntax hints.

What's collected (ContextBundle)

  • Project — first matching file from CLAUDE.md, AGENTS.md, .cursorrules, .clinerules, clarify.md, .clarify/rules.md. package.json plus sibling manifests (pyproject.toml, Cargo.toml, go.mod, Gemfile, composer.json, …) drive framework + language detection.

  • File — optional file_path / file_language / file_excerpt inputs.

  • Session — ring buffer (20 ops/session) of recent optimizations and outcomes. Accepted outputs get retrieved as few-shot examples for similar future prompts.

  • Target model — the LLM doing the rewrite, matched against a capability table.

  • User — locale, preferred mode, pinned instructions (highest-priority grounding).

Inspecting what the engine sees

Use the inspect_context tool to preview the full bundle without running an optimization. Same shape as optimize_prompt returns when include_bundle: true.

Extending context

Drop an AGENTS.md / clarify.md / CLAUDE.md at your project root. Next optimization picks it up automatically. To feed accepted outputs back into future rewrites, call save_outcome after the user acts on the result.

Tracing

$CLARIFYPROMPT_HOME/traces/YYYY-MM-DD.jsonl

Every optimization writes one JSONL line capturing {id, ts, sessionId, category, platform, mode, input, bundleSummary, systemPrompt, output, model, strategy, latencyMs, shape, groundingSources, error}. Use list_traces for summaries and get_trace for full records.

Privacy posture:

  • Traces are strictly local. No outbound network calls to any ClarifyPrompt-owned infrastructure.

  • Only calls out to the LLM endpoint you configured (LLM_API_URL) and optional search provider (SEARCH_API_KEY).

  • Disable tracing entirely with CLARIFYPROMPT_TRACE=off.

  • There is no telemetry in this release. When a telemetry option ships it will be opt-in, anonymous, and documented before the build includes it.

Known limitations & roadmap

Session memory is in-memory only (today)

The save_outcome + few-shot retrieval loop writes into a per-process ring buffer. Restarting the MCP server clears session state; two servers don't share memory. The MCP tool surface is deliberately stable — the interface won't change in 1.3. The upgrade is purely a backend swap to SQLite + sqlite-vec for disk persistence and richer similarity. Ship target: 1.3.

Intent quality scales with the model running the analyzer

The analyzer runs on the same LLM_MODEL that does the rewrite. In the integration battery:

  • Qwen 2.5 7B and 14B → correct on every well-formed prompt tested.

  • Llama 3.2 3B → occasionally over-commits on ambiguous prompts (e.g. tagged "make it better" as brand-voice/high when unknown/low is the right answer). Larger models on the same prompt correctly returned unknown/low.

Guidance: prefer a 7B+ local model (or any frontier hosted model) as LLM_MODEL. Latency-sensitive callers can set skip_intent_resolution: true to skip the analyzer; the engine falls back to user-hint category and default mode, losing intent-driven mode + overlay but keeping grounding + shape. A systematic eval harness with a public fixture set lands in 1.3 (Day 3) so you can score the analyzer against your own fixtures and detect regressions across model or classifier changes.

Capability table is not exhaustive

Entries today: Claude, GPT-4/o-series, Gemini, Grok, DeepSeek (chat + reasoning), Qwen, Llama, Mistral/Codestral, Mixtral, Gemma, Phi, Cohere Command, Aya, Kimi, GLM, Minimax, GPT-OSS, Yi, Nemotron. Unknown models fall back to capabilities: {} and standard prompt-shape — still functional, just without model-aware sizing. Adding entries is a data-only edit to src/engine/context/targetModelSignals.ts.

Reasoning / chain-of-thought models

Supported as a first-class case. The engine auto-detects reasoners at family level (o1/o3/o4, deepseek-reasoner, gpt-oss) and at variant level (anything whose ID matches /\b(thinking|reasoner|reasoning)\b/ or /\br[12]\b/: kimi-k2-thinking:cloud, qwen3-thinking:72b, qwen-r1-distill, etc.). For these, maxTokens is automatically bumped to ≥ 8192 so the model has room to think AND produce content. The reasoning field is never surfaced as the optimized prompt — only content is.

Architecture

clarifyprompt-mcp/
  src/
    index.ts                           MCP server entry point (11 tools, 1 resource)
    engine/
      config/
        categories.ts                  7 categories, 58 platforms, 7 modes
        paths.ts                       Unified $CLARIFYPROMPT_HOME resolver (1.2.0)
        persistence.ts                 ConfigStore — JSON config + .md file loading
        registry.ts                    PlatformRegistry — merges built-in + custom
      context/                         Context Engine (1.2.0)
        types.ts                       ContextBundle + signal types + AnalysisSignal
        projectSignals.ts              CLAUDE.md / AGENTS.md / .cursorrules / manifests scan
        fileSignals.ts                 Active-file path + language + excerpt
        sessionSignals.ts              In-memory per-session ring buffer + outcome retrieval
        targetModelSignals.ts          Model → capabilities mapping
        promptAnalyzer.ts              Unified analyzer: category + intent + recommendedMode
        bundle.ts                      Bundle orchestrator
      trace/                           Local tracing (1.2.0)
        types.ts                       TraceEntry schema (shape, groundingSources, error)
        writer.ts                      JSONL + OTel-stub writer, reader, lookup
      llm/client.ts                    Multi-provider LLM client (OpenAI + Anthropic)
      search/client.ts                 Web search (6 providers; results merge into Grounding Context)
      optimization/
        engine.ts                      Core orchestrator — analyzer, shape, grounding, retrieval, trace
        groundingContext.ts            Priority-ordered context assembly + mode/shape helpers
        types.ts                       OptimizationContext + result shape
        strategies/
          base.ts                      Bundle-aware base strategy (intent overlay + shape-aware sizing)
          chat.ts                      9 platforms
          image.ts                     10 platforms
          video.ts                     11 platforms
          voice.ts                     7 platforms
          music.ts                     4 platforms
          code.ts                      9 platforms
          document.ts                  8 platforms

Docker

docker build -t clarifyprompt-mcp .
docker run -e LLM_API_URL=http://host.docker.internal:11434/v1 -e LLM_MODEL=qwen2.5:7b clarifyprompt-mcp

Development

git clone https://github.com/LumabyteCo/clarifyprompt-mcp.git
cd clarifyprompt-mcp
npm install
npm run build

Test with MCP Inspector:

npx @modelcontextprotocol/inspector node dist/index.js

Set environment variables in the Inspector's "Environment Variables" section before connecting.

License

Apache-2.0

Available Tools

23 tools
clarify_with_userA

Given an ambiguous draft prompt, return 1–3 targeted clarifying questions instead of guessing. Each question carries a suggested_answer you can accept verbatim to keep moving, an optional 2–4 quick-pick options list, and a dimension tag (audience/scope/format/length/tone/constraints/goal/platform). When the analyzer is highly confident AND the prompt is non-trivially long, the tool short-circuits with clarificationNeeded: false so callers can pipeline this in front of optimize_prompt without paying a latency tax on every call. Pass force: true to always generate questions.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesThe draft prompt the user is unsure about.
categoryNoCategory hint. Will skip questions about category/platform if you pass it.
cwdNoWorking directory to pull workspace rules (CLAUDE.md / AGENTS.md / .cursorrules) from. Defaults to server cwd.
file_pathNoActive file path — informs the clarifier's defaults.
file_languageNoExplicit language override for the active file.
file_excerptNoShort excerpt of the active file to ground the questions.
user_localeNo
forceNoAlways generate questions even when the analyzer is highly confident. Useful for UIs that want to surface clarification on every call.
max_questionsNoCap on returned questions. Default 3, hard max 5.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully carries the burden of behavioral disclosure. It explains the short-circuit behavior (clarificationNeeded: false), the structure of each question (suggested_answer, options, dimension), and the effect of the force flag. It also notes that passing a category skips questions about category/platform. This is comprehensive for a non-destructive tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is thorough but not overly verbose. It front-loads the core purpose and then expands on behavior and structure. Every sentence contributes useful information. Could be slightly more compact, but it remains efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 9 parameters (1 required), high schema coverage, and no output schema, the description provides complete context. It explains the tool's behavior, response format, short-circuit logic, and ties to sibling tools (optimize_prompt). No critical gaps are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 89%, so the schema already documents most parameters. The description adds value by explaining the structure of the generated questions (suggested_answer, options, dimension) and the effect of force: true. It also clarifies how category can skip certain questions. These details enhance understanding beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Given an ambiguous draft prompt, return 1–3 targeted clarifying questions instead of guessing.' It specifies the verb (return), resource (clarifying questions), and distinguishes from alternative behaviors (short-circuiting). This is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use and when-not-to-use guidance: it explains the short-circuit behavior when the analyzer is highly confident and the prompt is non-trivially long, and mentions pipelining in front of optimize_prompt. It also describes the force parameter for overriding the short-circuit. This clearly differentiates usage contexts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compose_promptA

Run the canonical ClarifyPrompt pipeline in ONE call: clarify (optional pre-stage) → ground OR optimize (core) → critique (optional post-stage) → optional auto-revise. Use this when you want the four-tool happy path without orchestrating five round-trips. Short-circuits if pre_clarify surfaces questions — caller answers and re-calls. When sources is non-empty the chain takes the strict ground_prompt branch; otherwise it goes through optimize_prompt. When auto_revise is true and critique returns a non-accept verdict with an improved rewrite, final_prompt is the rewrite. The stages array is a per-call audit log so callers can see exactly what ran.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesThe prompt to compose.
pre_clarifyNo'auto' = run clarify only if analyzer confidence is low / prompt is short. 'always' = force clarify. 'never' = skip. When clarification questions surface, the chain stops; caller answers and re-calls.auto
max_questionsNo
sourcesNoWhen non-empty, the chain takes the strict ground_prompt branch (caller-provided sources pinned at highest priority).
post_critiqueNoRun the critique judge against the optimized output. Adds ~3-5s on a local model.
revise_thresholdNo
critique_criteriaNoOverride the default 5 critique criteria.
auto_reviseNoWhen true AND post_critique is true AND verdict !== 'accept' AND there's an improvedPrompt: `final_prompt` becomes the rewritten version instead of the raw optimization.
max_iterationsNoMax revise-loop iterations. With `auto_revise: true` AND `post_critique: true`, the engine can feed each iteration's improvedPrompt back through optimize+critique up to this cap. Stops early at verdict=accept or when there's no improvedPrompt. Default 1 (single-shot, no loop). Hard max 5 to prevent cost runaways.
clarify_modelNoOverride the LLM model for the clarify pre-stage. Default: env LLM_MODEL. Useful for per-stage cost/quality routing — e.g. run clarify on a cheap model while critique runs on a frontier one.
optimize_modelNoOverride the LLM model for the optimize/ground core stage.
critique_modelNoOverride the LLM model for the critique judge AND rewrite.
categoryNo
platformNo
modeNo
enrich_contextNo
session_idNo
file_pathNo
file_languageNo
file_excerptNo
cwdNo
user_localeNo
user_pinned_instructionsNo
skip_intent_resolutionNo
include_bundleNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It details the pipeline flow, short-circuit behavior, branch conditions, and auto-revise loop. It mentions stages as an audit log and cost limits (max_iterations). However, it lacks disclosure on potential side effects, auth needs, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph that front-loads the purpose and then explains behaviors. It is dense but efficient for the complexity. Could be improved with bullet points for scannability, but remains concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 25 parameters, 40% schema coverage, and no output schema, the description falls short. It does not describe the output structure (e.g., final_prompt, stages) nor error conditions. Many contextual parameters (session_id, file_path, etc.) are undocumented, leaving gaps for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 40%, so the description must compensate. It adds value for core parameters (pre_clarify, sources, post_critique, auto_revise, max_iterations, model overrides) explaining their behavior. However, many parameters (category, platform, mode, file_path, etc.) are not described, relying solely on schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool runs the canonical ClarifyPrompt pipeline in one call, covering clarify, ground/optimize, critique, and auto-revise. It distinguishes from sibling tools by explicitly noting it replaces orchestrating five round-trips. The branching based on sources (ground vs optimize) is also specified.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use the tool: when wanting the four-tool happy path without orchestrating. It covers short-circuit behavior for pre_clarify, branching conditions, and auto-revise. However, it does not explicitly state when not to use it (e.g., for fine-grained control, use individual tools), though this is implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

critique_promptA

LLM-as-judge for a prompt. Scores it 0–10 across 5 default dimensions (clarity, specificity, intent_alignment, format_fitness, length_appropriateness) — or your own custom criteria — and returns per-dimension rationale + concrete suggestions, an overall score, and a verdict (accept / revise / reject). When the score is below revise_threshold (default 7.0), the tool also returns an improvedPrompt you can use as a drop-in replacement. Use it pre-flight (is this prompt good enough for the expensive model?), postmortem (was the prompt the cause of a bad output?), or to A/B-pick the best of N optimization variants. Pass original_prompt when critiquing an optimized version so the judge can verify intent was preserved.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesThe candidate prompt to critique.
original_promptNoIf `prompt` is an optimized version, the user's original ask. Used for the intent_alignment dimension.
categoryNo
cwdNo
file_pathNo
file_languageNo
file_excerptNo
user_localeNo
criteriaNoOverride the default 5 criteria. Up to ~8 dimensions; more bloats the judge call.
revise_thresholdNoOverall score below this triggers the rewrite pass. Default 7.0.
skip_rewriteNoSkip the rewrite pass even when below threshold (faster; just returns scores).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully discloses key behaviors: it returns per-dimension rationale, concrete suggestions, an overall score, and a verdict. It explains that when below 'revise_threshold' (default 7.0), it returns an 'improvedPrompt'. It also mentions custom criteria and skip_rewrite. It does not discuss side effects or costs, but for a critique tool the disclosure is thorough.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and results, then explains use cases and special parameters. Each sentence adds value without redundancy. It is appropriately sized for the complexity—neither too terse nor verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 11 parameters and no output schema, the description covers the main function, return values (rationale, suggestions, score, verdict, improvedPrompt), and key optional parameters. It does not explain every parameter, but the core functionality is well-documented. The output structure is sufficiently described for an agent to understand what to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 45%, so the description must add meaning. It does for key parameters: 'prompt' (candidate prompt), 'original_prompt' (intent preservation for optimized versions), 'criteria' (custom dimensions), 'revise_threshold', and 'skip_rewrite'. However, parameters like 'cwd', 'file_path', 'file_language', 'file_excerpt', and 'user_locale' are not explained in the description, leaving gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with 'LLM-as-judge for a prompt', clearly stating the tool's core purpose. It specifies it scores 0–10 across dimensions, returns rationale, suggestions, overall score, and a verdict (accept/revise/reject). The name and verb 'critique' align, and the description distinguishes from siblings by mentioning pre-flight, postmortem, and A/B testing use cases.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool: 'pre-flight', 'postmortem', or 'to A/B-pick the best of N optimization variants'. It also advises passing 'original_prompt' when critiquing an optimized version. However, it does not mention when not to use it or provide explicit alternatives among the sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

explain_last_curationA

Render a human-readable explanation of the Context Curator's decisions for the most recent (or a specified) optimization. Shows every candidate that was considered, whether it was selected or rejected, why, and how many tokens it used against the budget. Use this when an output felt off and you want to understand which grounding sources the engine chose.

ParametersJSON Schema
NameRequiredDescriptionDefault
optimization_idNoOptional trace ID. If omitted, explains the most recent trace.
lookback_daysNo

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided. The description does not explicitly state that the tool is read-only or has no side effects, but it describes the output in detail (candidates, selections, reasons, token usage). For a read-only explanation tool, this is adequate but leaves some ambiguity about mutability.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences: the first states the main function, the second adds detail and a use case. It is front-loaded with no extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description adequately explains the return content (candidates, selection status, reasons, token usage). Parameters are mostly covered, and the use case is clear. Minor omission: doesn't specify if the tool queries stored data or triggers a new analysis.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: 'optimization_id' has a clear description; 'lookback_days' is only described via constraints. The tool description explains the main purpose but does not elaborate on 'lookback_days' beyond the schema, so it partially compensates for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool renders a human-readable explanation of curation decisions, specifying the resource (Context Curator's decisions for an optimization) and action (render explanation). It distinguishes from siblings like 'get_trace' or 'inspect_context' by focusing on curation decisions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear scenario: 'Use this when an output felt off and you want to understand which grounding sources the engine chose.' It implies a debugging use case but does not explicitly list when not to use it or compare to alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_traceB

Fetch the full trace for an optimization ID, including system prompt + output. Looks back 7 days by default.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
lookback_daysNo

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden for behavioral disclosure. It only reveals the default lookback period and that the trace includes system prompt and output. It does not mention whether the operation is read-only, error handling, rate limits, or permission requirements, leaving significant gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with two sentences, front-loading the primary action and content, then adding the temporal scope. No unnecessary words or repetitions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (2 parameters, fetch operation), the description provides enough context: what it returns (full trace, system prompt, output) and default behavior. However, it lacks details on return format or error handling, which are not critical but would improve completeness. Since there is no output schema, the description could be slightly more explicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds meaning by stating the 'id' is an optimization ID and the 'lookback_days' defaults to 7. However, it does not explain the exact format of the id or provide further context for the parameters beyond what the schema already conveys (e.g., exclusiveMinimum, maximum).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Fetch', the resource 'full trace for an optimization ID', and includes what it contains ('system prompt + output'). It also provides the default lookback period, which helps differentiate from sibling tools like list_traces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives, such as list_traces or inspect_context. It does not mention prerequisites, exclusions, or scenarios where another tool would be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ground_promptA

Optimize a prompt against EXPLICIT caller-provided grounding sources (a spec, a transcript excerpt, an RFC, an internal doc, etc.). Each source is pinned at the highest priority — above project rules, above pinned instructions — and tracked individually in the trace. Use this when you want the rewrite to cite specific material rather than letting the curator decide what's relevant. Requires at least one non-empty source; will error rather than silently fall through to optimize_prompt. Sources are capped at 4000 chars each so a single large paste can't dominate the budget.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesThe prompt to optimize.
sourcesYesCaller-provided grounding sources. Must be non-empty.
categoryNo
platformNo
modeNo
cwdNo
file_pathNo
file_languageNo
file_excerptNo
session_idNo
user_localeNo
user_pinned_instructionsNo
enrich_contextNo
skip_intent_resolutionNo
include_bundleNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses key behaviors: sources are pinned at highest priority, tracked individually in trace, capped at 4000 chars, requires at least one non-empty source, and error behavior. While no annotations exist, it covers most relevant aspects for decision-making.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence is purposeful: states purpose, priority, usage guidance, constraints, and error behavior. No redundant or filler content. Well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with 15 params and no output schema or annotations, the description covers core functionality, usage scenario, and important constraints. Some optional params are left to schema descriptions, but the essential context for selection and invocation is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds some context beyond the schema (e.g., priority over project rules, error behavior), but schema coverage is only 13%, and many optional parameters remain unexplained. The description compensates partially for the required params but not fully for all 15.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool optimizes a prompt against explicit grounding sources, distinguishes it from optimize_prompt by mentioning error behavior and priority, and provides specific examples of sources (spec, transcript, RFC, etc.).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells when to use this tool ('when you want the rewrite to cite specific material') and when not to ('will error rather than silently fall through to optimize_prompt'), providing clear alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_contextA

Preview the ContextBundle (workspace rules, frameworks, target-model capabilities, resolved analysis, session history) without running optimization. Returns the same bundle that optimize_prompt would assemble.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYes
categoryNo
cwdNo
file_pathNo
file_languageNo
file_excerptNo
session_idNo
skip_intent_resolutionNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Given that no annotations are provided, the description carries the full burden. It discloses that the tool is non-destructive ('without running optimization') and what it returns, which is sufficient for a preview operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long with no superfluous information. It front-loads the core action and provides clear, efficient context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 8 parameters, no output schema, and no annotations, the description is too brief. It lacks details on parameter usage, output format, and behavioral edge cases, making it insufficient for complex calls.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not explain any of the 8 parameters despite 0% schema description coverage. It adds no value over the schema, leaving the agent to infer meaning from parameter names alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Preview' and identifies the resource 'ContextBundle' with details on its contents. It explicitly distinguishes itself from the sibling tool 'optimize_prompt' by noting that it runs without optimization.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states it returns the same bundle as optimize_prompt would assemble, implying it is for previewing. However, it does not explicitly state when to use this over alternatives or provide exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_categoriesA

List all available prompt optimization categories with platform counts including custom platforms

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries the burden of behavioral disclosure. It accurately implies a safe read operation but omits details such as ordering or error handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no extraneous words, efficiently conveying the tool's purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity (no parameters, no output schema), the description is largely complete. However, it could mention the read-only nature for completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With zero parameters, the baseline is 4. The description adds no parameter info because none exist, but it implies output structure without an output schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists prompt optimization categories with platform counts, using specific language that distinguishes it from sibling tools like list_platforms or list_modes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives; no context or exclusion criteria are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_modesA

List available output modes for prompt optimization

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It clearly indicates a read-only listing operation without side effects, which is adequate for a simple list tool. However, it could mention if any authentication or context requirements exist, but none are needed here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, no wasted words, front-loaded with the action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no parameters and no output schema, the description fully captures what it does. It is complete for a simple list operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters, and schema coverage is 100%. Baseline for 0 parameters is 4. The description does not need to add parameter info.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists available output modes for prompt optimization, using specific verb and resource. It distinguishes from siblings like list_categories and list_platforms by specifying 'for prompt optimization'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives like list_categories or list_packs. Usage is implied—when you need output modes for prompt optimization—but no exclusions or comparisons are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_packsB

List knowledge packs currently loaded in the persistent memory store.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeNoFilter by scope (e.g. 'user', 'project:myapp'). Omit to list all.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden for behavioral disclosure. It only states the read operation ('list') but lacks details on permissions, rate limits, side effects, or result format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence of 10 words, front-loaded with the action and resource, containing zero superfluous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given one optional parameter and no output schema, the description minimally covers the tool's purpose. However, it lacks usage guidance and behavioral details, leaving gaps for effective invocation among 22 siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a clear description of the 'scope' parameter including examples. The tool description adds the context 'currently loaded' but does not further enrich parameter semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'List knowledge packs currently loaded in the persistent memory store,' with a specific verb and resource. It is distinct from siblings like load_knowledge_pack or unload_pack, but does not explicitly differentiate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives such as memory_list_facts or list_categories. The description provides no context for selecting this tool over others.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_platformsB

List available platforms for a category, including custom registered platforms.

ParametersJSON Schema
NameRequiredDescriptionDefault
categoryYesCategory to list platforms for

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds that custom registered platforms are included, which is a behavioral detail beyond the bare function. But without annotations, more details like auth requirements or rate limits are missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence, efficient, front-loaded with the purpose. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with one parameter and no output schema, the description is fairly complete. It covers what and scope. Missing return format but acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. Description adds no extra meaning beyond what the schema already provides for the category parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states it lists platforms for a category, including custom ones. Verb 'list' and resource 'platforms' are specific. However, it does not explicitly differentiate from sibling list_categories, but context makes it clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives like register_platform or list_categories. Lacks usage context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_tracesA

List recent optimization traces from the local tracer. Summary only; use get_trace for full records.

ParametersJSON Schema
NameRequiredDescriptionDefault
dayNoUTC day YYYY-MM-DD; defaults to the most recent day with data
limitNo

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description alone must convey behavioral traits. It notes 'recent' (though undefined) and 'from the local tracer,' but does not disclose ordering, pagination behavior, or whether the operation is read-only. The suggestion to use get_trace for full records adds some context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the primary purpose and then usage guidance. Every sentence adds value without redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with two optional parameters and no output schema, the description covers the main purpose and links to a more detailed sibling. It could be improved by clarifying what 'recent' means, but overall it's adequately complete given the tool's low complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not add any meaning beyond the input schema. Schema coverage is 50% (only 'day' has a description). The parameter 'limit' lacks a description in both schema and tool description, leaving its purpose and constraints unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states 'List recent optimization traces' (specific verb+resource) and distinguishes from sibling 'get_trace' by noting 'Summary only; use get_trace for full records.' This explicitly differentiates the tool's purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Description provides clear guidance to use 'get_trace' for full records, indicating when to use this tool vs. an alternative. However, it does not specify when not to use this tool (e.g., if more than recent data is needed) or other contextual triggers.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

load_knowledge_packA

Load a knowledge pack — a markdown document with optional YAML frontmatter — into the persistent memory store. The pack is chunked by heading, each chunk embedded, and made available for semantic retrieval during subsequent optimize_prompt calls. Packs can come from a local file path, an HTTPS URL, or be passed inline as raw markdown. Community pack registry: https://github.com/LumabyteCo/clarifyprompt-packs

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYesLocal file path, HTTPS URL, or inline markdown body (auto-detected).
source_typeNoOverride source-type detection. `registry` marks a pack as community-sourced.auto
scopeNoScope to load under (e.g. 'user', 'project:myapp'). Defaults to pack frontmatter or 'user'.
nameNoOverride the pack name (else pulled from frontmatter).
versionNoOverride the pack version (else pulled from frontmatter or '0.0.0').

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description partially covers behavior: chunking, embedding, and retrieval usage during optimize_prompt. However, it omits details on overwriting existing packs, error handling, performance implications, or side effects like data persistence.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loading the main action and then explaining chunking and source options. No wasted words, though slightly more structured formatting could improve readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers core functionality (loading, chunking, sources), but lacks details on overwrite behavior, size limits, or unload mechanism. Given no output schema and basic complexity, it is adequate but not thorough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds marginal value beyond schema: it explains auto-detection of source types and mentions the community pack registry, but mostly repeats schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool loads a knowledge pack (markdown with YAML frontmatter) into persistent memory for semantic retrieval, specifying chunking by heading and embedding. It distinguishes from siblings like list_packs (listing) and unload_pack (unloading) by focusing on loading for retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when needing to load documents for semantic retrieval, lists source types and a community registry, but does not explicitly exclude alternatives (e.g., memory_remember for facts) or provide when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_forgetA

Invalidate (soft-delete) a fact by its id. The fact is marked invalidated_at = now and won't appear in future memory_search or grounding, but its history is preserved (bi-temporal soft-delete). Use memory_list_facts first to find the id you want to forget.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesFact id (from memory_remember response, memory_search result, or memory_list_facts row).

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully discloses behavior: it's a soft-delete, marks invalidated_at, removes from future searches, and preserves history. This is comprehensive for a single-parameter tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the action, no wasted words, and essential information is efficiently presented.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one required parameter and no output schema, the description is complete. It explains behavior, prerequisite, and effect on future operations. No gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and already explains the id's sources. The main description does not add new parameter semantics beyond what the schema provides, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Invalidate' and resource 'fact by its id'. It clearly distinguishes from siblings like memory_remember and memory_search by describing the bi-temporal soft-delete behavior, and it sets the context for how to obtain the id.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly instructs to use memory_list_facts first to find the id, providing a clear prerequisite. However, it does not mention when not to use this tool or any alternatives, though for a simple delete this is adequate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_list_factsA

List live (non-invalidated) facts in persistent memory, optionally filtered by scope and predicate. Sorted by most-recently-observed first. Useful for inspecting what the engine knows, or finding fact ids to forget.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeNoMemory scope to filter by. Default 'user'. Examples: 'user', 'project:myapp', 'session:abc'.user
predicateNoOptional predicate filter (e.g., only 'prefers' facts).
limitNoMax facts to return. Default 50, hard max 100.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses core behaviors: lists live/non-invalidated facts, sorted by recency, optional filters. With no annotations, description carries full burden; missing details like pagination behavior, empty result handling, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two efficient sentences: first states operation and sorting, second provides use cases. No redundant information, well front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, filtering, sorting, and use cases. Lacks description of return format (e.g., fields like fact_id, predicate, value) but no output schema exists; would benefit from a brief hint about output structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline 3. Description mentions filtering by scope and predicate but adds no new meaning beyond schema descriptions which already detail defaults and examples. Limit parameter is not explicitly mentioned in description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states verb 'List' and resource 'live facts in persistent memory', with optional filtering by scope and predicate. Distinguishes from sibling tools like memory_search and memory_forget by specifying 'non-invalidated' facts and sorting order.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit use cases: 'inspecting what the engine knows' and 'finding fact ids to forget'. Implicitly excludes mutation or search operations, but does not explicitly state when not to use or name alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_rememberA

Explicitly add a fact to persistent memory. Use when the user says something the engine should remember across sessions (preferences, conventions, project facts). Complements save_outcome reflection, which extracts facts implicitly — this is the explicit, user-driven path. Returns the new fact id, which can be passed to memory_forget later.

ParametersJSON Schema
NameRequiredDescriptionDefault
subjectYesWho/what the fact is about. Examples: 'user', 'project', 'this codebase', a person's name.
predicateYesShort verb phrase. Examples: 'prefers', 'uses', 'avoids', 'requires', 'is'.
objectYesThe concrete value. Example: 'TypeScript with strict mode'.
scopeNoMemory scope. Default 'user' (cross-session, cross-project). Use 'project:<name>' for project-local memory, 'session:<id>' for ephemeral session-only memory.user
confidenceNo0-1 confidence. Default 1.0 for explicit user remember. Reflection-extracted facts use 0.6-0.8.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden. It reveals the tool returns a fact id and implies persistence, but does not disclose potential side effects (e.g., overwrite behavior), required permissions, error handling, or whether it can fail silently.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences front-load the purpose and usage context, with no redundant phrases. Every sentence serves a distinct function: purpose, usage guidance, sibling differentiation, and return value note.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the parameter set (5 params, 3 required) and absence of output schema, the description covers the basic lifecycle (add, return id, forget). However, it omits details like whether adding duplicate facts creates duplicates or updates, and doesn't discuss scope isolation or session behavior beyond the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already explains all parameters. The description adds minimal meaning beyond the schema (only linking to `save_outcome` and `memory_forget`). Baseline 3 is appropriate as the description does not materially enhance parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Explicitly add a fact to persistent memory' with a specific verb ('add') and resource ('fact'). It distinguishes from the sibling tool `save_outcome` by contrasting explicit vs implicit extraction, and notes the return of a new fact id for use with `memory_forget`.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use: 'when the user says something the engine should remember across sessions' and provides examples. It mentions the complementary role of `save_outcome` and tees up `memory_forget` for the returned id, but does not explicitly list non-usage cases or other alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

optimize_promptA

Optimize a prompt for a specific AI platform. Context-aware: auto-gathers workspace signals (CLAUDE.md / AGENTS.md / .cursorrules / package.json), resolves intent + category + recommended mode in a single analysis step, shapes the system prompt to the target model's capabilities, and grounds the rewrite in a priority-ordered Grounding Context. Supports 58+ platforms across 7 categories, plus custom registered platforms. Category, platform, and mode are all optional — the engine chooses sane defaults from the analysis.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesThe prompt to optimize
categoryNoPrompt category. Auto-detected via the analyzer when omitted. When provided, the analyzer can still override if it's confident the hint is wrong.
platformNoTarget platform ID (e.g. midjourney, dall-e, sora, suno, claude, cursor, or a custom platform ID). Uses category default when omitted.
modeNoOutput mode. When omitted, the engine uses the analyzer's intent-derived recommendation (e.g. production-code → technical, quick-draft → concise). When passed, user choice wins.
enrich_contextNoUse web search for context enrichment (Tavily/Brave/Serper/SerpAPI/Exa/SearXNG). Results merge into the single Grounding Context block.
session_idNoSession ID to stitch related optimizations so the engine can reuse accepted prior outputs as few-shot examples. Auto-generated when omitted.
file_pathNoActive file path — infers language and grounds the rewrite
file_languageNoExplicit language override for the active file
file_excerptNoShort excerpt (≤2 KB) of the active file to ground the rewrite
cwdNoWorking directory to scan for CLAUDE.md / AGENTS.md / .cursorrules / package.json. Defaults to server cwd.
user_localeNoUser locale hint (e.g. en-US, ar-EG)
user_pinned_instructionsNoPinned, always-applied user instructions (highest-priority grounding)
include_bundleNoInclude the full resolved ContextBundle in the response (same shape as inspect_context returns)
skip_intent_resolutionNoSkip the analyzer LLM call (faster; loses intent/category/mode recommendations)

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden. It discloses auto-gathering of workspace signals, intent resolution, mode recommendation, and grounding. It mentions support for many platforms and optional parameters with sensible defaults. It does not mention side effects, auth, or rate limits, but the coverage is good. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single paragraph of about 6 sentences, each adding meaningful information. It is front-loaded with the main purpose and logically flows through features. Slightly long but still efficient; no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 14 parameters and no output schema, the description covers the analysis pipeline, default behaviors, optional features, and even mentions response structure via include_bundle. It could explicitly state that the response is an optimized prompt string, but the inference is clear from context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by explaining how parameters interact (e.g., category auto-detected, mode chosen from intent, session_id for few-shot). This goes beyond the schema descriptions and helps the agent understand the tool's behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Optimize a prompt for a specific AI platform.' It uses a specific verb (optimize) and resource (prompt) and distinguishes itself from siblings like compose_prompt or critique_prompt by emphasizing context-awareness, auto-analysis, and multi-platform support.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains that category, platform, and mode are optional with intelligent defaults, and that the engine auto-gathers workspace signals. However, it does not explicitly state when to use this tool vs alternatives like ground_prompt or compose_prompt, nor does it provide exclusions. The context is clear but lacks explicit when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

register_platformC

Register a new custom AI platform for prompt optimization.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesUnique platform ID (lowercase, alphanumeric with hyphens)
categoryYesCategory this platform belongs to
labelYesHuman-readable platform name
descriptionYesShort description
syntax_hintsNo
instructionsNo
instructions_fileNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It only says 'register a new platform' without mentioning side effects (e.g., overwriting an existing ID), authentication needs, or any implications. This is insufficient for safe invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, which is concise. It covers the core purpose without extra words. However, given the tool's complexity, slightly more structure (e.g., listing key prerequisites) could improve clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description lacks important context such as return values (no output schema), error conditions, and post-registration effects. For a tool with 7 parameters and no annotations, this is insufficient to ensure correct usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not mention any parameters or their roles. With a schema coverage of 57%, the description adds no value beyond what the schema provides. The three undocumented parameters (syntax_hints, instructions, instructions_file) are left entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (register) and resource (new custom AI platform for prompt optimization). This distinguishes it from sibling tools like update_platform and unregister_platform, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as update_platform or unregister_platform. There is no context about prerequisites or suitable scenarios, leaving the agent uncertain about selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_outcomeA

Tell ClarifyPrompt whether an optimization's output was accepted, edited, or rejected. Feeds two loops: (1) the session ring buffer so accepted prior outputs are injected as few-shot examples into future similar prompts, and (2) the persistent memory layer via reflection — on accept/edit, ClarifyPrompt extracts atomic facts from the interaction and stores them; on reject, recent reflection facts from this session are invalidated. Reflection uses the same LLM you've configured; expect a 1–3s latency on local models.

ParametersJSON Schema
NameRequiredDescriptionDefault
optimization_idYesThe `id` returned from optimize_prompt
session_idYesThe `sessionId` returned from optimize_prompt. Required so the outcome lands in the right session bucket.
verdictYesaccepted = user used the output as-is; edited = user kept it with edits; rejected = user threw it away
diffNoOptional: the user's edited version or a diff. Helps reflection extract better facts.
skip_reflectionNoSkip the LLM-based fact extraction pass (faster, no facts learned)

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, but the description fully discloses the behavior: feeding the session ring buffer, triggering reflection for fact extraction/invalidation, and the latency impact on local models. It covers all significant side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise yet comprehensive. It front-loads the core purpose, then efficiently explains the two feedback loops and the reflection latency. Every sentence serves a purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers all essential aspects: the tool's function, its integration into two loops, behavior on each verdict, and a performance caveat. No output schema exists, but the side effects are fully described.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already has 100% description coverage, but the tool description adds operational context (e.g., how 'diff' helps reflection, the effect of 'skip_reflection') that provides additional meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it records the verdict of an optimization output (accepted/edited/rejected) and explains its role in two feedback loops, distinguishing it from sibling tools that handle other aspects of the optimization process.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage after obtaining an optimization output, but does not explicitly state when not to use it or list alternatives. It provides clear context for when it is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unload_packA

Remove a loaded knowledge pack (and all its chunks + embeddings) from the memory store.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesPack id (as returned by list_packs).

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description adds moderate behavioral context by noting that unloading removes both chunks and embeddings. However, it does not disclose other important traits like destructiveness, reversibility, or authorization requirements, which would be expected for a tool with no annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-formed sentence that is concise and to the point. Every word adds value, with no extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of the tool (1 parameter, no output schema), the description is mostly complete. It explains the primary action and scope (pack, chunks, embeddings). Minor gaps remain, such as failure scenarios or state changes, but for a straightforward removal operation, it is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the parameter description in the schema is clear ('Pack id (as returned by list_packs)'). The tool description adds no additional semantic information beyond the schema, resulting in a baseline score of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (remove) and the resource (a loaded knowledge pack, including its chunks and embeddings). It effectively distinguishes this tool from siblings like 'load_knowledge_pack' and 'list_packs' by specifying its unique purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no explicit guidance on when to use this tool versus alternatives (e.g., memory_forget for individual facts). It does not specify prerequisites or contraindications, leaving the agent to infer usage context from the tool name and schema alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unregister_platformB

Remove a custom platform, or clear instruction overrides on a built-in.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
categoryYes
remove_override_onlyNo

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description should fully disclose behavior. It mentions removal and clearing overrides but omits side effects, permission requirements, or consequences.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence, but it could add more detail without becoming overly verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and 0% schema coverage, the description is insufficient. It lacks information on return values, error conditions, and post-removal effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description provides no explanation of the three parameters (id, category, remove_override_only) or their roles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool removes a custom platform or clears instruction overrides on a built-in platform, distinguishing its purpose from siblings like register_platform and update_platform.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use (removing or clearing overrides) but lacks explicit guidance on when not to use or alternatives among the listed siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_platformC

Update a custom platform or add/override instructions on a built-in platform.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
categoryYes
labelNo
descriptionNo
syntax_hintsNo
syntax_hints_appendNo
instructionsNo
instructions_fileNo

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description only says 'update' and 'add/override', indicating mutation but lacking details on side effects, authorization, or what happens to unspecified fields. Behavioral transparency is minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence covering both use cases without unnecessary words. It is well-structured for its length, though breaking it into two sentences could improve clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 8 parameters and no output schema, the description is insufficient. It omits key details like required fields, partial vs full update behavior, and the meaning of complex parameters (e.g., instructions_file vs instructions, syntax_hints vs syntax_hints_append).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain any of the 8 parameters (e.g., id, instructions, syntax_hints). The agent has no insight into how to correctly populate the fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (update/add/override) and resource (platform), and distinguishes between custom and built-in platforms. However, it does not explicitly differentiate from sibling tools like register_platform, leaving some ambiguity about when to use each.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies two modes (update custom, add/override built-in) but provides no explicit guidance on when to use this tool versus alternatives like register_platform or unregister_platform.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 20 tool updatesv1.6.8
    • Addedclarify_with_user
    • Addedcompose_prompt
    • Addedcritique_prompt
    • Addedexplain_last_curation
    • Addedget_trace
    • Addedground_prompt
    • Addedinspect_context
    • Addedlist_packs
    • Addedlist_traces
    • Addedload_knowledge_pack
    • Addedmemory_forget
    • Addedmemory_list_facts
    • Addedmemory_remember
    • Addedmemory_search
    • Changedoptimize_prompt13 fields changed
      • changedInput schema / properties / category / description
        Previous value: -"Prompt category. Auto-detected from prompt content when omitted."New value: +"Prompt category. Auto-detected via the analyzer when omitted. When provided, the analyzer can still override if it's confident the hint is wrong."
      • addedInput schema / properties / cwd
        Added value: +{
        +  "description": "Working directory to scan for CLAUDE.md / AGENTS.md / .cursorrules / package.json. Defaults to server cwd.",
        +  "type": "string"
        +}
      • changedInput schema / properties / enrich_context / description
        Previous value: -"Use web search for context enrichment (supports Tavily, Brave, Serper, SerpAPI, Exa, SearXNG)"New value: +"Use web search for context enrichment (Tavily/Brave/Serper/SerpAPI/Exa/SearXNG). Results merge into the single Grounding Context block."
      • addedInput schema / properties / file_excerpt
        Added value: +{
        +  "description": "Short excerpt (≤2 KB) of the active file to ground the rewrite",
        +  "type": "string"
        +}
      • addedInput schema / properties / file_language
        Added value: +{
        +  "description": "Explicit language override for the active file",
        +  "type": "string"
        +}
      • addedInput schema / properties / file_path
        Added value: +{
        +  "description": "Active file path — infers language and grounds the rewrite",
        +  "type": "string"
        +}
      • addedInput schema / properties / include_bundle
        Added value: +{
        +  "default": false,
        +  "description": "Include the full resolved ContextBundle in the response (same shape as inspect_context returns)",
        +  "type": "boolean"
        +}
      • removedInput schema / properties / mode / default
        Removed value: -"detailed"
      • changedInput schema / properties / mode / description
        Previous value: -"Output mode"New value: +"Output mode. When omitted, the engine uses the analyzer's intent-derived recommendation (e.g. production-code → technical, quick-draft → concise). When passed, user choice wins."
      • addedInput schema / properties / session_id
        Added value: +{
        +  "description": "Session ID to stitch related optimizations so the engine can reuse accepted prior outputs as few-shot examples. Auto-generated when omitted.",
        +  "type": "string"
        +}
      • addedInput schema / properties / skip_intent_resolution
        Added value: +{
        +  "default": false,
        +  "description": "Skip the analyzer LLM call (faster; loses intent/category/mode recommendations)",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / user_locale
        Added value: +{
        +  "description": "User locale hint (e.g. en-US, ar-EG)",
        +  "type": "string"
        +}
      • addedInput schema / properties / user_pinned_instructions
        Added value: +{
        +  "description": "Pinned, always-applied user instructions (highest-priority grounding)",
        +  "type": "string"
        +}
    • Changedregister_platform6 fields changed
      • changedInput schema / properties / description / description
        Previous value: -"Short description of the platform"New value: +"Short description"
      • changedInput schema / properties / id / description
        Previous value: -"Unique platform ID (lowercase, alphanumeric with hyphens, e.g. 'my-llm')"New value: +"Unique platform ID (lowercase, alphanumeric with hyphens)"
      • removedInput schema / properties / instructions / description
        Removed value: -"Inline instructions for prompt optimization on this platform"
      • removedInput schema / properties / instructions_file / description
        Removed value: -"Path to a .md file with detailed instructions (relative to config dir's instructions/ folder, or absolute)"
      • changedInput schema / properties / label / description
        Previous value: -"Human-readable platform name (e.g. 'My Custom LLM')"New value: +"Human-readable platform name"
      • removedInput schema / properties / syntax_hints / description
        Removed value: -"Platform-specific syntax hints (e.g. ['system prompts', 'JSON mode'])"
    • Addedsave_outcome
    • Addedunload_pack
    • Changedunregister_platform3 fields changed
      • removedInput schema / properties / category / description
        Removed value: -"Category the platform belongs to"
      • removedInput schema / properties / id / description
        Removed value: -"Platform ID to remove"
      • removedInput schema / properties / remove_override_only / description
        Removed value: -"If true, only remove instruction overrides (for built-in platforms)"
    • Changedupdate_platform8 fields changed
      • removedInput schema / properties / category / description
        Removed value: -"Category the platform belongs to"
      • removedInput schema / properties / description / description
        Removed value: -"Updated description (custom platforms only)"
      • removedInput schema / properties / id / description
        Removed value: -"Platform ID to update"
      • removedInput schema / properties / instructions / description
        Removed value: -"Inline instructions (replaces existing)"
      • removedInput schema / properties / instructions_file / description
        Removed value: -"Path to .md instructions file (replaces existing)"
      • removedInput schema / properties / label / description
        Removed value: -"Updated display name (custom platforms only)"
      • removedInput schema / properties / syntax_hints / description
        Removed value: -"Replace syntax hints (custom platforms only)"
      • removedInput schema / properties / syntax_hints_append / description
        Removed value: -"Additional syntax hints to append (works for both built-in and custom)"
  2. 7 tool updates
    • Addedlist_categories
    • Addedlist_modes
    • Addedlist_platforms
    • Addedoptimize_prompt
    • Addedregister_platform
    • Addedunregister_platform
    • Addedupdate_platform

TDQS

A3.7/5.0

Scored across 23 tools

Disambiguation4/5

Most tools have clearly distinct purposes, but compose_prompt overlaps slightly with clarify_with_user, optimize_prompt, ground_prompt, and critique_prompt, as it bundles their functionality. The memory and platform tools are well-separated.

Naming Consistency5/5

All tool names follow a consistent verb_noun snake_case pattern, with clear verbs like 'list', 'create', 'optimize', 'memory_', etc. There are no mixed conventions or vague names.

Tool Count5/5

23 tools is appropriate for the comprehensive feature set of prompt optimization, memory management, platform registration, and inspection. Each tool serves a distinct function without unnecessary bloat.

Completeness4/5

The tool surface covers the full lifecycle of prompt optimization: clarification, grounding, optimization, critique, memory, and platform management. Minor gaps exist (e.g., no bulk optimization), but the core workflows are well-covered.

Maintenance

ActivityInactive
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    An MCP server that uses Claude 3.5 Sonnet to transform ordinary prompts into structured, professionally engineered instructions for any LLM. It enhances AI interactions by adding context, requirements, and structural clarity to raw user inputs.
    1
    3
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    An MCP server for deterministic prompt optimization in Claude Code. Score prompts across 7 quality dimensions, auto-select from 11 Anthropic techniques, and return a structural scaffold.
    1
    21 npm
    2
    MIT