Converse MCP Server
Converse MCP Server lets Claude talk to other AI models — single or multi-model chats, consensus/roundtable discussions, background jobs, and calibrated typed decisions.
Chat with one or more models —
chatmode runs 1..N models in parallel, each answering independently (OpenAI, Google, Anthropic, X.AI, Mistral, DeepSeek, OpenRouter, Abliteration, Codex, Claude Agent SDK, Antigravity/Gemini CLI, GitHub Copilot).Get consensus —
consensusmode has ≥2 models answer, then refine after seeing each other's responses.Run a roundtable —
roundtablemode has models speak sequentially in a given order, each seeing the running transcript.Attach context — pass
files(with line ranges likefile.txt{10:50}) andimages(paths or base64) so code/files are shared directly instead of pasted.Hold multi-turn conversations — reuse a
continuation_idto continue a thread; you can switch modes or models on resume.Run long tasks in the background — set
async: trueto get an ID immediately, then poll withcheck_status.Monitor and manage jobs —
check_statusreports progress, AI-generated titles/summaries, and full history;cancel_jobstops a running or queued job.Get calibrated decisions instead of text —
decideasks a System One model (TypeSafe Jev) batched typed questions about a state: yes/no probability (noul), one-of-many choice (choice), or rubric position (score).Control reasoning depth —
reasoning_effortfromnonetomax, clamped to what each model supports.Export conversations —
export: truewrites numbered request/response files and metadata to disk.Pick models flexibly —
auto, a bare provider name,provider:model(e.g.openai:gpt-6-astra,copilot:sonnet), or a bare model alias routed to the first configured provider.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Converse MCP Serveruse consensus to evaluate startup ideas"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Converse MCP Server
An MCP (Model Context Protocol) server that lets Claude talk to other AI models. Use it to chat with models from OpenAI, Google, Anthropic, X.AI, Mistral, DeepSeek, OpenRouter, or Abliteration. You can either talk to one model at a time or get multiple models to weigh in on complex decisions.
📋 Requirements
Node.js: Version 20 or higher
Package Manager: npm (or pnpm/yarn)
API Keys: At least one from any supported provider
Related MCP server: Zen MCP Server
🚀 Quick Start
Step 1: Get Your API Keys
You need at least one API key from these providers:
Provider | Where to Get | Example Format |
OpenAI |
| |
Google/Gemini |
| |
X.AI |
| |
Anthropic |
| |
Mistral |
| |
DeepSeek |
| |
OpenRouter |
| |
Abliteration |
| |
Codex | ChatGPT login (system-wide) | Local agentic assistant |
Note: Codex uses your ChatGPT login (not an API key). If you have an active ChatGPT session, Codex will work automatically. For headless/server deployments, set CODEX_API_KEY in your environment.
Step 2: Add to Claude Code or Claude Desktop
For Claude Code (Recommended)
# Add the server with your API keys
claude mcp add converse \
-e OPENAI_API_KEY=your_key_here \
-e GEMINI_API_KEY=your_key_here \
-e XAI_API_KEY=your_key_here \
-e ANTHROPIC_API_KEY=your_key_here \
-e MISTRAL_API_KEY=your_key_here \
-e DEEPSEEK_API_KEY=your_key_here \
-e OPENROUTER_API_KEY=your_key_here \
-e ABLITERATION_API_KEY=ak_your_key_here \
-e ENABLE_RESPONSE_SUMMARIZATION=true \
-e SUMMARIZATION_MODEL=gpt-5-nano \
-s user \
npx converse-mcp-serverFor Claude Desktop
Add this configuration to your Claude Desktop settings:
{
"mcpServers": {
"converse": {
"command": "npx",
"args": ["converse-mcp-server"],
"env": {
"OPENAI_API_KEY": "your_key_here",
"GEMINI_API_KEY": "your_key_here",
"XAI_API_KEY": "your_key_here",
"ANTHROPIC_API_KEY": "your_key_here",
"MISTRAL_API_KEY": "your_key_here",
"DEEPSEEK_API_KEY": "your_key_here",
"OPENROUTER_API_KEY": "your_key_here",
"ABLITERATION_API_KEY": "ak_your_key_here",
"ENABLE_RESPONSE_SUMMARIZATION": "true",
"SUMMARIZATION_MODEL": "gpt-5-nano"
}
}
}
}Windows Troubleshooting: If npx converse-mcp-server doesn't work on Windows, try:
{
"command": "cmd",
"args": ["/c", "npx", "converse-mcp-server"],
"env": {
"ENABLE_RESPONSE_SUMMARIZATION": "true",
"SUMMARIZATION_MODEL": "gpt-5-nano"
// ... add your API keys here
}
}Step 3: Start Using Converse
Once installed, you can:
Chat with a specific model: Ask Claude to use the chat tool with your preferred model
Get consensus: Ask Claude to use the chat tool with
mode: "consensus"when you need multiple perspectivesRun tasks in background: Use
async: truefor long-running operations that you can check laterMonitor progress: Use the check_status tool to monitor async operations with AI-generated summaries
Cancel jobs: Use the cancel_job tool to stop running operations
Smart summaries: Get auto-generated titles and summaries for better context understanding
Get help: Type
/converse:helpin Claude
🛠️ Available Tools
1. Chat Tool
One tool, three modes. Pass a models array and choose a mode. Supports files, images, conversation history, and background execution. The tool routes each model to the right provider by name; "auto" picks the first available provider. When AI summarization is enabled, it generates smart titles and summaries.
// mode "chat" (default) — 1..N models answer independently, in parallel
{
"prompt": "How should I structure the authentication module for this Express.js API?",
"models": ["gemini-2.5-flash"], // Routes to Google
"files": ["/path/to/src/auth.js", "/path/to/config.json"],
"images": ["/path/to/architecture.png"],
"reasoning_effort": "medium"
}
// mode "consensus" — ≥2 models answer, then refine after seeing each other
{
"prompt": "Should we use microservices or a monolith for our e-commerce platform?",
"models": ["gpt-5.6", "gemini-2.5-flash", "grok-4.5"],
"mode": "consensus",
"files": ["/path/to/requirements.md"]
}
// mode "roundtable" — models speak SEQUENTIALLY in the given order, each seeing
// the running transcript. One call = one lap; pass continuation_id for more laps.
{
"prompt": "Critique this caching strategy and propose improvements.",
"models": ["codex", "gemini", "claude"], // ORDER MATTERS
"mode": "roundtable"
}
// Asynchronous execution (for long-running tasks) — any mode
{
"prompt": "Analyze this large codebase and provide optimization recommendations",
"models": ["gpt-5.6"],
"files": ["/path/to/large-project"],
"async": true, // Enables background processing
"continuation_id": "my-analysis-task" // Optional: custom ID for tracking
}Codex Notes:
Uses thread-based sessions in
chatmode (context persists withcontinuation_id)Responses typically take 6-20 seconds (complex tasks may take minutes)
Accesses files directly from your working directory
Configure sandbox mode via
CODEX_SANDBOX_MODEenvironment variable
2. Check Status Tool
Monitor the progress and retrieve results from asynchronous operations. When AI summarization is enabled, provides intelligent summaries of ongoing and completed tasks.
// Check status of a specific job
{
"continuation_id": "my-analysis-task"
}
// List recent jobs (shows last 10)
// With summarization enabled, displays titles and final summaries
{}
// Get full conversation history for completed job
{
"continuation_id": "my-analysis-task",
"full_history": true
}3. Cancel Job Tool
Cancel running asynchronous operations when needed.
// Cancel a running job
{
"continuation_id": "my-analysis-task"
}4. Decide Tool
Ask a System One decision model (TypeSafe's Jev) typed questions about a state and get calibrated answers instead of text: a yes probability (noul), a chosen option with per-option probabilities (choice), or a rubric position with per-level probabilities (score). Batch many questions into one call; they are judged in parallel against the same state. Needs TYPESAFE_API_KEY or OPENROUTER_API_KEY (TypeSafe first, falling back to OpenRouter).
{
"state": { "message": "I was charged twice for order A-104. Please fix this ASAP." },
"questions": {
"urgent": { "type": "noul", "instructions": "Does the message convey urgency?" },
"team": { "type": "choice", "instructions": "Which team should handle this?",
"criteria": { "billing": "Payments, refunds", "technical": "Bugs, outages" } },
"frustration": { "type": "score", "instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Frustrated", "Very angry"] }
}
}See docs/API.md for the full schema, model routing, and usage guidance.
🤖 AI Summarization Feature
When enabled, the server automatically generates intelligent titles and summaries for better context understanding:
Automatic Title Generation: Creates descriptive titles (up to 60 chars) for each request
Streaming Summaries: Status check returns an up-to-date summary of the progress based on the partially streamed response
Final Summaries: Concise 1-2 sentence summaries of completed responses
Smart Status Display: Enhanced check_status tool shows titles and summaries in job listings
Persistent Context: Summaries are stored with async jobs for better progress tracking
Configuration:
# Enable in your environment
ENABLE_RESPONSE_SUMMARIZATION=true # Default: false
SUMMARIZATION_MODEL=gpt-5-nano # Default: gpt-5-nanoBenefits:
Quickly understand what each async job is doing without reading full responses
Better context when reviewing multiple ongoing operations
Improved job management with at-a-glance understanding of task progress
Graceful fallback to text snippets when summarization is disabled or fails
📊 Supported Models
OpenAI Models
gpt-6.1-sol (default; aliases:
gpt-6,gpt-6.1,gpt-5,sol): GPT-6.1 Sol (1M context, 128K output) - Near-Astra coding, computer use, and professional work at a fifth of the Astra price; effortlow–maxgpt-6-sol: Previous GPT-6 Sol (1M context, 128K output), reachable by versioned name only; effort
none–maxgpt-6-luna (alias:
luna): Most efficient GPT-6 (1M context, 128K output) - Focused, high-volume tasks; effortnone–maxgpt-6-astra (alias:
astra): Frontier GPT-6 flagship (1M context, 128K output) - Hardest end-to-end work; effortlow–max(EXPENSIVE: 5x Sol)gpt-5.6-sol (alias:
gpt-5.6): Previous flagship GPT-5.6 (1M context, 128K output)gpt-5.6-terra (alias:
terra): Lower-cost GPT-5.6 (400K context, 128K output) - Performance competitive with GPT-5.5 at half the pricegpt-5.6-luna: Fastest, most affordable GPT-5.6 (400K context, 128K output)
gpt-5.4: Flagship-class reasoning (1M context, 128K output)
gpt-5.4-pro (alias:
gpt-5-pro): Maximum-performance reasoning (1M context, 272K output) - Hardest problems, extended compute time (EXPENSIVE)gpt-5-mini, gpt-5-nano: Faster, cost-efficient GPT-5 tiers (400K context, 128K output)
gpt-5.4-mini, gpt-5.4-nano: Fast, efficient GPT-5.4 tiers (400K context, 128K output)
o3, o3-pro, o4-mini: Advanced reasoning models (200K context)
gpt-4.1: Large context (1M tokens, 32K output)
o3-deep-research (30-90 min runtime), o4-mini-deep-research (15-60 min runtime): Deep research models (200K context)
Google/Gemini Models
API Key Options:
GEMINI_API_KEY: For Gemini Developer API (recommended)
GOOGLE_API_KEY: Alternative name (GEMINI_API_KEY takes priority)
Vertex AI: Use
GOOGLE_GENAI_USE_VERTEXAI=truewith project/location settings
Supported Models:
gemini-3.1-pro-preview (aliases:
pro,gemini-pro): Most advanced reasoning with expanded thinking levels (1M context, 64K output)gemini-3.5-flash (aliases:
gemini-3.5,flash-3.5): Frontier-level agentic and coding performance at Flash speed (1M context, 65K output)gemini-3.8-flash (aliases:
gemini-3.8,flash-3.8): Current-generation Flash with stronger long-horizon agentic performance (1M context, 65K output; thinking levels low/medium/high — no minimal)gemini-2.5-pro (alias:
pro 2.5): Deep reasoning with thinking budget (1M context, 65K output)gemini-2.5-flash (alias:
flash): Ultra-fast (1M context, 65K output)gemini-2.5-flash-lite (alias:
flash-lite): Lightweight fast model (1M context, 65K output)
Note: The Antigravity CLI provider serves gemini-3.8-flash and gemini-3.1-pro-preview under the same names and aliases, so bare pro, gemini-pro, flash and those IDs go to Antigravity first when it is installed (see Model Selection). Use google:<model> to always use the Google API.
X.AI/Grok Models
grok-4.5 (default; aliases:
grok,grok-4.5-latest,grok-build-latest): Flagship model with image input, reasoning content, and native web/X search (500K context). Reasoning maps tolow/medium/highand cannot be disabled; web search is automatic.
Anthropic Models
claude-opus-5-5 (default; aliases:
opus,opus-5.5,claude-opus): Flagship Opus for complex agentic coding and deep reasoning; thinking always on, effortlow–max(1M context, 128K output)claude-fable-5 (alias:
fable): Most capable model for demanding reasoning and long-horizon agentic work (1M context, 128K output)claude-opus-5 (alias:
opus-5): Previous Opus generation (1M context, 128K output)claude-opus-4-8 / claude-opus-4-7 / claude-opus-4-6: Earlier Opus generations with adaptive thinking (200K context, 1M via beta, 128K output)
claude-opus-4-5 / claude-opus-4-1: Legacy Opus models with extended thinking (64K / 32K output)
claude-sonnet-5-5 (aliases:
sonnet,sonnet-5.5): Current Sonnet with adaptive thinking and effort (1M context, 128K output)claude-sonnet-4-6 (alias:
sonnet-4.6): Previous Sonnet generation with adaptive thinking (64K output)claude-haiku-4-5 (alias:
haiku): Fast and intelligent for simple queries (64K output)
Mistral Models
mistral-medium-3-5 (default; aliases:
mistral,mistral-medium): Frontier-class multimodal model with adjustable reasoning (256K context)mistral-small-2603 (alias:
mistral-small): Hybrid multimodal model unifying instruct, reasoning, and coding (256K context)mistral-large-2512 (alias:
mistral-large): Open-weight MoE flagship, image-capable, no adjustable reasoning (256K context)
Reasoning maps to high (enabled) or none (disabled) on Medium 3.5 and Small; Large has no adjustable reasoning.
DeepSeek Models
deepseek-v4-pro (default; aliases:
deepseek,deepseek-pro): Flagship MoE model with thinking mode (1M context, 384K max output, text-only)deepseek-v4-flash (alias:
deepseek-flash): Faster, lower-cost V4 tier with thinking mode (1M context, 384K max output, text-only)
Thinking mode maps reasoning_effort to none (off), high (enabled levels up to high), or max (xhigh and max).
OpenRouter Models
z-ai/glm-5.2 (default; aliases:
glm,glm-5.2): Large-scale reasoning model, text-only (1M context)deepseek/deepseek-v4-pro, deepseek/deepseek-v4-flash: DeepSeek V4 reasoning tiers, text-only (1M context)
qwen/qwen3.7-max, qwen/qwen3.7-plus: Flagship Qwen tiers (1M context;
plusis image-capable)moonshotai/kimi-k2.7-code, moonshotai/kimi-k2.6: Image-capable Moonshot models (256K context;
k2.7-codealways reasons)openrouter/auto: Auto-selects the best model for your prompt
Any other model works via its full provider/model slug or the openrouter: namespace — no extra configuration. Append :online to a slug to opt into web search (adds a real per-request cost).
Abliteration Models
Abliteration provides uncensored ("abliterated") reasoning models through its OpenAI-compatible Chat Completions API at https://api.abliteration.ai/v1; see its documentation. Use the abliteration or ablit namespace; set ABLITERATION_DEFAULT_MODEL to override the default. Web search is not supported.
abliterated-model-large-v2 (default; aliases:
abliterated-large,abliterated-large-v2): GLM-5.3-derived, text-only (1M context, 999,990 max output)abliterated-model-large (alias:
abliterated-large-v1): Previous large model, GLM-5.2-derived, text-only (1M context, 999,990 max output)abliterated-model (aliases:
abliterated,abliterated-base): Multimodal with image input (256K context, 262,134 max output)
All models reason by default and return reasoning traces (reasoning_content), streamed as thinking. reasoning_effort is clamped to each model's supported levels: large-v2 runs low/high/max and cannot disable reasoning (none/minimal/low → low, medium/high → high, xhigh/max → max); large runs high/max and can disable (none → none, minimal–high → high, xhigh/max → max); base accepts every level, with none disabling reasoning. Only abliterated-model accepts images; image requests to the text-only large models are rejected before sending.
Codex Models
OpenAI Codex agentic coding assistant. codex uses its default model (GPT-6 Astra, or CODEX_DEFAULT_MODEL); codex:<model> picks one (e.g. codex:luna, codex:astra, codex:gpt-5.6-terra):
gpt-6.1-sol (
sol,gpt-6), gpt-6-sol, gpt-6-luna (luna), gpt-6-astra (astra)gpt-5.6-sol (
gpt-5.6), gpt-5.6-terra (terra), gpt-5.6-luna, gpt-5.5, gpt-5.3-codex-spark (spark)reasoning_effortmaps onto the tiers the chosen backend accepts (versioned GPT-6 Sol, Luna, and GPT-5.6:nonethroughmax; GPT-6.1 Sol and GPT-6 Astra:lowthroughmax, nonone)Thread-based sessions with persistent context
Direct filesystem access from working directory
Typical response time: 6-20 seconds (longer for complex tasks)
Requires ChatGPT login or CODEX_API_KEY
See Configuration for sandbox and approval settings
Claude Agent SDK Models
Claude via the Claude Agent SDK. claude uses its default model (Opus 5.5, or CLAUDE_DEFAULT_MODEL); claude:<model> picks one:
claude-opus-5-5 (default; aliases:
opus,claude-opus,opus-5.5), claude-opus-5 (opus-5)claude-fable-5-1 (aliases:
fable,claude-fable,fable-5.1), claude-fable-5 (fable-5)claude-sonnet-5-5 (aliases:
sonnet,claude-sonnet,sonnet-5.5)Uses Claude Code CLI authentication (
claude login) - no API key neededDirect filesystem access from working directory
GitHub Copilot SDK Models
Reach these only with the copilot: namespace (e.g. copilot:gpt-6.1-sol) — Copilot never serves bare model names. copilot alone uses GPT-6.1 Sol, or COPILOT_DEFAULT_MODEL. Uses your GitHub Copilot subscription (gh auth login) - no API key needed:
OpenAI:
gpt-6.1-sol(aliases:gpt-6,gpt-6.1,gpt-5,sol),gpt-6-sol,gpt-6-luna(alias:luna),gpt-5.6-sol(alias:gpt-5.6),gpt-5.6-terra,gpt-5.6-luna(all supportreasoning_effort)Anthropic:
claude-opus-5.5(aliases:opus,claude),claude-fable-5(alias:fable),claude-sonnet-5.5(alias:sonnet),claude-sonnet-5,claude-opus-5,claude-opus-4.8Google:
gemini-3.1-pro-preview(aliases:gemini,gemini-3.1-pro),gemini-3.8-flash(aliases:gemini-3.8,flash-3.8),gemini-3.5-flash(alias:gemini-flash)
📚 Help & Documentation
Built-in Help
Type these commands directly in Claude:
/converse:help- Full documentation/converse:help tools- Tool-specific help (includes async features)/converse:help models- Model information/converse:help parameters- Configuration details/converse:help examples- Usage examples (sync and async)/converse:help async- Async execution guide
Additional Resources
API Reference: docs/API.md
Architecture Guide: docs/ARCHITECTURE.md
Integration Examples: docs/EXAMPLES.md
⚙️ Configuration
Environment Variables
Create a .env file in your project root:
# Required: At least one API key
OPENAI_API_KEY=sk-proj-your_openai_key_here
GEMINI_API_KEY=your_gemini_api_key_here # Or GOOGLE_API_KEY (GEMINI_API_KEY takes priority)
XAI_API_KEY=xai-your_xai_key_here
ANTHROPIC_API_KEY=sk-ant-your_anthropic_key_here
MISTRAL_API_KEY=your_mistral_key_here
DEEPSEEK_API_KEY=your_deepseek_key_here
OPENROUTER_API_KEY=sk-or-your_openrouter_key_here
ABLITERATION_API_KEY=ak_your_key_here
# Optional: Server configuration
PORT=3157
LOG_LEVEL=info
# Optional: AI Summarization (Enhanced async status display)
ENABLE_RESPONSE_SUMMARIZATION=true # Enable AI-generated titles and summaries
SUMMARIZATION_MODEL=gpt-5-nano # Model to use for summarization (default: gpt-5-nano)
# Optional: OpenRouter attribution (for ranking credit; both optional)
OPENROUTER_REFERER=https://github.com/FallDownTheSystem/converse
OPENROUTER_TITLE=Converse
# Optional: Codex configuration
CODEX_API_KEY=your_codex_api_key_here # Optional if ChatGPT login available
CODEX_SANDBOX_MODE=read-only # read-only (default), workspace-write, danger-full-access
CODEX_SKIP_GIT_CHECK=true # true (default), false
CODEX_APPROVAL_POLICY=never # never (default), untrusted, on-failure, on-request
# Optional: per-provider default model, used for a bare provider name (`codex`,
# `openai`, ...) and for "auto". Must be a model or alias from that provider's
# list; startup fails with suggestions otherwise.
CODEX_DEFAULT_MODEL=gpt-6-astra # CODEX_MODEL still works as a fallback
CLAUDE_DEFAULT_MODEL=claude-opus-5-5
AGY_DEFAULT_MODEL=gemini-3.8-flash # Antigravity CLI (gemini)
COPILOT_DEFAULT_MODEL=gpt-6.1-sol # COPILOT_MODEL still works as a fallback
OPENAI_DEFAULT_MODEL=gpt-6.1-sol
GOOGLE_DEFAULT_MODEL=gemini-3.1-pro-preview
XAI_DEFAULT_MODEL=grok-4.5
ANTHROPIC_DEFAULT_MODEL=claude-opus-5-5
MISTRAL_DEFAULT_MODEL=mistral-medium-3-5
DEEPSEEK_DEFAULT_MODEL=deepseek-v4-pro
OPENROUTER_DEFAULT_MODEL=z-ai/glm-5.2 # any vendor/model slug is accepted
ABLITERATION_DEFAULT_MODEL=abliterated-model-large-v2Configuration Options
Server Environment Variables (.env file)
Variable | Description | Default | Example |
| Server port |
|
|
| Logging level |
|
|
Claude Code Environment Variables (System/Global)
These must be set in your system environment or when launching Claude Code, NOT in the project .env file:
Variable | Description | Default | Example |
| Token response limit |
|
|
| Wall-clock limit per tool call (ms) | ~28 hours when unset |
|
| Idle window (ms) — aborts a call that produces no output for this long | 30 min (stdio) / 5 min (HTTP) |
|
# Example: Set globally before starting Claude Code
export MAX_MCP_OUTPUT_TOKENS=200000
export CLAUDE_CODE_MCP_TOOL_IDLE_TIMEOUT=3600000 # 60 min for long silent agentic calls
claude # Then start Claude CodeOr persist them in ~/.claude/settings.json:
{
"env": {
"MAX_MCP_OUTPUT_TOKENS": "200000",
"CLAUDE_CODE_MCP_TOOL_IDLE_TIMEOUT": "3600000"
}
}The idle timeout is usually what kills long calls. Agentic models (Codex, Claude Agent SDK) can work silently for 30+ minutes; Converse holds one MCP request open the whole time, and Claude Code aborts it after the idle window with an error like "failed after 30 minutes of silence. The idle timeout aborted it." Progress-notification heartbeats can't prevent this — Claude Code doesn't send a progressToken on tools/call, so an MCP server has no spec-compliant way to emit them (claude-code#58687). Raising the idle window is the only fix.
Codex CLI Tool Timeout
If you register Converse in OpenAI's Codex CLI, note that Codex enforces its own hard 300-second default per MCP tool call (tool_timeout_sec, undocumented — it exists only in Codex's config schema). Progress notifications don't extend it. Raise it in ~/.codex/config.toml:
[mcp_servers.converse]
# ... command/env ...
tool_timeout_sec = 3600 # default 300 kills long calls at 5 minutesModel Selection
Every entry in models takes one of four forms:
// Auto-selection (recommended): the first available provider's default model
"auto";
// A provider: its default model (hardcoded, or <PROVIDER>_DEFAULT_MODEL)
"codex"; // -> Codex (GPT-6 Astra)
"claude"; // -> Claude Agent SDK (Claude Opus 5.5)
"gemini"; // -> Antigravity CLI (Gemini 3.8 Flash); `agy` works too
"openai"; // -> OpenAI API (GPT-6.1 Sol)
// provider:model — that model on that provider only
"codex:astra"; // -> Codex (GPT-6 Astra)
"openai:gpt-6-astra"; // -> OpenAI API, even when Codex is available
"gemini:pro"; // -> Antigravity CLI (Gemini 3.1 Pro)
"google:gemini-3.1-pro-preview"; // -> Google API
"copilot:sonnet"; // -> GitHub Copilot (Claude Sonnet 5.5)
"openrouter:z-ai/glm-5.2:online"; // -> OpenRouter with web search opt-in
// A bare model ID or alias — the first configured provider that offers it
"gpt-6-astra"; // -> Codex, else OpenAI API
"opus"; // -> Claude Agent SDK, else Anthropic API
"pro"; // -> Antigravity CLI, else Google API
"grok"; // -> grok-4.5 (XAI)
"z-ai/glm-5.2"; // -> OpenRouter (any vendor/model slug)Bare model names go to the first provider, in the order below, whose model list contains the name and that is set up (API key present; for Codex and Claude, a login file or token; for Antigravity, the agy binary). If that provider then fails with an authentication or availability error, the next provider that serves the same model takes over — a provider whose alias of that name points at a different model is never substituted (bare fable is Fable 5.1 on the Claude Agent SDK and Fable 5 on the Anthropic API, so it does not fail over between them). Copilot is never picked for bare names; use copilot:<model>.
Unknown names are rejected, never guessed: a typo or an unlisted model returns an error with up to three close matches, e.g. Unknown model "gtp-6-astra". Did you mean: gpt-6-astra? or Unknown openai model "spark" in "openai:spark". Did you mean: codex:spark?. OpenRouter is the exception for full vendor/model slugs, which are checked against OpenRouter's live catalog.
Auto Model Behavior:
chat mode:
["auto"]selects the first available provider and uses its default model, failing over down the listconsensus mode:
["auto"]automatically expands to the first 3 available providersroundtable mode:
["auto"]uses the first available provider
Provider priority order (subscription-based local providers first, then API-key providers), used by both auto and bare model names:
Codex (
codex→ GPT-6 Astra)Gemini via Antigravity CLI (
gemini/agy→ Gemini 3.8 Flash)Claude Agent SDK (
claude→ Claude Opus 5.5)Copilot (
copilot→ GPT-6.1 Sol;autoonly, never bare names)OpenAI (
openai→ GPT-6.1 Sol)Google (
google→ Gemini 3.1 Pro)XAI (
xai→ Grok 4.5)Anthropic (
anthropic→ Claude Opus 5.5)Mistral (
mistral→ Mistral Medium 3.5)DeepSeek (
deepseek→ DeepSeek V4 Pro)OpenRouter (
openrouter→ GLM 5.2)Abliteration (
abliteration/ablit→abliterated-model-large-v2)
Local agent permissions: Bare model names and auto reach the local agent providers whenever they are set up, not only when named explicitly. The Antigravity CLI runs agy with --dangerously-skip-permissions because headless calls cannot prompt for tool approval — every tool request is auto-approved, including shell commands and file writes. The Claude Agent SDK runs with bypassPermissions. Codex uses CODEX_SANDBOX_MODE (read-only by default). A read-only prompt is not an enforced security boundary for these providers, so use them only with trusted prompts and context. To keep a request on a plain API, name the provider: google:pro, anthropic:opus, openai:gpt-6-astra.
Advanced Configuration
Manual Installation Options
Option A: Direct Node.js execution
If you've cloned the repository locally:
{
"mcpServers": {
"converse": {
"command": "node",
"args": [
"C:\\Users\\YourUsername\\Documents\\Projects\\converse\\src\\index.js"
],
"env": {
"OPENAI_API_KEY": "your_key_here",
"GEMINI_API_KEY": "your_key_here",
"XAI_API_KEY": "your_key_here",
"ANTHROPIC_API_KEY": "your_key_here",
"MISTRAL_API_KEY": "your_key_here",
"DEEPSEEK_API_KEY": "your_key_here",
"OPENROUTER_API_KEY": "your_key_here",
"ABLITERATION_API_KEY": "ak_your_key_here"
}
}
}
}Option B: Local HTTP Development (Advanced)
For local development with HTTP transport (optional, for debugging):
First, start the server manually with HTTP transport:
# In a terminal, navigate to the project directory cd converse MCP_TRANSPORT=http npm run dev # Starts server on http://localhost:3157/mcpThen configure Claude to connect to it:
{ "mcpServers": { "converse-local": { "url": "http://localhost:3157/mcp" } } }
Important: HTTP transport requires the server to be running before Claude can connect to it. Keep the terminal with the server open while using Claude.
Configuration File Locations
The Claude configuration file is typically located at:
Windows:
%APPDATA%\Claude\claude_desktop_config.jsonmacOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonLinux:
~/.config/Claude/claude_desktop_config.json
For more detailed instructions, see the official MCP configuration guide.
💻 Running Standalone (Without Claude)
You can run the server directly without Claude for testing or development:
# Quick run (no installation needed)
npx converse-mcp-server
# Alternative package managers
pnpm dlx converse-mcp-server
yarn dlx converse-mcp-serverFor development setup, see the Development section below.
🐛 Troubleshooting
Common Issues
Server won't start:
Check Node.js version:
node --version(needs v20+)Try a different port:
PORT=3001 npm start
API key errors:
Verify your .env file has the correct format
Test with:
npm run test:real-api
Module import errors:
Clear cache and reinstall:
npm run clean
Long tool calls aborted mid-run (idle/timeout errors):
The abort almost always comes from the MCP client, not Converse — Converse's own limits are 30 min per provider call and 90 min per async job.
Claude Code: raise
CLAUDE_CODE_MCP_TOOL_IDLE_TIMEOUT(idle window, default 30 min stdio / 5 min HTTP) and checkMCP_TOOL_TIMEOUT(wall-clock). See Claude Code Environment Variables.Codex CLI: set
tool_timeout_secunder[mcp_servers.converse]in~/.codex/config.toml— the undocumented default is 300 seconds.Immune alternative: run the call with
async: trueand pollcheck_status— each poll is a fresh short request, so no client timeout applies.
Debug Mode
# Enable debug logging
LOG_LEVEL=debug npm run dev
# Start with debugger
npm run debug
# Trace all operations
LOG_LEVEL=trace npm run dev🔧 Development
Getting Started
# Clone the repository
git clone https://github.com/FallDownTheSystem/converse.git
cd converse
npm install
# Copy environment file and add your API keys
cp .env.example .env
# Start development server
npm run devScripts Available
# Server management
npm start # Start server (auto-kills existing server on port 3157)
npm run start:clean # Start server without killing existing processes
npm run start:port # Start server on port 3001 (avoids port conflicts)
npm run dev # Development with hot reload (auto-kills existing server)
npm run dev:clean # Development without killing existing processes
npm run dev:port # Development on port 3001 (avoids port conflicts)
npm run dev:quiet # Development with minimal logging
npm run kill-server # Kill any server running on port 3157
# Testing
npm test # Run all tests
npm run test:unit # Unit tests only
npm run test:integration # Integration tests
npm run test:e2e # End-to-end tests (requires API keys)
# Integration test subcategories
npm run test:integration:mcp # MCP protocol tests
npm run test:integration:tools # Tool integration tests
npm run test:integration:providers # Provider integration tests
npm run test:integration:performance # Performance tests
npm run test:integration:general # General integration tests
# Other test categories
npm run test:mcp-client # MCP client tests (HTTP-based)
npm run test:providers # Provider unit tests
npm run test:tools # Tool tests
npm run test:coverage # Coverage report
npm run test:watch # Run tests in watch mode
# Code quality
npm run lint # Check code style
npm run lint:fix # Fix code style issues
npm run format # Format code with ESLint (alias for lint:fix)
npm run validate # Full validation (lint + test)
# Utilities
npm run build # Build for production
npm run debug # Start with debugger
npm run check-deps # Check for outdated dependencies
npm run kill-server # Kill any server running on port 3157Development Notes
Port conflicts: The server uses port 3157 by default. If you get an "EADDRINUSE" error:
Run
npm run kill-serverto free the portOr use a different port:
PORT=3001 npm start
Transport Modes:
Stdio (default): Works automatically with Claude
HTTP: Better for debugging, requires manual start (
MCP_TRANSPORT=http npm run dev)
Testing with Real APIs
After setting up your API keys in .env:
# Run end-to-end tests
npm run test:e2e
# Test specific providers
npm run test:integration:providers
# Full validation
npm run validateValidation Steps
After installation, run these tests to verify everything works:
npm start # Should show startup message
npm test # Should pass all unit tests
npm run validate # Full validation suiteProject Structure
converse/
├── src/
│ ├── index.js # Main server entry point
│ ├── config.js # Configuration management
│ ├── router.js # Central request dispatcher
│ ├── continuationStore.js # State management
│ ├── systemPrompts.js # Tool system prompts
│ ├── providers/ # AI provider implementations
│ │ ├── index.js # Provider registry
│ │ ├── interface.js # Unified provider interface
│ │ ├── openai.js # OpenAI provider
│ │ ├── xai.js # XAI provider
│ │ ├── google.js # Google provider
│ │ ├── anthropic.js # Anthropic provider
│ │ ├── mistral.js # Mistral AI provider
│ │ ├── deepseek.js # DeepSeek provider
│ │ ├── openrouter.js # OpenRouter provider
│ │ ├── abliteration.js # Abliteration provider
│ │ ├── openrouter-discovery.js # Request-local OpenRouter slug discovery
│ │ ├── openai-compatible.js # Base for OpenAI-compatible APIs
│ │ ├── codex.js # Codex agentic SDK provider
│ │ ├── claude.js # Claude Agent SDK provider
│ │ ├── gemini-cli.js # Gemini via Antigravity CLI provider
│ │ └── copilot.js # GitHub Copilot SDK provider
│ ├── tools/ # MCP tool implementations
│ │ ├── index.js # Tool registry
│ │ ├── chat.js # Unified chat tool (chat/consensus/roundtable modes)
│ │ └── modes/ # parallel.js + roundtable.js execution engines
│ └── utils/ # Utility modules
│ ├── contextProcessor.js # File/image processing
│ ├── errorHandler.js # Error handling
│ └── logger.js # Logging utilities
├── tests/ # Comprehensive test suite
├── docs/ # API and architecture docs
└── package.json # Dependencies and scripts📦 Publishing to NPM
Note: This section is for maintainers. The package is already published as
converse-mcp-server.
Quick Publishing Checklist
# 1. Ensure clean working directory
git status
# 2. Run full validation
npm run validate
# 3. Test package contents
npm pack --dry-run
# 4. Test bin script
node bin/converse.js --help
# 5. Bump version (choose one)
npm version patch # Bug fixes: 1.0.1 → 1.0.2
npm version minor # New features: 1.0.1 → 1.1.0
npm version major # Breaking changes: 1.0.1 → 2.0.0
# 6. Test publish (dry run)
npm publish --dry-run
# 7. Publish to npm
npm publish
# 8. Verify publication
npm view converse-mcp-server
npx converse-mcp-server --helpVersion Guidelines
Patch (
npm version patch): Bug fixes, documentation updates, minor improvementsMinor (
npm version minor): New features, new model support, new tool capabilitiesMajor (
npm version major): Breaking API changes, major architecture changes
Post-Publication
After publishing, update installation instructions if needed and verify:
# Test direct execution
npx converse-mcp-server
npx converse
# Test MCP client integration
# Update Claude Desktop config to use: "npx converse-mcp-server"Troubleshooting Publication
Git not clean: Commit all changes first
Tests failing: Fix issues before publishing
Version conflicts: Check existing versions with
npm view converse-mcp-server versionsPermission issues: Ensure you're logged in with
npm whoami
🤝 Contributing
Fork the repository
Create a feature branch:
git checkout -b feature/amazing-featureMake your changes
Run tests:
npm run validateCommit changes:
git commit -m 'Add amazing feature'Push to branch:
git push origin feature/amazing-featureOpen a Pull Request
Development Setup
# Fork and clone your fork
git clone https://github.com/yourusername/converse.git
cd converse
# Install dependencies
npm install
# Create feature branch
git checkout -b feature/your-feature
# Make changes and test
npm run validate
# Commit and push
git add .
git commit -m "Description of changes"
git push origin feature/your-feature🙏 Acknowledgments
This MCP Server was inspired by and builds upon the excellent work from BeehiveInnovations/zen-mcp-server.
📄 License
MIT License - see LICENSE file for details.
🔗 Links
Available Tools
4 toolscancel_jobA
Cancel a running async job by its continuation_id.
Terminates queued or running jobs with graceful cleanup. Preserves partial results when available.
| Name | Required | Description | Default |
|---|---|---|---|
| continuation_id | Yes | The continuation_id of the job to cancel |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses graceful cleanup and partial results preservation, but lacks detail on side effects, reversibility, or error handling. No annotations provided.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with main action, no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers cancellation and cleanup, but lacks output description and error scenarios; no output schema to compensate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage 100% and description merely restates the parameter purpose; adds little beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states verb 'Cancel', resource 'async job', and identifier 'continuation_id'. Differentiates from siblings like check_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies use when needing to stop a job, but no explicit guidance on when to use vs check_status or when not to use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chatA
UNIFIED CHAT — talk to one or more AI models. mode "chat" (default): 1..N models answer independently in parallel. mode "consensus": ≥2 models answer, then refine after seeing each other. mode "roundtable": models answer sequentially, each building on the running transcript. Supports files, images, and continuation_id for multi-turn threads (you may switch modes on resume). Use model "auto" for automatic selection. IMPORTANT: use the "files" parameter to share code/file content instead of pasting into the prompt.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Execution mode. "chat" (default): independent parallel answers. "consensus": ≥2 models answer then refine via cross-feedback. "roundtable": sequential turn-based dialogue in the given model order. Default: "chat". | |
| async | No | Execute in the background. When true, returns a continuation_id immediately and processes the request asynchronously; poll with check_status. Default: false | |
| files | No | File paths to include as context (absolute or relative). Supports line ranges: file.txt{10:50}, file.txt{100:}. Example: ["./src/utils/auth.js{50:100}", "./config.json"]. IMPORTANT: Always use this parameter to share file content instead of copying code into the prompt. | |
| export | No | Export the conversation to disk. Creates a folder named for the continuation_id with numbered request/response files and metadata. Default: false | |
| images | No | Image paths for visual context (absolute or relative paths, or base64 data). Example: ["C:\Users\username\diagram.png", "./screenshot.jpg", "data:image/jpeg;base64,/9j/4AAQ..."] | |
| models | No | Models to use. Examples: ["auto"] (recommended), ["codex"], ["codex", "gemini", "claude"], ["codex:astra"], ["gpt-6-astra"]. Forms: "provider" (its default model), "provider:model" (that provider only), or a bare "model" (served by the first configured provider that offers it, local CLI providers first, failing over to the next). Providers: codex, gemini (agy), claude, copilot, openai, google, xai, anthropic, mistral, deepseek, openrouter, abliteration. Unknown names are rejected with suggestions. In mode "chat" each model answers independently; in "consensus" they refine after seeing each other; in "roundtable" they speak in the given ORDER, each seeing the transcript. Default: ["auto"]. | |
| prompt | Yes | Your question, topic, or task with relevant context. More detail enables better responses. Example: "How should I structure the authentication module for this Express.js API?" | |
| continuation_id | No | Continuation ID for a persistent multi-turn thread. Auto-generated in the first response; pass it back to continue. You MAY change the mode or models on a resuming turn. | |
| reasoning_effort | No | Reasoning depth for thinking models, weakest to strongest: "none" (reasoning off, where the model allows it), "minimal", "low", "medium" (balanced), "high", "xhigh", "max". Passed through by name when the model accepts it, otherwise clamped to the nearest tier it does. Default: "medium" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it explains that async returns a continuation_id immediately for background processing, that export writes numbered files and metadata to disk, and that continuation_id enables multi-turn threads with mode switching on resume. It omits any mention of cost, latency, or per-model failure behavior, which is a real gap for a multi-provider fan-out tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Long but front-loaded: the unified purpose and mode taxonomy come first, then supporting capabilities. The 'IMPORTANT: use the files parameter' instruction is duplicated verbatim in the schema's files description, which is minor redundancy rather than bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no annotations and no output schema, the description covers modes, asynchronous flow, file/image context, and thread continuity well. Its main omission is return shape — the agent is not told how per-model answers are presented or how consensus/roundtable output is structured, and with no output schema that must come from prose.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 9 parameters in detail, including mode, async, files, models, and reasoning_effort. The description largely restates that schema content (modes, files-vs-prompt advice, model 'auto') rather than adding syntax or edge-case meaning beyond it, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('talk to one or more AI models') and immediately enumerates the three execution modes with their distinct semantics. An agent can distinguish this from siblings cancel_job, check_status, and decide without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear routing guidance for mode selection, tells the agent to use the files parameter instead of pasting code, and links async=true to polling via check_status. It never contrasts this tool with the sibling 'decide', which appears to be an adjacent model-consultation tool, so the sibling boundary is left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_statusA
Check the status and progress of async jobs. Query specific jobs by continuation_id or list the 10 most recent jobs. Returns job status with start time and progress information.
| Name | Required | Description | Default |
|---|---|---|---|
| full_history | No | When used with continuation_id, returns the full conversation history for that continuation ID. Only use when there are multiple turns and you need the full conversation. | |
| continuation_id | No | Optional job continuation ID to query. If not provided, returns the 10 most recent jobs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must cover behavioral traits. It mentions return of job status, start time, and progress information, implying read-only behavior. However, it lacks details on error handling, rate limits, or authentication requirements, which are uncovered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise with two sentences, front-loading the main purpose. Every sentence adds value without redundancy, achieving maximum efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description adequately explains return values (status, start time, progress). It covers both usage modes. However, it lacks details on possible status values or error conditions, leaving minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with good parameter descriptions. The description adds minimal value beyond the schema, merely restating the two modes. Baseline 3 is appropriate as the schema already provides necessary semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'check' and resource 'async jobs', defining two specific usage modes: query by continuation_id or list recent jobs. This distinguishes it from sibling tools 'chat' and 'cancel_job', which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use each mode (by continuation_id or not), providing clear context. However, it does not explicitly state when not to use it or mention alternatives like 'cancel_job', though sibling names are present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
decideA
DECIDE — ask a System One decision model (TypeSafe Jev) typed questions about a state and get calibrated answers, not text. Question types: "noul" (yes/no → probability 0..1), "choice" (pick one of 2–255 named options → choice, per-option probabilities, confidence), "score" (ordered rubric of 2–10 levels → weighted position, per-level probabilities, confidence). Batch independent questions over the same state into one call: they are judged in parallel and in isolation, so none sees another's answer; each extra question adds its own input tokens. Best for fast semantic judgments (classify, route, select, verify, rank). Ask one narrow, coherent judgment per question, with its full meaning in the question; split independently useful dimensions, but a bounded action choice or contextual interpretation is a valid single question. Do counting, arithmetic, and date comparison in code. confidence measures how concentrated the distribution is, not permission to act: take the top option to pick a best, and treat a noul near 0.5 as "yes and no equally likely". Text only, no explanations are returned. Limits: ~64k tokens per request, ~32k for state plus the longest question.
| Name | Required | Description | Default |
|---|---|---|---|
| files | No | Text files added to the state as { "files": { "<path>": "<content>" } }; a given state moves to "input". Supports line ranges: file.txt{10:50}. Images are rejected. | |
| model | No | Decision model. "auto" (default): TypeSafe, falling back to OpenRouter. "jev-latest", "jev-1.13": first configured provider that serves it, with fallback. "typesafe:jev-1.13.0", "openrouter:~typesafe/jev-latest": that provider only. Providers: typesafe (TYPESAFE_API_KEY), openrouter (OPENROUTER_API_KEY). | |
| state | No | The material to judge: plain text, or JSON (an object with descriptively named fields is best; an array for sequences such as messages). Optional when "files" is given. | |
| questions | Yes | Named questions, all answered against the same state. The name is your own label and is returned as the answer key. Each question: { "type": "noul"|"choice"|"score", "instructions": string|object|array, "criteria": ... }. criteria — noul: optional { "true": "...", "false": "..." }; choice: required { "<option>": "description" | null } (2–255 options); score: required ordered array of level descriptions, lowest first (2–10 levels). instructions may be an object bundling the question with reference data, referenced by `name` in the text. Example: { "team": { "type": "choice", "instructions": "Which team should handle this?", "criteria": { "billing": "Payments, refunds", "technical": "Bugs, outages" } } } |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations at all, the description carries the full burden and fully delivers: parallel and isolated judging, token accounting per question, confidence semantics, "Text only, no explanations are returned," and token limits. This is rich behavioral disclosure beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place; it is front-loaded with the core purpose and then efficiently layers usage rules, output shape, batching behavior, and limits. The density is justified by the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter nested tool with no output schema, the description fully covers inputs, question construction, return semantics, model selection, batching, confidence interpretation, and limits. An agent has enough information to invoke and interpret results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema coverage is 100%, the description adds substantial meaning: it defines each question type's answer shape (noul → probability, choice → per-option probabilities, score → weighted position), explains model fallback behavior, and shows how to structure questions with an example. This significantly exceeds the baseline for schema-covered parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: "ask a System One decision model (TypeSafe Jev) typed questions about a state and get calibrated answers, not text." It clearly distinguishes itself from the sibling chat tool by emphasizing non-text, calibrated answers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong when-to-use guidance: "Best for fast semantic judgments (classify, route, select, verify, rank)" and explicit when-not-to-use instruction: "Do counting, arithmetic, and date comparison in code." It stops short of naming sibling tools as alternatives, so it doesn't fully meet the 5-level bar for explicit alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v4.5.0- Changed
chat1 field changed- changed
Input schema / properties / models / descriptionPrevious value: -"Models to use. Examples: [\"auto\"] (recommended), [\"codex\"], [\"codex\", \"gemini\", \"claude\"], [\"codex:astra\"], [\"gpt-6-astra\"]. Forms: \"provider\" (its default model), \"provider:model\" (that provider only), or a bare \"model\" (served by the first configured provider that offers it, local CLI providers first, failing over to the next). Providers: codex, gemini (agy), claude, copilot, openai, google, xai, anthropic, mistral, deepseek, openrouter. Unknown names are rejected with suggestions. In mode \"chat\" each model answers independently; in \"consensus\" they refine after seeing each other; in \"roundtable\" they speak in the given ORDER, each seeing the transcript. Default: [\"auto\"]."New value: +"Models to use. Examples: [\"auto\"] (recommended), [\"codex\"], [\"codex\", \"gemini\", \"claude\"], [\"codex:astra\"], [\"gpt-6-astra\"]. Forms: \"provider\" (its default model), \"provider:model\" (that provider only), or a bare \"model\" (served by the first configured provider that offers it, local CLI providers first, failing over to the next). Providers: codex, gemini (agy), claude, copilot, openai, google, xai, anthropic, mistral, deepseek, openrouter, abliteration. Unknown names are rejected with suggestions. In mode \"chat\" each model answers independently; in \"consensus\" they refine after seeing each other; in \"roundtable\" they speak in the given ORDER, each seeing the transcript. Default: [\"auto\"]."
1 tool update
v4.1.2- Added
decide
1 tool update
v4.0.0- Changed
chat1 field changed- changed
Input schema / properties / models / descriptionPrevious value: -"Models to use. Examples: [\"auto\"] (recommended), [\"codex\"], [\"codex\", \"gemini\", \"claude\"]. In mode \"chat\" each model answers independently; in \"consensus\" they refine after seeing each other; in \"roundtable\" they speak in the given ORDER, each seeing the transcript. Default: [\"auto\"]."New value: +"Models to use. Examples: [\"auto\"] (recommended), [\"codex\"], [\"codex\", \"gemini\", \"claude\"], [\"codex:astra\"], [\"gpt-6-astra\"]. Forms: \"provider\" (its default model), \"provider:model\" (that provider only), or a bare \"model\" (served by the first configured provider that offers it, local CLI providers first, failing over to the next). Providers: codex, gemini (agy), claude, copilot, openai, google, xai, anthropic, mistral, deepseek, openrouter. Unknown names are rejected with suggestions. In mode \"chat\" each model answers independently; in \"consensus\" they refine after seeing each other; in \"roundtable\" they speak in the given ORDER, each seeing the transcript. Default: [\"auto\"]."
1 tool update
v3.6.0- Changed
chat2 fields changed- changed
Input schema / properties / reasoning_effort / descriptionPrevious value: -"Reasoning depth for thinking models. Examples: \"none\" (no reasoning, fastest - GPT-5.1+ only), \"minimal\", \"low\", \"medium\" (balanced), \"high\", \"max\". Default: \"medium\""New value: +"Reasoning depth for thinking models, weakest to strongest: \"none\" (reasoning off, where the model allows it), \"minimal\", \"low\", \"medium\" (balanced), \"high\", \"xhigh\", \"max\". Passed through by name when the model accepts it, otherwise clamped to the nearest tier it does. Default: \"medium\"" - changed
Input schema / properties / reasoning_effort / enumPrevious value: -[ - "none", - "minimal", - "low", - "medium", - "high", - "max" -]New value: +[ + "none", + "minimal", + "low", + "medium", + "high", + "xhigh", + "max" +]
2 tool updates
v3.2.4- Added
cancel_job - Added
check_status
1 tool update
v3.1.0- Removed
check_status
1 tool update
v3.0.2- Removed
cancel_job
3 tool updates
v3.0.1- First observed
cancel_job - First observed
chat - First observed
check_status
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: chat handles model conversations, decide handles typed decision questions, and cancel_job/check_status manage async job lifecycle from opposite angles. Boundaries are reinforced by explicit modes and output types, so misselection is unlikely.
All names use lowercase snake_case and readable verbs, but the pattern is not fully uniform: cancel_job and check_status follow verb_noun, while chat and decide are bare verbs. This is a minor deviation rather than a confusing mix.
Four tools is well-scoped for a unified chat/decision server: chat, one async job control pair, and decide. Each tool has a clear role, and the surface does not feel thin or bloated.
Core capabilities are covered: multi-mode model chat, decision queries, and async job cancellation/status. A minor possible gap is explicit retrieval of completed async job results, though this may be handled through chat continuation_id.
Maintenance
Related MCP Connectors
Convene a panel of expert AI personas to debate any decision from every side.
Multi-model AI debates: GPT-4o, Claude, Gemini & 200+ models discuss, then synthesize insight.
Multiple Google accounts (Gmail, Calendar, Drive, Contacts, Tasks) in one Claude connector.
Multiple Google accounts (Gmail, Calendar, Drive, Contacts, Tasks) in one Claude connector.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceConnects Claude Code with multiple AI models (Gemini, Grok-3, ChatGPT, DeepSeek) simultaneously, allowing users to get diverse AI perspectives, conduct AI debates, and leverage each model's unique strengths.153MIT
- -licenseNot gradedqualityNot gradedmaintenanceGives Claude access to multiple AI models (Gemini, OpenAI, OpenRouter, Ollama) for enhanced development capabilities including extended reasoning, collaborative development, code review, and advanced debugging.-
- AlicenseNot gradedqualityAmaintenanceEnables Claude to consult over 17 AI platforms and 800,000+ models to provide alternative perspectives, code reviews, and diverse feedback. It features a unique personality system and supports multi-AI group discussions and debates directly within the chat interface.45MIT
- AlicenseNot gradedqualityNot gradedmaintenanceEnables Claude to orchestrate multi-model AI roundtable discussions between GPT-4o, Gemini, Grok, and DeepSeek through independent responses and cross-critique rounds. It supports various discussion presets such as debate, brainstorm, and consensus to facilitate diverse perspectives and synthesis.MIT