Agent Exchange
Allows chatting with OpenAI models via the Chat Completions API, with configurable model and base URL settings and persistent multi-turn conversation history.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Agent Exchangeask claude to plan a 3-day trip to kyoto, then get grok's take"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Agent_Exchange_MCP
An MCP server that lets an AI client (Claude Code, Claude Desktop, …) talk to other AI models through their APIs. Supports Anthropic, OpenAI and xAI, with multi-turn conversations saved to disk.
Features
Multiple providers — Anthropic (Claude), OpenAI, xAI (Grok). Choose the provider from the client UI, and switch it per call, even mid-conversation.
Multi-turn conversations — history is kept per
conversation_id.Persistent — each conversation is a JSON file on disk and survives server restarts.
Simple config — API keys and models live in a
.envfile.
Related MCP server: Multi-CLI MCP
Installation
Requires Python 3.11+.
git clone https://github.com/cat-people-82/Agent_Exchange_MCP.git
cd Agent_Exchange_MCP
pip install -e .
cp .env.example .env # then add your API keysConfiguration
Settings are read from .env in the project root (or the current directory, or the file named by AGENT_EXCHANGE_ENV_FILE). Real environment variables take precedence over .env. .env is git-ignored — never commit it.
Variable | Default | Purpose |
|
| Where conversations are saved |
| – / | Anthropic credentials and default model |
| – / | OpenAI settings |
| – / | xAI settings |
Provider selection is not configured here — it happens in the client (see below). .env only holds API keys, models and the data directory. You only need keys for the providers you use. Adjust the default OpenAI/xAI model IDs to whatever your account offers.
Register with an MCP client
Claude Code:
claude mcp add agent-exchange -- agent-exchange-mcpClaude Desktop — edit claude_desktop_config.json (Settings → Developer → Edit Config). command must be an executable (not the path to server.py), so point it at the Python interpreter of the environment where you ran pip install -e . and run the package as a module:
{
"mcpServers": {
"agent-exchange": {
"command": "/absolute/path/to/Agent_Exchange_MCP/.venv/bin/python",
"args": ["-m", "agent_exchange_mcp"]
}
}
}Tips:
Create the environment first:
python3 -m venv .venv && .venv/bin/pip install -e .(find an existing interpreter's path withwhich pythonwhile its environment is active).Claude Desktop doesn't inherit your shell environment; keys are read from the project's
.env, or pass them via an"env": { "ANTHROPIC_API_KEY": "..." }entry.Fully quit and reopen Claude Desktop after editing the config. Logs are in
~/Library/Logs/Claude/mcp-server-agent-exchange.log.
Tools
Tool | Description |
| Send a message and get the reply. Omit |
| Configured providers, default models, and whether a key is set. |
| Saved conversations (id, turns, system prompt). |
| Delete a conversation. |
chat returns {conversation_id, provider, model, reply}.
Example flow: ask Claude a question, then continue the same conversation with provider: "xai" to get Grok's take with the full history.
Choosing a provider
The provider is picked in your MCP client, not in config:
chathas aproviderparameter (anthropic|openai|xai), which clients show as a choice list and which the calling model or you can set.If it is omitted when starting a conversation and more than one provider has an API key, the server asks you to choose through the client's prompt (MCP elicitation, supported by clients such as Claude Code).
If the client doesn't support elicitation, or only one provider has a key, the first provider with a key is used.
A continued conversation keeps its last provider unless you pass a different one.
Storage
Conversations are stored as <conversation_id>.json (system prompt + message list) in the data directory. Files are plain text and contain your prompts and replies, so protect the directory accordingly. Delete a conversation with reset_conversation or by removing its file.
Notes
If a provider call fails, the turn is not saved, so history stays consistent.
Anthropic calls use adaptive thinking and streaming; OpenAI and xAI use the Chat Completions API.
Available Tools
4 toolschatA
Send a message to another AI model and get its reply.
Omit conversation_id to start a new conversation; pass the returned conversation_id to continue it (history is saved to disk, and you may switch provider between turns).
Args: message: The user message. conversation_id: Continue an existing conversation. provider: Provider to use. If omitted: a continued conversation keeps its last provider; a new one asks the user to choose (via the client UI when supported). model: Model ID override for this turn. system: System prompt (only applied when starting a conversation). max_tokens: Maximum tokens in the reply. effort: Thinking effort (Anthropic only; ignored by other providers).
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| effort | No | medium | |
| system | No | ||
| message | Yes | ||
| provider | No | ||
| max_tokens | No | ||
| conversation_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does much of it: discloses that history is persisted to disk, that the provider may be switched between turns, that an omitted provider prompts the user via the client UI, and that 'effort' is Anthropic-only and ignored elsewhere. It omits auth/permission needs, cost/rate-limit implications of max_tokens, and error behavior on invalid conversation_id.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core usage rule is front-loaded in the first two sentences, before the exhaustive Args list. The Args block is somewhat redundant for trivially named parameters, but each line carries actionable context, so it mostly earns its space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, no-annotation, no-output-schema tool, the description covers conversation lifecycle, provider resolution, and per-parameter constraints, and even hints at the returned conversation_id. It stops short of describing the reply payload shape or failure modes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it documents all seven parameters with meaningful semantics beyond the raw schema — notably the conditional behavior of 'provider' (inherits on continuation, prompts on new), the start-only application of 'system', and the Anthropic-only scope of 'effort'. 'max_tokens' and 'message' are near restatements, so this is not a full 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Send a message to another AI model and get its reply'), which is concrete and unambiguous. It is clearly distinct from the sibling tools (list_providers, list_conversations, reset_conversation), though it never names them to reinforce the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit operating instructions: omit conversation_id to start new, pass the returned id to continue, provider is auto-resolved for continuations, and 'system' only applies at conversation start. It lacks explicit when-to-use-this-vs-siblings routing (e.g., reset_conversation to clear history), which keeps it short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_conversationsA
List active conversations (id, turn count, system prompt).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It adds one behavioral fact — that only 'active' conversations are returned, implying archived ones are excluded — but is silent on ordering, pagination, or read-only safety profile. The listed return fields are redundant with the existing output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no waste. The action verb leads and the parenthetical detail follows efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and an output schema already documenting the return shape, the description needs little. Restating the returned fields is mildly redundant, and it omits ordering or scope details, but nothing critical to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a no-param tool applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List active conversations') and even names the fields returned, so the agent knows what it gets back. It does not differentiate itself against siblings like reset_conversation or list_providers, but the operations are distinct enough that confusion is unlikely.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the name — an agent would call this to discover existing conversations — but the description gives no explicit when-to-use, no prerequisites, and no mention of alternatives such as chat. Adequate but with a clear gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_providersA
List configured providers, their default models, and whether an API key is set.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the behavioral burden. 'List' strongly implies a read-only, non-destructive operation, and it discloses that API key presence is returned, but it does not explicitly state permissions, side effects, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence that front-loads the primary verb and lists only meaningful return attributes. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the zero-parameter schema and the presence of an output schema, the description is nearly complete. It adds enough context to distinguish the tool, though it could mention read-only safety explicitly since annotations are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is no parameter information in the description because none is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specifies the verb (List), the resource (configured providers), and the exact returned fields (default models, API key status). It is clearly distinguishable from the chat/conversation siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative guidance is present. The description only states what the tool does, leaving the agent to infer the use case from the operation itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reset_conversationC
Delete a conversation and its history.
| Name | Required | Description | Default |
|---|---|---|---|
| conversation_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose that history is destroyed alongside the conversation, which is valuable, but says nothing about irreversibility, required permissions, confirmation steps, or effects on related resources.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with zero filler and the key scope (history included) front-loaded. It is arguably too terse for a destructive operation, but every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. However, for a destructive, annotation-free operation the description omits irreversibility, permissions, and post-deletion state, leaving significant gaps for a mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the schema documents nothing about conversation_id beyond its type. The description only implies a single conversation is targeted; it adds no format, source, or constraint information to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (delete) and resource (a conversation), and usefully clarifies that the reset operation also removes history, which the name alone doesn't convey. It does not differentiate itself from siblings like list_conversations or chat, but the core action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the sibling chat or list_conversations tools, nor any stated prerequisites or consequences to consider before invoking. The agent is left to infer that this is the destructive counterpart to chat.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
chat - First observed
list_conversations - First observed
list_providers - First observed
reset_conversation
TDQS
Scored across 4 tools
chat is the sole action tool, while list_providers, list_conversations, and reset_conversation are clearly scoped to different resources. There is no meaningful overlap between messaging and conversation/provider management.
Three tools follow a clear verb_noun pattern (list_providers, list_conversations, reset_conversation), but chat is a single-word noun/verb that breaks the pattern slightly. The set remains readable and predictable overall.
Four tools cleanly cover the server's narrow scope: sending messages, discovering providers, inspecting conversations, and deleting a conversation. No tool feels redundant, and the set is not bloated.
Core conversation lifecycle is covered (start/continue/delete/list) and providers can be listed. However, there is no get_conversation or history-retrieval tool, and provider/model management is read-only, leaving minor gaps.
Maintenance
Related MCP Connectors
MCP server for AI dialogue using various LLM models via AceDataCloud
Hosted MCP server connecting claude.ai, ChatGPT and other AI apps to your own computer
Cloud-hosted MCP server for durable AI memory
- QuallaaOAuthcom.quallaa
Talk to your public-facing AI from any MCP client — Claude, ChatGPT, Cursor, Cline, Windsurf.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceAn MCP server that functions as an intelligent gateway for multiple LLM backends including OpenAI, Claude, and Ollama. It supports automatic provider fallback, streaming responses via Server-Sent Events, and real-time monitoring for robust AI integration.MIT
- FlicenseAqualityDmaintenanceAn MCP server that bridges multiple AI clients (Claude, Gemini, Codex, OpenCode) so they can call each other as tools.138 npm73-
- -licenseNot gradedqualityNot gradedmaintenanceAn MCP server that enables multiple AI models (Claude, ChatGPT, Gemini) to share and persist context via a local-first vault of markdown files and SQLite index, allowing seamless cross-AI memory.-
- AlicenseNot gradedqualityDmaintenanceA production-ready MCP server for persistent AI memory across LLMs like Claude and ChatGPT. Provides automatic conversation backup, multi-user support, and multi-storage (PostgreSQL, Redis, Qdrant).8 npm14MIT