Custom MCP Server
Provides a proxy layer for interacting with OpenAI-compatible chat completion APIs, enabling session-based conversational memory, context management, and resilient agent calls.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Custom MCP ServerAsk the agent: "What's the weather like today?""
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Custom MCP Server
A custom MCP (Model Control Proxy) server that sits between clients and an OpenAI-compatible chat-completion backend. It provides:
Input validation, normalization, and chunking
Session-based conversational memory (in-memory store, swappable)
Sliding-window context management with rough token capping
Prompt construction with injectable system instructions
Resilient async agent calls (timeout + single retry on 429/5xx)
This is the MVP scope (milestones M1 + M2). Streaming, real summarization, persistent stores, and auth are deferred.
Requirements
Python 3.11+
uv (package & environment manager)
Related MCP server: MCP Server with LLM Integration
Install
uv sync --all-extras
# activate the environment
source .venv/bin/activateConfigure
cp .env.example .env
# edit .env and set AGENT_API_KEY, AGENT_API_BASE_URL, AGENT_MODELRun
uv run uvicorn app.main:app --reload --port 8000Endpoints
POST /agent/input
curl -s http://localhost:8000/agent/input \
-H 'Content-Type: application/json' \
-d '{"user_input": "Hello!"}'GET /session/{id}
curl -s http://localhost:8000/session/<session_id>DELETE /session/{id}
curl -X DELETE http://localhost:8000/session/<session_id>GET /health
curl -s http://localhost:8000/healthTest
uv run pytest -qLayout
app/
main.py FastAPI app, logging, health
config.py Settings (pydantic-settings)
api/ HTTP routers
models/schemas.py Pydantic request/response models
core/ input_processor, context_manager, prompt_builder, agent_client
storage/ SessionStore protocol + in-memory implementation
utils/ logging + token estimation
tests/ pytest suite (agent backend mocked)
docs/ Design documentation (architecture, data model, API, ADRs)Design documentation
See docs/ for the full design package:
docs/architecture.md— system context, container, component, sequence, and deployment diagrams (Mermaid)docs/data-model.md— Session / Message ER diagram and lifecycledocs/api.md— HTTP API spec and error modeldocs/adr/— Architecture Decision Records
This server cannot be deployed
Maintenance
Related MCP Connectors
Pass messages between AI agents with cleaning, metadata enrichment, and metered billing.
Drop-in proxy keeping OpenAI Assistants API calls working past the August 26, 2026 sunset, plus a Th
Real-time chat hub for AI agents — Claude Code, Cursor, Cline, Codex over MCP or REST.
Related MCP Servers
- FlicenseNot gradedqualityNot gradedmaintenanceHigh-performance Model Context Protocol server supporting multiple LLM providers (OpenRouter, OpenAI, Groq) with WebSocket API and conversation history persistence.-
- AlicenseNot gradedqualityDmaintenanceEnables chat with multiple LLM providers (OpenAI and Anthropic) while maintaining persistent conversation memory. Provides extensible tool framework for various operations including echo functionality and conversation storage/retrieval.MIT
- AlicenseAqualityCmaintenanceBridges MCP clients (like Claude Desktop) to A2A agents, enabling message sending and agent card retrieval via a stateless, non-persistent server.3MIT
- FlicenseNot gradedqualityCmaintenanceEnables coordinating multiple AI agents over HTTP with authenticated messaging, cached read-only Notion context, and safe proxying to registered endpoints.-