polymnemo
Uses Cloudflare R2 (S3-compatible) for object storage to store multimedia files (images, video, etc.) via presigned URLs, with only searchable descriptions embedded.
Provides persistent long-term memory storage using PostgreSQL with pgvector, enabling semantic vector search and retrieval across memories.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@polymnemoRemember that I prefer Python over JavaScript"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
polymnemo
A shared long-term memory across any LLM, over MCP.
Point Claude Desktop, an MCP-capable IDE, or any MCP client at one polymnemo endpoint and they share the same memories โ stored in your own Postgres. Store a fact with one assistant, recall it from another; save whole sessions and reload them; even attach files, images, or video. Embeddings run locally (no embedding API key), and the server makes no generative-LLM calls.
Status: active development. Semantic memory + pgvector store, session save/reload, and multimedia memories all work; deployable to Azure Container Apps (or Cloud Run) via Terraform. Wiki ยท Issues
Features
๐ Cross-LLM shared โ point any MCP client at one endpoint; they share the same memory.
๐ง Semantic recall โ vector search over Postgres + pgvector, not keyword matching.
๐ฌ Sessions โ save a full transcript and reload it verbatim, or recall across it.
๐ผ๏ธ Multimedia โ attach files, images, or video; bytes go to object storage, only a searchable description is embedded.
๐ Local & private โ embeddings run locally (ONNX): no embedding API key, and no generative-LLM calls, ever.
๐ฅ Namespaces โ "born-shared" collections readable by everyone, vs. private-to-owner; writes are always owner-scoped.
๐งฉ Pluggable layers โ store, embedder, auth, retriever, blob store, and rate limiter are all swappable
Protocols.โ๏ธ Multi-cloud deploy โ one Terraform stack to Azure Container Apps or Cloud Run, scale-to-zero.
๐ฆ Rate limiting โ optional global token bucket.
Related MCP server: memento
How it works
flowchart LR
Clients["MCP clients<br/>(Claude Desktop, IDEs, โฆ)"] -->|"/mcp ยท Bearer key"| P["polymnemo<br/>(MCP server)"]
P --> DB[("Postgres + pgvector<br/>text + pointers")]
P -. "large files<br/>(presigned URLs)" .-> OS[("Object storage<br/>S3 / R2")]A request carries a bearer key (which resolves to a user_id); the tool passes
the rate-limit gate, then delegates to a MemoryService that chunks + embeds
text and stores the vectors in pgvector โ large files go to object storage via
presigned URLs, with only a searchable description embedded.
Every layer is a typing.Protocol, wired together by a composition root
(context.py), so you can swap an implementation
without touching the tools:
Layer | Default | Swap for |
Store |
|
|
Embedder |
|
|
Auth | per-user bearer keys | static single-user (dev) |
Retriever |
| your own ranker |
BlobStore | S3 / R2 | off |
RateLimiter | global token bucket | off |
The ping tool returns the active layers, so you can see how a running server is
wired.
Quickstart
Requires Python 3.11+.
1. Install
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e . # add ".[dev]" for the test + lint tooling2. Provision Postgres (pgvector)
The durable store is Postgres + pgvector; the easiest hosted option is Neon (use the pooled connection string). Apply the schema once:
psql "<your-connection-string>" -f scripts/schema.sql3. Configure
Copy .env.example to .env and set the database URL and at least one API key:
POLYMNEMO_DATABASE_URL=postgresql://user:pass@host/db?sslmode=require
POLYMNEMO_API_KEYS=sk-alice-secret:alice,sk-bob-secret:bob # "key:user_id" pairsEach key maps a bearer token to a user_id; writes are scoped to that user.
4. Run
polymnemo # Streamable HTTP at http://127.0.0.1:8000/mcpOr with Docker:
docker build -t polymnemo .
docker run -e POLYMNEMO_DATABASE_URL="..." -e POLYMNEMO_API_KEYS="sk-alice-secret:alice" \
-e PORT=8080 -p 8080:8080 polymnemo5. Connect an MCP client
Any client that supports remote (HTTP) MCP servers with custom headers needs two things:
Endpoint:
http://<host>:<port>/mcpHeader:
Authorization: Bearer <your-key>
For clients that read an mcpServers config:
{
"mcpServers": {
"polymnemo": {
"url": "http://127.0.0.1:8000/mcp",
"headers": { "Authorization": "Bearer sk-alice-secret" }
}
}
}Or verify with the inspector:
npx @modelcontextprotocol/inspector
# Transport: Streamable HTTP ยท URL: http://127.0.0.1:8000/mcp
# Header: Authorization: Bearer sk-alice-secret โ call ping / remember / recallConcepts
Users & keys โ each bearer key maps to a
user_id; writes are owner-scoped (you can only edit or delete your own memories).Namespaces โ memories live in namespaces. A shared namespace (default
shared) is readable by everyone ("born shared"); everything else is private to its owner. Sessions and media default to private namespaces.Chunking โ long content is split into chunks on write (one vector each), so
remembermay return several ids andrecallreturns the closest chunks.
Tools
polymnemo exposes MCP tools for storing, searching, and managing memories:
Memory โ
remember,recall,list_memories,get_memory,update,forgetSessions โ
save_session,load_sessionMedia โ
create_upload,confirm_upload,get_download_url
Plus a ping health check and a memory://{namespace} resource for
auto-injecting a collection. Media bytes go to object storage via presigned URLs
โ never through the MCP channel โ with only a searchable description embedded
(needs the blob extra).
Full arguments and return shapes live in the dedicated MCP tools reference (coming soon). See Concepts for how keys, namespaces, and chunking work.
Configuration
POLYMNEMO_* environment variables (or .env) โ full list in
.env.example. The essentials:
Variable | Default | Purpose |
| (unset) | Postgres+pgvector DSN. Required in production; unset โ in-memory (dev/tests). |
| (empty) |
|
|
| Namespaces readable by every user. |
|
| Transport. |
|
| Optional global rate limit (ops per minute). |
|
| Object storage for media memories; |
Development
pip install -e ".[dev]"
pytest # fast, offline (stub embedder, in-memory store)
ruff check . && ruff format --check . # lint + format (enforced in CI)Postgres tests run when TEST_DATABASE_URL points at a pgvector Postgres. CI
(.github/workflows/ci.yml) runs lint + the suite with coverage and posts a
pass/fail/coverage table to each run's summary.
Deploy
polymnemo is stateless (all state in Neon + object storage), so it runs on
Azure Container Apps (primary) or Google Cloud Run with scale-to-zero.
Everything is Terraform in deploy/terraform/: shared
neon/ (Postgres) and r2/
(media bucket) roots own the durable state, and a compute root deploys a service
that reads both โ so memories and media are shared across clouds.
The Azure path deploys from CI in one click: set the GitHub secrets
(scripts/setup-github-secrets.sh), bootstrap the state backend
(scripts/bootstrap-tfstate-azure.sh), then run the deploy (azure) workflow
(neon โ schema โ r2 โ build โ app). GCP is a manual failover. Full walkthrough
in docs/deploy.md.
License
MIT โ see LICENSE.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Shared, governed long-term memory for AI agents across tools and sessions via MCP and REST.
Private persistent memory for Claude, ChatGPT & Gemini via MCP - semantic search, zero-code setup.
Persistent AI memory shared across Claude, ChatGPT, coding agents, and compatible MCP clients.
Persistent memory for AI agents โ log and recall conversation context over MCP.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceProvides persistent, local-first AI memory across sessions via MCP tools for storing, searching, and retrieving context from past interactions.1MIT
- AlicenseNot gradedqualityAmaintenanceProvides persistent memory for AI coding agents via MCP, enabling agents to store and semantically recall facts, events, and lessons across sessions, all running locally without cloud dependencies.Apache 2.0
- AlicenseAqualityDmaintenanceProvides persistent memory with semantic search for MCP-based AI agents, enabling them to store and recall information across sessions using vector embeddings.41MIT
- AlicenseNot gradedqualityBmaintenanceProvides persistent semantic memory for AI agents via MCP, enabling them to remember, recall, list, update, and forget memories with vector-based similarity search.ISC