Skip to main content
Glama

polymnemo

CI License: MIT

Python MCP Postgres Neon SQLAlchemy Pydantic Cloudflare R2 Azure Cloud Run Terraform Docker pytest Ruff

A shared long-term memory across any LLM, over MCP.

Point Claude Desktop, an MCP-capable IDE, or any MCP client at one polymnemo endpoint and they share the same memories โ€” stored in your own Postgres. Store a fact with one assistant, recall it from another; save whole sessions and reload them; even attach files, images, or video. Embeddings run locally (no embedding API key), and the server makes no generative-LLM calls.

Status: active development. Semantic memory + pgvector store, session save/reload, and multimedia memories all work; deployable to Azure Container Apps (or Cloud Run) via Terraform. Wiki ยท Issues

Features

  • ๐Ÿ”— Cross-LLM shared โ€” point any MCP client at one endpoint; they share the same memory.

  • ๐Ÿง  Semantic recall โ€” vector search over Postgres + pgvector, not keyword matching.

  • ๐Ÿ’ฌ Sessions โ€” save a full transcript and reload it verbatim, or recall across it.

  • ๐Ÿ–ผ๏ธ Multimedia โ€” attach files, images, or video; bytes go to object storage, only a searchable description is embedded.

  • ๐Ÿ”’ Local & private โ€” embeddings run locally (ONNX): no embedding API key, and no generative-LLM calls, ever.

  • ๐Ÿ‘ฅ Namespaces โ€” "born-shared" collections readable by everyone, vs. private-to-owner; writes are always owner-scoped.

  • ๐Ÿงฉ Pluggable layers โ€” store, embedder, auth, retriever, blob store, and rate limiter are all swappable Protocols.

  • โ˜๏ธ Multi-cloud deploy โ€” one Terraform stack to Azure Container Apps or Cloud Run, scale-to-zero.

  • ๐Ÿšฆ Rate limiting โ€” optional global token bucket.

Related MCP server: memento

How it works

flowchart LR
    Clients["MCP clients<br/>(Claude Desktop, IDEs, โ€ฆ)"] -->|"/mcp ยท Bearer key"| P["polymnemo<br/>(MCP server)"]
    P --> DB[("Postgres + pgvector<br/>text + pointers")]
    P -. "large files<br/>(presigned URLs)" .-> OS[("Object storage<br/>S3 / R2")]

A request carries a bearer key (which resolves to a user_id); the tool passes the rate-limit gate, then delegates to a MemoryService that chunks + embeds text and stores the vectors in pgvector โ€” large files go to object storage via presigned URLs, with only a searchable description embedded.

Every layer is a typing.Protocol, wired together by a composition root (context.py), so you can swap an implementation without touching the tools:

Layer

Default

Swap for

Store

PostgresStore (pgvector)

InMemoryStore (dev/tests)

Embedder

fastembed (local ONNX)

StubEmbedder (offline)

Auth

per-user bearer keys

static single-user (dev)

Retriever

VectorRetriever

your own ranker

BlobStore

S3 / R2

off

RateLimiter

global token bucket

off

The ping tool returns the active layers, so you can see how a running server is wired.

Quickstart

Requires Python 3.11+.

1. Install

python -m venv .venv
source .venv/bin/activate            # Windows: .venv\Scripts\activate
pip install -e .                     # add ".[dev]" for the test + lint tooling

2. Provision Postgres (pgvector)

The durable store is Postgres + pgvector; the easiest hosted option is Neon (use the pooled connection string). Apply the schema once:

psql "<your-connection-string>" -f scripts/schema.sql

3. Configure

Copy .env.example to .env and set the database URL and at least one API key:

POLYMNEMO_DATABASE_URL=postgresql://user:pass@host/db?sslmode=require
POLYMNEMO_API_KEYS=sk-alice-secret:alice,sk-bob-secret:bob   # "key:user_id" pairs

Each key maps a bearer token to a user_id; writes are scoped to that user.

4. Run

polymnemo                            # Streamable HTTP at http://127.0.0.1:8000/mcp

Or with Docker:

docker build -t polymnemo .
docker run -e POLYMNEMO_DATABASE_URL="..." -e POLYMNEMO_API_KEYS="sk-alice-secret:alice" \
           -e PORT=8080 -p 8080:8080 polymnemo

5. Connect an MCP client

Any client that supports remote (HTTP) MCP servers with custom headers needs two things:

  • Endpoint: http://<host>:<port>/mcp

  • Header: Authorization: Bearer <your-key>

For clients that read an mcpServers config:

{
  "mcpServers": {
    "polymnemo": {
      "url": "http://127.0.0.1:8000/mcp",
      "headers": { "Authorization": "Bearer sk-alice-secret" }
    }
  }
}

Or verify with the inspector:

npx @modelcontextprotocol/inspector
# Transport: Streamable HTTP ยท URL: http://127.0.0.1:8000/mcp
# Header:    Authorization: Bearer sk-alice-secret  โ†’ call ping / remember / recall

Concepts

  • Users & keys โ€” each bearer key maps to a user_id; writes are owner-scoped (you can only edit or delete your own memories).

  • Namespaces โ€” memories live in namespaces. A shared namespace (default shared) is readable by everyone ("born shared"); everything else is private to its owner. Sessions and media default to private namespaces.

  • Chunking โ€” long content is split into chunks on write (one vector each), so remember may return several ids and recall returns the closest chunks.

Tools

polymnemo exposes MCP tools for storing, searching, and managing memories:

  • Memory โ€” remember, recall, list_memories, get_memory, update, forget

  • Sessions โ€” save_session, load_session

  • Media โ€” create_upload, confirm_upload, get_download_url

Plus a ping health check and a memory://{namespace} resource for auto-injecting a collection. Media bytes go to object storage via presigned URLs โ€” never through the MCP channel โ€” with only a searchable description embedded (needs the blob extra).

Full arguments and return shapes live in the dedicated MCP tools reference (coming soon). See Concepts for how keys, namespaces, and chunking work.

Configuration

POLYMNEMO_* environment variables (or .env) โ€” full list in .env.example. The essentials:

Variable

Default

Purpose

POLYMNEMO_DATABASE_URL

(unset)

Postgres+pgvector DSN. Required in production; unset โ†’ in-memory (dev/tests).

POLYMNEMO_API_KEYS

(empty)

"key1:alice,key2:bob" โ€” required for bearer auth.

POLYMNEMO_SHARED_NAMESPACES

shared

Namespaces readable by every user.

POLYMNEMO_HOST / POLYMNEMO_PORT / POLYMNEMO_MCP_PATH

127.0.0.1 / 8000 / /mcp

Transport.

POLYMNEMO_RATELIMIT_ENABLED / _PER_MIN

false / 600

Optional global rate limit (ops per minute).

POLYMNEMO_BLOB_BACKEND (+ _BUCKET / _ENDPOINT_URL / _ACCESS_KEY_ID / _SECRET_ACCESS_KEY)

none

Object storage for media memories; s3 = Cloudflare R2 / S3-compatible.

Development

pip install -e ".[dev]"
pytest                               # fast, offline (stub embedder, in-memory store)
ruff check . && ruff format --check .   # lint + format (enforced in CI)

Postgres tests run when TEST_DATABASE_URL points at a pgvector Postgres. CI (.github/workflows/ci.yml) runs lint + the suite with coverage and posts a pass/fail/coverage table to each run's summary.

Deploy

polymnemo is stateless (all state in Neon + object storage), so it runs on Azure Container Apps (primary) or Google Cloud Run with scale-to-zero. Everything is Terraform in deploy/terraform/: shared neon/ (Postgres) and r2/ (media bucket) roots own the durable state, and a compute root deploys a service that reads both โ€” so memories and media are shared across clouds.

The Azure path deploys from CI in one click: set the GitHub secrets (scripts/setup-github-secrets.sh), bootstrap the state backend (scripts/bootstrap-tfstate-azure.sh), then run the deploy (azure) workflow (neon โ†’ schema โ†’ r2 โ†’ build โ†’ app). GCP is a manual failover. Full walkthrough in docs/deploy.md.

License

MIT โ€” see LICENSE.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides persistent, local-first AI memory across sessions via MCP tools for storing, searching, and retrieving context from past interactions.
    1
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Provides persistent memory for AI coding agents via MCP, enabling agents to store and semantically recall facts, events, and lessons across sessions, all running locally without cloud dependencies.
    Apache 2.0
  • A
    license
    A
    quality
    D
    maintenance
    Provides persistent memory with semantic search for MCP-based AI agents, enabling them to store and recall information across sessions using vector embeddings.
    4
    1
    MIT