mcp-memory-server
Supports Perplexity as a remote client for the persistent knowledge-graph memory server, enabling memory storage and retrieval across sessions.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-memory-serverremember that the project deadline is Friday"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Sepia — Memory MCP Server
7 tools. 1 purpose: remember everything so your AI doesn't forget — and never needs to be reminded.
A personal, self-hosted remote knowledge-graph memory server for AI coding agents — a drop-in upgrade from the official local-file memory MCP server, with:
7 focused MCP tools over Streamable HTTP (not 17)
MCP server instructions — a usage contract auto-injected into the model's system prompt, so your AI recalls and persists without you asking
Always-on editor instructions — per-editor instruction files (VS Code prompts, Cursor rules, CLAUDE.md, AGENTS.md) injected into every session, so no editor can skip memory
A bundled Agent Skill (
SKILL.md, open standard) that teaches any editor the full usage guideA web dashboard with search, CRUD, stats, and an interactive knowledge-graph view
Out-of-the-box support for online AIs that speak MCP: Grok, ChatGPT, Claude, Gemini, Perplexity
$0/month on the free tiers of Fly.io + Neon Postgres + Netlify
Why this exists
The official MCP memory server is a single local JSONL file — no remote access, no search across sessions, no scaling. The remote alternatives are either overkill (17 tools, RBAC, audit trails, team workflows), gone (mem0 went hosted-only SaaS), or local-first (basic-memory, claude-mem).
Nobody ships a self-hosted, single-user, remote knowledge-graph memory server with a dashboard and a skill. That's the gap this project fills — the Goldilocks version, on infrastructure you control.
Related MCP server: MemoryVault MCP
Features
Feature | What it does |
🧠 Knowledge graph | Entities (nodes), weighted relations (edges), memories (facts/observations) with importance scoring, in isolated namespaces |
🔎 Search + traversal | Unified keyword search across everything; BFS graph traversal from any entity |
🧹 | Idempotent maintenance sweep: decay-scoring, dedup, purge — pure SQL, no LLM calls |
📋 Server instructions | A usage contract sent in the MCP |
⚡ Always-on instructions |
|
🛠️ Bundled Agent Skill |
|
� Conversation migration |
|
🖥️ Web dashboard | SvelteKit app on Netlify (SSR + remote functions): landing + pricing pages, search with URL-backed filters, graph view, conversations, stats, and a "Connect an AI" page — never wakes the API's scaled-to-zero machine |
🌐 Online AI support | Grok, ChatGPT, Claude web, Gemini (Spark), Perplexity, Le Chat all accept remote MCP connectors — your memory follows you to the web |
🔐 Two-phase auth | Phase 1: static Bearer token (local editors). Phase 2: OAuth 2.1 + PKCE live — built-in authorization server via |
Architecture
┌─────────────────────────────────┐
│ Browser (you) │
│ sepia.svelte-apps.me │
│ Dashboard (SvelteKit SSR, │
│ served from Netlify) │
└───────────────┬─────────────────┘
│ HTTPS + Bearer token / PKCE
▼
┌────────────────────────┐ ┌────────────────────────────────────┐
│ MCP Clients │ │ Fly.io App (Bun.serve) │
│ │ │ │
│ Local: Cursor, Zed, │──▶│ /mcp TMCP server (7 tools + │
│ Claude Code, Copilot, │ │ instructions) │
│ OpenCode │ │ /api/* REST (same auth, CORS │
│ Web: Grok, ChatGPT, │ │ allowlist) │
│ Claude.ai, Gemini, │ └───────────────┬────────────────────┘
│ Perplexity │ │ @neondatabase/serverless
└────────────────────────┘ ▼
┌────────────────────────────────────┐
│ Neon Postgres (Free Tier) │
│ namespaces · entities · relations │
└────────────────────────────────────┘Stack: Bun · TMCP (Valibot adapters, HttpTransport) · Neon Postgres · Drizzle ORM (type-safe query builder + sql template + migrations) · Svelte 5/SvelteKit (adapter-netlify, SSR + remote functions) · Tailwind CSS v4 · cytoscape.js
Key decision: the MCP endpoint and the REST API share one Bun process on one Fly.io machine — TMCP's HttpTransport mounts at /mcp inside an existing Bun.serve. The dashboard is a SvelteKit app on Netlify (SSR + remote functions): free tier, and it never wakes the Fly VM (which scales to zero) — the machine only spins up for real API calls from agents.
flowchart LR
subgraph Clients
L[Local editors<br/>Cursor · Zed · Claude Code<br/>Copilot · OpenCode]
W[Online AIs<br/>Grok · ChatGPT · Claude<br/>Gemini · Perplexity]
end
subgraph Fly["Fly.io (scale-to-zero)"]
B[Bun.serve]
M["/mcp — TMCP server<br/>7 tools + instructions"]
A["/api/* — REST<br/>CORS allowlist"]
end
N[(Neon Postgres<br/>free tier)]
D[Netlify<br/>Dashboard app]
L --> M
W --> M
B --> N
D -- "fetch /api/*" --> AThe 7 Tools
# | Tool | Actions | What it does |
1 |
| create, list, get, delete | Organize memory into isolated spaces |
2 |
| create, get, update, delete, find, batch_update | Knowledge graph nodes (people, concepts, projects, tools) |
3 |
| create, delete, list | Directed, weighted edges between entities |
4 |
| create, get, update, delete, query, batch_update, ingest | Facts/observations/preferences with importance scoring; conversation digests |
5 |
| — | Unified keyword + metadata search across all data |
6 |
| — | BFS walk of the knowledge graph from an entity |
7 |
| — | Decay sweep + dedup + purge (idempotent maintenance) |
Why 7 instead of 17: FlarelyLegal's 17 tools split entity search, memory queries, conversations, and admin into separate tools. By using action enums inside manage_* tools, the LLM surface stays clean while covering all capabilities — including conversation migration (manage_memory action=ingest) and bulk updates (batch_update). No RBAC, no audit trails — those are team features a personal server doesn't need. Semantic/vector search is a deliberate future upgrade; search ships keyword + metadata for v1.
Remember Without Being Asked
Three complementary channels, one contract (src/instructions.ts):
MCP
instructionsfield — the server sends a usage contract in theinitializehandshake; clients that support it (Claude Code, Codex, VS Code Copilot Chat, Goose, Claude Desktop) inject it into the model's system prompt. The model recalls before working and persists after learning — no reminder prompts.Always-on instruction files (
skills/sepia/always-on/) — the same condensed contract, installed into each editor's own instruction system: VS Code*.instructions.mdwithapplyTo: '**/*'(auto-attached to every chat request), Cursor.mdcwithalwaysApply: true(every session, unconditionally), a section in~/.claude/CLAUDE.md(loaded at session start), and anAGENTS.mdsection for Codex/other agents. Skills are on-demand by design in every platform, so this channel is what actually forces memory usage in editors that ignoreinstructions(Cursor).Bundled Agent Skill (
skills/sepia/SKILL.md) — the extended guide (tool-by-tool detail, examples, edge cases), delivered through the open Agent Skills standard. Loads when memory is relevant.
The contract teaches: search before meaningful work, persist durable facts (preferences, decisions, conventions), prefer update over duplicate, link memories to entities, score importance 0–1, never store credentials or ephemeral chat content — and migrate conversations between AIs via manage_memory action=ingest (handoff digests with status: active/paused/done).
The Dashboard
A SvelteKit app at sepia.svelte-apps.me (SSR + remote functions on Netlify), talking to the same database through /api/*:
🏠 Landing + pricing pages — what Sepia is, how to install it, and the hosted plan
🔍 Search all memories/entities; browse by namespace, type, importance — filters persist in the URL (back/forward works)
🕸️ Interactive knowledge-graph view (cytoscape.js + dagre)
✏️ CRUD on memories, entities, and relations from the browser
💬 Conversations — browse handoff digests by status (active/paused/done), resume or delete them
📊 Stats: counts, top entities, recent memories, decay/consolidation status
🔗 Connect an AI page: copy-paste configs for Grok, ChatGPT, Claude, Gemini, Perplexity, and the local editors
Project Structure
A Bun workspace monorepo: one repo, one lockfile, three deploy entries — Fly.io builds the server from the root Dockerfile, Netlify builds dashboard/ from netlify.toml, and the skill is installed by a script (no build).
sepia/ # Bun workspace monorepo
├── package.json # root scripts (dev, deploy:*)
├── bun.lock # ONE lockfile for the whole repo
├── Dockerfile # Fly.io entry — installs only the server's deps (--filter)
├── fly.toml # scale-to-zero config
├── netlify.toml # builds dashboard/, publishes dashboard/build
├── .env.example
├── src/ # SERVER (deployed by Fly.io)
│ ├── index.ts # Bun.serve: mounts /mcp + /api/* + /api/auth/* (CORS)
│ ├── instructions.ts # The memory contract (system-prompt injection)
│ ├── auth.ts # Bearer token (Phase 1) / OAuth guard (Phase 2)
│ ├── oauth.ts # OAuth 2.1 authorization server (@tmcp/auth)
│ ├── rate-limit.ts # per-user sliding-window rate limits
│ ├── db.ts # Drizzle client (lazy init) + MemoryError
│ ├── tools/ # 7 tools, one file each
│ ├── lib/ # CRUD + search + BFS + decay (shared by tools & API)
│ └── api.ts # /api/* router (same auth as /mcp)
├── drizzle/ # Drizzle migrations (generated by drizzle-kit)
│ ├── 0000_*.sql # baseline (introspected from the live schema)
│ └── 0001_*.sql # constraints + trigram indexes (see below)
├── packages/
│ └── shared/ # @sepia/shared — Valibot schemas + types, no build step
│ ├── src/{schemas,types}.ts # single source of truth for tools, API, and dashboard
│ └── src/db/ # Drizzle schema + owner-scoped CRUD libs (plans, users)
├── dashboard/ # DASHBOARD (deployed by Netlify)
│ └── src/routes/ # (public)/ landing + pricing, app/ search, memories,
│ # entities, conversations, graph, connect, settings
├── skills/
│ └── sepia/ # SKILL (static, installed by script)
│ ├── SKILL.md
│ ├── always-on/ # per-editor instruction files (vscode, cursor, claude…)
│ └── references/tools.md # generated from @sepia/shared schemas
├── sql/schema.sql # namespaces · entities · relations · memories · oauth_clients
└── scripts/
├── install-skill.sh # copies the skill into every editor dir it finds
└── gen-skill-ref.ts # regenerates references/tools.md from shared schemas@sepia/shared is imported as TypeScript directly (no build step) by both the server (Bun) and the dashboard (Vite) — the dashboard's forms validate against exactly what the server enforces, and the skill reference is generated from the same schemas: three consumers, one source of truth.
Getting Started
Prereqs: Bun 1.x.
bun install # one lockfile for the whole workspace
cp .env.example .env # set DATABASE_URL + MCP_BEARER_TOKEN (see below)
bun run dev # starts the server (MCP on /mcp, REST on /api/*)
bun run dev:dashboardEnvironment variables
Variable | Purpose |
| Neon Postgres pooled connection string ( |
| Phase 1 auth token for |
| OAuth 2.1 consent-page password — setting it enables the OAuth endpoints (single user; hosted accounts with plans are in progress). The dashboard itself signs in with the bearer token, not this password |
| OAuth issuer URL (defaults to |
| Dashboard build-time REST base (e.g. |
| Dashboard build-time MCP URL shown on |
Database
Single source of truth: src/db/schema.ts (Drizzle). No hand-written DDL for new changes.
bun run db:generate # schema.ts → new migration in drizzle/
bun run db:migrate # apply migrations (fresh DB: creates full schema in one go)
bun run db:push # dev: sync schema.ts diff directlyExisting live DB (has sql/schema.sql but no migration history):
bun run db:cleanup && bun run db:baseline && bun run db:migrate
# cleanup → repair data, baseline → mark 0000 applied, migrate → 0001 (constraints + pg_trgm)
db:migrate/pullneedpg(real TCP + transactions). The app itself uses@neondatabase/serverless(HTTP).sql/schema.sqlis the original baseline — keep it, new changes go through Drizzle.
Schema: namespaces → entities (cascade, UNIQUE(namespace_id, name)) → relations (UNIQUE(source, target, relation_type)) → memories (importance 0–1, archived) + memory_entity_links + oauth_clients.
Deployment
Server → Fly.io
fly apps create sepia
fly secrets set DATABASE_URL="postgresql://..." MCP_BEARER_TOKEN="$(openssl rand -hex 32)"
fly deployDockerfile runs
bun install --frozen-lockfile— SvelteKit never enters the image (all workspacepackage.jsonfiles must be copied before install; Bun validates the full workspace graph against the lockfile).fly.tomluses scale-to-zero (min_machines_running = 0): the free tier covers it, and cold starts (1–2s for a thin Bun process) are acceptable for personal use. Setmin_machines_running = 1($1–3/mo) if you want always-on.⚠️ Don't add a Fly HTTP smoke check — raw GETs confuse Streamable HTTP servers. If you want a health endpoint, expose
GET /healthzwith a TCP check.
Verify with curl -i https://sepia.fly.dev/mcp (expect 401 without a token — correct) or npx @modelcontextprotocol/inspector (Streamable HTTP, Authorization: Bearer <token>).
Dashboard → Netlify
SvelteKit app (SSR + remote functions), built from the repo root (the workspace install must happen at root), published from dashboard/build. Attach the sepia.svelte-apps.me subdomain, and add the origin to the API's CORS allowlist in src/index.ts. Remote functions run in Netlify Functions (Node runtime) and talk to Neon directly via @sepia/shared — no CORS, no exposed API keys.
Connect Clients
Local editors (Phase 1 — bearer token)
Claude Code:
claude mcp add --transport http sepia https://sepia.fly.dev/mcp \
--header "Authorization: Bearer YOUR_TOKEN"Cursor / VS Code Copilot (.cursor/mcp.json / .vscode/mcp.json):
{
"mcpServers": {
"sepia": {
"type": "http",
"url": "https://sepia.fly.dev/mcp",
"headers": { "Authorization": "Bearer YOUR_TOKEN" }
}
}
}Zed (Settings → Agent → MCP): same shape as above.
Older stdio-only clients: use Fly's shim — fly mcp proxy https://sepia.fly.dev/mcp (or npx mcp-remote --header "Authorization: Bearer ...").
Online AIs (Phase 2 — OAuth 2.1, verified mid-2026)
AI | Where | Gate |
Claude | Settings → Connectors → custom connector | Every plan (Free = 1 connector) |
Grok | grok.com/connectors → New Connector → Custom | Paid plans |
ChatGPT | Settings → Apps → Developer mode → Create | Plus+, web only |
Gemini | Settings → Connected Apps → Custom apps for Spark | Google AI Pro/Ultra (Spark) |
Perplexity | Settings → Connectors → Custom → Remote | Pro/Max/Enterprise |
Le Chat | Connectors → + Add Connector → Custom | Free/paid |
All connect from the provider's cloud, so the server must be publicly reachable (it is — Fly with force_https); Streamable HTTP is the universal transport.
✅ OAuth 2.1 is live. Paste the MCP URL into any of these connectors and you'll get a browser sign-in (password =
DASHBOARD_PASSWORD) instead of a manual credential form. Step-by-step for Grok: Connecting Sepia to Grok. Bearer-token clients (Claude Code, Cursor, Zed, Copilot) keep working unchanged.
Install the skill + always-on instructions
The skill and the always-on instruction files are served over HTTP from the same server as the MCP endpoint, so you can install them without cloning the repo:
# One-liner — fetches SKILL.md + references + always-on files from the server
# and installs into every editor dir it finds:
# skills: ~/.agents, .cursor, .claude, .codex, .opencode
# always-on: VS Code prompts folder, Cursor user rules, ~/.claude/CLAUDE.md, AGENTS.md
curl -fsSL https://sepia.fly.dev/install | bashOr via the skills.sh CLI (open agent skills ecosystem):
# From the GitHub repo (discovers skills/sepia/)
npx skills add Michael-Obele/sepia
# Or directly from the server's SKILL.md URL
npx skills add https://sepia.fly.dev/skillIf you have the repo cloned, the local installer works too:
bun run scripts/install-skill.sh # skill + always-on files, idempotentRestart your editor to pick it up. Claude Code users can also invoke the skill on demand with /sepia.
Roadmap
Milestone | Exit criteria |
M1 — Server on Fly.io, Bearer auth, 7 tools ✅ | Inspector connects; CRUD works end-to-end against Neon |
M2 — Server instructions + skill ✅ | New chat in Claude Code recalls a memory with zero reminder prompts; skill works in Zed + Cursor; always-on files installed in VS Code + Cursor |
M3 — REST API + dashboard on Netlify ✅ | Browse/search/graph/CRUD at |
M4 — OAuth 2.1 ( |
|
M5 — Online AI rollout ✅ | Grok + ChatGPT + Gemini connectors authorized; memory usable from web chats |
M6 — Hosted accounts 🚧 | Signup/sign-in, per-user namespaces, plan limits, API keys, per-user rate limits — built and smoke-tested, shipping soon |
Release gate: everything in M1–M3 works in a fresh chat with zero reminder prompts (verified via instructions + always-on files + skill), and the dashboard shows the same data the agents write.
Future enhancements
Semantic search — pgvector on Neon (paid) or a small embeddings service;
searchis already a single tool, so the engine swaps without schema changesMulti-user namespaces — per-person namespaces + shared read-only access
Memory ingestion API — browser extension or CLI to dump chat transcripts into memory
MCP resources — expose the graph as
memory://resources for subscription-capable clientsPublishing — the skill to skills.sh; the server to an MCP marketplace
Costs
Item | Cost |
Fly.io (shared-cpu 256MB VM, scale-to-zero) | $0 (free tier) |
Netlify (dashboard SPA) | $0 (~20–60 of 300 credits/mo) |
Neon Postgres free tier | $0 (0.5 GB, 100 CU-hours — fine for ~10K memories) |
Domains | $0–12/yr |
Total | $0/mo (always-on variant: ~$1–3/mo) |
License
AGPL-3.0 — GNU Affero General Public License v3.0. Copyright © 2026 Michael Obele.
Self-host free. You may run, modify, and redistribute Sepia for any purpose — personal or commercial — as long as modified versions offered as a network service publish their source under AGPL-3.0 (section 13).
Hosted service (optional, paid). The maintainers run a hosted, multi-account instance of Sepia on shared infrastructure (always-on availability, multiple machines for uptime and speed). Using that hosted service is a separate paid offering that covers the always-on infrastructure cost — the AGPL does not require hosted services to be free. Self-hosting remains free forever.
Contributions. By submitting a pull request, you agree that your contributions are licensed under AGPL-3.0-or-later, so the project can keep this license (and dual-license later if needed).
Security & Privacy
All traffic TLS (
force_https = true); secrets live infly secrets, never in the imageThe memory contract forbids storing credentials/secrets — the server is a memory, not a vault
OAuth consent screen (Phase 2) lists scopes (
memory:read,memory:write)consolidatepurges archived rows; retention rules can be added (e.g. importance < 0.2 and unaccessed 90 days → archive)
Documentation
The full design lives in plan/ — implementation specs, decision records, and research with citations:
plan/README.md— master plan: tool behavior spec, DB schema, deploy parts 1–7, milestonesplan/arch.md— architecture decisions (Bun vs Node, TMCP, Neon, Fly.io, instructions + skill, auth phases)plan/dashboard.md— web dashboard spec (REST API surface, pages, graph payload, Netlify deploy)plan/skill.md— the bundled Agent Skill (full draft + install matrix)plan/notes.md— research notes: competition analysis, client compatibility, citations (Aug 2026)
Status
M1–M5 shipped. The server, instructions, skill, dashboard, and OAuth rollout are all live and verified. Hosted accounts (signup, plans, API keys) are built and in final testing — coming soon. PRs, issues, and ideas welcome.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceA self-hosted MCP memory server that gives AI assistants persistent, semantic memory by storing facts as vector embeddings locally, supporting semantic search and swappable embedding models.48MIT
- AlicenseNot gradedqualityCmaintenanceA self-hosted, graph-aware memory server for AI assistants that provides persistent memory across sessions with hybrid search and knowledge graph capabilities.11MIT
- AlicenseNot gradedqualityBmaintenanceKnowledge-graph memory server for MCP-compatible AI tools, providing persistent, connected memory with typed relationships and auto-consolidation.77MIT
- FlicenseNot gradedqualityCmaintenanceA self-hosted MCP server that gives AI agents persistent, searchable memory with importance scoring, knowledge graphs, and autonomous memory consolidation.1
Related MCP Connectors
Persistent memory and knowledge management for AI agents with semantic search and 50+ tools.
Persistent memory and knowledge graphs for AI agents. Hybrid search, context checkpoints, and more.
Shared, governed long-term memory for AI agents across tools and sessions via MCP and REST.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/Michael-Obele/sepia'
If you have feedback or need assistance with the MCP directory API, please join our Discord server