KOPENG
Enables optional MinIO integration for S3-compatible artifact storage, allowing memory-related files to be stored externally.
Allows optional integration with Neo4j for graph and entity traversal, enabling relationship-aware queries over memories.
Integrates with a local Ollama instance to run an optional LLM reasoner for pair classification during memory consolidation, all locally without data egress.
Offers PostgreSQL with pgvector as an alternative storage backend, enabling scalable, vector-search-capable memory persistence.
Supports optional Redis usage for ephemeral context storage, providing fast temporary data caching.
Provides SQLite as a storage backend for memories, embeddings, and audit logs, with support for in-memory vector indexing.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@KOPENGwhat do we know about the database schema?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
KOPENG
Memory that curates itself instead of growing into landfill.
KOPENG is persistent, self-curating, local-first memory for coding agents (Claude Code, Codex CLI). It learns from how you actually work, recalls the right context on every prompt, and — unlike append-only RAG — cleans up after itself: duplicates collapse, stale facts decay, contradictions get routed to you for review. Every mutation is snapshot-first, audited, and reversible. All inference is 100% local: no per-query API cost, and your codebase context never leaves the box.
Install
One prerequisite: Node.js 20+. That's it — no Docker, no database server, no cloud account, no API keys, no admin rights.
npx kopeng@latest initOn Windows, also install the Visual C++ Redistributable first — the local embedding runtime is a native module that will not load without it (KOPENG degrades to keyword-only search rather than failing, but you want the semantic half).
The installer is dry-run-first: it shows you exactly what it will do, then waits for your yes. On your next Claude Code prompt, recall just fires.
What it puts on your machine (and what npx kopeng uninstall removes):
Item | Where |
KOPENG server + CLI (prebuilt — nothing compiles) |
|
Your memory data: SQLite database(s), logs, and |
|
Embedding model (all-MiniLM-L6-v2, ~30 MB) and Reranker model (ms-marco-MiniLM-L-6-v2, downloaded on first search) |
|
A user-level autostart entry (Scheduled Task / systemd --user unit / LaunchAgent) | your user session, recorded in |
MCP registration + 5 fail-open Claude Code hooks, merged minimally into your existing config (backed up first) |
|
Learning-profile flags (observation ingestion, discovery detection, dreaming — all off until you opt in) |
|
npx kopeng uninstall reverses all of it and keeps your memory data unless you pass --purge. Nothing phones home — there is no telemetry to opt out of.
Prefer to run from source, wire configs by hand, or install as a real system service? The clone-and-npm run wire path lives in SETUP.md, along with Codex CLI wiring and troubleshooting.
Related MCP server: exocortex
Watch it think
npx kopeng vizKOPENG ships a six-tab dashboard that makes the whole system observable while you work: a live feed of tool-use observations streaming in real time, a memory graph you can traverse, ops panels (confidence distribution, decay, dream history, corpus health), the dream review queue where you accept or reject the Librarian's proposed consolidations, session replay, and slots. This is the difference between "trust me, it's learning" and watching it learn. Runs at http://localhost:8780 in the foreground — Ctrl-C stops it.
Why KOPENG
Most "agent memory" is append-only store-and-retrieve. It remembers; it never prunes. The corpus drifts — duplicates pile up, stale facts outrank current ones, contradictory memories ("we use X" / "we switched to Y") both keep surfacing — and the operator becomes the garbage collector. KOPENG adds the missing half: curation.
It curates, not just recalls — the dreaming Librarian. An operator-gated nightly consolidation pass collapses duplicates, decays stale memories, and routes contradictions and supersessions for your review. The engine is deterministic-first: every mutation is deterministic code; the optional local LLM only classifies pairs and is structurally locked out of the write path. Supersession is a temporal chain, not a deletion — the history stays.
It learns passively from real tool use, at zero LLM cost. Six heuristic detectors turn observed behavior — repeated commands, error-then-fix patterns, hot files, cross-session sequences — into confidence-scored memories. You never have to remember to save anything.
Every write that changes a memory is reversible. Snapshot to revisions → mutate → append-only audit, with a compensation path if the audit append fails. Roll back any memory to any revision via one endpoint. The revision is a whole-row snapshot — content, embedding, tags, confidence, scope, type, the decay clock, and the is_locked Hard Anchor — so a rollback reverts protection state along with everything else. Scope of the claim, honestly: what is reversible is memory state. Automated (dream/promotion/maintenance) mutations additionally append a dream_audit_log row; operator edits and crystallization carry the revision alone as their reversibility record. Deletes are archives, never row removal — except the two admin revision-purge routes, which are the deliberate redaction escape hatch and do destroy history.
Fully local. Embeddings and reranking run as quantized ONNX in-process; the optional reasoner is Ollama on your own GPU. No cloud model sits in the retrieval path or the consolidation path.
How it works
Layer | What it does |
Retrieval | Hybrid search — RRF fusion of semantic + keyword (FTS5), optional cross-encoder rerank, confidence-blended ranking. A fast hook-optimized recall path skips reranking. |
Dreaming / consolidation | Deterministic-first nightly pass: duplicate collapse, durability-scaled decay, contradiction routing, supersession chains. Snapshot-first, audited, reversible. Off by default. |
Auto-discovery | Tool-use observations → 6 zero-cost detectors → confidence scoring → semantic dedup → memories. |
Static surfacing | Per-prompt injection of relevant tools, skills, and project conventions from your |
Observability | Live SSE event stream + the six-tab dashboard above. |
Optional reasoner | Local Ollama pair classifier — classify/extract only, never writes. Absent or down, behavior degrades byte-for-byte to deterministic-only. |
Storage | SQLite (supported preview backend) + in-process ONNX models. Optional, feature-flagged: Neo4j, Redis, MinIO. |
Interfaces | 19 MCP tools (thin stdio client) + Fastify REST API on |
Two entry points, one backend: src/server.ts is the real server (owns the database, index, and services); src/index.ts is the MCP stdio client that proxies tool calls to it. The REST contract never changes across backend swaps.
Memory types: user, feedback, project, reference, discovery. Scopes: global, project:<name>, client:<name>, with a PRIMARY_SCOPE routing rule so nothing silently lands in global.
Three commands answer the three operational questions, any time: wire (is it connected?), doctor (is the whole install correct?), canary (does recall actually work, end to end?).
Preview status
KOPENG is a developer preview: a self-hosted memory system for a single expert developer, on one machine, bound to loopback (127.0.0.1:3200).
Every autonomous layer ships OFF and is labeled advanced. Passive learning, dreaming, and the reasoner are opt-in flags; the installer (
init) andwireboth offer minimal / recommended / everything profiles at install time.Auto-apply is hard-restricted in code to exactly two change classes (exact duplicates and decay) — both also OFF by default. Everything else queues for operator review.
Reliability comes from engineering rigor, not scale: a zero-LLM pinned-clock replay regression net, adversarial reviews run against a copy of real data, idempotent consolidation passes, and fail-open behavior everywhere a hook or service could stall. ~1,800 tests run against in-memory SQLite with no server needed (
npm test).Security posture, in one paragraph: operator mutations require a generated admin key; observation ingestion uses its own separate optional key; reads are public by design on loopback. Remote deployment is unsupported — a non-loopback bind refuses to boot without both keys, and even then requires an outer boundary (VPN or authenticating reverse proxy). Full threat model: SECURITY.md.
Renamed 2026-07 from its previous codename. Everything uses
kopeng— hook env vars areKOPENG_*, client data lives in~/.kopeng/.
Evals
Retrieval quality is measurable, not vibes: npm run eval reports P@K, R@K, MRR, and NDCG@K, with baseline-vs-reranked comparisons, and the consolidation layer has its own zero-LLM regression and effectiveness harnesses (npm run dream:replay, npm run dream:effectiveness). One deliberate exception to local-only, clearly flagged: npm run eval:seed sends selected memory content to the Anthropic API to draft eval queries — the only shipped command that egresses corpus data. The runtime itself never does.
License
KOPENG is source-available under the Business Source License 1.1: read, run, modify, and self-host it freely — including production use — but you may not offer it to third parties as a commercial hosted memory service. On 2030-07-05 the license automatically converts to Apache 2.0.
Contact
Built by djy89 — a single-maintainer project.
More of my work: djy89.net
Bugs, questions, design discussion: open a GitHub issue — public discussion preferred; it helps the next person.
Anything else:
hello@kopeng.netSecurity vulnerabilities: use GitHub's private vulnerability reporting per SECURITY.md — not issues, not email.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- AlicenseAqualityAmaintenanceMCP-native, local-first memory for coding agents that turns real sessions into reusable decisions, gotchas, and domain knowledge.176MIT
- AlicenseNot gradedqualityBmaintenancePersonal unified memory system for AI coding agents, providing persistent memory with hybrid RAG retrieval via MCP integration, allowing agents to store, search, and manage memories locally.1MIT
- AlicenseNot gradedqualityAmaintenanceA persistent, local memory layer for AI coding agents that remembers decisions, bugs, and rules across sessions with three core MCP verbs (recall, remember, search).Apache 2.0
- AlicenseNot gradedqualityBmaintenancePersistent memory for AI coding agents that stores and recalls preferences, decisions, and conventions via semantic similarity, with zero cloud dependencies and plug-and-play MCP integration for Claude Code.Apache 2.0
Related MCP Connectors
Persistent memory and knowledge graphs for AI agents. Hybrid search, context checkpoints, and more.
Persistent memory and cross-session learning for AI coding assistants (hosted remote MCP).
Persistent memory and knowledge management for AI agents with semantic search and 50+ tools.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/djy89/kopeng'
If you have feedback or need assistance with the MCP directory API, please join our Discord server