ZaiMem
Used as the recommended package manager and runtime for installing dependencies and running the server.
Listed as an example reverse proxy for exposing the MCP endpoint publicly.
Cloudflare Tunnel listed as an example for exposing the MCP endpoint publicly.
Distributed as a Docker image published to GHCR, with support for docker run and docker compose deployment.
Mirrors all user data (memories, sessions, ledger pages, skills, stats) to a private GitHub repo as a human-readable cloud database. Supports GitHub PAT pairing, automatic private repo creation, idempotent sha-based pushes, debounced auto-sync, and scheduled daily backup snapshots.
Built as a Next.js 16 web application providing the dashboard, REST API, and MCP endpoint.
Listed as an example reverse proxy for exposing the MCP endpoint publicly.
Uses SQLite as the default local database backend for storing memories, sessions, and other user data.
Referenced as a blog channel link for the project.
Implemented in strict TypeScript.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ZaiMemRemember my project context and sync it to my private GitHub memory"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ZaiMem
๐ Launch the live demo โ โ no signup: open the app and a private
zm_โฆtoken is generated for you instantly. Paste the magic prompt into any chat.z.ai agent chat and watch the memory, context boost and token savings appear in your dashboard.
Session memory & context enhancer for AI agents like chat.z.ai โ an MCP-powered web service that gives any chat.z.ai agent persistent vector memory, automatic context enhancement, token saving, and smart-skill orchestration โ with every user's data mirrored to their own private GitHub repo as a human-readable cloud database.
โโโโโโโโโโโโโโโโ 1. visit โโโโโโโโโโโโโโโโโโโโโ 2. auto private token
โ ZaiMem web โ โโโโโโโโโโโโโบ โ token issued & โ โโโโโโโโโโโโโโโโโโโโโโโโโ
โ app โ โ stored locally โ โผ
โโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ 3. login โ get MAGIC PROMPT โ
โ (endpoint + key embedded) โ
โโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ 4. paste prompt into a new chat.z.ai AGENT-mode chat โ
โ โ session auto-synced: vector memory, context enhancement, token saving, skills โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โผ
5. pair GitHub PAT โ private repo auto-created
โ ALL data auto-synced (repo = user's cloud DB)Features
Automated private tokens โ no signup: visit the app, get a
zm_โฆtoken instantly, log in with it.Magic prompt (universal) โ the dashboard generates a ready-to-paste activation prompt with your MCP endpoint and key embedded. Works in any MCP-capable agent: chat.z.ai, Claude Code, Cursor, Cline, Windsurf, Trae, Antigravity, zcode, Koda, Pi, Grok and more.
MCP server (JSON-RPC 2.0, streamable HTTP) โ 33 tools in 8 togglable tool packs, 4 resources, prompt templates, batch calls, session ids, CORS.
Document ingestion โ
zaimem_ingest_fileMCP tool + dashboard drag-and-drop upload (PDF / DOCX / TXT / MD / CSV / code): text is auto-chunked into ~600-token overlapping pieces, every chunk is embedded as adocumentmemory tagged with its source filename, and re-ingestion is idempotent (same content hash โ no-op; changed file โ chunks replaced). Recall hits cite[doc:file.pdf ยท part i/N].Pinned memories โ
zaimem_remember {pinned: true}(or one click in the dashboard) marks a memory as always-in-force: it is injected into everyzaimem_enhance_contextblock and gets a recall ranking boost.Right to be forgotten โ
zaimem_forgetMCP tool with a two-phase preview โ confirm flow: match by id, semantic query, kind, source filename (purge a whole document) or created-before date; dashboard rows also support inline edit (re-embeds instantly) and pin toggle.Pre-created sessions & handoffs โ create a session before the work starts (title + brief + optional project), then copy its generated bootstrap prompt into a fresh agent chat in any IDE: the agent syncs onto that exact session with the brief, summary, key memories and open tasks baked in. Every existing session (even auto-created ones) has a Continue elsewhere prompt.
Projects โ agent teams โ a project is a shared workspace: team instructions, attached files/prompts and a shared memory namespace. Connect one or many agents (via the dashboard or
zaimem_project_brief {project, agent, role}); agents see the roster, the brief, the files and everything teammates shared โ and leave structuredzaimem_project_handoffnotes (done / in_progress / blocked) so the next agent picks up cleanly. Like a team of devs with one shared brain.Meeting intelligence (Tactiq-style, self-hosted) โ paste a Google Meet / Zoom / Teams transcript: ZaiMem chunks + embeds the full transcript, writes an LLM summary with every decision, extracts action items and pushes them onto the global tasks board. Ask across all meetings with
zaimem_meeting_search("what did we decide about X?").Universal tools โ the zero-config staples from the MCP ecosystem, wired into memory:
zaimem_web_search+zaimem_web_fetch(JS-rendered pages, optional auto-ingest as searchable memory),zaimem_calc(safe arithmetic, no code execution),zaimem_time(IANA timezones),zaimem_think(ledger-backed sequential-thinking scratchpad that survives compaction).HEADROOM compression mode โ togglable (dashboard switch or
zaimem_headroom {enabled}), inspired by headroomlabs-ai/headroom: when ON, every context injection is compressed harder (shorter excerpts, skill protocols withheld) to preserve context-window headroom, while originals stay full-fidelity in the store and remain retrievable viazaimem_doc_read/zaimem_recall. Lifetime tokens freed are counted per account.Local vector memory โ 384-dim hashed word/bigram/char-4gram embeddings with cosine recall; auto-dedupe (0.94 duplicate / 0.80 merge thresholds), recency + keyword boosts, context-block assembly.
Token saver โ LLM-powered digests (with extractive fallback) compress long context; token accounting per action.
Smart skills (zcode-smart-skill integration) โ 8 builtin SKILL.md skills (
smart,context-boost,token-frugal,meeting-notes,web-research,session-continuity,project-team,doc-memory), auto trigger detection, difficulty budgets (E5: 2/6/12), ledger pages (notes.md,tasks.json) with size budgets, handoff brief with TRUST clause, reflection schema. Every skill and every tool pack has an on/off switch in the Skills tab.GitHub Cloud DB (free hosting) โ pair a GitHub PAT (pre-filled token link, 2 minutes, free โ private repos cost nothing) and a private repo is auto-created in YOUR account; every session, memory, project & file is auto-synced in real time (sha-based idempotent pushes, 4s debounced, full audit trail). Your memory stops consuming ZaiMem's local storage entirely โ the dashboard keeps the recommendation visible until you pair.
Scheduled daily backup โ an optional heartbeat push of a full snapshot ~every 24h (configurable via
ZAIMEM_BACKUP_HOURS), even when nothing changed โ proof the backup pipeline is alive. Toggle it in the Cloud DB panel.One-PAT account rescue โ fresh token / new account, but your old ZaiMem already synced to a private GitHub repo? Point the Cloud DB tab's rescue card at that repo (
owner/name+ PAT): memories are re-imported through the dedupe engine (safe to run twice), sessions come back with titles and summaries, and skills you don't have yet are added. Pairing not required, nothing deleted.Global search (โK) โ one query across all sessions, memories (vector + substring), ledger pages and skills, with jump-to-result navigation. Filter results by kind (sessions / memories / ledger / skills) and by date range (24 h โ 1 year) right from the command bar.
Related MCP server: memorix
Quick start
โก Fastest path: open the hosted instance โ nothing to install. The steps below are for running your own copy.
Clone & run locally
# 1. clone the repo
git clone https://github.com/romangalaxys10-spec/zaimem.git
cd zaimem
# 2. install dependencies (bun โฅ 1.2 recommended; npm/pnpm work too)
bun install
# 3. configure env
cp .env.example .env # SQLite by default
# 4. create the database schema
bun run db:push
# 5. run
bun run dev # http://localhost:3000Production build: bun run build โ bun run start (standalone output on localhost:3000).
Deploy with Docker (public demo / self-host)
Every push to main publishes a fresh image to GHCR via Actions (docker.yml):
# pull the prebuilt image and run โ http://localhost:3000
docker pull ghcr.io/romangalaxys10-spec/zaimem:latest
docker run -d -p 3000:3000 -v zaimem-db:/app/db ghcr.io/romangalaxys10-spec/zaimem:latest
# or build & run from source in one command
docker compose up --build -d
# or without compose
docker build -t zaimem .
docker run -d -p 3000:3000 -v zaimem-db:/app/db zaimemImage tags: latest (default branch), vX.Y.Z / vX.Y (git tags), short sha-*, and the branch name. Browse all tags on the package page.
The SQLite database lives in the zaimem-db volume (mounted at /app/db) and survives rebuilds. Healthcheck, restart policy and env knobs are pre-configured in docker-compose.yml.
Env var | Default | Purpose |
|
| SQLite location (inside the volume) |
|
| Scheduled cloud-DB backup interval (hours, 0.02โ168) |
|
| Set |
Expose port 3000 through your reverse proxy / tunnel (Caddy, nginx, Cloudflare Tunnel) for a public demo โ all agent traffic goes through the single MCP endpoint /api/mcp.
First run
Then open the app โ Get my private token โ log in โ copy the Magic Prompt โ paste into a new chat.z.ai agent-mode chat.
MCP endpoint
POST /api/mcp
Authorization: Bearer <api-key> # or ?token=<api-key>
Content-Type: application/jsonThe endpoint implements the Model Context Protocol over streamable HTTP (initialize, tools/list, tools/call, resources/*, prompts/*, batch arrays, MCP-Session-Id).
Tools
Tool | Purpose |
| Register/sync the current agent session |
| Store a memory (auto-dedupe + merge, optional |
| Right-to-be-forgotten: preview โ confirm deletion by id/query/kind/source/date |
| Ingest a whole document: chunk + embed + hash-dedupe, source citations |
| Vector + keyword recall with recency/importance boosts |
| Build an injectable context block for the current message |
| Digest/compress long content and bank the savings |
| Detect a matching skill for the user's request |
| List registered skills (SKILL.md convention) |
| Fetch one skill's full SKILL.md protocol body |
| Write a smart-skill ledger page ( |
| Read a ledger page |
| Distill an end-of-session summary |
| Cross-session handoff brief (TRUST clause) |
| Batch-store up to 25 memories in one round-trip |
| Progressive document loading: outline or one chunk |
| History pressure + activity report for a session |
| "What's new since you left" cross-session digest |
| Resume a session: summary + open tasks + checkpoint diff |
| Global task board: pick the next open task |
| Pre-create a custom session + get its bootstrap prompt |
| Bootstrap prompt to continue any session elsewhere |
| Join a project team + get the full brief (roster, files, shared memory) |
| Structured end-of-shift handoff to teammates |
| Meeting transcript โ chunks + summary + action items |
| List ingested meetings with summaries |
| "Ask my meetings" โ semantic search across transcripts |
| Web search (zero API keys) |
| Read a web page as text (+ optional auto-ingest) |
| Safe arithmetic calculator (shunting-yard, no code execution) |
| Current time / timezone info |
| Sequential-thinking scratchpad (ledger-backed chain) |
| HEADROOM compression mode: status + toggle |
Resources & prompts
zaimem://protocolโ operating protocol (markdown)zaimem://memoryโ recent memories (json)zaimem://skillsโ skill registry (json)Prompt template
zaimem-bootโ boot a ZaiMem-synced session with memory recall
Dashboard REST API
Route | Description |
| Issue an automated private token |
| Log in with the token (cookie session) |
| Current user + stats + config |
| Session management |
| Memory browsing & vector search |
| Skills + MCP tool packs + headroom mode: GET returns all three, PATCH toggles ( |
| Usage statistics (tokens saved, actions) |
| Cloud DB status, pair / unpair / sync / toggle / restore / import (new-account rescue) |
GitHub Cloud DB
Pair your PAT (classic, repo scope) from the dashboard's Cloud DB tab:
PAT is validated against
GET /userand stored AES-256-GCM encrypted (only last 4 chars are ever displayed).A private repo (default
zaimem-cloud-db) is auto-created in your account.Every mutation (memories, sessions, ledger, skills, stats) triggers a debounced auto-sync.
Pushes are idempotent โ unchanged files are skipped via git blob SHA comparison; your own files in the repo are never touched.
New account?
POST /api/github {action: "import", pat, repo, branch?}โ or the Cloud DB tab's rescue card โ re-syncs an existing cloud-DB repo into the current account (memories + sessions + skills, additive & dedupe-aware).
Repo layout:
index.json # snapshot index + counts
README.md # human-readable overview (regenerated)
memories.json # all memories
vectors.jsonl # 384-dim embeddings, one JSON per line
sessions/<slug>.json # one file per session (turns, summary, metadata)
skills.json # skill registry
stats.json # usage statistics
sync/log.json # sync audit trailTool packs โ capability switches
All 33 MCP tools belong to one of 8 packs. Flip a pack off in Skills โ MCP tool packs and its tools vanish from tools/list; a direct call answers with a clear "pack is switched OFF" hint instead of executing. core-memory is locked โ it is the product:
Pack | Tools |
Core memory (locked) | sync_session ยท remember ยท remember_many ยท recall ยท forget ยท enhance_context |
Session continuity | session_status ยท brief_me ยท resume ยท task_next ยท session_summary ยท handoff_brief ยท session_create ยท session_prompt |
Project agent teams | project_brief ยท project_handoff |
Meeting intelligence | ingest_meeting ยท meetings_list ยท meeting_search |
Document ingestion | ingest_file ยท doc_read |
Web & utilities | web_search ยท web_fetch ยท calc ยท time ยท think |
Smart skills & ledger | detect_skill ยท list_skills ยท get_skill ยท ledger_write ยท ledger_read |
Token saver & headroom | save_tokens ยท headroom |
Deployment
Vercel (auto-sync): every push to main deploys automatically via the Vercel deploy GitHub Actions workflow (prebuilt deploy with VERCEL_TOKEN / VERCEL_ORG_ID / VERCEL_PROJECT_ID repo secrets). Project: zaimem.vercel.app.
Database on Vercel: Vercel lambdas each own an ephemeral /tmp, so the default DATABASE_URL=file:/tmp/zaimem.db gives every instance a separate blank database (boot warning + schema bootstrap are built in: vercel-build.sh ships db/vercel-bootstrap.db and db.ts copies it into place on cold start). For real persistence and cross-instance consistency, set DATABASE_URL to any managed Postgres (Vercel Postgres / Neon / Supabase) โ the build detects a postgres:// URL and automatically generates + syncs prisma/schema.postgres.prisma. Zero raw SQL is used anywhere, so the provider switch is loss-free. scripts/vercel-smoke.sh sweeps production: HEALTH checks must always pass; state flows are db-dep until Postgres is connected.
Self-hosted / Docker: the long-running mode stays the reference deployment โ docker-compose.yml (or bun run build && bun start) keeps SQLite on a real disk, runs the GitHub backup scheduler in-process, and shares one rate-limit/session registry. GHCR semver images are built by CI on every release tag.
Architecture
src/
โโโ app/api/โฆ # 12 route handlers (auth, sessions, memories, skills, stats, github, mcp)
โโโ components/zaimem/โฆ # landing, dashboard, cloud-db, panels, magic-prompt builder
โโโ lib/zaimem/
โโโ vector.ts # hashed n-gram embeddings + cosine engine
โโโ memory.ts # remember/recall/enhance with dedupe & boosts
โโโ compress.ts # token saver (LLM digest + extractive fallback)
โโโ skills.ts # smart-skill engine (triggers, budgets, ledger, handoff)
โโโ mcp.ts # MCP tool/resource/prompt registry & dispatcher
โโโ github.ts # PAT validation, repo provisioning, idempotent sync, debounce queue
โโโ crypto.ts # AES-256-GCM at-rest encryption, blob sha
โโโ auth.ts # token issuance & session auth
โโโ seed.ts # builtin skills bootstrapStack: Next.js 16 (App Router) ยท React 19 ยท TypeScript ยท Tailwind CSS 4 ยท shadcn/ui ยท Prisma + SQLite ยท z-ai-web-dev-sdk.
Testing
npx tsx scripts/e2e-test.ts # 156 checks โ auth, MCP tools & packs, sessions, projects, meetings, export, dedupe, GitHub guards โฆ
npx tsx scripts/e2e-github-unit.ts # 26 engine checks โ full sync engine vs mock GitHub APIThe 9 live rescue-import checks skip hermetically unless
ZAIMEM_E2E_PAT/ZAIMEM_E2E_REPOare set โ no credentials are ever hard-coded.
Security notes
Login tokens are stored as SHA-256 hashes โ plaintext tokens are never persisted (v1.7.2 migration + backfill).
GitHub PATs are encrypted at rest (AES-256-GCM, scrypt-derived key from
ZAIMEM_SECRET) and never returned by the API; legacy rows are transparently re-encrypted on load.Rate limiting: auth init 60 req / 5 min per IP; MCP endpoint 1200 req / min per token.
JSON-RPC hardening: batches capped at 25 requests; MCP sessions evicted after 24 h idle (10 k cap); verbose error scrubbing.
Security headers (v1.8): Content-Security-Policy (self-locked), HSTS, X-Content-Type-Options, Referrer-Policy, Permissions-Policy and COOP on every response. Frame directives are deliberately omitted (v1.8.1) so the dashboard embeds in chat sidebars, IDE panels and preview gateways โ safe because ZaiMem holds no cookies; auth is PAT via explicit
Authorizationheaders that browsers never attach cross-site.Strict type-checked builds (v1.8):
next buildfails on TypeScript errors โ enforced in CI.Data portability (v1.8):
GET /api/exportdownloads the full account (memories, sessions, ledger, skills, tool packs, project teams) as one JSON archive โ no vectors, no credentials.The SQLite database and
.envare local-only and excluded from version control. A full security audit (v1.7.2) is reproduced byscripts/sec-report-*.py.
Community & links
โก Built with | |
๐ Lead by | |
โ๏ธ Telegram Blog | |
๐ข The Claw Blog | |
๐ Author of |
License
MIT ยฉ RyzenCode
This server cannot be deployed
Maintenance
Related MCP Connectors
Shared long-term memory vault for AI agents with 20 MCP tools.
Private, portable memory and reusable skills for AI agents.
Private persistent memory for Claude, ChatGPT & Gemini via MCP - semantic search, zero-code setup.
Git-backed platform for skills, tools, and context for AI agents
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables deployment of autonomous AI agents with memory and tool execution capabilities through a WebSocket-based MCP protocol. Provides production-ready infrastructure with REST API access, persistent state management, and extensible function registry for building self-hosted AI systems.-
- AlicenseAqualityAmaintenanceCross-agent memory bridge for AI coding assistants. Persistent knowledge graph shared across 10 IDEs (Cursor, Windsurf, Claude Code, Codex, Copilot, Kiro, Antigravity, OpenCode, Trae, Gemini CLI) via MCP. 22 tools including team collaboration, auto-cleanup, mini-skills, session management, and workspace sync. 100% local, zero API keys required.91,211 npm817Apache 2.0
- AlicenseNot gradedqualityNot gradedmaintenancePortable MCP memory server giving AI agents persistent, verified, cross-session memory. 30 tools, SQLite + cloud sync, Chrome Extension for every AI chat platform. The only JavaScript MCP memory server. Includes behavioral learning engine, semantic search, knowledge scoping, session quality scoring, and web dashboard.-
- AlicenseCqualityAmaintenanceEnables AI agents to maintain persistent, searchable two-layer memory with 37 tools, hybrid search, knowledge graphs, and enterprise features like authentication and backups.636 PyPIMIT