Skip to main content
Glama
jagoff

MEMO MCP Server

by jagoff

memo — local memory for AI

memo

Your coding agent starts every session with amnesia. memo fixes that — 100% on your own machine.

Persistent, searchable memory for Claude Code, Codex, Cursor, Cline, Devin, and OpenCode. No cloud, no API keys, no Ollama, no vector DB to run. And it spends fewer tokens, not more.

PyPI Downloads License: MIT MCP

Save a fact once — every later session recalls it automatically, all stored locally.


Install

curl -fsSL https://raw.githubusercontent.com/jagoff/memo/v4.5.0/install.sh | bash

Prefer a package manager? uv tool install mlx-memo · pipx install mlx-memo · brew tap jagoff/memo && brew install mlx-memo

Then:

memo doctor                                   # self-check
memo save 'we use Postgres, not Mongo'        # save a decision
memo search 'what database did we pick?'      # search by meaning

That's it. Your agents pick it up over MCP automatically — the installer wires every client it finds.

New Mac:

curl -fsSL https://raw.githubusercontent.com/jagoff/memo/v4.5.0/install.sh | bash
memo sync bootstrap git@github.com:yourname/memo-sync.git

Agent-managed setup:

curl -fsSL https://raw.githubusercontent.com/jagoff/memo/v4.5.0/install.sh | bash
memo doctor --strict-runtime

On Linux or just want to look around first?

docker run --rm ghcr.io/jagoff/memo:latest memo doctor

Related MCP server: neuromcp

Why this saves you money

Most memory servers add context. memo is built to remove it.

Profile

Tools

Schema tokens

agent (default)

38

~3.8k

core / slim

55

~5.0k

full / default

159

~18k

The default MCP surface is 38 tools, not 159: about 79% fewer tool schemas. It exposes 38 tools / ~3.8k schema tokens versus 159 tools / ~18k tokens on the full surface — overhead paid every session, in every client.

Ambient recall injects one relevant memory before the model answers. The bundled Claude Code hook caps that injection at ~160 tokens. memo roi reads the real grounding and re-ask ledgers, then estimates accumulated savings with disclosed defaults (350 tokens per grounded recall and 900 per avoided re-ask).

memo roi       # value from grounded recalls and avoided re-asks
memo tokens    # usage-savings ledger

Three things nothing else does

🕰️ Time-machine — query your knowledge as it was

memo as-of ask "what was the deploy strategy?" --date 2026-02-01
memo diff --from 2026-01-01 --to 2026-03-01

Full historical reconstruction by reverse-replaying history.db. Useful when you need to know why past-you made a call, not just what past-you decided.

⚡ Contradiction radar — memory that notices when you change your mind

memo contradict scan      # find conflicting facts corpus-wide
memo contradict triage    # resolve: fuse / newer-wins / dismiss

Change a decision and memo flags the now-stale version, so the agent stops reintroducing what you already threw out.

🔮 Dream — it optimizes itself while you sleep

memo dream run

A 7-phase nightly pipeline: inventory → mine signals → resolve conflicts → prune stale → synthesize cross-cluster insights → optimize → pre-warm the top-100 query embeddings so tomorrow's recall stays under 200 ms. Every run writes a receipt you can audit. Zero intervention.


How it works

Hybrid retrieval. A vector leg (MLX on Apple Silicon, sentence-transformers on CPU) and a BM25 leg (FTS5, diacritic-folding for Spanish) run in parallel, fuse via Reciprocal Rank Fusion, then go through an optional MLX cross-encoder rerank.

vector + keyword search in parallel, fused, reranked, top memory injected

Markdown is the source of truth. Every memory is a plain .md file you can read, grep, and version-control. SQLite is a derived index that rebuilds from the files at any time — hand-edit in Obsidian and your edit wins on the next memo reindex. Nothing is locked in a database you can't open.

Prompts and memories stay on your machine. Embedder, reranker, and LLM all run in-process. No telemetry. Memory travels only if you point memo sync at a git remote you own. Normal startup is fully offline; remote update checks and auto-update require an explicit opt-in. → Privacy and network policy

Also in the box: cross-agent memo resume (reopen any session from any agent), cross-Mac git sync, a knowledge graph with optional codegraph symbol edges, encrypted secret storage, OCR/audio ingestion, evidence packs, outcome learning, and signed federation. → Full feature reference


How it compares

Verified July 2026 against each project's own docs. Corrections welcome — open an issue and I'll fix the table.

memo

mem0

letta

cognee

basic-memory

cipher

100% local, no cloud API

⚠️

⚠️

⚠️

⚠️

Time-machine (rewind to any date)

⚠️

⚠️

⚠️

Contradiction detection + resolution

⚠️

⚠️

Autonomous nightly maintenance

Token-economy MCP profiles

⚠️

Markdown / Obsidian as source of truth

⚠️

✅ first-class · ⚠️ partial, config-gated, or add-on · ❌ absent

Closest comparators are basic-memory (local-first + Obsidian + MCP — same thesis) and cipher (memory for coding agents).


Requirements

Support

macOS, Apple Silicon (M1–M4)

Full — MLX embedder + reranker + ask/synthesize/dream

Linux / Ubuntu

Standalone CPU backend — search, recall, save. pipx install "mlx-memo[cpu]" · docs/ubuntu.md

Intel Mac

Unsupported — current PyTorch releases do not ship Python 3.13 wheels for this platform

Docker

Cross-platform, CPU backend · docs/docker.md

Python ≥ 3.13 (the installer handles this via uv if you don't have it). First install pulls ~8 GB of models, 5–15 min. Optional: an Obsidian vault — without one, memo uses ~/Documents/memo/.


Docs

Install detail, installer knobs, new-Mac migration

reference.md › Install

Per-client MCP setup (Claude Desktop, Cursor, Cline, Continue)

reference.md › MCP setup

Ambient recall, capture, and tuning

reference.md › Ambient memory

Full CLI reference (137 commands) + memo tui

reference.md › CLI

All MEMO_* flags and model profiles

reference.md › Configuration

Architecture and design notes

reference.md › Design

Privacy and network policy

PRIVACY.md

All 138 top-level CLI commands

Core: save search ask get edit rename delete list

Recall & Hooks: recall recall-hook context briefing continuity prewarm capture-tick capture-stop interject ask-gaps guard digest

Session & History: history as-of diff record-history session resume reflect mine-history episodes chronicle

Maintenance: reindex maintain review dream consolidate synthesize dedupe retier contradict invalidate temporal compress-context ops

Analysis & Quality: health stats doctor journey-check lint drift analytics eval roi tokens token-savings usefulness gaps outcome profile confidence graduation hype definitive evidence

Knowledge Graph: graph entities entity extract-entities links version related

Advanced Search: embed rerank contextual retrieve context-pack chat chat-ask repo

Import / Export / Sync: import export backup restore sync ingest federation

Visualization: tui dashboard map logs hook-log

Setup & Config: init setup config install-mcp install-watcher uninstall-watcher install-slash install-statusline install-recall-hook install-shell-wrapper install-shims startup-banner migrate migrate-vault migrate-independence update upgrade self-update watch release onboard

Daemons: recall-daemon ingest-daemon maint-daemon embed-daemon idle-daemon

Other: backend-native collaborative feedback query mandate drift sleep-cycle operational ocr-image provenance secret verbatim mcp-command codex-badge debug-recall http-api mine-git token-gate fix undo code-facts code-nudge code-health


Contributing

git clone https://github.com/jagoff/memo && cd memo
uv pip install -e '.[dev]'

Issues and PRs welcome — see CONTRIBUTING.md. If memo is useful to you, a ⭐ genuinely helps other people find it.

MIT licensed. Built on Apple MLX, sqlite-vec, and codegraph.

Install Server
A
license - permissive license
A
quality
A
maintenance

Maintenance

Maintainers
Response time
1dRelease cycle
46Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • A
    license
    B
    quality
    B
    maintenance
    Basic Memory is a knowledge management system that allows you to build a persistent semantic graph from conversations with AI assistants. All knowledge is stored in standard Markdown files on your computer, giving you full control and ownership of your data. Integrates directly with Obsidan.md
    Last updated
    17
    3,539
    AGPL 3.0
  • A
    license
    -
    quality
    A
    maintenance
    Semantic memory for AI agents — local-first MCP server with hybrid search, knowledge graph, contradiction detection, and plan-then-commit consolidation.
    Last updated
    794
    4
    AGPL 3.0
  • A
    license
    -
    quality
    D
    maintenance
    Self-hosted semantic memory for AI agents. Save worklogs, decisions, and notes via MCP, then recall them across sessions by meaning rather than keyword. Backed by Postgres + pgvector with local embeddings (multilingual-e5-base).
    Last updated
    1
    MIT

View all related MCP servers

Related MCP Connectors

  • User-owned memory for AI agents, Copilot, Claude, IDEs, CLIs, and chat apps over remote MCP.

  • Persistent memory and knowledge graphs for AI agents. Hybrid search, context checkpoints, and more.

  • Local-first RAG engine with MCP server for AI agent integration.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/jagoff/memo'

If you have feedback or need assistance with the MCP directory API, please join our Discord server