engram
Supports ingesting notes and content from Obsidian vaults into the semantic memory store.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@engramremember that my dog's name is Max"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
What is Engram?
Engram is a personal knowledge base that stores your notes, conversations, documents, and web pages as vector embeddings — then lets any MCP-compatible AI assistant (Claude, Cursor, Windsurf, etc.) search and recall them by meaning, not just keywords. Everything lives in a PostgreSQL database you control, runs locally or in the cloud, and costs near-zero to self-host.
Related MCP server: mcp-recall
Features
Hybrid Search — Combines vector similarity (pgvector HNSW) with PostgreSQL full-text BM25, fused via Reciprocal Rank Fusion
MCP Server — Any MCP-compatible client can store, search, and manage memories over HTTP
REST API — Full CRUD + search, with OpenAPI docs at
/api/docs/Multi-format Ingestion — Ingest PDFs, DOCX, TXT, Markdown, URLs, and Obsidian vaults
Auto-Enrichment — Optional LLM-powered tagging, entity extraction, and memory decay
React Dashboard — Browse, search, and visualize your memory graph
Privacy-First — Runs 100% locally with Ollama; no data leaves your machine
Pluggable Embeddings — Ollama (default, free), OpenRouter, or bring your own provider
Architecture
┌──────────────────────────────┐
│ AI Clients │
│ (Claude, Cursor, Windsurf) │
└─────────────┬────────────────┘
│ MCP / REST
▼
┌──────────────────────────────────────┐
│ MCP Server (FastMCP :8080) │
│ REST API (Django DRF :8000) │
│ Dashboard (React + Vite :5173) │
└─────────────┬────────────────────────┘
│
┌──────────┴──────────┐
│ Application Layer │
│ Embedder ─→ Ollama │
│ Auto-Tagger │
│ Entity Extractor │
│ Memory Decay │
└──────────┬──────────┘
│
┌──────────▼──────────┐
│ PostgreSQL 16 │
│ + pgvector (768d) │
│ + Full-text search │
│ + HNSW index │
└──────────────────────┘Quick Start
Prerequisites: Python 3.12+, PostgreSQL 16 with pgvector, Ollama (or Docker)
The supported install path is source + Docker:
# 1. Clone the repository
git clone https://github.com/jblacketter/engram.git && cd engram
# 2. Start database + Ollama via Docker
docker compose up -d
# 3. Pull the embedding model into the Ollama container
docker compose exec ollama ollama pull nomic-embed-text
# (Alternative: if Ollama already runs natively on :11434, start only the
# database and pull on the host instead:
# docker compose up -d db
# ollama pull nomic-embed-text)
# 4. Install Python dependencies
pip install -e ".[dev]"
# 5. Copy environment config (set DJANGO_SECRET_KEY)
cp .env.example .env
# 6. Run migrations and start the server
python manage.py migrate
python manage.py runserverNote on PyPI: the
engram-semanticwheel ships the Python packages only — it does not includemanage.py, the React dashboard, or a console entry point, so it cannot run the full stack on its own. Use it when you need engram's modules as a library; install from source for everything else.
The API is now live at http://localhost:8000/api/ and docs at http://localhost:8000/api/docs/.
Start the MCP server (separate terminal):
python -m mcp_serverStart the frontend (separate terminal):
cd frontend && npm install && npm run devDashboard at http://localhost:5173.
Connecting AI Clients
Claude Desktop
Add to ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"engram": {
"url": "http://localhost:8080/mcp"
}
}
}Claude Code
claude mcp add --transport http engram http://localhost:8080/mcpCursor / Windsurf
Add to your MCP settings:
{
"mcpServers": {
"engram": {
"url": "http://localhost:8080/mcp"
}
}
}See docs/connecting-clients.md for authenticated setups and advanced config.
API Reference
REST Endpoints
Method | Endpoint | Description |
|
| Health check |
|
| List memories (paginated) |
|
| Create a memory |
|
| Get memory by UUID |
|
| Update memory |
|
| Delete memory |
|
| Hybrid semantic + keyword search |
|
| Memory statistics |
|
| List all tags |
|
| Ingest a file (PDF, DOCX, TXT, etc.) |
|
| Scrape and ingest a URL |
|
| Batch ingest files and URLs |
Auth: Authorization: Bearer <REST_API_KEY> (disabled when env var is empty).
MCP Tools
Tool | Description |
| Store a new memory with optional tags and importance |
| Retrieve a memory by UUID |
| Update content, tags, or importance |
| Delete a memory |
| Hybrid search with configurable semantic/keyword weight, filterable by |
| Find semantically similar memories, filterable by |
| List recent memories, optionally filtered by |
| Fetch and ingest a URL |
| Ingest a base64-encoded file |
| System statistics |
Scoping memories across domains
A single Engram instance can host memories from multiple domains (e.g. QA-tool output and personal notes) without cross-contamination, using tag-based soft scoping. The convention is documented here so any client can follow it.
The convention
Every write should set two things:
A scoping tag in the
tagsfield —domain:<name>, optionally refined withproject:<slug>. Examples:tags=["domain:qa", "project:my-app"]tags=["domain:personal"]
The
sourcefield — already part of the schema — to identify the writer (e.g."aegis","qaagent","mcp","manual").sourceis a distinct top-level field, not a tag; the two are complementary.
Reading with scope
The MCP tools and REST search honor these on read. Examples:
# MCP
await search_brain("flaky test", tags=["domain:qa"])
await find_related(memory_id, tags=["domain:qa"])
await list_recent_memories(tags=["domain:qa"])
# REST
POST /api/search/ {"query": "flaky test", "tags": ["domain:qa"]}Scoped read surfaces (filterable)
The following read surfaces accept tags and/or source filters, so callers
that pass them can stay inside one domain:
Surface |
|
|
MCP | yes | yes |
MCP | yes | yes |
MCP | yes | yes |
REST | yes | yes |
Per-agent keys (enforced scoping)
Each tool/agent can hold its own key bound to the domains it may touch:
python manage.py agent_keys create claude-code --default-domain engram --allow tagteam
python manage.py agent_keys list
python manage.py agent_keys revoke claude-codeThe plaintext key (egk_…) is shown exactly once; only its SHA-256 hash is
stored. Use it as a bearer token on REST (Authorization: Bearer egk_…)
and MCP. For agent principals the scoping is enforced server-side,
identically on both surfaces:
Writes with no
domain:tag get the key's default domain injected; adomain:tag outside the key's allowed list is rejected (403 / tool error) — never silently rewritten.Scoped reads default to the key's default domain; other allowed domains must be named explicitly; disallowed domains are rejected.
Formerly unscoped surfaces are closed for agents:
GET /api/memories/,/api/memories/<id>/,/stats/,/tags/, MCPget_memory/get_stats/list_domainsare filtered to the key's allowed domains; out-of-scope ids return 404 (no existence oracle); memories with nodomain:tag are invisible to agent keys.
Note: the MCP server checks for agent keys at startup — restart it after creating the first key.
Owner surface (intentionally unscoped)
The global REST_API_KEY/MCP_API_KEY and the dashboard session remain
the single-user owner surface with full visibility across all domains
— including untagged memories. With no agent keys configured, behavior is
exactly the pre-scoping single-user system.
Limits of soft scoping
The project model is still tag-based: no
workspace/tenantcolumn exists (a decision checkpoint tracks whether to harden it — seedocs/decision_log.md). Hard multi-tenant isolation would mean separate instances or a schema column.Owner-surface writes are still discipline-based: the owner can create untagged memories (visible only to the owner surface).
Project Structure
engram/
├── api/ # REST API (DRF views, serializers, auth)
├── core/ # Data models, services (memory CRUD, search)
├── embeddings/ # Embedding providers (Ollama, OpenRouter)
├── intelligence/ # Auto-tagger, entity extraction, decay, reports
├── ingestion/ # File, URL, batch, and Obsidian importers
├── mcp_server/ # FastMCP server with tool definitions
├── engram/ # Django project settings and URL config
├── frontend/ # React + TypeScript + Tailwind SPA
├── docker/ # Entrypoint scripts, pgvector init SQL
├── nginx/ # Reverse proxy config (production)
├── docs/ # Extended documentation
├── tests/ # Test suite
├── Dockerfile # Django production image
├── Dockerfile.mcp # MCP server image
├── docker-compose.yml # Dev: PostgreSQL + Ollama
└── pyproject.toml # Python package configConfiguration
Key environment variables (see .env.example for the full list):
Variable | Default | Description |
| — | Django secret key |
|
| Settings module |
|
| Database name |
|
| Database host |
|
| Ollama API URL |
|
| Embedding model |
|
| Chat model for enrichment |
| — | Fallback embedding provider |
| — | MCP auth token (empty = open) |
| — | REST auth token (empty = open) |
|
| Auto-tag and extract entities |
Production Deployment
docker compose -f docker-compose.prod.yml up -dThis starts PostgreSQL, Ollama (with GPU support), Django (gunicorn), the MCP server, and nginx with TLS. See docs/setup-docker.md for full production setup and docs/setup-supabase.md for cloud-hosted PostgreSQL.
Tech Stack
Backend: Django 5.1 • Django REST Framework • FastMCP 2.0 • PostgreSQL 16 + pgvector • Ollama Frontend: React 18 • TypeScript • Vite • Tailwind CSS • D3 • Recharts Infra: Docker • nginx • gunicorn • uvicorn
Documentation
Guide | Description |
Claude Desktop, Claude Code, Cursor setup | |
Ollama, OpenRouter, OpenAI, Cohere comparison | |
Local and production Docker guide | |
Cloud PostgreSQL with Supabase | |
Database schema deep dive | |
Daily capture and search patterns | |
Adding providers, tools, and importers | |
Deploy on a Windows server for home LAN access | |
Common issues and fixes | |
Feature roadmap |
License
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Persistent personal memory for AI assistants — save, search, and recall across every MCP client.
An MCP memory server. One memory your agents share — across models, devices and apps.
Persistent, portable memory for AI assistants — your private memory graph, from any MCP client.
Private persistent memory for Claude, ChatGPT & Gemini via MCP - semantic search, zero-code setup.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceA local MCP server for AI assistants to store and retrieve personal memories on disk, with optional semantic search using embeddings.-
- AlicenseNot gradedqualityCmaintenanceA self-hosted MCP memory server that gives AI assistants persistent, semantic memory by storing facts as vector embeddings locally, supporting semantic search and swappable embedding models.75 npmMIT
- AlicenseNot gradedqualityCmaintenanceA self-hosted MCP memory server with hybrid semantic and keyword search, providing persistent memory for AI coding assistants like Claude Code, Cursor, and Windsurf.15 npm4MIT
- AlicenseNot gradedqualityDmaintenanceA local-first memory MCP server that enables storing, searching, and managing personal memories with hybrid keyword and semantic recall, all on-device.20 npmMIT