structured-memory-mcp
# structured-memory-mcp
**An MCP server that gives an LLM a persistent, structured brain it reads at session start and writes to as it works, so you never re-explain yourself.**
Architecture reference implementation. Python, Postgres + pgvector, about 20 MCP tools, no LLM calls inside the server. Everything in the demo is synthetic.
---
## The problem
Every conversation with an LLM starts cold. The usual fixes all fail in a predictable way:
| Fix | Why it breaks |
|---|---|
| Paste chat history into context | Cost and noise grow forever. Old, superseded statements sit next to current ones with equal authority. |
| A bigger context window | Same pile, bigger. Nothing says which line is true *now*. |
| Plain vector search over everything | Great at "what's similar", blind to "what's current", and it always returns *something*, even when the honest answer is "nothing". |
What's missing isn't capacity. It's **structure**: knowing which kind of fact you're holding (a rule, a current state, a historical event, a hard number), which one wins when they disagree, and when to say "I don't know."
This repo is one answer: split memory into layers with different jobs, different write rules and different staleness rules, serve them over MCP, and teach the model the protocol on connect.
## Architecture
```mermaid
flowchart TB
subgraph Client["MCP client (any host)"]
Model["LLM"]
end
subgraph Server["MCP server — stdio, or http + bearer token"]
Tools["tool surface (20 tools)"]
Service["service layer: validate → repository → format"]
Repo["repository: the only SQL"]
Tools --> Service --> Repo
end
subgraph DB["Postgres + pgvector"]
subgraph Present["Present tense: state, overwritten in place"]
direction LR
Constraints["constraints<br/>standing rules"]
Records["records<br/>hard data"]
Pins["pins<br/>state per lane"]
Wiki["wiki<br/>compiled truth"]
Threads["threads<br/>open loops"]
end
subgraph Past["Past tense: the trail, append-only"]
Memory["vector memory<br/>hybrid search + honesty floor"]
end
subgraph Meta["Meta"]
Friction["friction<br/>model failure log"]
end
end
Model -- "session start: get_session_continuation()" --> Tools
Model -- "mid-session: read / write" --> Tools
Model -- "session close: set_pin, write_wiki, write_memory" --> Tools
Repo --> DB
Friction -. "human triage only" .-> Human(("human"))
```
### The seven layers
| Layer | Job | Tense | Write rule |
|---|---|---|---|
| **constraints** | Standing rules, injected every session | present | Added/retired only when the user says so. Never deleted. |
| **pins** | One "where things stand + what's next" slot per lane (`work`, `personal`, `school`) | present | Rewritten in place at session close. State, not history. |
| **threads** | Open loops tracked below the lane level; quiet ones surface at session start | present | Status changes record only what the user *said*. Silence never moves a thread. |
| **wiki** | One compiled page per entity; `CORE` pages are always-on operating instructions | present | Rewritten in place with a version bump. Never appended, never dated snapshots. |
| **vector memory** | What happened: dated facts and decisions. Hybrid search (vector cosine + BM25, rank-fused) with an honesty floor | past | Append-only, idempotent by key. One fact per write. |
| **records** | Hard structured data: invoices, deadlines, measurements | present | Upsert on (kind, title, date). Beats recall when they disagree. |
| **friction** | The model's own failure log, evidence-gated | meta | Written by the model, read and acted on by a human only. |
Details: [`docs/architecture.md`](docs/architecture.md).
### Trust order, in one line
`constraints > records > pins > wiki > vector memory > the model's own recall`: and if a *lower* layer is **newer** than a higher state layer and disagrees, the answer is **contested**: the model says so and asks, instead of silently picking. Executable version in [`brain/trust.py`](brain/trust.py); rationale and staleness rules in [`docs/trust-and-staleness.md`](docs/trust-and-staleness.md).
### Session lifecycle
1. **Start**: `get_session_continuation()` returns, deterministically and with no vector search: hard rules → one pin per lane → QUIET THREADS → full text of CORE wiki pages → wiki index.
2. **During**: topic-specific history from `read_memory` / `walk_trail`; hard data in `log_record`; the model touches threads, sets constraints, and logs its own friction when it has a concrete artifact.
3. **Close**: `set_pin` for the lane, `write_wiki` for changed entities, `write_memory` for durable facts.
Full protocols: [`docs/protocols.md`](docs/protocols.md). Honest failure modes and what the design does about each: [`docs/failure-modes.md`](docs/failure-modes.md).
## Run it locally (about 2 minutes, no API keys)
Prerequisites: Docker, Python 3.11+.
```bash
git clone <this-repo> && cd structured-memory-mcp
make setup # writes .env with a generated DB password (gitignored), creates .venv
make demo # starts Postgres+pgvector, migrates, seeds a fictional user, runs the walkthrough
make test # unit tests + integration tests against the live database
```
`make demo` plays a scripted session for a **fictional** freelance analyst (every name, client and number is invented): the session-start payload, a quiet-thread check-in, a search that hits, a search that correctly finds *nothing*, both trust-order conflicts, the friction evidence gate, and a session close.
An excerpt of real output:
```text
QUIET THREADS — my picture of these is stale; ask to refresh it (silence is missing data, never a verdict):
Half-marathon training (personal) — quiet 12d (cadence 7d) · Paused after a knee tweak; ...
Kestrel dashboard handoff (work) — quiet 6d (cadence 3d) · Handoff doc and walkthrough video ...
5. SEARCH THAT SHOULD FIND NOTHING — the honesty floor
No relevant memories found for this query.
6. TRUST ORDER 1 — a record beats a recalled memory
winner: record -> $1,450
flag: memory disagrees with record; record wins — flag the memory for correction.
7. TRUST ORDER 2 — a NEWER memory contradicts an older pin
contested: True (no silent winner)
flag: memory is newer than pin and disagrees — the pin may be stale; ask before asserting.
```
### Connect it to an MCP client
After `make up migrate seed` (or `make demo`):
```bash
# Claude Code
claude mcp add structured-memory -- /absolute/path/to/structured-memory-mcp/.venv/bin/python \
/absolute/path/to/structured-memory-mcp/server.py
```
Any MCP host that can spawn a stdio server works: point it at the venv's Python and `server.py`. The server reads its secrets by name from `.env`; nothing is baked into the config.
### Configuration
All by environment variable; the names are in [`.env.example`](.env.example). Highlights:
| Variable | Purpose |
|---|---|
| `POSTGRES_PASSWORD` / `DATABASE_URL` | Database access (`make setup` generates a password) |
| `EMBED_PROVIDER` | `local-hash` (default; offline, demo-grade) · `voyage` · `ollama` |
| `VOYAGE_API_KEY` | Only for `EMBED_PROVIDER=voyage` (`pip install '.[voyage]'`) |
| `MIN_SIMILARITY` | Honesty-floor override (per-provider defaults in `brain/config.py`) |
| `TRANSPORT` | `stdio` (default) or `http` (requires `MCP_AUTH_TOKEN`) |
> **Embeddings:** `local-hash` is a hashed bag-of-words projection. It tracks lexical overlap, not meaning. It exists so the demo and tests run offline. Use a real embedding model for anything real, and recalibrate the floor when you do ([`docs/failure-modes.md`](docs/failure-modes.md)).
## What's in the repo
```
brain/ the server: layers (pure logic), service, tools, repository, search, trust order
migrations/ schema only. No data rows ship in this repo.
demo/ synthetic data for a fictional user + the scripted walkthrough
tests/ pure-logic unit tests, plus integration + MCP-protocol tests (need the DB)
docs/ architecture, trust & staleness, protocols, failure modes, blog draft
```
Design choices worth a look: pure formatting/selection logic is separated from SQL so every layer's rules are unit-testable without a database; the server makes no LLM calls (retrieval and continuation are deterministic); the honesty floor and the multiplicative re-rank exist because of specific failures documented in [`docs/failure-modes.md`](docs/failure-modes.md).
## Not included (on purpose)
A web dashboard, OAuth, multi-user isolation, automatic contradiction detection across the store, and background "dreaming" jobs.
## Writeup
[`docs/blog/why-structured-memory.md`](docs/blog/why-structured-memory.md): why structured memory beats dumping chat history into context.
## License
MIT
TDQS
Scored across 20 tools
Most tools target distinct resources and actions, but read_memory and search_memory overlap heavily (both hybrid search the trail), and write_memory vs write_wiki could be confused about where durable information belongs. The detailed descriptions help, but the boundaries are not perfectly crisp.
All tool names use snake_case with a consistent verb_noun pattern (load_wiki, retire_constraint, write_memory, list_constraints, track_thread, etc.). The convention is predictable throughout with no mixed styles.
20 tools is heavy for a single MCP server and sits in the borderline range. While each tool covers a distinct facet (memory, wiki, constraints, threads, friction, records, session state), the surface could likely be consolidated to reduce selection overhead.
The server covers core lifecycle operations across its subdomains: memory read/write/search/walk, wiki load/write/list, constraints list/set/retire, threads list/track/touch, friction log/list/triage, and records log/list. Missing hard delete/update for memories, records, and wiki pages may be intentional by design, but still leaves minor gaps.