Skip to main content
Glama
README.md
# structured-memory-mcp

**An MCP server that gives an LLM a persistent, structured brain it reads at session start and writes to as it works, so you never re-explain yourself.**

Architecture reference implementation. Python, Postgres + pgvector, about 20 MCP tools, no LLM calls inside the server. Everything in the demo is synthetic.

---

## The problem

Every conversation with an LLM starts cold. The usual fixes all fail in a predictable way:

| Fix | Why it breaks |
|---|---|
| Paste chat history into context | Cost and noise grow forever. Old, superseded statements sit next to current ones with equal authority. |
| A bigger context window | Same pile, bigger. Nothing says which line is true *now*. |
| Plain vector search over everything | Great at "what's similar", blind to "what's current", and it always returns *something*, even when the honest answer is "nothing". |

What's missing isn't capacity. It's **structure**: knowing which kind of fact you're holding (a rule, a current state, a historical event, a hard number), which one wins when they disagree, and when to say "I don't know."

This repo is one answer: split memory into layers with different jobs, different write rules and different staleness rules, serve them over MCP, and teach the model the protocol on connect.

## Architecture

```mermaid
flowchart TB
    subgraph Client["MCP client (any host)"]
        Model["LLM"]
    end

    subgraph Server["MCP server — stdio, or http + bearer token"]
        Tools["tool surface (20 tools)"]
        Service["service layer: validate → repository → format"]
        Repo["repository: the only SQL"]
        Tools --> Service --> Repo
    end

    subgraph DB["Postgres + pgvector"]
        subgraph Present["Present tense: state, overwritten in place"]
            direction LR
            Constraints["constraints<br/>standing rules"]
            Records["records<br/>hard data"]
            Pins["pins<br/>state per lane"]
            Wiki["wiki<br/>compiled truth"]
            Threads["threads<br/>open loops"]
        end
        subgraph Past["Past tense: the trail, append-only"]
            Memory["vector memory<br/>hybrid search + honesty floor"]
        end
        subgraph Meta["Meta"]
            Friction["friction<br/>model failure log"]
        end
    end

    Model -- "session start: get_session_continuation()" --> Tools
    Model -- "mid-session: read / write" --> Tools
    Model -- "session close: set_pin, write_wiki, write_memory" --> Tools
    Repo --> DB
    Friction -. "human triage only" .-> Human(("human"))
```

### The seven layers

| Layer | Job | Tense | Write rule |
|---|---|---|---|
| **constraints** | Standing rules, injected every session | present | Added/retired only when the user says so. Never deleted. |
| **pins** | One "where things stand + what's next" slot per lane (`work`, `personal`, `school`) | present | Rewritten in place at session close. State, not history. |
| **threads** | Open loops tracked below the lane level; quiet ones surface at session start | present | Status changes record only what the user *said*. Silence never moves a thread. |
| **wiki** | One compiled page per entity; `CORE` pages are always-on operating instructions | present | Rewritten in place with a version bump. Never appended, never dated snapshots. |
| **vector memory** | What happened: dated facts and decisions. Hybrid search (vector cosine + BM25, rank-fused) with an honesty floor | past | Append-only, idempotent by key. One fact per write. |
| **records** | Hard structured data: invoices, deadlines, measurements | present | Upsert on (kind, title, date). Beats recall when they disagree. |
| **friction** | The model's own failure log, evidence-gated | meta | Written by the model, read and acted on by a human only. |

Details: [`docs/architecture.md`](docs/architecture.md).

### Trust order, in one line

`constraints > records > pins > wiki > vector memory > the model's own recall`: and if a *lower* layer is **newer** than a higher state layer and disagrees, the answer is **contested**: the model says so and asks, instead of silently picking. Executable version in [`brain/trust.py`](brain/trust.py); rationale and staleness rules in [`docs/trust-and-staleness.md`](docs/trust-and-staleness.md).

### Session lifecycle

1. **Start**: `get_session_continuation()` returns, deterministically and with no vector search: hard rules → one pin per lane → QUIET THREADS → full text of CORE wiki pages → wiki index.
2. **During**: topic-specific history from `read_memory` / `walk_trail`; hard data in `log_record`; the model touches threads, sets constraints, and logs its own friction when it has a concrete artifact.
3. **Close**: `set_pin` for the lane, `write_wiki` for changed entities, `write_memory` for durable facts.

Full protocols: [`docs/protocols.md`](docs/protocols.md). Honest failure modes and what the design does about each: [`docs/failure-modes.md`](docs/failure-modes.md).

## Run it locally (about 2 minutes, no API keys)

Prerequisites: Docker, Python 3.11+.

```bash
git clone <this-repo> && cd structured-memory-mcp
make setup     # writes .env with a generated DB password (gitignored), creates .venv
make demo      # starts Postgres+pgvector, migrates, seeds a fictional user, runs the walkthrough
make test      # unit tests + integration tests against the live database
```

`make demo` plays a scripted session for a **fictional** freelance analyst (every name, client and number is invented): the session-start payload, a quiet-thread check-in, a search that hits, a search that correctly finds *nothing*, both trust-order conflicts, the friction evidence gate, and a session close.

An excerpt of real output:

```text
QUIET THREADS — my picture of these is stale; ask to refresh it (silence is missing data, never a verdict):
  Half-marathon training (personal) — quiet 12d (cadence 7d) · Paused after a knee tweak; ...
  Kestrel dashboard handoff (work) — quiet 6d (cadence 3d) · Handoff doc and walkthrough video ...

5. SEARCH THAT SHOULD FIND NOTHING — the honesty floor
No relevant memories found for this query.

6. TRUST ORDER 1 — a record beats a recalled memory
   winner:  record -> $1,450
   flag:    memory disagrees with record; record wins — flag the memory for correction.

7. TRUST ORDER 2 — a NEWER memory contradicts an older pin
   contested: True  (no silent winner)
   flag:      memory is newer than pin and disagrees — the pin may be stale; ask before asserting.
```

### Connect it to an MCP client

After `make up migrate seed` (or `make demo`):

```bash
# Claude Code
claude mcp add structured-memory -- /absolute/path/to/structured-memory-mcp/.venv/bin/python \
  /absolute/path/to/structured-memory-mcp/server.py
```

Any MCP host that can spawn a stdio server works: point it at the venv's Python and `server.py`. The server reads its secrets by name from `.env`; nothing is baked into the config.

### Configuration

All by environment variable; the names are in [`.env.example`](.env.example). Highlights:

| Variable | Purpose |
|---|---|
| `POSTGRES_PASSWORD` / `DATABASE_URL` | Database access (`make setup` generates a password) |
| `EMBED_PROVIDER` | `local-hash` (default; offline, demo-grade) · `voyage` · `ollama` |
| `VOYAGE_API_KEY` | Only for `EMBED_PROVIDER=voyage` (`pip install '.[voyage]'`) |
| `MIN_SIMILARITY` | Honesty-floor override (per-provider defaults in `brain/config.py`) |
| `TRANSPORT` | `stdio` (default) or `http` (requires `MCP_AUTH_TOKEN`) |

> **Embeddings:** `local-hash` is a hashed bag-of-words projection. It tracks lexical overlap, not meaning. It exists so the demo and tests run offline. Use a real embedding model for anything real, and recalibrate the floor when you do ([`docs/failure-modes.md`](docs/failure-modes.md)).

## What's in the repo

```
brain/          the server: layers (pure logic), service, tools, repository, search, trust order
migrations/     schema only. No data rows ship in this repo.
demo/           synthetic data for a fictional user + the scripted walkthrough
tests/          pure-logic unit tests, plus integration + MCP-protocol tests (need the DB)
docs/           architecture, trust & staleness, protocols, failure modes, blog draft
```

Design choices worth a look: pure formatting/selection logic is separated from SQL so every layer's rules are unit-testable without a database; the server makes no LLM calls (retrieval and continuation are deterministic); the honesty floor and the multiplicative re-rank exist because of specific failures documented in [`docs/failure-modes.md`](docs/failure-modes.md).

## Not included (on purpose)

A web dashboard, OAuth, multi-user isolation, automatic contradiction detection across the store, and background "dreaming" jobs.

## Writeup

[`docs/blog/why-structured-memory.md`](docs/blog/why-structured-memory.md): why structured memory beats dumping chat history into context.

## License

MIT

TDQS

B3.4/5.0

Scored across 20 tools

Disambiguation4/5

Most tools target distinct resources and actions, but read_memory and search_memory overlap heavily (both hybrid search the trail), and write_memory vs write_wiki could be confused about where durable information belongs. The detailed descriptions help, but the boundaries are not perfectly crisp.

Naming Consistency5/5

All tool names use snake_case with a consistent verb_noun pattern (load_wiki, retire_constraint, write_memory, list_constraints, track_thread, etc.). The convention is predictable throughout with no mixed styles.

Tool Count3/5

20 tools is heavy for a single MCP server and sits in the borderline range. While each tool covers a distinct facet (memory, wiki, constraints, threads, friction, records, session state), the surface could likely be consolidated to reduce selection overhead.

Completeness4/5

The server covers core lifecycle operations across its subdomains: memory read/write/search/walk, wiki load/write/list, constraints list/set/retire, threads list/track/touch, friction log/list/triage, and records log/list. Missing hard delete/update for memories, records, and wiki pages may be intentional by design, but still leaves minor gaps.

Maintenance

ActivityMaintained
ResponsivenessNo issues