Skip to main content
Glama
yashraz23

memory-immune-system

by yashraz23
README.md
# A Memory Immune System

### Preventing memory-induced hallucination in LLM assistants

**An auditing layer for LLM long-term memory** — it catches when an assistant's stored
knowledge becomes self-contradictory, absorbs user-introduced falsehoods, goes stale, or
contains unsourced fabrications, *before* those corrupt entries produce hallucinated answers.

Every assistant with long-term memory (ChatGPT, Claude, Gemini) has this failure mode today:
memory gets things wrong, and wrong memories silently drive hallucinated answers later. This
project is an "immune system" that runs over a memory store, flags corruption, and —
critically — is measured by a benchmark showing it prevents the hallucinations a naive store
lets through.

**Where this sits in the hallucination landscape.** Hallucination splits into *factuality*
(fabricating facts) and *faithfulness / grounding* (answering unfaithfully to the provided
context). A corrupted memory is corrupted context, so an answer faithfully drawn from a false
memory is a **grounding hallucination**. This project attacks hallucination through the
**memory channel** — keeping the assistant's long-term memory clean so it has less to
hallucinate from. It does not claim to fix generation-time fabrication with no memory
involved; the scope is *memory-induced* hallucination, which is exactly what keeps the claim
defensible.

Shipped as an **MCP server** so any client can install it, plus a **benchmark** that reports
precision/recall per corruption type.

### How this differs from prior work

Prior work audits *adversarially-poisoned* agent memory *post-hoc* (e.g. MemAudit), or
benchmarks hallucination attribution over agent trajectories (e.g. AgentHallu). This project
builds a **real-time defensive layer for _benign_ memory corruption in personal-assistant
agents, shipped as an installable MCP server** — the target is a well-meaning user who is
simply wrong, not an attacker. Three distinctions carry the novelty: benign vs. adversarial,
real-time defensive vs. post-hoc audit, installable tool vs. research benchmark. The
**user-induced false belief** case is the headline contribution. Full positioning and
citations in [`docs/related-work.md`](docs/related-work.md).

---

## See it work

```bash
python examples/demo.py          # 40 seconds, offline, no API key
```

```
You tell your assistant a few things:
   "My manager is Dana"
   "My standup is at 9am"
   "My skills are Python"

Weeks later, something changes:
   "My standup is at 10am"

The audit runs on write:
   ! self_contradiction — '9am' superseded by '10am' on user.standup

Asked about the standup, it now gets ONE answer, not two:
   -> My standup is at 10am
   (the 9am version is no longer offered)

The old memory isn't deleted — it's held for review:
   [flagged] My standup is at 9am
            unresolved conflict, needs review: '9am' vs '10am' (distinct)

And you have the last word. If the audit got it wrong, overrule it:
   restore("My standup is at 9am")  ->  [active]
```

Meanwhile `My skills are Python` is never touched — skills are multi-valued, so a second skill
isn't a contradiction. That distinction is the difference between an immune system and an
autoimmune one.

## Results

Three scenarios: two scripted, and one built on a **real multi-session LoCoMo transcript** with
corruptions planted among 60 turns of genuine human dialogue. Each includes clean facts and
near-miss discriminators (restatements, refinements, multi-valued additions) so false positives
are measured, not hidden.

```bash
python -m mis.benchmark                  # offline, deterministic
python -m mis.benchmark --locomo --llm   # + real transcript, + Tier-3 LLM hooks
```

### Detection is not the bottleneck — visibility is

**Precision 1.00 and recall 1.00 on every corruption type, in every scenario, at both tiers**,
with zero false positives — including zero false alarms on real dialogue. The detectors are
accurate on whatever they can see. The interesting question turned out to be *how much they can
see*, and that is what the expensive tier actually buys:

| scenario | parse coverage: regex only | + LLM extractor | resolved outright: regex only | + LLM |
|---|---|---|---|---|
| `personal_assistant` (13 memories) | 100% | 100% | 5/5 | 5/5 |
| `messy_realistic` (14 memories) | 86% | **100%** | 1/3 | **3/3** |
| `locomo[0]` (66 memories, real dialogue) | **12%** | **55–59%** | 4/4 | 4/4 |

An unparsed memory is invisible to every detector, so parse coverage is a hard ceiling on
recall. On real human conversation the regex tier sees **12%** — people say *"Just moved over to
the Denver office last month"*, not *"My employer is Acme"*. The LLM extractor lifts that to
roughly **55–59%** (it's a range because LLM extraction is non-deterministic: repeat runs
structure slightly different sentences). Separately, *resolved outright* counts corruptions that
got the ideal action rather than merely being surfaced for the user; adjudication takes
`messy_realistic` from 1/3 to 3/3.

### Downstream effect

On questions whose answers depend on a corruption — does the store still assert something wrong?

| scenario | naive store | audited store |
|----------|-------------|---------------|
| `personal_assistant` | 0/5 correct | **5/5** |
| `messy_realistic` | 1/4 correct | **4/4** |
| `locomo[0]` (real transcript) | 0/4 correct | **4/4** |

### Honest caveats

- **Small scenarios.** These show the detectors behave correctly and that the naive-vs-audited
  gap is real. They do not yet map where the system breaks.
- **41% of real dialogue is still invisible** even with the LLM extractor. Much of that is
  genuinely factless (greetings, questions), but not all of it.
- **Organic dialogue carries no gold labels** — corruptions planted in LoCoMo are labeled;
  the surrounding real utterances aren't, so flags on them are reported separately for manual
  review rather than scored.
- **LLM-tier numbers vary between runs.** Extraction is non-deterministic, so parse coverage
  moves by a few points run to run. The offline tier is fully deterministic (fixed clock, fixed
  scenarios) and reproduces exactly.

## Install it (MCP server)

Give any MCP client an audited memory instead of a plain one.

```bash
pip install -e ".[server]"
```

Then add it to Claude Desktop's config (`claude_desktop_config.json`):

```json
{
  "mcpServers": {
    "memory-immune-system": {
      "command": "C:\\path\\to\\memoryimmune system\\.venv\\Scripts\\python.exe",
      "args": ["-m", "mis.server"]
    }
  }
}
```

Restart the client and the assistant gains these tools:

| tool | what it does |
|------|--------------|
| `remember` | store a fact — audited on write, reports anything it flags |
| `recall` | retrieve memories that are safe to reason from (excludes quarantined/superseded) |
| `review_flagged` | show what the immune system flagged, and why |
| `restore` | overrule the audit when the user says a memory was right |
| `quarantine` | retire a memory the user says is wrong (retained, not deleted) |
| `audit_sweep` | periodic staleness check |
| `reliability` | explainable 0–1 trust score for a memory |
| `memory_stats` | store overview |

Memory persists at `~/.memory-immune-system/memory.db` (override with `MIS_DB_PATH`). If
`ANTHROPIC_API_KEY` is set the expensive LLM audit tier turns on automatically; without it,
everything still runs on the offline tiers.

## What it catches (taxonomy)

**Tier 1 (first):** self-contradiction over time; user-induced false belief.
**Tier 2 (if time):** staleness; fabrication / source-less memory.

Full operational definitions and discriminators (the near-miss cases we must *not* flag) live in
[`docs/definitions.md`](docs/definitions.md).

## Architecture (three pieces)

1. **Memory layer** — store/retrieve/update with per-entry provenance metadata. Shippable as an
   MCP server. *(This layer: built.)*
2. **Immune system** — a separate auditing process over the store. Cheap pre-filters first
   (provenance, recency, embedding contradiction search), expensive LLM adjudication only on
   candidates. *(Weeks 5–7.)*
3. **Benchmark** — scripted long-horizon sessions with injected corruptions; measures whether the
   auditor catches them and whether a naive baseline gets corrupted where this one doesn't.
   Reports precision/recall per type. *(Weeks 8–9.)*

## Quickstart

```bash
python examples/quickstart.py     # multi-turn session, no API keys, no network
```

```python
from mis import MemoryStore, MemoryEntry, Source, SourceKind

store = MemoryStore("memory.db")                     # SQLite file (or ":memory:")
store.store(MemoryEntry(
    content="Standup is at 9am",
    confidence=0.7,
    source=Source(kind=SourceKind.USER, session_id="s1", turn_id=4),
))

for entry, sim in store.search("when is standup", k=3):
    print(sim, entry.content)
```

By default the store uses a dependency-free hashing embedder so it runs offline. For real
semantics install the extra and pass a `LocalEmbedder`:

```bash
pip install -e ".[embeddings]"
```

```python
from mis import MemoryStore, LocalEmbedder
store = MemoryStore("memory.db", embedder=LocalEmbedder())   # all-MiniLM-L6-v2, local
```

## Tests

```bash
pip install -e ".[dev]"
pytest
```

## Tech stack

- **Storage:** SQLite (rows + metadata) + in-process vector search. No external services.
- **Embeddings:** local sentence-transformers for the contradiction pre-filter (no API cost).
- **LLM adjudication:** Anthropic API, on candidates only (Weeks 5+).
- **Interface:** MCP (Python SDK) — the memory layer as an installable server (Weeks 3–4).
- **Orchestration:** LangGraph for the host agent + auditing process (Weeks 3+).

## Layout

```
docs/definitions.md   # testable Tier-1 corruption definitions  ← the intellectual core
docs/schema.md        # memory schema rationale
src/mis/schema.py     # MemoryEntry + enums
src/mis/store.py      # SQLite store + vector search + audit log
src/mis/embeddings.py # Embedder protocol; HashEmbedder (default), LocalEmbedder
src/mis/agent.py      # minimal multi-turn host agent
examples/quickstart.py
tests/
```