Aokun Memory MCP Server
by pangak959
README.md
# Aokun Memory MCP Server
English | [简体中文](README.zh-CN.md)
**A persistent long-term memory MCP server for AI assistants.** It distills conversations into three retrievable layers of memory (documents / memory atoms / a knowledge graph), recalls them through a multi-channel hybrid of **BM25 + vector + graph** retrieval, and adds human-like forgetting, atom lifecycles, and crash recovery — so any MCP-capable client gets long-term memory that can fade, strengthen, and evolve.
[](LICENSE)


> The system is optimized for **Chinese text** (word segmentation uses [jieba](https://github.com/fxsjy/jieba)); the architecture and all algorithms are language-agnostic.
---
## Table of Contents
- [Features](#features)
- [Architecture](#architecture)
- [Quick Start](#quick-start)
- [Memory Model](#memory-model)
- [Write Pipeline](#write-pipeline)
- [Retrieval Pipeline (core)](#retrieval-pipeline-core)
- [Forgetting & Lifecycle](#forgetting--lifecycle)
- [Reliability](#reliability)
- [MCP Tools](#mcp-tools)
- [Configuration](#configuration)
- [Maintenance Scripts & Config Panel](#maintenance-scripts--config-panel)
- [Tech Stack](#tech-stack)
- [Design Origins](#design-origins)
- [License](#license)
---
## Features
- **Three-layer memory model**
- **Document memory**: a complete, structured memory distilled from a conversation (summary / topics / key facts / sentiment / importance)
- **Memory atoms**: key facts broken into fine-grained facts, each with a type, a shelf life (TTL), and its own decay curve
- **Knowledge graph**: people, topics, and facts become nodes; *describes / mentioned_in / co_occurs_with* become edges — every node and edge maps back to its source memory
- **Parallel multi-channel hybrid retrieval**: SQLite FTS5 full-text (BM25), FAISS vectors (cosine), graph keyword and graph-vector channels, all fused with **RRF (Reciprocal Rank Fusion)**
- **Dual-route retrieval + dynamic routing**: the document route and graph route are fused with weights that are adjusted in real time by query intent (relational / temporal / factual)
- **MMR diversity re-ranking**: prevents the top results from being near-duplicate phrasings of the same topic
- **Human-like forgetting**: daily importance decay, archival of low-value old memories, and per-type atom TTLs with a full state machine; **memories that are recalled repeatedly decay more slowly** (access reinforcement)
- **Zero extra tokens for structuring**: atom classification and graph construction are entirely **deterministic rules** (regex + jieba segmentation) — no extra LLM calls
- **Crash recoverable**: every write is tracked by a `write_op` journal; if the process dies at any stage, a restart back-fills the missing indexes / atoms / graph
- **Local web config panel**: double-click a `.bat` to tune parameters in a browser, no manual YAML editing required
---
## Architecture
```
┌──────────────────────────────────────────────────────────────────────┐
│ MCP client (Claude Code / any MCP host) │
└───────────────────────────────┬──────────────────────────────────────┘
│ stdio / MCP
┌───────────────────────────────▼──────────────────────────────────────┐
│ server.py tools/memory_tools.py │
│ MCP server entrypoint registers memorize/recall/list/edit │
└───────────────────────────────┬──────────────────────────────────────┘
│
┌───────────────────────────────▼──────────────────────────────────────┐
│ core/memory_engine.py (orchestrator) │
│ add_memory() write pipeline search_memories() recall pipeline │
│ write_op crash-recovery journal search cache (TTL + LRU + epoch) │
└───────┬────────────────┬──────────────────┬───────────────┬──────────┘
│ │ │ │
┌───────▼──────┐ ┌───────▼────────┐ ┌───────▼────────┐ ┌────▼─────────┐
│ processors/ │ │ summarizer.py │ │ embedding.py │ │ retrieval/ │
│ segmentation │ │ LLM chat │ │ OpenAI-compat │ │ dual-route │
│ atom rules │ │ summarization │ │ embedding client│ │ hybrid search│
│ graph rules │ │ (LLM used here │ │ L2 normalized │ │ (see below) │
│ entity resolve│ │ only) │ │ │ │ │
└───────┬──────┘ └────────────────┘ └────────────────┘ └────┬─────────┘
│ │
┌───────▼────────────────────────────────────────────────────▼─────────┐
│ storage/ (persistence) │
│ DocumentStore AtomStore GraphStore (SQLite / FTS5) │
└───────┬───────────────────────────┬──────────────────────────────────┘
│ │
┌───────▼─────────────────┐ ┌───────▼─────────────────────────────────┐
│ Main DB aokun_memory.db │ │ Graph DB aokun_memory_graph_documents │
│ · documents │ │ · graph entry text + vector metadata │
│ · memory_atoms │ └───────┬─────────────────────────────────┘
│ · graph_nodes/edges/ │ │
│ entries │ ┌───────▼─────────────────────────────────┐
│ · 3 FTS5 virtual tables│ │ FAISS index files (under data/) │
│ · memory_write_ops │ │ · faiss_index (document vectors) │
│ (recovery journal) │ │ · graph_faiss_index (graph-entry vecs) │
└─────────────────────────┘ └──────────────────────────────────────────┘
Sidecars: decay_scheduler.py (daily decay) · atom_lifecycle_manager.py (TTL state machine)
backup_manager.py (version + daily hot backups) · config_panel.py (local web UI)
```
**Module responsibilities**
| Path | Responsibility |
|---|---|
| `server.py` | MCP stdio server entrypoint; wires the engine and tools together |
| `tools/memory_tools.py` | Exposed MCP tool definitions and parameter validation |
| `core/memory_engine.py` | Orchestrator: write pipeline, recall pipeline, decay, cache, crash recovery |
| `core/config.py` | YAML config loading and dotted-path access (with defaults); anchors relative paths to the config directory |
| `core/embedding.py` | OpenAI-compatible `/embeddings` client, batch embedding + L2 normalization |
| `core/summarizer.py` | Calls an LLM to summarize a conversation into structured memory JSON |
| `core/processors/text_processor.py` | jieba segmentation, custom words, stop-word management |
| `core/processors/atom_classifier.py` | **Pure-rule** classification of key facts into memory atoms |
| `core/processors/graph_extractor.py` | **Pure-rule** extraction of graph nodes / edges / entries from metadata |
| `core/processors/entity_resolver.py` | Entity-name canonicalization (alias merging) |
| `core/retrieval/` | Retrievers: BM25, vector, hybrid, RRF, three graph channels, dual-route, atom |
| `core/decay_scheduler.py` | Daily scheduled document-importance decay (catches up after downtime) |
| `core/managers/atom_lifecycle_manager.py` | Atom expiry / forgetting / physical-purge state machine |
| `core/managers/graph_memory_manager.py` | Graph write orchestration (delete old → extract → persist → embed) |
| `core/managers/backup_manager.py` | Full backup on version change + daily online SQLite backup |
| `storage/` | All SQLite / FTS5 / FAISS read-and-write wrappers |
---
## Quick Start
### 1. Requirements
- Python 3.10–3.14 (64-bit; 3.11 / 3.12 recommended)
- Two OpenAI-compatible APIs:
- **Embedding API** (defaults to SiliconFlow's `Qwen/Qwen3-Embedding-8B`, 4096 dimensions)
- **Chat/summarization API** (defaults to DeepSeek; used only by the `summarize_memory` tool or your own external conversation pipeline — everyday `memorize` / `recall` consume **no** LLM tokens)
### 2. Install
```bash
git clone https://github.com/pangak959/Aokun_memory.git
cd Aokun_memory
pip install -r requirements.txt
```
### 3. Configure
```bash
cp config.example.yaml config.yaml # on Windows, copy and rename the file manually
```
Edit `config.yaml` and fill in at least the two `api_key` values. Database and index paths default to the `data/` directory inside the project and are created automatically. Relative paths are always resolved against the directory containing `config.yaml`, so it does not matter which working directory the MCP client launches the server from.
### 4. Connect an MCP client
See `mcp_config.example.json` and replace the path with the real path on your machine:
```json
{
"mcpServers": {
"aokun_memory": {
"command": "python",
"args": ["C:\\path\\to\\Aokun_memory\\server.py"],
"env": {}
}
}
}
```
The MCP server is launched automatically by the client when needed — you do not start it yourself. To sanity-check initialization, run `python server.py` from the project directory (or double-click `start.bat`); seeing `Aokun Memory MCP Server starting...` means the config and databases initialized successfully. The process then waits on stdio, so just close the window.
### 5. Tune parameters (optional)
Double-click `配置面板.bat` (or run `python config_panel.py`); your browser opens <http://127.0.0.1:8765>, where you can adjust recall counts, fusion weights, decay rates, and more, grouped by function. Save and reconnect the MCP client for changes to take effect.
---
## Memory Model
The system maintains three granularities of memory at once; they are generated together on write and complement each other at recall time.
### 1. Document memory (`documents`)
A complete memory produced by summarizing one conversation — one row in the `documents` table:
| Field | Description |
|---|---|
| `text` | Memory body (the code prepends an `【原始时间 / original time: …】` marker) |
| `metadata.topics` | List of topic terms |
| `metadata.participants` | Participants (people) |
| `metadata.key_facts` | Key facts (the raw material for memory atoms) |
| `metadata.sentiment` | `positive` / `neutral` / `negative` |
| `metadata.importance` | 0.0–1.0; used by both decay and ranking |
| `metadata.pinned` | Pinned memories are immune to decay and archival |
| `metadata.session_id / persona_id` | Session and persona isolation |
| `metadata.access_count / last_access_time` | Inputs to access reinforcement |
### 2. Memory atoms (`memory_atoms`)
Key facts are split into fine-grained facts **with shelf lives**, stored, retrieved, and expired independently.
- **Six atom types** (decided by `atom_classifier.py` with regex rules — zero LLM calls):
| Type | Meaning | Default TTL | Decay curve |
|---|---|---|---|
| `episodic` | Episodes: things that happened | 7 days | exponential |
| `planned` | Plans: things to do in the future (with Chinese time-expression parsing) | 2 days | step (no decay until the deadline, then instant expiry) |
| `factual` | Factual knowledge | 180 days | exponential |
| `relational` | Inter-person relationships | 90 days | linear |
| `preference` | Preferences / likes and dislikes | 60 days | exponential |
| `unknown` | Fallback for unclassified facts | 30 days | linear |
- **The effective TTL is modulated by importance and reinforcement count**:
```
ttl = base_ttl × (0.5 + importance) × (1 + 0.5 × reinforcement_count)
```
- **Atom state machine**: `active` → (on expiry) `expired` → (soft-deleted after 7 days, removed from FTS) `forgotten` → (physically deleted after 30 days); there are also `dormant` and `superseded` states.
- **Automatic reinforcement**: when a new atom has a bag-of-words Jaccard similarity ≥ 0.6 with an existing atom, it is treated as the same fact being mentioned again — the old atom gets `reinforcement_count +1` and its confidence and TTL are refreshed. **Facts that recur live longer.**
### 3. Knowledge graph (`graph_nodes` / `graph_edges` / `graph_entries`)
The graph is also **built deterministically, without an LLM** (`graph_extractor.py`):
- **Nodes (`GraphNode`)**, keyed by `node_type:node_value:canonical_value`:
- `topic`: from `topics`
- `person`: from `participants`
- `fact`: from each `key_fact` / each atom
- **Edges (`GraphEdge`)**, three relation types:
- `topic/person` —`describes`→ `fact`: which person/topic a fact is about
- `fact` —`mentioned_in`→ `fact`: co-occurrence of facts within one memory
- `topic/person` pairs —`co_occurs_with`: associative co-occurrence
- **Entries (`GraphEntry`)**: both nodes and edges produce a **searchable text** (e.g. `"Fact: Alice likes iced americano. Summary: <memory body>"`), and each entry maps back to its source document via `source_memory_id`. This is the key design that lets the graph participate in full-text and vector retrieval.
- **Entity resolution**: `EntityResolver.canonicalize()` normalizes case, whitespace, and common aliases so the same person/topic never splits into multiple nodes.
- Edges carry `confidence` (inherited from atom-classification confidence) and `weight`, both used by the graph-route score.
### 4. On-disk layout
| Storage | Contents |
|---|---|
| `data/aokun_memory.db` (SQLite) | `documents`, `memory_atoms`, `graph_nodes`, `graph_edges`, `graph_entries`, 3 FTS5 virtual tables, and the `memory_write_ops` recovery journal |
| `data/aokun_memory_graph_documents.db` (SQLite) | Graph-entry text and vector metadata |
| `data/faiss_index` | Document vectors (`IndexIDMap2(IndexFlatIP)`) |
| `data/graph_faiss_index` | Graph-entry vectors (same structure) |
All vectors are **L2-normalized and indexed with inner product**, so the inner-product score equals cosine similarity.
---
## Write Pipeline
The `memorize` tool calls `MemoryEngine.add_memory()`. The whole pipeline is tied together by one `write_op` journal entry, making it crash-recoverable:
```
memorize(content, metadata)
│
├─ 1. classify_atoms(metadata.key_facts) rules only: key facts → atoms (type/TTL/confidence)
│
├─ 2. write write_op journal (status=pending) ── crash-recovery anchor
│
├─ 3. Document layer (docs stage)
│ ├─ insert into documents, obtain integer doc_id
│ ├─ jieba segmentation → write FTS5 full-text index (BM25)
│ └─ call embedding API → L2 normalize → write FAISS
│ (a vector failure does not block: the document is already persisted;
│ rebuild_index can back-fill later)
│ → write_op advances to docs_done
│
├─ 4. Atom layer (atoms stage)
│ ├─ batch-insert memory_atoms + memory_atoms_fts
│ ├─ on failure, fall back to row-by-row inserts; bad rows are flagged
│ │ needs_repair instead of rolling back the whole batch
│ └─ same-fact auto reinforcement (Jaccard ≥ 0.6)
│ → write_op advances to atoms_done
│
├─ 5. Graph layer (graph stage)
│ ├─ delete this memory's previous graph output (supports rebuilding after an edit)
│ ├─ GraphExtractor rule-extracts nodes / edges / entries
│ ├─ EntityResolver canonicalizes, then persist graph DB and graph FTS
│ └─ batch-embed entries → write graph vector DB, write back vector_doc_id
│ → write_op advances to graph_done → done
│
└─ 6. search-cache epoch +1 (a search immediately after a write never hits stale cache)
```
---
## Retrieval Pipeline (core)
The `recall` tool calls `MemoryEngine.search_memories()`. When the graph is enabled it runs **dual-route retrieval (`DualRouteRetriever`)**, and each route is itself a multi-channel hybrid.
```
query
│
search-cache hit? (TTL 45s / LRU 256)
┌──────────────┴────────────────┐
yes no
│ │
return directly ┌──────────▼──────────┐
│ DualRouteRetriever │
└──────┬──────┬──────┘
doc route 0.65 ◄┴──────┴► 0.35 graph route
┌─────────────────────┐ ┌──────────────────────────┐
│ HybridRetriever │ │ GraphRetriever │
│ │ │ │
│ BM25 ch. Vector ch.│ │ graph-keyword graph-vec │
│ (FTS5) (FAISS) │ │ (FTS+traversal) (FAISS) │
└────┬────────┬───────┘ └──────┬───────────┬────────┘
│ RRF fuse (k=60) │ │ RRF fuse (k=60) │
▼ ▼ ▼ ▼
3-dim weighted rank + MMR dedup 4-dim weighting (confidence × atom TTL)
└────────────────┬────────────┴───────────────┬─────────┘
│ cross-route bonus + dynamic intent weighting
▼ │
fused ranking, take Top-K │
│ │
async touch: refresh last_access_time / access_count (reinforce) │
▼ │
recall returns (persona_summary preferred, text fallback)
```
### I. Document route: `HybridRetriever`
#### 1. BM25 channel (`BM25Retriever`)
- Both queries and documents go through **jieba segmentation + stop-word removal**, are joined with spaces, and written into a SQLite **FTS5** virtual table.
- Query terms are joined with `OR` (favor high recall; leave precision to the fusion layer). FTS5's built-in `bm25()` gives the raw score, which a sigmoid maps into `(0,1)`:
```
bm25_norm = 1 / (1 + exp(-raw_bm25 / 8))
```
- Session/persona filtering happens in Python (hence the vector channel pre-fetches `k×10` candidates before filtering).
#### 2. Vector channel (`VectorRetriever`)
- The query is embedded, L2-normalized, and searched with FAISS `IndexFlatIP` inner product — inner product is cosine similarity — then linearly mapped into `(0,1)`:
```
vec_score = (cosine + 1) / 2
```
#### 3. Parallelism, degradation, and RRF fusion
- The two channels run **in parallel** with `asyncio.gather`.
- If either channel fails and `fallback_enabled=true`, retrieval automatically **degrades to the surviving channel** without interruption.
- The two rankings are fused with **RRF (Reciprocal Rank Fusion, Cormack et al., 2009)**, which looks only at ranks and is immune to differing score scales:
```
RRF(d) = Σ_channels 1 / (k + rank(d) + 1) # k = rrf_k, paper recommends 60
```
#### 4. Three-dimensional weighting
The normalized RRF score captures relevance only; the final score layers on importance and recency (default weights 0.5 / 0.25 / 0.25):
```
final(d) = α · rrf_norm(d) # α = score_alpha = 0.5
+ β · importance(d) # β = score_beta = 0.25
+ γ · recency(d) # γ = score_gamma = 0.25
recency(d) = exp( - age_days(d) / 30 )
age_days = now - max(create_time, last_access_time)
```
Note that age uses the **later of creation time and last-access time** — an old memory that was recalled recently gets its freshness pulled back up. This is the first mechanism behind "the more you use it, the stronger it gets".
#### 5. MMR diversity re-ranking
After fusion, **MMR (Maximal Marginal Relevance)** iteratively selects results, with bag-of-words Jaccard as the similarity measure:
```
MMR = λ · relevance(d) − (1 − λ) · max_sim(d, already-selected) # λ = mmr_lambda = 0.7
```
Higher `λ` favors relevance; lower `λ` favors topic diversity, keeping the top-K from being duplicate phrasings of one event.
### II. Graph route: `GraphRetriever`
The graph route is itself a "keyword + vector" hybrid, and can expand across graph structure.
#### 1. Graph-keyword channel (`GraphKeywordRetriever`)
After segmentation, the query hits graph entries through four entry points, all traced back to `source_memory_id`:
| Entry point | Match method | Weight |
|---|---|---|
| Entry full text | Direct FTS5 BM25 hit on `graph_entries` | 1.0 |
| Node match | Query terms hit a node token (asking about a specific person/topic) | 0.7 |
| **One-hop neighbors** | Expand from a hit node along edges to adjacent entries | 0.7 |
| **Two-hop neighbors** (optional, `graph_expansion_hops=2`) | Neighbors of neighbors | 0.4 |
When one source memory is hit through multiple entry points, scores aggregate (each extra entry point adds a ×0.35 boost); expansion is capped by `graph_expansion_limit` (default 24).
> This layer answers questions BM25/vector are bad at — e.g. "what is the relationship between Alice and Bob?". Through structures like `Alice ─ co_occurs_with ─ Bob`, memories that never literally contain the word "relationship" are still recalled.
#### 2. Graph-vector channel (`GraphVectorRetriever`)
- Graph-entry text (`"Fact: … Summary: …"`) is embedded independently and stored in a **separate FAISS index and a separate SQLite DB**.
- Vector hits trace back to source memories via `source_memory_id` in the metadata.
- The separate index keeps graph semantic retrieval from polluting the document vector space when there are far more entries than documents.
#### 3. Fusion and four-dimensional weighting
The two channels are first RRF-fused, then weighted across four dimensions (defaults 0.55 / 0.20 / 0.15 / 0.10):
```
graph_final(d) = 0.55 · rrf_norm(d)
+ 0.20 · importance(d)
+ 0.15 · recency(d)
+ 0.10 · graph_confidence(d) # extraction confidence (edge/node weight)
graph_final(d) ×= temporal_factor(atoms of d) # then multiplied by the TTL decay
# factor of that memory's atoms
```
### III. Dual-route fusion: `DualRouteRetriever`
1. The document and graph routes are each normalized to the same scale;
2. Weighted sum (defaults: document 0.65, graph 0.35):
```
score(d) = 0.65 · doc_score(d) + 0.35 · graph_score(d)
```
3. **Cross-route bonus**: a memory hit by both routes gets an extra `+0.08` (capped at 1.0) — memories corroborated by both routes are more trustworthy;
4. **Dynamic intent routing** (`dynamic_route_weighting=true`): weights are adjusted in real time from surface query features:
| Query signature | Example terms | Weight adjustment |
|---|---|---|
| Relational | who, relationship, friend, know, family, colleague… | graph route **+0.20** |
| Temporal | last time, yesterday, before, what time, which day… | graph route **+0.10** |
| Factual | what is, definition, difference, principle, why… | document route **+0.15** |
### IV. Atom retrieval: `search_atoms` (separate tool)
- Runs BM25 directly over `memory_atoms_fts`;
- Each atom's score is multiplied by its own **TTL time factor**: `exp(-age/ttl)` while unexpired, and 0 once `expired/forgotten`;
- Only `active` atoms are returned — ideal for "one specific fact, not the whole memory".
### V. Post-recall reinforcement and caching
- **Async touch**: after each recall hit, the memory's `last_access_time` and `access_count` are updated in the background without blocking the response. This feeds three mechanisms: slower document-route freshness decay, a reduced daily importance-decay rate, and atom reinforcement.
- **Search cache**: keyed by `(query, k, session_id, persona_id)`, default TTL 45 seconds, LRU 256 entries; any write/delete/edit bumps the cache epoch and invalidates the whole cache immediately.
---
## Forgetting & Lifecycle
### Document level: daily importance decay (`DecayScheduler`)
- Runs every day at `decay_hour:decay_minute` (default 00:05); state is written to `decay_state.json`, and if the machine was off at run time, startup **catches up by the number of missed days**.
- Pinned memories (`pinned`) are immune.
- **Access frequency modulates the decay rate**: the more a memory was accessed within the last 30 days (`access_decay_window_days`), the slower it decays that day:
```
effective_rate = decay_rate × max(0.3, 1 − 0.1 × min(recent_access_count, 10))
importance_new = max(0.05, importance × (1 − effective_rate))
```
- After each daily decay, the access count is multiplied by 0.5 (`access_count_decay_multiplier`), so reinforcement fades over time.
- Optional **auto-archival** (`auto_cleanup_enabled`): memories older than `cleanup_days_threshold` with importance below `cleanup_importance_threshold` are flagged `archived` — they are **not deleted**; they leave automatic recall but can still be found by an explicit `recall`.
### Atom level: the TTL state machine (`AtomLifecycleManager`)
- Runs a sweep every 24 hours (`atom_maintenance_interval_hours`);
- Atoms past their effective TTL become `expired`;
- `atom_forget_delay_days` (default 7) after expiry they are soft-deleted as `forgotten` (rows kept, removed from FTS, no longer retrieved);
- `atom_purge_delay_days` (default 30) after soft deletion they are physically purged;
- `planned` atoms use step decay: fully valid until the deadline, invalid the instant it passes — an expired plan has no "vaguely remember" value.
---
## Reliability
- **Write crash recovery (the `write_op` journal)**: each memory write first persists a `pending` journal row, advancing its status as the three stages complete. If the process crashes at any stage, a restart scans unfinished write_ops and back-fills missing vectors, atoms, or graph; entries past `write_op_max_retries` are marked failed with an alert. You never get the silent corruption of "document exists but indexes are incomplete".
- **Index consistency self-check**: at startup it compares documents / BM25 / vector counts; missing BM25 rows are **back-filled at zero cost with jieba**; missing vectors are only reported, and you can run `rebuild_index.py` when convenient (back-filling vectors needs the embedding API).
- **Atomic FAISS writes**: index changes are first written to `*.tmp` and then `os.replace`d atomically; a failed deletion rolls back, so an index file is never half-written.
- **Embedding-service degradation**: when the embedding API is unavailable (wrong key, network outage), writes do not fail — documents and BM25 indexes persist normally and only vectors are skipped; retrieval automatically falls back to BM25 + graph keyword. Once the service is back, run `python rebuild_index.py` to back-fill all vectors offline.
- **Backups**: a full backup is taken automatically when the server version changes; an online SQLite hot backup (`backup` API, no database lock) runs daily and is kept for `backup.keep_days` days.
---
## MCP Tools
By default only the high-frequency tools are placed in the model's context (to save tokens); the other tools are implemented and can be exposed via `EXPOSED_TOOLS` in `tools/memory_tools.py`.
| Tool | Public by default | Purpose |
|---|---|---|
| `memorize` | ✅ | Write one complete memory (body + topics / key facts / sentiment / importance and other metadata); atom splitting and graph construction happen internally |
| `recall` | ✅ | Hybrid retrieval of relevant memories, returning Top-K (prefers the `persona_summary`, falls back to text) |
| `list_memories` | ✅ | List memories newest-first, with session/persona filtering and pagination |
| `edit_memory` | ✅ | Edit a memory's body; automatically rebuilds the affected BM25 / graph indexes |
| `delete_memory` | ❌ | Delete a memory (cascade cleanup of documents, atoms, graph, vectors) |
| `pin_memory` | ❌ | Pin / unpin a memory (immune to decay and archival) |
| `memory_stats` | ❌ | View memory-store statistics |
| `summarize_memory` | ❌ | Ask an LLM to summarize a conversation into structured memory JSON |
| `search_atoms` | ❌ | Retrieve fine-grained memory atoms only |
| `memory_maintenance` | ❌ | Manually trigger decay / archival / atom maintenance |
---
## Configuration
See `config.example.yaml` for the full configuration. Common options:
| Option | Default | Description |
|---|---|---|
| `memory.recall_top_k` / `recall_max_k` | 5 / 15 | Default and maximum number of `recall` results |
| `memory.rrf_k` | 60 | RRF smoothing constant |
| `advanced.score_alpha/beta/gamma` | 0.5/0.25/0.25 | Document-route relevance / importance / recency weights |
| `advanced.graph_score_alpha/beta/gamma/delta` | 0.55/0.2/0.15/0.1 | Graph-route four-dimension weights |
| `advanced.document_route_weight` / `graph_route_weight` | 0.65 / 0.35 | Dual-route fusion weights |
| `advanced.cross_route_bonus` | 0.08 | Bonus when both routes hit the same memory |
| `advanced.dynamic_route_weighting` | true | Dynamic relational/temporal/factual intent weighting |
| `advanced.mmr_lambda` | 0.7 | MMR relevance–diversity balance |
| `advanced.graph_expansion_hops` | 1 | Graph neighbor expansion hops (1 or 2) |
| `memory.decay_rate` | 0.05 | Daily document-importance decay rate (the example config uses 0.01) |
| `memory.atom_enabled` / `graph_memory_enabled` | true | Atom-layer / graph-route switches |
| `embedding.dimensions` | 4096 | **Must be changed when you switch embedding models** |
---
## Maintenance Scripts & Config Panel
| File | Purpose |
|---|---|
| `config_panel.py` (`配置面板.bat`) | Local web config panel (127.0.0.1:8765) for grouped parameter tuning and memory-store statistics |
| `rebuild_index.py` | Fully rebuild the FAISS vector indexes (after switching embedding models / if an index is corrupted) |
| `repair_associations.py` | Scan and repair missing memory atoms and graph entries; by default it only scans and reports — pass `--apply` to actually repair |
| `start.bat` | Launch the MCP server |
---
## Tech Stack
- **MCP**: Model Context Protocol (stdio transport)
- **SQLite + FTS5**: structured storage and BM25 full-text retrieval (async access via `aiosqlite`)
- **FAISS**: `IndexIDMap2(IndexFlatIP)` vector index (`faiss-cpu`)
- **jieba**: Chinese segmentation, custom dictionary, and stop words
- **httpx**: async calls to OpenAI-compatible embedding / chat endpoints
- Pure-Python async implementation (`asyncio`); no external service dependencies; all data stays local
---
## Design Origins
The storage structure, memory-atom model, and hybrid-retrieval architecture are aligned with the design of the LivingMemory plugin from the AstrBot ecosystem, re-architected here as a **standalone MCP server**: it depends on no bot framework, and any MCP-capable client can connect. Graph construction and atom classification remain deterministic-rule implementations, and the graph and documents use independent vector indexes and SQLite databases.
---
## License
[MIT License](LICENSE) © 2026 pangak959
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues