Skip to main content
Glama
pangak959

Aokun Memory MCP Server

by pangak959

Aokun Memory MCP Server

English | 简体中文

A persistent long-term memory MCP server for AI assistants. It distills conversations into three retrievable layers of memory (documents / memory atoms / a knowledge graph), recalls them through a multi-channel hybrid of BM25 + vector + graph retrieval, and adds human-like forgetting, atom lifecycles, and crash recovery — so any MCP-capable client gets long-term memory that can fade, strengthen, and evolve.

License: MIT Python MCP

The system is optimized for Chinese text (word segmentation uses jieba); the architecture and all algorithms are language-agnostic.


Table of Contents


Related MCP server: Memsolus MCP Server

Features

  • Three-layer memory model

    • Document memory: a complete, structured memory distilled from a conversation (summary / topics / key facts / sentiment / importance)

    • Memory atoms: key facts broken into fine-grained facts, each with a type, a shelf life (TTL), and its own decay curve

    • Knowledge graph: people, topics, and facts become nodes; describes / mentioned_in / co_occurs_with become edges — every node and edge maps back to its source memory

  • Parallel multi-channel hybrid retrieval: SQLite FTS5 full-text (BM25), FAISS vectors (cosine), graph keyword and graph-vector channels, all fused with RRF (Reciprocal Rank Fusion)

  • Dual-route retrieval + dynamic routing: the document route and graph route are fused with weights that are adjusted in real time by query intent (relational / temporal / factual)

  • MMR diversity re-ranking: prevents the top results from being near-duplicate phrasings of the same topic

  • Human-like forgetting: daily importance decay, archival of low-value old memories, and per-type atom TTLs with a full state machine; memories that are recalled repeatedly decay more slowly (access reinforcement)

  • Zero extra tokens for structuring: atom classification and graph construction are entirely deterministic rules (regex + jieba segmentation) — no extra LLM calls

  • Crash recoverable: every write is tracked by a write_op journal; if the process dies at any stage, a restart back-fills the missing indexes / atoms / graph

  • Local web config panel: double-click a .bat to tune parameters in a browser, no manual YAML editing required


Architecture

┌──────────────────────────────────────────────────────────────────────┐
│                  MCP client (Claude Code / any MCP host)              │
└───────────────────────────────┬──────────────────────────────────────┘
                                │  stdio / MCP
┌───────────────────────────────▼──────────────────────────────────────┐
│  server.py                        tools/memory_tools.py               │
│  MCP server entrypoint            registers memorize/recall/list/edit │
└───────────────────────────────┬──────────────────────────────────────┘
                                │
┌───────────────────────────────▼──────────────────────────────────────┐
│                     core/memory_engine.py (orchestrator)              │
│  add_memory() write pipeline      search_memories() recall pipeline   │
│  write_op crash-recovery journal  search cache (TTL + LRU + epoch)    │
└───────┬────────────────┬──────────────────┬───────────────┬──────────┘
        │                │                  │               │
┌───────▼──────┐ ┌───────▼────────┐ ┌───────▼────────┐ ┌────▼─────────┐
│ processors/  │ │ summarizer.py  │ │ embedding.py   │ │ retrieval/   │
│ segmentation │ │ LLM chat       │ │ OpenAI-compat  │ │ dual-route   │
│ atom rules   │ │ summarization  │ │ embedding client│ │ hybrid search│
│ graph rules  │ │ (LLM used here │ │ L2 normalized  │ │ (see below)  │
│ entity resolve│ │  only)         │ │                │ │              │
└───────┬──────┘ └────────────────┘ └────────────────┘ └────┬─────────┘
        │                                                    │
┌───────▼────────────────────────────────────────────────────▼─────────┐
│                          storage/ (persistence)                        │
│        DocumentStore   AtomStore   GraphStore (SQLite / FTS5)          │
└───────┬───────────────────────────┬──────────────────────────────────┘
        │                           │
┌───────▼─────────────────┐ ┌───────▼─────────────────────────────────┐
│ Main DB aokun_memory.db │ │ Graph DB aokun_memory_graph_documents    │
│  · documents            │ │  · graph entry text + vector metadata   │
│  · memory_atoms         │ └───────┬─────────────────────────────────┘
│  · graph_nodes/edges/   │         │
│    entries              │ ┌───────▼─────────────────────────────────┐
│  · 3 FTS5 virtual tables│ │ FAISS index files (under data/)          │
│  · memory_write_ops     │ │  · faiss_index (document vectors)        │
│    (recovery journal)   │ │  · graph_faiss_index (graph-entry vecs)  │
└─────────────────────────┘ └──────────────────────────────────────────┘

  Sidecars: decay_scheduler.py (daily decay) · atom_lifecycle_manager.py (TTL state machine)
            backup_manager.py (version + daily hot backups) · config_panel.py (local web UI)

Module responsibilities

Path

Responsibility

server.py

MCP stdio server entrypoint; wires the engine and tools together

tools/memory_tools.py

Exposed MCP tool definitions and parameter validation

core/memory_engine.py

Orchestrator: write pipeline, recall pipeline, decay, cache, crash recovery

core/config.py

YAML config loading and dotted-path access (with defaults); anchors relative paths to the config directory

core/embedding.py

OpenAI-compatible /embeddings client, batch embedding + L2 normalization

core/summarizer.py

Calls an LLM to summarize a conversation into structured memory JSON

core/processors/text_processor.py

jieba segmentation, custom words, stop-word management

core/processors/atom_classifier.py

Pure-rule classification of key facts into memory atoms

core/processors/graph_extractor.py

Pure-rule extraction of graph nodes / edges / entries from metadata

core/processors/entity_resolver.py

Entity-name canonicalization (alias merging)

core/retrieval/

Retrievers: BM25, vector, hybrid, RRF, three graph channels, dual-route, atom

core/decay_scheduler.py

Daily scheduled document-importance decay (catches up after downtime)

core/managers/atom_lifecycle_manager.py

Atom expiry / forgetting / physical-purge state machine

core/managers/graph_memory_manager.py

Graph write orchestration (delete old → extract → persist → embed)

core/managers/backup_manager.py

Full backup on version change + daily online SQLite backup

storage/

All SQLite / FTS5 / FAISS read-and-write wrappers


Quick Start

1. Requirements

  • Python 3.10–3.14 (64-bit; 3.11 / 3.12 recommended)

  • Two OpenAI-compatible APIs:

    • Embedding API (defaults to SiliconFlow's Qwen/Qwen3-Embedding-8B, 4096 dimensions)

    • Chat/summarization API (defaults to DeepSeek; used only by the summarize_memory tool or your own external conversation pipeline — everyday memorize / recall consume no LLM tokens)

2. Install

git clone https://github.com/pangak959/Aokun_memory.git
cd Aokun_memory
pip install -r requirements.txt

3. Configure

cp config.example.yaml config.yaml       # on Windows, copy and rename the file manually

Edit config.yaml and fill in at least the two api_key values. Database and index paths default to the data/ directory inside the project and are created automatically. Relative paths are always resolved against the directory containing config.yaml, so it does not matter which working directory the MCP client launches the server from.

4. Connect an MCP client

See mcp_config.example.json and replace the path with the real path on your machine:

{
  "mcpServers": {
    "aokun_memory": {
      "command": "python",
      "args": ["C:\\path\\to\\Aokun_memory\\server.py"],
      "env": {}
    }
  }
}

The MCP server is launched automatically by the client when needed — you do not start it yourself. To sanity-check initialization, run python server.py from the project directory (or double-click start.bat); seeing Aokun Memory MCP Server starting... means the config and databases initialized successfully. The process then waits on stdio, so just close the window.

5. Tune parameters (optional)

Double-click 配置面板.bat (or run python config_panel.py); your browser opens http://127.0.0.1:8765, where you can adjust recall counts, fusion weights, decay rates, and more, grouped by function. Save and reconnect the MCP client for changes to take effect.


Memory Model

The system maintains three granularities of memory at once; they are generated together on write and complement each other at recall time.

1. Document memory (documents)

A complete memory produced by summarizing one conversation — one row in the documents table:

Field

Description

text

Memory body (the code prepends an 【原始时间 / original time: …】 marker)

metadata.topics

List of topic terms

metadata.participants

Participants (people)

metadata.key_facts

Key facts (the raw material for memory atoms)

metadata.sentiment

positive / neutral / negative

metadata.importance

0.0–1.0; used by both decay and ranking

metadata.pinned

Pinned memories are immune to decay and archival

metadata.session_id / persona_id

Session and persona isolation

metadata.access_count / last_access_time

Inputs to access reinforcement

2. Memory atoms (memory_atoms)

Key facts are split into fine-grained facts with shelf lives, stored, retrieved, and expired independently.

  • Six atom types (decided by atom_classifier.py with regex rules — zero LLM calls):

Type

Meaning

Default TTL

Decay curve

episodic

Episodes: things that happened

7 days

exponential

planned

Plans: things to do in the future (with Chinese time-expression parsing)

2 days

step (no decay until the deadline, then instant expiry)

factual

Factual knowledge

180 days

exponential

relational

Inter-person relationships

90 days

linear

preference

Preferences / likes and dislikes

60 days

exponential

unknown

Fallback for unclassified facts

30 days

linear

  • The effective TTL is modulated by importance and reinforcement count:

ttl = base_ttl × (0.5 + importance) × (1 + 0.5 × reinforcement_count)
  • Atom state machine: active → (on expiry) expired → (soft-deleted after 7 days, removed from FTS) forgotten → (physically deleted after 30 days); there are also dormant and superseded states.

  • Automatic reinforcement: when a new atom has a bag-of-words Jaccard similarity ≥ 0.6 with an existing atom, it is treated as the same fact being mentioned again — the old atom gets reinforcement_count +1 and its confidence and TTL are refreshed. Facts that recur live longer.

3. Knowledge graph (graph_nodes / graph_edges / graph_entries)

The graph is also built deterministically, without an LLM (graph_extractor.py):

  • Nodes (GraphNode), keyed by node_type:node_value:canonical_value:

    • topic: from topics

    • person: from participants

    • fact: from each key_fact / each atom

  • Edges (GraphEdge), three relation types:

    • topic/person —describes→ fact: which person/topic a fact is about

    • fact —mentioned_in→ fact: co-occurrence of facts within one memory

    • topic/person pairs —co_occurs_with: associative co-occurrence

  • Entries (GraphEntry): both nodes and edges produce a searchable text (e.g. "Fact: Alice likes iced americano. Summary: <memory body>"), and each entry maps back to its source document via source_memory_id. This is the key design that lets the graph participate in full-text and vector retrieval.

  • Entity resolution: EntityResolver.canonicalize() normalizes case, whitespace, and common aliases so the same person/topic never splits into multiple nodes.

  • Edges carry confidence (inherited from atom-classification confidence) and weight, both used by the graph-route score.

4. On-disk layout

Storage

Contents

data/aokun_memory.db (SQLite)

documents, memory_atoms, graph_nodes, graph_edges, graph_entries, 3 FTS5 virtual tables, and the memory_write_ops recovery journal

data/aokun_memory_graph_documents.db (SQLite)

Graph-entry text and vector metadata

data/faiss_index

Document vectors (IndexIDMap2(IndexFlatIP))

data/graph_faiss_index

Graph-entry vectors (same structure)

All vectors are L2-normalized and indexed with inner product, so the inner-product score equals cosine similarity.


Write Pipeline

The memorize tool calls MemoryEngine.add_memory(). The whole pipeline is tied together by one write_op journal entry, making it crash-recoverable:

memorize(content, metadata)
   │
   ├─ 1. classify_atoms(metadata.key_facts)     rules only: key facts → atoms (type/TTL/confidence)
   │
   ├─ 2. write write_op journal (status=pending)      ── crash-recovery anchor
   │
   ├─ 3. Document layer (docs stage)
   │     ├─ insert into documents, obtain integer doc_id
   │     ├─ jieba segmentation → write FTS5 full-text index (BM25)
   │     └─ call embedding API → L2 normalize → write FAISS
   │        (a vector failure does not block: the document is already persisted;
   │         rebuild_index can back-fill later)
   │     → write_op advances to docs_done
   │
   ├─ 4. Atom layer (atoms stage)
   │     ├─ batch-insert memory_atoms + memory_atoms_fts
   │     ├─ on failure, fall back to row-by-row inserts; bad rows are flagged
   │     │     needs_repair instead of rolling back the whole batch
   │     └─ same-fact auto reinforcement (Jaccard ≥ 0.6)
   │     → write_op advances to atoms_done
   │
   ├─ 5. Graph layer (graph stage)
   │     ├─ delete this memory's previous graph output (supports rebuilding after an edit)
   │     ├─ GraphExtractor rule-extracts nodes / edges / entries
   │     ├─ EntityResolver canonicalizes, then persist graph DB and graph FTS
   │     └─ batch-embed entries → write graph vector DB, write back vector_doc_id
   │     → write_op advances to graph_done → done
   │
   └─ 6. search-cache epoch +1 (a search immediately after a write never hits stale cache)

Retrieval Pipeline (core)

The recall tool calls MemoryEngine.search_memories(). When the graph is enabled it runs dual-route retrieval (DualRouteRetriever), and each route is itself a multi-channel hybrid.

                                 query
                                   │
                     search-cache hit? (TTL 45s / LRU 256)
                    ┌──────────────┴────────────────┐
                   yes                              no
                    │                               │
             return directly             ┌──────────▼──────────┐
                                          │ DualRouteRetriever │
                                          └──────┬──────┬──────┘
                                   doc route 0.65 ◄┴──────┴► 0.35 graph route
                          ┌─────────────────────┐   ┌──────────────────────────┐
                          │  HybridRetriever    │   │      GraphRetriever       │
                          │                     │   │                           │
                          │ BM25 ch.  Vector ch.│   │ graph-keyword  graph-vec  │
                          │ (FTS5)    (FAISS)   │   │ (FTS+traversal) (FAISS)   │
                          └────┬────────┬───────┘   └──────┬───────────┬────────┘
                               │  RRF fuse (k=60) │         │  RRF fuse (k=60)      │
                               ▼                  ▼         ▼                       ▼
                     3-dim weighted rank + MMR dedup    4-dim weighting (confidence × atom TTL)
                          └────────────────┬────────────┴───────────────┬─────────┘
                                           │  cross-route bonus + dynamic intent weighting
                                           ▼                                      │
                                    fused ranking, take Top-K                       │
                                           │                                       │
                  async touch: refresh last_access_time / access_count (reinforce)  │
                                           ▼                                       │
                          recall returns (persona_summary preferred, text fallback)

I. Document route: HybridRetriever

1. BM25 channel (BM25Retriever)

  • Both queries and documents go through jieba segmentation + stop-word removal, are joined with spaces, and written into a SQLite FTS5 virtual table.

  • Query terms are joined with OR (favor high recall; leave precision to the fusion layer). FTS5's built-in bm25() gives the raw score, which a sigmoid maps into (0,1):

bm25_norm = 1 / (1 + exp(-raw_bm25 / 8))
  • Session/persona filtering happens in Python (hence the vector channel pre-fetches k×10 candidates before filtering).

2. Vector channel (VectorRetriever)

  • The query is embedded, L2-normalized, and searched with FAISS IndexFlatIP inner product — inner product is cosine similarity — then linearly mapped into (0,1):

vec_score = (cosine + 1) / 2

3. Parallelism, degradation, and RRF fusion

  • The two channels run in parallel with asyncio.gather.

  • If either channel fails and fallback_enabled=true, retrieval automatically degrades to the surviving channel without interruption.

  • The two rankings are fused with RRF (Reciprocal Rank Fusion, Cormack et al., 2009), which looks only at ranks and is immune to differing score scales:

RRF(d) = Σ_channels  1 / (k + rank(d) + 1)        # k = rrf_k, paper recommends 60

4. Three-dimensional weighting

The normalized RRF score captures relevance only; the final score layers on importance and recency (default weights 0.5 / 0.25 / 0.25):

final(d) = α · rrf_norm(d)                         # α = score_alpha = 0.5
         + β · importance(d)                       # β = score_beta  = 0.25
         + γ · recency(d)                          # γ = score_gamma = 0.25

recency(d) = exp( - age_days(d) / 30 )
age_days   = now - max(create_time, last_access_time)

Note that age uses the later of creation time and last-access time — an old memory that was recalled recently gets its freshness pulled back up. This is the first mechanism behind "the more you use it, the stronger it gets".

5. MMR diversity re-ranking

After fusion, MMR (Maximal Marginal Relevance) iteratively selects results, with bag-of-words Jaccard as the similarity measure:

MMR = λ · relevance(d) − (1 − λ) · max_sim(d, already-selected)      # λ = mmr_lambda = 0.7

Higher λ favors relevance; lower λ favors topic diversity, keeping the top-K from being duplicate phrasings of one event.

II. Graph route: GraphRetriever

The graph route is itself a "keyword + vector" hybrid, and can expand across graph structure.

1. Graph-keyword channel (GraphKeywordRetriever)

After segmentation, the query hits graph entries through four entry points, all traced back to source_memory_id:

Entry point

Match method

Weight

Entry full text

Direct FTS5 BM25 hit on graph_entries

1.0

Node match

Query terms hit a node token (asking about a specific person/topic)

0.7

One-hop neighbors

Expand from a hit node along edges to adjacent entries

0.7

Two-hop neighbors (optional, graph_expansion_hops=2)

Neighbors of neighbors

0.4

When one source memory is hit through multiple entry points, scores aggregate (each extra entry point adds a ×0.35 boost); expansion is capped by graph_expansion_limit (default 24).

This layer answers questions BM25/vector are bad at — e.g. "what is the relationship between Alice and Bob?". Through structures like Alice ─ co_occurs_with ─ Bob, memories that never literally contain the word "relationship" are still recalled.

2. Graph-vector channel (GraphVectorRetriever)

  • Graph-entry text ("Fact: … Summary: …") is embedded independently and stored in a separate FAISS index and a separate SQLite DB.

  • Vector hits trace back to source memories via source_memory_id in the metadata.

  • The separate index keeps graph semantic retrieval from polluting the document vector space when there are far more entries than documents.

3. Fusion and four-dimensional weighting

The two channels are first RRF-fused, then weighted across four dimensions (defaults 0.55 / 0.20 / 0.15 / 0.10):

graph_final(d) = 0.55 · rrf_norm(d)
               + 0.20 · importance(d)
               + 0.15 · recency(d)
               + 0.10 · graph_confidence(d)      # extraction confidence (edge/node weight)
graph_final(d) ×= temporal_factor(atoms of d)    # then multiplied by the TTL decay
                                                 #     factor of that memory's atoms

III. Dual-route fusion: DualRouteRetriever

  1. The document and graph routes are each normalized to the same scale;

  2. Weighted sum (defaults: document 0.65, graph 0.35):

score(d) = 0.65 · doc_score(d) + 0.35 · graph_score(d)
  1. Cross-route bonus: a memory hit by both routes gets an extra +0.08 (capped at 1.0) — memories corroborated by both routes are more trustworthy;

  2. Dynamic intent routing (dynamic_route_weighting=true): weights are adjusted in real time from surface query features:

Query signature

Example terms

Weight adjustment

Relational

who, relationship, friend, know, family, colleague…

graph route +0.20

Temporal

last time, yesterday, before, what time, which day…

graph route +0.10

Factual

what is, definition, difference, principle, why…

document route +0.15

IV. Atom retrieval: search_atoms (separate tool)

  • Runs BM25 directly over memory_atoms_fts;

  • Each atom's score is multiplied by its own TTL time factor: exp(-age/ttl) while unexpired, and 0 once expired/forgotten;

  • Only active atoms are returned — ideal for "one specific fact, not the whole memory".

V. Post-recall reinforcement and caching

  • Async touch: after each recall hit, the memory's last_access_time and access_count are updated in the background without blocking the response. This feeds three mechanisms: slower document-route freshness decay, a reduced daily importance-decay rate, and atom reinforcement.

  • Search cache: keyed by (query, k, session_id, persona_id), default TTL 45 seconds, LRU 256 entries; any write/delete/edit bumps the cache epoch and invalidates the whole cache immediately.


Forgetting & Lifecycle

Document level: daily importance decay (DecayScheduler)

  • Runs every day at decay_hour:decay_minute (default 00:05); state is written to decay_state.json, and if the machine was off at run time, startup catches up by the number of missed days.

  • Pinned memories (pinned) are immune.

  • Access frequency modulates the decay rate: the more a memory was accessed within the last 30 days (access_decay_window_days), the slower it decays that day:

effective_rate  = decay_rate × max(0.3, 1 − 0.1 × min(recent_access_count, 10))
importance_new  = max(0.05, importance × (1 − effective_rate))
  • After each daily decay, the access count is multiplied by 0.5 (access_count_decay_multiplier), so reinforcement fades over time.

  • Optional auto-archival (auto_cleanup_enabled): memories older than cleanup_days_threshold with importance below cleanup_importance_threshold are flagged archived — they are not deleted; they leave automatic recall but can still be found by an explicit recall.

Atom level: the TTL state machine (AtomLifecycleManager)

  • Runs a sweep every 24 hours (atom_maintenance_interval_hours);

  • Atoms past their effective TTL become expired;

  • atom_forget_delay_days (default 7) after expiry they are soft-deleted as forgotten (rows kept, removed from FTS, no longer retrieved);

  • atom_purge_delay_days (default 30) after soft deletion they are physically purged;

  • planned atoms use step decay: fully valid until the deadline, invalid the instant it passes — an expired plan has no "vaguely remember" value.


Reliability

  • Write crash recovery (the write_op journal): each memory write first persists a pending journal row, advancing its status as the three stages complete. If the process crashes at any stage, a restart scans unfinished write_ops and back-fills missing vectors, atoms, or graph; entries past write_op_max_retries are marked failed with an alert. You never get the silent corruption of "document exists but indexes are incomplete".

  • Index consistency self-check: at startup it compares documents / BM25 / vector counts; missing BM25 rows are back-filled at zero cost with jieba; missing vectors are only reported, and you can run rebuild_index.py when convenient (back-filling vectors needs the embedding API).

  • Atomic FAISS writes: index changes are first written to *.tmp and then os.replaced atomically; a failed deletion rolls back, so an index file is never half-written.

  • Embedding-service degradation: when the embedding API is unavailable (wrong key, network outage), writes do not fail — documents and BM25 indexes persist normally and only vectors are skipped; retrieval automatically falls back to BM25 + graph keyword. Once the service is back, run python rebuild_index.py to back-fill all vectors offline.

  • Backups: a full backup is taken automatically when the server version changes; an online SQLite hot backup (backup API, no database lock) runs daily and is kept for backup.keep_days days.


MCP Tools

By default only the high-frequency tools are placed in the model's context (to save tokens); the other tools are implemented and can be exposed via EXPOSED_TOOLS in tools/memory_tools.py.

Tool

Public by default

Purpose

memorize

✅

Write one complete memory (body + topics / key facts / sentiment / importance and other metadata); atom splitting and graph construction happen internally

recall

✅

Hybrid retrieval of relevant memories, returning Top-K (prefers the persona_summary, falls back to text)

list_memories

✅

List memories newest-first, with session/persona filtering and pagination

edit_memory

✅

Edit a memory's body; automatically rebuilds the affected BM25 / graph indexes

delete_memory

❌

Delete a memory (cascade cleanup of documents, atoms, graph, vectors)

pin_memory

❌

Pin / unpin a memory (immune to decay and archival)

memory_stats

❌

View memory-store statistics

summarize_memory

❌

Ask an LLM to summarize a conversation into structured memory JSON

search_atoms

❌

Retrieve fine-grained memory atoms only

memory_maintenance

❌

Manually trigger decay / archival / atom maintenance


Configuration

See config.example.yaml for the full configuration. Common options:

Option

Default

Description

memory.recall_top_k / recall_max_k

5 / 15

Default and maximum number of recall results

memory.rrf_k

60

RRF smoothing constant

advanced.score_alpha/beta/gamma

0.5/0.25/0.25

Document-route relevance / importance / recency weights

advanced.graph_score_alpha/beta/gamma/delta

0.55/0.2/0.15/0.1

Graph-route four-dimension weights

advanced.document_route_weight / graph_route_weight

0.65 / 0.35

Dual-route fusion weights

advanced.cross_route_bonus

0.08

Bonus when both routes hit the same memory

advanced.dynamic_route_weighting

true

Dynamic relational/temporal/factual intent weighting

advanced.mmr_lambda

0.7

MMR relevance–diversity balance

advanced.graph_expansion_hops

1

Graph neighbor expansion hops (1 or 2)

memory.decay_rate

0.05

Daily document-importance decay rate (the example config uses 0.01)

memory.atom_enabled / graph_memory_enabled

true

Atom-layer / graph-route switches

embedding.dimensions

4096

Must be changed when you switch embedding models


Maintenance Scripts & Config Panel

File

Purpose

config_panel.py (配置面板.bat)

Local web config panel (127.0.0.1:8765) for grouped parameter tuning and memory-store statistics

rebuild_index.py

Fully rebuild the FAISS vector indexes (after switching embedding models / if an index is corrupted)

repair_associations.py

Scan and repair missing memory atoms and graph entries; by default it only scans and reports — pass --apply to actually repair

start.bat

Launch the MCP server


Tech Stack

  • MCP: Model Context Protocol (stdio transport)

  • SQLite + FTS5: structured storage and BM25 full-text retrieval (async access via aiosqlite)

  • FAISS: IndexIDMap2(IndexFlatIP) vector index (faiss-cpu)

  • jieba: Chinese segmentation, custom dictionary, and stop words

  • httpx: async calls to OpenAI-compatible embedding / chat endpoints

  • Pure-Python async implementation (asyncio); no external service dependencies; all data stays local


Design Origins

The storage structure, memory-atom model, and hybrid-retrieval architecture are aligned with the design of the LivingMemory plugin from the AstrBot ecosystem, re-architected here as a standalone MCP server: it depends on no bot framework, and any MCP-capable client can connect. Graph construction and atom classification remain deterministic-rule implementations, and the graph and documents use independent vector indexes and SQLite databases.


License

MIT License © 2026 pangak959

Related MCP Connectors

Related MCP Servers