Skip to main content
Glama
udjin-labs
by udjin-labs

mnemostack

PyPI Python versions CI License: Apache 2.0

Self-hosted hybrid memory & retrieval for AI apps.

mnemostack is a durable retrieval layer over your own Qdrant (and optional Memgraph): semantic, keyword (BM25), temporal, and graph recall, fused with Reciprocal Rank Fusion and refined by an 8-stage ranking pipeline — with payload filters for multi-tenant isolation, optional LLM answer synthesis (confidence + citations), and an ingest path that enriches and projects structured fields. One recall(query) call, usable as a Python library, an HTTP service, or an MCP server.

Flagship use case — durable memory for AI agents. Long-running agents hit the same wall: context gets compacted, sessions restart, useful decisions disappear, and the next run pays the re-orientation tax again. mnemostack gives them a persistent memory layer to query when the context window is not enough — durable, searchable, scoped, and explainable, not just embedded and hoped for.

The same engine backs other retrieval-heavy work: RAG over mixed corpora, multi-tenant or per-user knowledge stores, and time-aware search backends — anywhere pure vector similarity falls short on its own.

Status: Actively developed — public API is stable; new functionality lands additively in minor releases. Breaking changes are rare and called out in CHANGELOG.md.

Quickstart: agent memory over MCP

The fastest on-ramp is the MCP server — it gives Claude Desktop, Claude Code, Cursor, ChatGPT, or another MCP-capable agent durable memory in a few commands. Building an app instead of wiring up an agent? Use the HTTP API or the Python library over the same collection.

1. Install

pip install 'mnemostack[mcp]'

Run a local Qdrant for the vector store:

# tag rule: match your installed qdrant-client - same major, minor within 1
docker run -p 6333:6333 qdrant/qdrant:v1.18.3

Optional: run Memgraph for graph-backed memory:

docker run -p 7687:7687 memgraph/memgraph:latest

2. Start the MCP server

export GEMINI_API_KEY=your-key-here
mnemostack mcp-serve --provider gemini --collection my-memory

Claude Desktop config example:

{
  "mcpServers": {
    "mnemostack": {
      "command": "mnemostack",
      "args": ["mcp-serve", "--provider", "gemini", "--collection", "my-memory"],
      "env": {
        "GEMINI_API_KEY": "your-key-here"
      }
    }
  }
}

Claude will then be able to call mnemostack_search, mnemostack_answer, and graph tools.

3. Store memory

Index a folder of notes, docs, transcripts, or project context:

mnemostack index ./my-notes/ --provider gemini --collection my-memory --recreate

--recreate drops the existing collection, so it asks for confirmation first; pass --yes to skip the prompt (required in scripts/CI — non-interactive runs without it exit with code 2).

For a running app or assistant, use the streaming Ingestor API shown below to store messages as they arrive.

4. Recall memory

From an agent, ask a memory-style question and let the MCP tools retrieve the right facts.

From the shell, test the same collection directly:

mnemostack search "what did we decide about auth" --provider gemini --collection my-memory
mnemostack answer "what did we decide about auth" --provider gemini --collection my-memory

Related MCP server: Second Brain

Why hybrid memory?

Vector search answers "what sounds similar?" Real retrieval over a growing corpus needs to answer "what actually matters for this query?" — which takes exact matches, semantic similarity, relationship tracking, recency, user/project scope, and feedback from past recalls. Mnemostack uses hybrid retrieval so recall is reliable instead of embedding roulette. (Agent memory is the most demanding version of this problem, which is why it's the flagship use case.)

Use cases

Agent and chatbot memory (flagship):

  • Long-running coding agents that need to survive compaction and session restarts.

  • Chat assistants and conversational bots that remember a user's earlier messages, preferences, and decisions across sessions — scope each user's memory with payload filters so one user never sees another's history.

  • Personal assistant memory for preferences, recurring tasks, and long-horizon context.

  • Multi-agent context sharing through one durable memory backend.

  • Session compaction recovery when the useful details no longer fit in the prompt.

Beyond agents — the same engine as a retrieval backend:

  • A searchable knowledge base in memory: ingest your docs/notes/FAQ once, then serve hybrid search or grounded answer() (with confidence and source citations) over it from the CLI, HTTP, or library.

  • RAG over mixed corpora (code, docs, transcripts) where exact-token and temporal recall beat pure vector similarity.

  • Multi-tenant or per-user knowledge stores — payload filters isolate each tenant's data inside every retriever, or enable service-key auth (serve --auth) for a hard, key-resolved tenant boundary with optional per-tenant quotas (see the HTTP API).

  • Time-aware search and knowledge bases — "what changed last week", point-in-time graph facts, freshness-weighted ranking.

  • Team or project knowledge recall across docs, notes, tickets, and chat history.

Three ways to use Mnemostack

  1. MCP server — for agent users. Start mnemostack mcp-serve, connect your agent, and use memory tools from the chat/runtime you already use.

  2. HTTP API — for app developers. Run mnemostack serve and call /recall, /answer, /feedback, the memory write/lifecycle endpoints (/memories, /invalidate, /triples), /health, /metrics, or /docs from any language.

  3. Python SDK — for library users. Compose retrievers, stores, rerankers, graph tools, and the streaming ingest API inside your own Python application.

Architecture

Mental model

Think of it as a storage hierarchy for agent memory:

  • Context window = RAM. Fast, limited (typically 100K–200K tokens for many agent models; larger windows exist, but usable working context is often much smaller after tools, instructions, MCP output, and context rot — ~45K usable tokens on a 200K window is a realistic working number in long-running agent sessions. Clears on session restart.

  • mnemostack corpus = Disk. Persistent, searchable, grows forever — every fact the agent has ever seen, queryable on demand.

  • recall(query) = page fault handler. When the agent needs something that isn't in the current context, it pulls the exact fact from storage with a single hybrid query — not a grep, not a reload of the whole corpus.

The practical effect: you stop re-explaining your project to the agent after every /compact. You stop losing momentum to the re-orientation tax that shows up in any agent with session compaction. mnemostack solves it at the library level, not tied to any single agent runtime.

How it works, in one paragraph

On each recall(query): the configured retrievers (Vector and Temporal by default, with BM25 and Memgraph when configured) run in parallel and return ranked lists. Reciprocal Rank Fusion merges them. The optional 8-stage pipeline can reweight results using query classification, exact-token rescue, gravity/hub dampening, freshness, inhibition-of-return, curiosity boosts, Q-learning weights supplied through its state store, and graph resurrection. An optional LLM reranker does a final ordering pass. You get a list of RecallResult with source, score, and provenance — ready to hand to a model. The list order is authoritative: score has no single scale — many stages and fallback paths write it, and a rerank changes the order without rewriting the numbers — so re-sorting by them undoes it. It is not a similarity, not a confidence, and not comparable across queries; see what score is not.

mnemostack architecture

Where mnemostack fits

Most memory tools in the agent ecosystem pick one axis and optimize for it: simple vector similarity for RAG, framework-bound memory tied to a specific agent library, platform-level runtimes with audit and compliance features, or CLI wrappers over a single vendor's session store. Each makes sense for its scope.

mnemostack takes a different slice: it is a recall quality layer, offered as a plain Python package. Four retrievers (Vector + BM25 + Memgraph + Temporal), RRF fusion, an 8-stage pipeline, and an optional LLM reranker — composed to handle mixed workloads on the same corpus: exact-token lookups, semantic queries, temporal questions, and multi-hop reasoning, without forcing you to choose one mode over another.

We are not a replacement for your agent framework and not a full platform runtime. We are the piece that actually finds the right fact in a growing corpus. Drop mnemostack into your own Python agent or application, or let a higher-level service call recall() over a plain function boundary. The retrievers, pipeline, and reranker are individually composable — take only the parts you need.

Design

See ARCHITECTURE.md for detailed design: pipeline stages, Qdrant schema, Memgraph temporal model, consolidation runtime, MCP tools.

Storage, index, and retrieval layers

  • Storage/index layer: Qdrant stores vector points and payloads; BM25 indexes exact-token corpora; Memgraph stores temporal graph facts; the Temporal retriever handles time-aware vector recall.

  • Fusion layer: Reciprocal Rank Fusion merges ranked lists from Vector, BM25, Memgraph, and Temporal retrievers, with optional static or adaptive weights.

  • Recall pipeline: the 8-stage pipeline can classify the query, rescue exact tokens, dampen gravity/hubs, blend freshness, apply inhibition-of-return, add curiosity boosts, use Q-learning state, and resurrect graph-linked memories.

  • Feedback loop: HTTP and MCP recall can apply existing state; explicit /feedback or mnemostack feedback updates usefulness signals without silently training on every response.

  • Inference layer: optional LLM reranking and answer generation sit on top of recall, so retrieval still works when the LLM is unavailable.

Pipeline state

The 8-stage pipeline can use a small state store between calls (Q-learning weights, inhibition-of-return history, per-document gravity/hub counters). FileStateStore(path) persists it to a JSON file. HTTP recall applies existing state and can record inhibition-of-return exposure with --auto-record-ior; Q-learning updates only through explicit /feedback calls. CLI/MCP recall still apply existing state but do not collect feedback automatically. For deterministic benchmarks, call build_full_pipeline(enable_stateful_stages=False) so IoR/Q-learning/curiosity state cannot affect scores. For multi-process servers, implement your own StateStore (three methods: get(), set(), update()) backed by Redis or your database.

Graceful degradation

Any retriever can fail (Memgraph down, Qdrant unreachable, BM25 corpus empty). Recaller logs and continues with the remaining sources. The LLM reranker is wrapped in try/except by convention — if the LLM is rate-limited, the pre-rerank order is returned. This is deliberate: a memory stack that goes dark because one component hiccuped is worse than a slightly degraded one.

One exception: query expansion runs before retrieval, so a misconfigured expansion step (query_expansion=True without an expansion_llm, or a provider error inside it) surfaces as an error instead of degrading silently — see ARCHITECTURE.md for the full fail-open contract. Degradations themselves are visible, not silent: every HTTP/MCP response carries degraded tags, and the full per-retriever trace is available opt-in via include_trace.

Comparison and benchmarks

On LoCoMo, Mnemostack reaches 82.9% strict accuracy in our evaluation setup. The table below includes our baseline runs and externally reported numbers for context. Results depend on dataset version, configuration, judge model, scoring rules, and query type. Treat externally reported numbers as directional unless they were run with the same harness and settings.

Benchmarks

Full LoCoMo runs use the official SNAP-Research dataset (10 samples / 1986 QA) from a clean state. Across the tables below: Strict = exact match, Combined = strict + partial. Counts in cells are correct / total.

Some LoCoMo cat_5 questions have empty ground-truth answers. Under the current scorer, these are counted as correct because there is no expected answer to match. To avoid overstating recall quality, we also report signal-only scores with those questions removed. Signal-only scores are computed on the 1,540 questions with non-empty ground-truth answers.

LoCoMo, current judge (gemini-3-flash-preview)

Run

Strict (full)

Combined (full)

Strict (signal-only)

Combined (signal-only)

Baseline v0.3.0 (Vector + BM25 + 8-stage pipeline)

76.7% (1524 / 1986)

88.1% (1750 / 1986)

70.0% (1078 / 1540)

84.7% (1304 / 1540)

Retrieval improvements (window_size=3, query expansion, top-K 25)

82.5% (1639 / 1986)

92.2% (1832 / 1986)

77.5% (1193 / 1540)

90.0% (1386 / 1540)

v0.4.5 + photo captions (same config as above)

82.9% (1647 / 1986)

92.7% (1842 / 1986)

78.0% (1201 / 1540)

90.6% (1396 / 1540)

Honest numbers disclaimer. (full) is the headline aggregate across all 1986 questions, the format vendors typically report — some publish only their strongest sub-category, we publish the full aggregate because it's what actually predicts behavior on mixed workloads. (signal-only) strips the cat_5 auto-pass artifact described above, so what you read there is the real recall quality on questions that have a ground-truth answer.

Per-category breakdown (v0.4.5 + photo captions run):

Category

Strict

Combined

cat_1 single-hop lists

51.4%

88.3%

cat_2 temporal

79.8%

85.7%

cat_3 open-domain reasoning

62.5%

79.2%

cat_4 multi-hop reasoning

88.0%

94.6%

cat_5 adversarial open-domain

100.0%

100.0%

Notes:

  • Judge model matters: gemini-3-flash-preview is more accurate than the previous Gemini Flash judge on synonyms, partial matches, and empty ground truth.

  • cat_5 questions have empty ground truth in this new run and are auto-scored as correct by the benchmark harness. That makes the new cat_5 strict score (446 / 446, 100.0%) useful for aggregate harness accounting, but not directly comparable to the historical cat_5 strict score (89.7%) from the older adversarial-question evaluation.

  • Pipeline: Vector retrieval with Gemini embeddings + BM25 + RRF + 8-stage reranking pipeline. The LLM reranker is not part of the benchmark loop (it is a runtime/server feature), so rerank_mode does not affect these numbers.

  • The v0.4.5 run additionally ingests the photo captions (blip_caption) that LoCoMo attaches to image-sharing turns — 697 of the 1540 signal questions cite image turns as evidence, and earlier runs silently dropped that content. Answer prompts also show the time of day of each memory since v0.4.5.

Historical LoCoMo results (gemini-2.5-flash judge)

Metric

First full run

mnemostack 0.2.1

Strict

66.4% (1319 / 1986)

67.8% (1346 / 1986)

Partial

12.8% (254 / 1986)

12.6% (250 / 1986)

Wrong

20.8% (413 / 1986)

19.6% (390 / 1986)

Combined

79.2% (1573 / 1986)

80.4% (1596 / 1986)

By question category (combined not tracked for the first full run):

Category

First run Strict

0.2.1 Strict

0.2.1 Combined

Δ Strict

cat_1 single-hop lists

34.8%

34.4%

74.1%

−0.4pp

cat_2 temporal

64.5%

69.8%

77.9%

+5.3pp

cat_3 open-domain reasoning

31.2%

41.7%

49.0%

+10.5pp

cat_4 multi-hop reasoning

69.2%

69.6%

82.0%

+0.4pp

cat_5 adversarial open-domain

90.1%

89.7%

89.7%

−0.4pp

Last historical run: 2026-04-27, mnemostack 0.2.1, same dataset, judged by gemini-2.5-flash.

Comparison with reported numbers from other systems

Caveat: different judges, evaluation protocols, and in some cases category cherry-picking. Vendor numbers below are taken at face value from their published material.

System

LoCoMo correct

Hindsight (reported range)

78–85%

Memobase (temporal subset)

85%

mnemostack

82.9%

Letta filesystem agent

74%

Mem0 graph variant

~68.5%

Zep (independently replicated)

58.4%

Real-corpus needle benchmark

LoCoMo measures generic long-term dialogue recall. We also run a private needle-in-haystack benchmark on the production workload that drove the original design — a ~17k-point memory stack indexed from a long-running assistant. Queries mix exact tokens (IP addresses, tickers), telegram IDs, paraphrased facts, and temporal probes.

Metric

Value

recall@1

90% (9/10)

recall@5

100% (10/10)

recall@10

100% (10/10)

Query latency p50

1.26 s

Query latency max

1.70 s

Honest numbers disclaimer. Reporting only recall@5 = 100% would look impressive, but it would also hide the harder top-1 behavior. recall@1 = 90% is what an agent reading only the top hit actually experiences, and the gap between @1 and @5 is where reranker quality (or the lack of it) shows up. We publish all three so you can read the metric that matches your downstream usage.

Useful because LoCoMo's failure modes (list exhaustion, open-domain reasoning) are orthogonal to what production memory stacks actually spend time on (find the specific fact the user mentioned weeks ago). This benchmark is not in the public repo; its methodology is in benchmarks/synthetic_longhorizon.py, which is the closest reproducible approximation.

Reproduce LoCoMo from a fresh clone

pip install -e '.[dev]'
bash benchmarks/download_locomo.sh   # fetches SNAP Research's public dataset
export GEMINI_API_KEY=...
bash benchmarks/run_locomo.sh        # full 10-sample run, writes results/ts.{json,log}

Details, category definitions, and notes on the judge protocol: benchmarks/README.md.

Who is this for?

Build it in if you need:

  • Long-lived agent memory that survives session restarts and doesn't drift into irrelevance as the corpus grows.

  • Recall quality on mixed workloads — exact-token lookups (IDs, tickers, error strings), semantic queries, temporal questions, multi-hop reasoning — not just one of them.

  • A stack you can plug into your own infrastructure: bring your own embedding model, LLM, vector store, or graph DB.

Not the best fit if you only need a single call to text-embedding-3-small + cosine similarity — something simpler will do. mnemostack earns its complexity on mixed, long-horizon workloads.

Features

  • 🧠 4-source hybrid retrieval — Vector (Qdrant) + BM25 (exact tokens) + Memgraph (knowledge graph) + Temporal (time-aware vector), all fused via Reciprocal Rank Fusion. Pluggable Retriever abstraction — add your own sources.

  • ⚖️ Weighted & adaptive RRF fusion — reciprocal_rank_fusion(weights=[...]) lets you lift sources you trust more; Recaller(adaptive_weights=True) picks a per-query-shape profile (exact-token / person / temporal / general). See the honest write-up below for where this helps and where it doesn't.

  • 🧪 HyDE retriever (opt-in) — embeds a hypothetical answer instead of the query. Useful for query↔document vocabulary gaps in documentation-style corpora; does not reliably help on dialogue-backed memory and always costs one extra LLM roundtrip per search(). Not included in the default Recaller.

  • 🪜 8-stage recall pipeline — ClassifyQuery → ExactTokenRescue → GravityDampen → HubDampen → FreshnessBlend → InhibitionOfReturn → CuriosityBoost → QLearningReranker. Opt-in; stateful HTTP feedback is explicit via /feedback, and recall exposure logging is off unless --auto-record-ior is enabled.

  • 🔁 Reranking — Gemini Flash (or any LLM) reorders top-K by relevance, or plug a cross-encoder / hosted rerank service through the score-based ScoringReranker. See docs/recipes.md for a runnable bge-reranker-v2-m3 example.

  • 🔤 Pluggable BM25 analyzer — the default is lowercase + Unicode word split (great for exact tokens); pass BM25Retriever(tokenizer=...) for stemming / lemmatization / language routing. Core stays dependency-free; docs/recipes.md has per-language recipes.

  • ⚡ Async API — every blocking surface has a signature-stable async mirror: Recaller.recall_async, recall_flow_async, Ingestor.ingest_async / ingest_one_async, AnswerGenerator.generate_async, synthesize_async, plus AsyncVectorStore over the native async Qdrant client. Retrievers dispatch in parallel; five concurrent HTTP recalls finish in roughly one single-recall wall-clock.

  • 🕰️ Stale-fact invalidation — mark superseded memories stale without deleting or re-embedding them: store.invalidate(ids, valid_until=...) sets bi-temporal payload keys (invalidated_at system-time, valid_until/valid_from world-time) via a cheap merge write. Recall hides invalidated facts by default; include_invalidated=True shows them and as_of="<iso>" reconstructs what was valid at a past instant from valid_from/valid_until — each optional, and invalidated_at is not read there at all, so a point-in-time view can be a superset of the default one (contract). The vector-side twin of the graph's valid_until model. CLI mnemostack invalidate <id>..., MCP mnemostack_invalidate, and — since 2.2 — HTTP POST /invalidate (with DELETE /memories for irreversible erasure), selecting either an id list or a whole source.

  • 🌍 Unicode-aware entity resolution — Memgraph retriever probes by telegram_id, handle, and precomputed name_lower so non-ASCII names match correctly (Memgraph's toLower() lower-cases ASCII only).

  • 📥 Streaming Ingestor API — batched, idempotent, LRU-cached ingest from any Python code. Lazy iterator means large corpora ingest with bounded memory. Same (source, offset, text) → same deterministic UUID-shaped content id, so re-runs are no-ops.

  • 📝 Markdown indexer — mnemostack index-markdown <dir> indexes a folder of markdown with structure: YAML frontmatter → payload filters, header-aware chunking with heading paths, and [[wikilinks]] / [text](note.md) → File -[LINKS_TO]-> File graph edges (with a Memgraph URI). Generic for any markdown folder; Obsidian vaults work as a side effect. Depends only on the already-present pyyaml.

  • 🌐 HTTP API (optional) — pip install 'mnemostack[server]' gives you /recall, /answer, the write/lifecycle surface (POST/GET/DELETE /memories, /invalidate, /triples), /health, /docs, plus /metrics in Prometheus text format. See the HTTP server section below.

  • 🔌 Pluggable embeddings — Gemini, Ollama, or HuggingFace (local GPU), via provider registry

  • 🤖 Pluggable LLM — Gemini Flash / Ollama for answer generation and reranking

  • 📚 Temporal knowledge graph — facts have valid_from/valid_until, query point-in-time state; graph resurrection stage recovers evicted-but-relevant memories.

  • 💬 Answer mode — inference layer synthesizes concise factual answers with source citations and confidence. Category-aware prompts (lists / temporal / multi-hop / inference / adversarial), specificity resolver, and cat_3 inference retry with query decomposition are on by default.

  • 📋 Knowledge synthesis — synthesize(entity) rolls up everything memory knows about a person, project, or topic into a structured profile (SynthesisFact / SynthesisResult, markdown or JSON). CLI: mnemostack synthesize <entity>. Optional related-entities expansion via graph and LLM summarization pass.

  • 📏 Progressive Tiers API — search --tier {1,2,3} and answer --tier {1,2,3} bound output size (~50 / ~200 / ~500 tokens) so agents can pay only for the detail they actually need. Omit --tier for unchanged full output.

  • ✂️ Chunkers + sliding window — plain, fixed-size, and MessagePairChunker for chat transcripts (keeps user↔assistant pairs together). The new vector.window_size config carries adjacent-turn context inside each chunk; window_size=3 was worth +5.8pp strict / +4.1pp combined on LoCoMo (v0.4.0).

  • 🔎 Query expansion + smart retry — Recaller(expansion_llm=...) widens recall with reformulated queries; AnswerGenerator(retry_with_expansion=True) retries low-confidence answers with the expanded query and a HyDE-style hypothetical before giving up. Opt-in via --query-expansion on mnemostack answer.

  • ⚙ Consolidation runtime — phase orchestrator for nightly memory lifecycle

  • 🔌 MCP server — expose memory tools to Claude Desktop, ChatGPT, Cursor, etc.

  • 🛡 Graceful degradation — retrieval keeps working if graph or any retriever is down

  • 🔐 Multi-tenancy — soft filters or a hard auth boundary — filters={"tenant": "a"} applies inside every retriever (exact match + ranges) on HTTP/MCP/CLI/library; results never include points outside the scope, verified by adversarial isolation tests. Filters are caller-supplied, so for a real trust boundary run the server with service-key auth (serve --auth / mcp-serve --auth): the tenant is resolved from the key (a client can't assert another's), enforced across the vector store, the knowledge graph, and per-tenant learning state. Optional per-tenant storage quotas apply at ingest, and request-rate quotas on the authenticated HTTP surface (serve --auth). Off by default. See the HTTP API section.

  • 🧩 Ingest enrichment + answer projection — Ingestor(enrich=callable) extracts structured facts into payloads at ingest (fail-open, --refresh-payloads updates existing collections without re-embedding); context_fields=[...] shows them to the answer LLM; rewrite_followup() resolves conversational follow-ups before recall.

  • 🧠 Reasoning-model friendly — Ollama think is off by default (reasoning models otherwise burn the whole token budget on thoughts and return empty text); options={...} passes any generation option through.

Recall tuning: fusion weights & HyDE

Some of the newer knobs help in specific workloads and do nothing (or mildly hurt) in others. Measured, not promised — both are opt-in by design, and the default Recaller stays classical equal-weight RRF over Vector + BM25 (+ Memgraph + Temporal when supplied).

Recaller(adaptive_weights=True) — picks a weight profile per query shape:

Query shape

Detection

Profile (bm25 / memgraph / vector / temporal)

exact_token

IPv4 / port / version / UUID / API-style tokens

1.4 / 1.4 / 1.0 / 0.9

person

"who is", @handle, username, contact, etc.

1.0 / 1.5 / 1.0 / 0.9

temporal

"when", "yesterday", "today", dates

1.0 / 1.0 / 1.0 / 1.4

general

everything else

classical equal-weight RRF

Measured on a real production corpus with 10 needle probes: recall@1 went 50% → 60%, recall@5 stayed at 90% (zero regression). On LoCoMo (pure dialogue questions, all classified general), adaptive weights had no effect — the profile simply isn't triggered. Rule of thumb: turn it on for production ops-style workloads (IPs, tickers, IDs, named entities); leave it off, or don't expect a lift, for dialogue benchmarks. Static retriever_weights={...} always wins over adaptive when both are set.

HyDERetriever — generates a short hypothetical answer via your LLM and embeds that instead of the raw query, then fuses alongside the other retrievers. Useful when the question and the stored answer use very different vocabulary (documentation corpora, FAQ-style content). On our LoCoMo cat_3 smoke (conv-43, 14 open-domain reasoning questions) it moved accuracy from 14.3% to 21.4% (+1 correct answer); on dialogue-backed memory overall it's roughly a wash. It always costs one extra LLM call per search(), so budget accordingly and treat it as a tool for specific workloads rather than a default.

Utilities

Agent runtimes often wrap transcript messages in metadata envelopes before the real body, which can dominate embeddings and make unrelated turns look similar. Clean messages before chunking/indexing with strip_metadata_blocks():

from mnemostack.utils import strip_metadata_blocks

clean = strip_metadata_blocks(raw_message)

Built-in profiles cover OpenClaw webchat and Telegram envelopes; pass profiles= or extra_patterns= to tune the cleanup for your runtime.

Environment

Variable

Purpose

Required for

GEMINI_API_KEY

Google Generative AI key

Gemini embedding + Gemini Flash LLM

OLLAMA_HOST

Ollama server URL (default http://localhost:11434)

Ollama embeddings / LLM

MNEMOSTACK_COLLECTION

Qdrant collection name (default mnemostack)

CLI convenience

MNEMOSTACK_QDRANT_URL

Qdrant URL (default http://localhost:6333)

Remote Qdrant

MNEMOSTACK_GRAPH_URI / MNEMOSTACK_MEMGRAPH_URI

Memgraph bolt URI

Graph retriever / GraphStore

MNEMOSTACK_TENANT

Tenant every recall is scoped to when the call names none (recall.tenant in the config, --tenant on the CLI)

Recall on every surface

MNEMOSTACK_ALLOW_CROSS_TENANT

Deliberately allow tenantless recall over a multi-tenant collection (tooling inside the trust boundary; not a security control)

Recall / serve startup

MNEMOSTACK_QUANTIZATION_RESCORE / MNEMOSTACK_QUANTIZATION_OVERSAMPLING

Qdrant quantization search parameters for dense queries (vector.quantization_* in the config); unset = not sent. See the recipe

Dense search on a quantized collection

MNEMOSTACK_LLM_HOST / MNEMOSTACK_LLM_TIMEOUT

LLM endpoint (ollama: default inherits the embedding --ollama-host; openai: required base URL) and LLM request timeout

Answer / reranker / expansion LLM

MNEMOSTACK_LLM_API_KEY

Bearer token for the openai LLM provider; unset or none = no auth header (keyless vLLM / llama.cpp)

Answer / reranker / expansion LLM

MNEMOSTACK_PROVIDER / MNEMOSTACK_EMBEDDING_PROVIDER

Embedding provider

CLI / HTTP / MCP

MNEMOSTACK_LLM / MNEMOSTACK_LLM_PROVIDER

LLM provider

Answer generation / reranking

MNEMOSTACK_BM25_PATHS

BM25 corpus paths separated by os.pathsep (: on Unix)

CLI / HTTP / MCP BM25 retriever

MNEMOSTACK_AUTO_RECORD_IOR

true/false toggle for HTTP recall exposure logging

HTTP stateful pipeline

MNEMOSTACK_EMBEDDING_MODEL / MNEMOSTACK_LLM_MODEL

Override the embedding / LLM model name

CLI / HTTP / MCP

MNEMOSTACK_VECTOR_HOST / MNEMOSTACK_VECTOR_COLLECTION

Aliases for the Qdrant URL / collection

CLI / HTTP / MCP

MNEMOSTACK_VECTOR_FLOOR

Keep top-N raw vector hits in results even when fusion/rerank would drop them (0 = off)

Recall tuning

MNEMOSTACK_RERANK_MODE

LLM reranker mode: relevant_only (default) or full_reorder

HTTP / MCP runtime reranker

MNEMOSTACK_TOKEN_BUDGET

Default recall token budget — cut results to the ranked prefix that fits (unset = off)

CLI / HTTP / MCP recall surfaces

MNEMOSTACK_GRAPH_TIMEOUT / MNEMOSTACK_GRAPH_HEALTH_TIMEOUT

Memgraph query / health-check timeouts in seconds

Graph retriever

MNEMOSTACK_CONFIG

Path to the YAML config file

All entry points

Only the providers you actually use need their keys. HuggingFace local-GPU embeddings need no keys at all. mnemostack init writes the same settings as YAML; explicit CLI flags override config/env defaults.

Setup and usage details

Try it in 30 seconds (Docker)

Fastest way to kick the tyres. No Python install, no manual Qdrant / Memgraph setup.

git clone https://github.com/udjin-labs/mnemostack && cd mnemostack
cp README.md examples/notes/              # any markdown will do
GEMINI_API_KEY=your-key docker compose -f examples/docker-compose.yml up -d --build

# Index the notes volume and ask a question over HTTP
docker compose -f examples/docker-compose.yml exec mnemostack \
    mnemostack index /data --provider gemini --collection demo

curl -s http://localhost:8000/recall \
    -H 'content-type: application/json' \
    -d '{"query":"what is this about","limit":5}' | jq

The mnemostack container runs the HTTP API on port 8000 by default. Interactive docs are at http://localhost:8000/docs. Use docker compose exec mnemostack mnemostack <cmd> for CLI-style operations (index, search, health) against the same stack.

Tear down with docker compose -f examples/docker-compose.yml down -v (the -v wipes Qdrant + Memgraph state).

Prefer Ollama (no cloud key needed)? Run Ollama on the host and pass --provider ollama everywhere instead of gemini. The endpoint resolves as: --ollama-host flag > MNEMOSTACK_OLLAMA_HOST env / embedding.ollama_host config > the native OLLAMA_HOST variable > http://localhost:11434 — so a client running in a container or VM can reach a remote Ollama daemon directly. An ollama LLM follows the same chain and inherits the embedding host by default; set llm.host / MNEMOSTACK_LLM_HOST only when generation lives on a different box:

mnemostack index-markdown memory/ \
    --provider ollama \
    --embedding-model qwen3-embedding:8b \
    --ollama-host http://192.0.2.10:11434 \
    --embedding-timeout 180 \
    --embedding-batch-size 64

Embedding uses the batch POST /api/embed endpoint (one request per batch; servers too old for it are detected once and served per-item with a loud warning). The embedding timeout (--embedding-timeout / MNEMOSTACK_EMBEDDING_TIMEOUT, default 180s) is independent of the short Qdrant liveness timeout — cold loads of larger local models are legitimately slow. Vector dimensions come from the model tables (quantization-suffix aware) or, for unknown models, a one-shot probe of the live model — there is no blind fallback dimension, so a wrong-size collection can't be created.

Behind an OpenAI-compatible endpoint (LiteLLM proxy, vLLM, llama.cpp server, an API gateway)? Use the openai LLM provider — it speaks POST {base}/v1/chat/completions, which all of them accept:

MNEMOSTACK_LLM_API_KEY=sk-... mnemostack serve \
    --llm openai --llm-model team-llm

with llm.host: http://gateway:4000 in the config (or MNEMOSTACK_LLM_HOST). Both the base URL and the model name are required — gateways have no meaningful defaults, so a missing one is a loud, actionable error (serve logs it and disables /answer) instead of a silent dial to the wrong place. A base URL already ending in /v1 (the OpenAI SDK convention) works too. Leave the key unset (or set it to none) for keyless vLLM / llama.cpp deployments; embeddings are unaffected and keep their own provider. Redirects are refused outright — a gateway 3xx becomes a normal error instead of carrying the bearer token to another origin. Reasoning models pointed straight at the cloud OpenAI endpoint (o1 family) reject the classic fields; via the SDK, get_llm("openai", token_param="max_completion_tokens", options={"temperature": None}) renames the budget field and drops the fields they refuse (gateways normally translate this themselves).

Reasoning models (qwen3, deepseek-r1 and similar): mnemostack disables thinking by default (think=False in OllamaLLM) — with thinking on, these models spend the whole token budget on thoughts and return empty text, silently degrading reranking, expansion and extraction. Pass get_llm("ollama", think=None) to keep the model's own default, or think=True to force it on models that support thinking. Extra generation options go through options={...} (e.g. {"num_ctx": 8192}).

Installation

# From PyPI
pip install mnemostack

# Optional extras
pip install 'mnemostack[huggingface]'  # local GPU embeddings
pip install 'mnemostack[mcp]'          # MCP server
pip install 'mnemostack[dev]'          # tests + linters

Run a local Qdrant for the vector store:

# tag rule: match your installed qdrant-client - same major, minor within 1
docker run -p 6333:6333 qdrant/qdrant:v1.18.3

Optionally a Memgraph for the knowledge graph:

docker run -p 7687:7687 memgraph/memgraph:latest

CLI quick start

# Health check
mnemostack health --provider ollama

# Index a directory of notes
mnemostack index ./my-notes/ --provider gemini --collection my-memory --recreate

# Hybrid recall
mnemostack search "what did we decide about auth" --provider gemini --collection my-memory

# Synthesize answer
mnemostack answer "what is the capital of France" --provider gemini --collection my-memory

# Record explicit feedback into the same state file used by the HTTP/MCP pipeline
mnemostack feedback <hit-id> --signal clicked --query "what did we decide about auth" \
  --source-list vector --source-list bm25

# MCP server (for Claude Desktop, Cursor, etc.)
mnemostack mcp-serve --provider gemini --collection my-memory

Progressive tiers — pay only for the detail you need

search and answer accept an optional --tier {1,2,3} flag that bounds how much output a call produces. Useful when a recall is called from a long-running agent loop where full recall output would burn context unnecessarily.

# Tier 1 (~50 tokens) — just "is there anything in memory about this?"
# Returns id, score, source labels; no text.
mnemostack search "VPN failover" --tier 1 --provider gemini

# Tier 2 (~200 tokens) — triage with short snippets (~40 chars each)
mnemostack search "VPN failover" --tier 2 --provider gemini

# Tier 3 (~500 tokens) — full 200-char previews, up to 10 results
mnemostack search "VPN failover" --tier 3 --provider gemini

Omit --tier to get the full, uncapped output (backward compatible). Rule of thumb for agents: tier 1 for navigation / existence checks, tier 3 only when you actually need to read the memories. answer is already compressed, so it needs a tier less often — use --tier 1 there to drop the SOURCES: block when only the answer text is wanted.

Streaming ingest API

When you want to feed items into mnemostack from code — a chatbot that logs every message, a scraper, a daemon tailing a log — use the Ingestor. It handles batching, deduplication, and idempotency for you.

from mnemostack.embeddings import get_provider
from mnemostack.vector import VectorStore
from mnemostack import Ingestor, IngestItem

emb = get_provider("gemini")
store = VectorStore(collection="my-memory", dimension=emb.dimension)
store.ensure_collection()

ing = Ingestor(embedding=emb, vector_store=store, batch_size=64)

stats = ing.ingest([
    IngestItem(text="alice joined acme on 2024-03-01", source="notes/alice.md",
               timestamp="2024-03-01T09:00:00Z"),  # event time — drives temporal recall
    IngestItem(text="alice left acme on 2025-06-15", source="notes/alice.md", offset=100),
])
print(stats)  # IngestStats(seen=2, embedded=2, upserted=2, skipped=0, failed=0)

Guarantees:

  • Idempotent. Each item gets a deterministic UUID-shaped content id computed from (source, offset, text). Re-running with the same input is a no-op: Qdrant upsert replaces the point onto itself, and an in-process LRU cache skips even the embedding call for items already seen in this session.

  • Batched. Items are embedded in batches of batch_size, so provider HTTP overhead amortises across many items.

  • Dated. Every payload records indexed_at (UTC). Pass timestamp= (or metadata={"timestamp": ...}) to set the event time the temporal retriever filters on. With window_size > 1, sliding-window chunks also carry the window's temporal range as window_start_ts / window_end_ts payload keys.

Images in the input (optional)

A memory stack that indexes only text answers "Not in memory" to questions whose answer lived in a photo. If your data contains images, describe them at ingest time and index the description:

from mnemostack.llm import get_llm

llm = get_llm("gemini")                 # or get_llm("ollama", model="llava") with a local vision model
desc = llm.describe_image(photo_bytes, mime_type="image/jpeg")  # one vision call per image
caption = f" [shared a photo: {desc.text}]" if desc.ok and desc.text else ""
item = IngestItem(text=f"{message_text}{caption}", source=..., timestamp=...)

describe_image is fully opt-in — nothing in the ingest or recall paths calls it, and text-only pipelines are unaffected. It works with any provider that has vision support (Gemini; Ollama vision models such as llava, llama3.2-vision, qwen2.5-vl) — providers without it return a normal fail-open error response. The default prompt produces a dense, index-oriented description (objects, any text/signs verbatim, setting, actions); pass prompt= to customize.

  • Streaming-friendly. ing.stream(item_iter) yields per-batch stats so long feeds can be monitored without waiting for the whole stream to drain.

  • Graceful. If a single item fails to embed, it is counted as failed but the rest of the batch still lands.

# Long-lived feed (e.g. inside a FastAPI or Celery worker)
for item in your_firehose():
    ing.ingest_one(IngestItem(text=item.body, source=item.channel, metadata={
        "user_id": item.user_id,
        "ts": item.ts.isoformat(),
    }))

Python API

from mnemostack.embeddings import get_provider
from mnemostack.vector import VectorStore
from mnemostack.recall import Recaller, AnswerGenerator
from mnemostack.llm import get_llm

emb = get_provider("gemini")
store = VectorStore(collection="my-memory", dimension=emb.dimension)
store.ensure_collection()

# ... index data here ...

recaller = Recaller(embedding_provider=emb, vector_store=store)
results = recaller.recall("what did we decide", limit=10)

# Each result: .id .text .score .source ("vector" | "bm25" | "memgraph" | "temporal") .metadata
# The list order is authoritative — .score need not follow it, and is not a
# confidence; re-sorting by it undoes reranking. docs/api-stability.md#what-score-is-not

# Optional: synthesize a concise answer
gen = AnswerGenerator(llm=get_llm("gemini"))
answer = gen.generate("what did we decide", results)
print(answer.text, answer.confidence, answer.sources)

Count and "list all X" questions need set completeness, which similarity top-K does not guarantee — the model counts what it sees and undercounts, returning a subset. For those, retrieve a wide candidate pool and enable the two-pass extract-and-aggregate mode:

gen = AnswerGenerator(llm=get_llm("gemini"), list_extract_mode=True)
pool = recaller.recall("how many trips did user A take", limit=150)   # wide pool, not top-10
answer = gen.generate("how many trips did user A take", pool)

list_extract_mode routes count/list questions through an extract pass (pulls every matching item as JSON) and a finalize pass (formats the list or count); other question categories are unaffected. The extract pass walks the whole pool you pass in, in batches of list_extract_batch_size (default 40), merging items across batches — so pool order does not decide whether a memory is seen, and the cost is one LLM call per batch plus finalize. An empty extract over a non-empty pool is retried once before abstaining. For guaranteed exhaustiveness on a bounded slice, build the pool from a full scan (VectorStore.scroll) filtered to the relevant slice. To evaluate it on your own data, the benchmark harness exposes the same knobs: benchmarks/locomo_single.py --list-extract --pool 150.

list_finalize="verbatim" skips the finalize LLM pass and assembles the answer deterministically from the extracted items (the count for count questions, the comma-joined items otherwise). Recommended for non-English corpora: an LLM finalize pass can paraphrase or distort items instead of repeating them verbatim. The default "llm" keeps the formatting pass.

Enriching payloads at ingest. Ingestor(enrich=callable) calls your function for every final item (including assembled window chunks) and merges the returned dict into the chunk payload — the mechanism is core, the extractor is yours (content extraction is corpus- and language-specific, so mnemostack ships none). Fail-open: a raising hook logs a warning and the item is indexed without enrichment; text/source/offset and an explicit item timestamp can't be overridden. From the CLI: mnemostack index docs/ --enrich mypkg.extractors:invoice_fields. Enriched fields combine with the rest of the stack: scope recall with filters={"amount": {"gte": 100}} and show them to the answer LLM with context_fields=["amount"].

def invoice_fields(item):          # yours — any language, any domain
    amounts = AMOUNT_RE.findall(item.text)
    return {"amounts": amounts} if amounts else {}

ing = Ingestor(embedding=emb, vector_store=store, enrich=invoice_fields)

Already-indexed collections don't need re-embedding to pick up enrichment: mnemostack index docs/ --enrich ... --refresh-payloads rewrites the payloads of existing chunks in place (Qdrant set_payload, vectors untouched) — only genuinely new chunks pay for embedding.

Structured payload fields in the answer prompt. By default the answer context shows each memory's timestamp, source and text. AnswerGenerator(context_fields=["author", "amount"]) additionally projects the named payload fields into each memory's context line (author=…, lists comma-joined, long values truncated; memories without the field render without it). Use it for structured facts the answer needs — who said it, amounts, your own ingest-time enrichments. Note the boundary: projection only changes what the answer prompt shows — retrieval ranks by text, so content that must be findable (image captions and similar) belongs in the text itself, not in a payload field.

Conversational follow-ups. "And who wrote that?" carries none of the conversation, so recall misses. rewrite_followup(query, history, llm) resolves pronouns and ellipses into a standalone question before recall — mnemostack holds no dialog state, you pass the history ((question, answer) pairs or plain lines, oldest first). One LLM call; the prompt instructs the model to return a self-contained question unchanged, and any failure falls back to the original query. To skip the call entirely for queries you already know are standalone, pass needs_rewrite=callable — that trigger heuristic is language-dependent, so core ships none (same boundary as question_classifier).

from mnemostack.recall import rewrite_followup

standalone = rewrite_followup("а кто это написал?", history, llm)
results = recall_flow(recaller, standalone, limit=10, pipeline=pipeline)

Non-English corpora. The built-in answer prompts and the question classifier are English; on other languages the extract/finalize passes degrade instead of helping. Both are pluggable:

gen = AnswerGenerator(
    llm=llm,
    list_extract_mode=True,
    prompt_overrides={                       # any subset; templates in YOUR corpus language
        "list_extract": MY_EXTRACT_TEMPLATE,    # must contain {context} and {query}
        "list_finalize": MY_FINALIZE_TEMPLATE,  # must contain {query} and {items}
        "temporal": MY_TEMPORAL_TEMPLATE,       # category prompts: {context} and {query}
    },
)
# the classifier's patterns are English too — route question classes yourself:
answer = gen.generate(query, pool, category="count")

Override names: the seven category prompts (general, list, count, temporal, multihop, inference, adversarial) plus list_extract / list_finalize. Required placeholders are validated at construction. mnemostack ships no translations by design — prompt quality is corpus- and domain-specific, so you own the templates.

Full stack: 4-source retrieval + 8-stage pipeline + reranker

This is the full runtime configuration. The LoCoMo numbers above are produced by a subset of it: the benchmark loop runs Vector + BM25 retrieval, the 8-stage pipeline, window_size=3, query expansion, and top-K 25 — the LLM reranker and the graph retriever are runtime-only features and are not part of the benchmark methodology (see benchmarks/run_locomo.sh for the exact reproduction path).

from mnemostack.embeddings import get_provider
from mnemostack.llm import get_llm
from mnemostack.vector import VectorStore
from mnemostack.recall import (
    Recaller, Reranker,
    VectorRetriever, BM25Retriever,
    MemgraphRetriever, TemporalRetriever,
    build_full_pipeline,
)
from mnemostack.recall.pipeline import FileStateStore, default_state_path

emb = get_provider("gemini")
store = VectorStore(collection="my-memory", dimension=emb.dimension)

retrievers = [
    VectorRetriever(embedding=emb, vector_store=store),
    BM25Retriever(docs=bm25_docs),                       # see "Building a BM25 corpus" below
    MemgraphRetriever(uri="bolt://localhost:7687"),      # optional
    TemporalRetriever(embedding=emb, vector_store=store),
]
recaller = Recaller(retrievers=retrievers)
raw = recaller.recall("what did we decide", limit=30)

pipeline = build_full_pipeline(state_store=FileStateStore(default_state_path()))
reranked = pipeline.apply("what did we decide", raw)
reranker = Reranker(llm=get_llm("gemini"), max_items=20)
final = reranker.rerank("what did we decide", reranked)[:10]

Reranker is generative: it asks an LLM to return candidate IDs. If you have a backend that returns numeric relevance scores instead (a local cross-encoder or a hosted rerank service), use ScoringReranker:

from mnemostack.recall import ScoringReranker

scoring_reranker = ScoringReranker(scorer=my_relevance_scorer, max_items=100)
final = scoring_reranker.rerank("retention policy", reranked)[:10]

The scorer object only needs score(query, documents) -> Iterable[float]. Scores are relative; no absolute threshold is applied by default. Generative LLMs can be wrapped as scorers, but dedicated rerank models/services are the more stable default because they avoid ID-format parsing.

Building a BM25 corpus

BM25Retriever needs a list of BM25Doc. Each doc is the atomic unit BM25 will rank — typically a paragraph or chunk of one of your source files:

from mnemostack.recall import BM25Doc
from pathlib import Path

docs = []
for i, path in enumerate(Path("my-notes/").rglob("*.md")):
    text = path.read_text()
    # chunk however you like — here: 800-char windows
    for j in range(0, len(text), 800):
        chunk = text[j : j + 800]
        if chunk.strip():
            docs.append(BM25Doc(
                id=f"{path.name}:{j}",
                text=chunk,
                payload={"source": str(path), "offset": j},
            ))

For transcript-like inputs (adjacent user and assistant turns), prefer MessagePairChunker so related turns stay in the same chunk. See mnemostack.chunking.

If your canonical memory corpus is already stored in Qdrant payloads, build the BM25 corpus from the same collection instead of maintaining a separate markdown export. This keeps exact-token lookup aligned with vector search (IDs, commit hashes, filenames, quoted phrases):

from qdrant_client import QdrantClient
from qdrant_client.models import FieldCondition, Filter, MatchValue
from mnemostack.recall import BM25Retriever

client = QdrantClient(host="localhost", port=6333)
bm25 = BM25Retriever.from_qdrant(
    client,
    "memory",
    scroll_filter=Filter(
        must=[FieldCondition(key="chunk_type", match=MatchValue(value="transcript"))]
    ),
    limit=40_000,
)

hits = bm25.search("api_key_rotation", limit=5)

You can also call bm25_docs_from_qdrant(...) directly if you want to combine Qdrant payload chunks with local BM25Docs before constructing BM25Retriever.

For morphologically rich languages or domain-specific normalization, pass a custom tokenizer/analyzer. The same analyzer is applied to corpus and query text; the default exact-token behavior is unchanged when omitted.

from mnemostack.recall import BM25Retriever

def analyzer(text: str) -> list[str]:
    # Normalize only what your corpus needs; preserve IDs, hashes and paths.
    ...

bm25 = BM25Retriever.from_qdrant(client, "memory", tokenizer=analyzer)

If you pre-tokenize BM25Doc objects yourself, pass retokenize=False when constructing BM25/BM25Retriever with the same analyzer. The BM25Retriever.from_qdrant(...) helper does this automatically.

HTTP server (optional)

If you want mnemostack available to callers that aren't Python — any service written in Node, Go, Rust, or a plain curl from a shell script — install the server extra and expose it over HTTP:

pip install 'mnemostack[server]'
export GEMINI_API_KEY=...
mnemostack serve --provider gemini --collection memory --port 8000

mnemostack serve binds to 127.0.0.1 by default. Use --host 0.0.0.0 only behind your own auth/rate-limit layer.

Endpoints:

Method

Path

Purpose

GET

/health

Qdrant + Memgraph reachability + config summary

GET

/healthz

Liveness probe — 200 whenever the process is up (no backend checks)

GET

/readyz

Readiness probe — 503 when Qdrant is unreachable (graph is fail-soft, never gates)

GET

/status

Operator snapshot — config, live dependency reachability, headline counters

POST

/recall

Hybrid recall with optional 8-stage pipeline

POST

/answer

Recall + LLM answer synthesis with citations

GET

/resolve/{chunk_id}

Verify a citation — resolve a chunk id back to its source document — read

POST

/feedback

Explicit click/usefulness feedback for stateful learning

POST

/memories

Create memories (server-side embedding, store-backed dedup) — write

GET

/memories

List what the tenant holds from one source (ids + integrity metadata, no text) — read

DELETE

/memories

Irreversible erasure by id list or source — write

POST

/invalidate

Non-destructive retraction by id list or source — write

POST

/triples

Write knowledge-graph facts — write

GET

/metrics

Prometheus scrape endpoint (counters + summary histograms)

GET

/docs

Interactive OpenAPI UI

curl -s http://localhost:8000/recall \
    -H 'content-type: application/json' \
    -d '{"query": "what did we decide about auth", "limit": 10}' | jq

Response shape (abridged):

{
  "query": "what did we decide about auth",
  "results": [
    { "id": "...", "text": "...", "score": 0.72, "source": "notes/...md", "metadata": {} }
  ],
  "degraded": [],  // components that ACTUALLY fell back, e.g. "retriever:bm25:failed", "reranker:fallback"; empty when healthy
  "notes": [],     // routine signals for stages that did not apply, e.g. "temporal:no_parse" on a date-less query; never a fault
  "tokens_estimate": 512   // estimated text tokens of the returned results
}

The order of results is authoritative — do not re-sort by score: many stages and fallback paths write that number on different scales, and a rerank changes the order without rewriting it. See what score is not.

Pass "include_trace": true in the request body to additionally get a trace object with per-retriever ranked lists, the fused order, the order after each ranking-pipeline stage (stages), and the post-rerank order — useful when debugging why a memory did or didn't surface. In Python, RecallTrace.loss_report(expected_ids) turns a trace into the position of each expected id at every step and the steps where it entered or left the top 1/5/10/20/30.

Pass "token_budget": 2000 to cap how much prompt space the results may occupy: the final ranking is cut to the prefix whose total text tokens fit the budget (a hard cap — never overshot, so an oversized top hit yields an empty list rather than a blown prompt). tokens_estimate in the response is the value the budget is enforced against; counting uses a dependency-free heuristic (≈4 chars/token for ASCII, ≈2 for non-ASCII scripts), so leave yourself margin rather than budgeting to the exact context limit. A server-wide default can be set with recall.token_budget in the config file (or MNEMOSTACK_TOKEN_BUDGET); per-request values override it. The same parameter is available on /answer (caps the memories fed to the LLM), on MCP mnemostack_search / mnemostack_answer, on the CLI as --token-budget, and in the library as recall_flow(..., token_budget=...) — where you can also pass an exact token_counter= (e.g. a tiktoken encoder) instead of the heuristic.

Pass "filters": {...} to scope recall by payload fields — exact match ({"tenant": "a"}) or inclusive ranges ({"timestamp": {"gte": "2026-01-01"}}). Filters apply inside every retriever, not as a post-filter on the output: the candidate pool itself is restricted, so top-K stays full and results never include points outside the scope — this is the isolation contract for multi-tenant and per-user memory. Sources that cannot attribute their results to the scope contribute nothing rather than leak. The knowledge-graph retriever attributes its hits where it can: a filter key the hit's own node metadata carries (e.g. index_root) is checked in place, the rest is proven through the hit's vector chunks — a graph file hit (with a recorded root, pinning the probe to its exact document) passes the filter exactly when at least one of its chunks does; entity nodes and anything else without pinnable chunks are still excluded, never leaked. The same filters parameter is available on /answer (the answer is generated only from in-scope memories, including retry sub-recalls), on MCP mnemostack_search / mnemostack_answer, on the CLI as --filters '{"tenant": "a"}', and in the library as recaller.recall(query, filters=...) / recall_flow(..., filters=...).

The /answer endpoint adds { answer, confidence, sources } alongside the memories and carries the same degraded / notes / opt-in trace fields, plus tokens_used — the LLM provider's reported token usage for the generation call that produced the answer (provider-specific semantics; null when the provider reports nothing). If the LLM isn't configured, /answer returns 503 and /recall still works — graceful degradation applies at the HTTP layer too.

Start the server with --retry-on-weak if you want a recall that comes back nearly empty to be paraphrased by the answer LLM and asked again, fusing the rounds by reciprocal rank — so a memory that two phrasings both find outranks one that only a single phrasing did, and a later paraphrase can beat an earlier one. What counts as "weak" is a COUNT (--retry-weak-below, default 1 — only a recall that returned nothing), not a score: fused scores are RRF values encoding rank, not confidence, so a threshold on them would measure nothing. One extra round, at most two paraphrases, and every retry carries the caller's tenant, filters, validity view and budget unchanged. Budget for it accordingly: each variant repeats your recall in full, reranker included, so a weak recall costs up to three LLM calls (one paraphrase plus one rerank per variant) and two extra retrieval rounds — not the single call the name suggests. Two consequences worth knowing before you switch it on: a retried response fuses over several rounds where an unretried one fuses over a single one, so the rank basis behind score differs and the two are never comparable — threshold on rank, not on the number, and see what score is not for what the number holds on each path; and a retry never returns fewer memories than it was given — if the fused list plus your token budget would hand back less than the recall already had, the original results come back untouched. That is a guarantee about the size of the response, not its membership: at a fixed limit a hit that a paraphrase ranks first will take the slot of one your original phrasing ranked last, which is what asking again is for. Off by default, and a request can only opt out of it — the server pays for the LLM call, so enabling it is the operator's call.

Stateful learning is explicit. Start the server with --auto-record-ior if you want /recall and /answer responses to update inhibition-of-return state, and with --record-access (or MNEMOSTACK_RECORD_ACCESS) if you want them to stamp access_count/last_accessed on every point they return — the reinforcement the freshness stage reads, recorded where the retrieval actually happens instead of in each client. Both are off by default: they turn reads into writes. Access recording is fail-open (a failed write is logged and counted, never raised), best-effort on the count (no atomic increment exists; concurrent recalls of one point can record one increment, and the reader clamps reinforcement at 10 anyway), and scoped to the caller's tenant.

It also changes what ranking means, so know which direction it moves things: recording these keys makes the freshness stage's access term live, and that term is a bonus, not a decay. A memory that has been retrieved gets a multiplier in [1.0, 1 + access_bonus_max] (0.25 by default, reached at 10 accesses) which fades back toward 1.0 as the access ages and stops there — use can raise a memory's rank, never lower it. A memory nothing has ever retrieved sits at exactly 1.0, so a deployment that records no accesses ranks exactly as it did before. Set --access-bonus-max 0 (or MNEMOSTACK_ACCESS_BONUS_MAX=0, honoured by serve, search, answer and mcp-serve alike) to take the access signal out of ranking entirely — the switch to reach for if your clients stamp last_accessed themselves, since leaving --record-access off does not help there: the stage reads those keys whoever wrote them. Values are clamped to [0, 1]. Because this is the one term a recall's own output feeds back into, it is bounded on purpose — small ceiling, saturating counter, and only the points actually handed to the caller are recorded. Send user actions to /feedback to update Q-learning:

curl -s http://localhost:8000/feedback \
    -H 'content-type: application/json' \
    -d '{"hit_id":"...","signal":"clicked","query":"what did we decide about auth","sources":["vector","bm25"]}' | jq

signal is one of useful, clicked, or irrelevant; pass the retrievers list returned by /recall as sources so Q-learning can update the right source weights. The same state update is available from CLI as mnemostack feedback ... and from MCP as mnemostack_feedback.

Tenantless recall over multi-tenant data is refused. A recall that names no tenant — no tenant= argument, no Recaller(default_tenant=...), no MNEMOSTACK_TENANT — is refused with CrossTenantRecallError when what it would read holds more than one tenant_id — every collection behind the recaller is asked, and so is a configured graph, and one tenant each in two stores counts as two (two small server-side queries per collection, exact with or without a payload index; the answer is cached and re-checked every minute), instead of silently searching every tenant at once. An unauthenticated serve over such data (collections and graph together) refuses to start, and the MCP mnemostack_graph_query tool refuses an unscoped query over a graph holding several tenants. Collections with a single tenant — including one tenant plus legacy points that carry no tenant_id — are unaffected, and allow_cross_tenant=True / --allow-cross-tenant / MNEMOSTACK_ALLOW_CROSS_TENANT restores the old behavior for deliberate cross-tenant tooling. A configured scope applies to the whole surface: on an unauthenticated serve or mcp-serve it scopes writes as well as reads (list, invalidate and delete included), the CLI's search, answer, synthesize, feedback and resolve default to it, and a file-backed BM25 corpus is stamped as that tenant's; mnemostack index takes its own explicit --tenant. Scoping hides points that carry no tenant_id — before switching an existing collection to a scope, stamp its points with mnemostack tenant-migrate --tenant <id> (a scoped serve warns at startup when it finds unstamped points). This guard catches misconfiguration; it is not a security boundary. It refuses only data verifiably holding several tenants: a probe that cannot answer (store down, timeout) lets the request through (the recall guard logs a warning once; the MCP graph query does not). A verdict is re-checked after a minute; a refresh that cannot answer keeps a known multi-tenant verdict (it keeps refusing until a probe answers again), while a single-tenant one turns undetermined and the graph sits out meanwhile. And whoever holds the store credentials can query the store directly. With auth off, a tenant scope configured on the server (MNEMOSTACK_TENANT) confines every request to that one tenant but is not access control (with auth on, the key decides the tenant and the configured scope does not restrict it), and a tenant= a caller passes itself is not verified at all. Untrusted consumers therefore need serve --auth below (the key decides the tenant) or an authentication layer in front.

Multi-tenant auth. By default the server is unauthenticated (single-tenant; put it behind your own auth layer). For a hard, per-tenant boundary, start it with --auth and issue service keys:

mnemostack keys add --tenant acme --scopes read,write   # prints the key once
mnemostack serve --auth                                  # default-deny on every data endpoint
curl -s http://localhost:8000/recall \
    -H 'X-API-Key: msk_...' -H 'content-type: application/json' \
    -d '{"query":"..."}'                                  # or: Authorization: Bearer msk_...

The tenant is resolved from the key (a client can't assert another's), enforced across the vector store, the tenant-scoped knowledge graph, and per-tenant learning state — so it's a real authorization boundary, unlike the caller-supplied filters above. /recall, /answer and GET /memories require read; /feedback, POST/DELETE /memories, /invalidate and /triples require write; a missing/invalid key is 401, insufficient scope 403. Cap each tenant with mnemostack quota set --tenant <id> --max-points N --max-rps R (storage enforced at ingest, rate on the HTTP surface → 429). The operator endpoints (/health, /healthz, /readyz, /status, /metrics) stay unauthenticated — protect them at your proxy if sensitive. Auth is off by default; you can still front the server with your own reverse proxy (nginx, Caddy, Traefik) either way. See docs/deployment.md and docs/api-stability.md.

Knowledge graph (optional)

from mnemostack.graph import GraphStore

graph = GraphStore(uri="bolt://localhost:7687")
graph.add_triple("alice", "works_on", "project-x", valid_from="2024-01-01")
graph.add_triple("alice", "works_on", "project-y", valid_from="2024-07-01")

# Who was alice working on in March?
march_facts = graph.query_triples(subject="alice", as_of="2024-03-15")

Current graph facts use the explicit valid_until="current" marker. If you created graph data with an older release, run mnemostack graph-migrate-current --dry-run first, then mnemostack graph-migrate-current to backfill legacy NULL markers.

MCP server for Claude Desktop

Add the entry below to your Claude Desktop config file:

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

  • Windows: %APPDATA%\Claude\claude_desktop_config.json

  • Linux: ~/.config/Claude/claude_desktop_config.json

{
  "mcpServers": {
    "mnemostack": {
      "command": "mnemostack",
      "args": ["mcp-serve", "--provider", "gemini", "--collection", "my-memory"],
      "env": {
        "GEMINI_API_KEY": "your-key-here"
      }
    }
  }
}

Claude will then be able to call mnemostack_search, mnemostack_answer, and graph tools.

Custom embedding provider

from mnemostack.embeddings import EmbeddingProvider, register_provider

class MyProvider(EmbeddingProvider):
    @property
    def name(self): return "my-provider"
    @property
    def dimension(self): return 512
    def embed(self, text): ...
    def embed_batch(self, texts): ...

register_provider("my-provider", MyProvider)

embed/embed_batch are the neutral primitives and stay the required surface. Ingestion and retrieval call the inherited role methods — embed_document(s) for chunks, embed_query/embed_queries for queries — which apply the model family's embedding profile exactly once before delegating to your primitives. Built-in profiles cover asymmetric families (Qwen3-Embedding's query instruction; E5's query: /passage: prefixes; the multilingual-E5-instruct variant's instruction format); symmetric models resolve to the identity profile and behave exactly as before. Transforms touch only inference input — stored text, chunk ids, lexical search and citations keep the original document text. To teach mnemostack a new family without patching it:

from mnemostack.embeddings import EmbeddingProfile, register_embedding_profile

register_embedding_profile("ollama", EmbeddingProfile(
    name="my-family",
    version=1,
    model_patterns=("my-family-embed:*",),
    query_transform={"kind": "prefix", "prefix": "query: "},
))

Indexed points are stamped with a document-space fingerprint (provider + model + document transform + dimension + provider inference knobs — profile name/version are metadata and never hashed; upgrading from 1.x? see docs/migration-1.x-to-2.0.md); mnemostack index / index-markdown refuse a collection embedded under a different space, and doctor reports the active profile and fingerprint.

Install for an OpenClaw-style agent

If you run an OpenClaw-style assistant with its own scripts/recall.sh or similar recall entry point, you don't have to wire mnemostack in by hand. Ask the assistant to do it and it can follow the steps below verbatim.

(These steps are written against the markdown-backed memory layout popularised by OpenClaw / OpenClaw — the reasoning applies to any assistant whose memory lives in MEMORY.md + daily notes.)

Assistant, read this section and follow the steps if your user asks you to install mnemostack.

  1. Install the package into whichever Python env your recall scripts use.

    pip install mnemostack

    Confirm with python3 -c "import mnemostack; print(mnemostack.__version__)".

  2. Make sure Qdrant and (optionally) Memgraph are reachable. If either is missing, start them:

    # tag rule: match your installed qdrant-client - same major, minor within 1
    docker run -d -p 6333:6333 --name qdrant qdrant/qdrant:v1.18.3
    docker run -d -p 7687:7687 --name memgraph memgraph/memgraph:latest   # optional

    Check with curl -sf http://localhost:6333/collections and echo 'RETURN 1;' | cypher-shell -a bolt://localhost:7687.

  3. Populate the vector store with the assistant's existing memory. Almost every OpenClaw-style assistant stores memory as markdown files — MEMORY.md, memory/*.md daily notes, maybe USER.md / TOOLS.md. Those files are the ground truth; Qdrant is just a searchable index on top of them. So "migration" here really means a first-time ingest.

    Two cases:

    Case A — user has no Qdrant yet (default OpenClaw setup, markdown-only memory).

    Stand up Qdrant (see step 2), then index the workspace in one command:

    mnemostack index ~/.openclaw/workspace/ \
        --provider gemini \
        --collection memory \
        --chunk-size 800

    This walks every *.md / *.txt under the workspace, chunks them, embeds each chunk with the chosen provider, and upserts into Qdrant. For a typical workspace (a few megabytes of notes) this finishes in a minute or two and stays well inside free-tier embedding quotas.

    Re-running is safe. mnemostack index uses a deterministic id per (source, offset, content) chunk, so a second run skips unchanged chunks (no duplicates, no wasted embedding calls) and only re-embeds chunks whose content changed or moved. Use --recreate if you deliberately want to wipe and rebuild from scratch.

    Memgraph is optional — skip it for a pure markdown-backed memory and MemgraphRetriever just won't be part of the retriever list.

    Case B — user already has Qdrant / Memgraph populated by another plugin or a prior assistant run. Don't re-index. mnemostack reads existing collections and graphs as-is; payload like chat_id, memory_class, valid_from is preserved. Just point VectorStore(collection=...) at the existing collection name and MemgraphRetriever(uri=...) at the existing graph.

    Critical for both cases: use the same embedding model for the query that was used for the stored points. Mixing providers/models across ingest and query silently returns garbage — no error, just bad results. Indexing commands now enforce this for documents: points carry a document-space fingerprint and index/index-markdown refuse a collection embedded under a different provider/model/profile.

  4. Point your recall entry at mnemostack. Swap your search helper to use the 4-retriever Recaller from the "Full stack" example above. If your recall script has a --unified flag, route to mnemostack first and fall back to the legacy code path on any exception.

  5. Keep the fallback wired. Don't delete the existing recall implementation. Route mnemostack as the primary, fall back to the old recall on any exception. A bad upgrade becomes a one-minute rollback, not an outage.

  6. Smoke-test on real queries from the user's memory, not toy strings. Ask at least one exact-token query (an IP, a ticker, an error string), one semantic query, one temporal query ("what did we decide last March"), and one list query. If any regresses vs the old recall, keep both running in parallel until you understand why.

  7. Write a short note in the user's daily memory describing what changed: which command was used to ingest, which embedding provider, which collection name, and where the fallback lives. Future sessions need this to avoid re-indexing with a different model and trashing recall quality.

If you're a human operator and you'd rather do this yourself, the same steps apply; treat them as a checklist.

OpenClaw auto-recall companion plugin

This repository also includes a self-contained OpenClaw companion plugin in openclaw-plugin/. mnemostack remains the recall backend — vector + BM25 + graph + temporal retrieval, fusion, reranking, and answer synthesis — while the plugin connects that backend to OpenClaw's before_prompt_build hook.

Zero-config path: install mnemostack, run the daemon on the default local port, install/enable the plugin, and OpenClaw will automatically inject bounded recall answers for recall-style questions:

mnemostack serve --host 127.0.0.1 --port 18793
cd openclaw-plugin
npm test

The plugin defaults to http://127.0.0.1:18793/answer, supports English/Russian trigger defaults with extensible language-agnostic trigger lists, and can fall back to a Script backend such as recall-selfeval.sh when you are not running the daemon.

Roadmap

  • Embedding provider registry (Gemini / Ollama / HuggingFace)

  • LLM provider registry (Gemini Flash / Ollama / any OpenAI-compatible endpoint)

  • Qdrant wrapper

  • BM25 + RRF recall pipeline

  • Answer mode with confidence + citations

  • LLM-based reranker

  • Memgraph wrapper with temporal validity

  • Consolidation runtime (phase orchestrator)

  • CLI (mnemostack health/doctor/inspect/search/answer/index/mcp-serve)

  • MCP server (Model Context Protocol)

  • Text → graph triple extractor helpers (mnemostack.graph.TripleExtractor)

  • Config file support YAML/JSON (mnemostack.config, mnemostack init/config CLI)

  • Async variants for high-throughput servers (mnemostack.vector.AsyncQdrantStore)

  • Docker compose examples (examples/docker-compose.yml)

  • Reproducible LoCoMo benchmark harness in-tree (benchmarks/run_locomo.sh)

  • First-class FastAPI/Starlette service wrapper (pip install 'mnemostack[server]', mnemostack serve)

  • Async Recaller.recall_async and parallel retriever dispatch (proven: 5 concurrent HTTP recalls complete in ~1x single-request wall-clock)

  • Benchmarks on longer-horizon synthetic corpora (benchmarks/synthetic_longhorizon.py)

  • Streaming Ingestor API (mnemostack.ingest)

  • Prometheus /metrics endpoint on the HTTP server

  • Unicode-aware MemgraphRetriever probes (telegram_id, handle, name_lower)

  • Community health: Code of Conduct, Security policy, issue/PR templates

  • Per-retriever latency in /metrics (mnemostack_recall_<name>_latency_ms)

  • Weighted RRF fusion (reciprocal_rank_fusion(weights=[...]))

  • Adaptive per-query-shape weights in Recaller (adaptive_weights=True)

  • HyDERetriever (opt-in, not in default Recaller)

  • Two-pass graph extraction (full + detail) in the agent's graph-sync pipeline

  • Progressive Tiers API on search/answer (--tier {1,2,3}, backward-compatible)

  • MCP integration guides for Claude Desktop/Code, Cursor, OpenClaw (integrations/)

  • Silent-zero fix in TemporalRetriever + dispatch-by-type filter builder (0.2.0a1)

  • Canonical recall_flow() — CLI/HTTP/MCP rank identically (0.5.0)

  • Chunk lifecycle: index --prune (root-scoped) + --refresh-payloads without re-embedding

  • Recall filters= on every surface with adversarially-tested tenant isolation

  • Ingest-time enrichment hook (Ingestor(enrich=...)) + context_fields answer projection

  • Ollama think control (off by default) + generation options passthrough

  • Follow-up question rewriting (rewrite_followup)

  • Access reinforcement instead of punitive decay — being used raises rank, going unused never lowers it (access_boost, --access-bonus-max) (2.3.1)

  • Half-open circuit breaker for an unreachable graph store + MNEMOSTACK_GRAPH_URI="" off switch (2.3.2)

  • llm.host / llm.timeout reaching the LLM on every construction surface (2.3.2)

  • openai LLM provider — LiteLLM proxies, vLLM, llama.cpp server, API gateways (2.4.0)

  • hermes-agent memory provider — hermes-mnemostack on PyPI, listed in the official Hermes plugin catalog

  • OpenClaw companion plugin published on ClawHub (@udjin79/mnemostack-auto-recall)

Contributing

Issues and PRs are welcome. Public APIs are intended to remain stable; new functionality should land additively where possible.

License

Apache 2.0 — see LICENSE.

Available Tools

7 tools
mnemostack_answerMnemostack AnswerA

Answer a question using retrieved memories.

Read-only, no side effects, no authentication required. Use this when you want a concise factual answer synthesized from memory search results instead of the raw matches returned by mnemostack_search. Returns a JSON object with ok, query, answer text, confidence (0.0-1.0), sources, notes (the AUTHORITATIVE routine signals — e.g. temporal:no_parse on any query without a date; NOT a fault), degraded (components that actually fell back, plus a deprecated back-compat duplicate of the routine tags until the next major), fallback_recommended, tokens_estimate (estimated text tokens of the context memories), tokens_used (LLM-provider-reported usage for the answer call; null when unreported), and error. Stale facts are hidden by default; use include_invalidated or as_of to see them.

ParametersJSON Schema
NameRequiredDescriptionDefault
as_ofNoPoint-in-time recall (ISO-8601); same contract as mnemostack_search.
limitNoMaximum number of results to return (default 10)
queryYesNatural language question or keyword to search memories for
filtersNoPayload filters applied inside every retriever (exact match or gte/lte ranges); the answer is generated only from memories inside the filtered scope.
token_budgetNoHard cap on the total (estimated) text tokens of the memories fed to the answer LLM (same contract as mnemostack_search). Unset uses the server-wide default.
include_invalidatedNoInclude facts marked stale (default false; same as mnemostack_search).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden and does so: read-only, no side effects, no authentication required, stale facts hidden by default, plus the important quirk that 'notes' are authoritative routine signals (e.g. temporal:no_parse) and NOT a fault, while 'degraded' means real fallback. This prevents an agent from misreading benign signals as errors.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and usage are front-loaded, but a large block enumerates return fields (ok, query, answer, confidence, sources, error) that an output schema already provides, which is redundant length. The parenthetical clarifications of 'notes' and 'degraded' earn their place; the field listing largely does not.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter read tool with a full output schema, the description covers the safety profile, the usage decision, and the semantically ambiguous return fields (notes/degraded/fallback_recommended). It omits nothing an agent needs to call it correctly, though it could tie the token_budget/token_estimate relationship together more tightly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all six parameters, establishing the baseline of 3. The description adds only marginal meaning beyond that ('Stale facts are hidden by default; use include_invalidated or as_of to see them'), restating contracts already present in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Answer a question using retrieved memories') and immediately differentiates itself from the sibling mnemostack_search by contrasting a synthesized answer with raw matches. An agent can pick this tool without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative ('instead of the raw matches returned by mnemostack_search') and the condition that selects it ('when you want a concise factual answer synthesized from memory search results'). It also flags the stale-data default and the include_invalidated/as_of escape hatch.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemostack_feedbackMnemostack FeedbackB

Record explicit feedback for stateful recall learning.

Use signal='clicked' to also record inhibition-of-return exposure. Pass retriever labels from mnemostack_search results as sources so Q-learning can update source weights.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNoNatural language question or keyword associated with the feedback
hit_idYes
rewardNo
signalYes
sourceNo
sourcesNo
query_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose non-obvious downstream behavior: Q-learning updates of source weights and inhibition-of-return exposure. It does not say what happens to a hit after feedback, whether it is reversible, or what reward/hit_id semantics govern the mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, then the two highest-value operational hints. No filler, though the phrasing is slightly fragmented.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, and the description captures the key learning/reinforcement loop. For a 7-parameter tool with 14% schema coverage and zero annotations, though, it leaves the reward and identity parameters under-explained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 14%, so the description must compensate, and it meaningfully clarifies two opaque parameters: 'signal' (the clicked value triggers IOR) and 'sources' (retriever labels from mnemostack_search). The remaining five parameters (hit_id, reward, query, source, query_type) get no added meaning beyond their names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Record explicit feedback for stateful recall learning.' This distinguishes it from read-oriented siblings like mnemostack_search and the write-oriented mnemostack_remember, though it never names a sibling to contrast against.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Offers concrete conditional guidance for one case ('Use signal=clicked to also record inhibition-of-return exposure') and tells the agent where to source the 'sources' value (mnemostack_search results). However, it gives no general when-to-use vs. when-not guidance relative to the other mnemostack tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemostack_healthMnemostack HealthA

Check health of all mnemostack components.

Read-only, no side effects, no authentication required. Returns a JSON object with ok (bool) and per-component status for the embedding provider, Qdrant vector store, and optional Memgraph graph database. Use this to verify the memory backend is reachable before issuing recall queries.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and discharges most of it: read-only, no side effects, no authentication required, plus a statement of the return shape (ok plus per-component status). It stops short of failure modes (what a degraded vs unreachable component looks like, timeouts/latency), so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the verb and resource, then the safety/return profile, then the usage trigger. Nothing is repeated from the schema and every sentence adds information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-argument health-check tool this is complete: the safety profile substitutes for missing annotations, the output shape is summarized, and the output schema exists so return values need not be detailed further.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters (schema is an empty object with additionalProperties: false), so there are no parameter semantics to explain; baseline 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Check health') and resource ('all mnemostack components'), and enumerates the components it covers (embedding provider, Qdrant, Memgraph). This is unmistakably distinct from every sibling (search, answer, resolve, invalidate, remember, feedback), which are all memory-content operations rather than diagnostics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use guidance: 'verify the memory backend is reachable before issuing recall queries.' There is no explicit when-not-to-use condition or named alternative, but the pre-flight check framing makes the trigger condition unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemostack_invalidateMnemostack InvalidateA

Mark memories stale by id, non-destructively.

A write tool (parallel to mnemostack_graph_add_triple). Sets invalidated_at (and optionally valid_until) on each point's payload without deleting or re-embedding it; invalidated facts drop out of default recall but stay reachable via include_invalidated / as_of. Points that do not exist are skipped. Pass index_root in multi-root collections to avoid marking another root's chunks stale. Returns ok, requested, and invalidated (the number of points actually updated).

ParametersJSON Schema
NameRequiredDescriptionDefault
idsYesPoint id(s) to mark stale (string or integer)
index_rootNoOwner guard: when set, points owned by a different index_root are skipped, so one root cannot invalidate another's chunks in a shared collection.
valid_untilNoWorld-time the fact stopped being true (ISO-8601); optional, separate from the system-time invalidation stamp.
invalidated_atNoSystem-time stamp (ISO-8601); default: now (UTC)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it states payloads are mutated in place (invalidated_at, optionally valid_until) with no deletion and no re-embedding, that records leave default recall but remain reachable via include_invalidated/as_of, that missing points are silently skipped, and that index_root guards against cross-root staleness. This is a rich, accurate disclosure of a mutation's side effects and edge cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with a one-line purpose followed by tightly packed behavioral detail; almost every clause carries information. The closing 'Returns ok, requested, and invalidated' sentence is redundant given an output schema exists, and the parenthetical about a non-sibling tool adds marginal value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter write tool with no annotations and an output schema, the description covers the mutation semantics, safety profile, edge cases, and scoping guard that an agent needs before calling. Nothing material is left to the annotations, which are absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents ids, index_root, valid_until, and invalidated_at precisely. The description reinforces the intent of index_root ('avoid marking another root's chunks stale') and the valid_until/invalidated_at split, but adds no syntax or format detail beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line gives a specific verb and resource ('Mark memories stale by id') plus the crucial qualifier 'non-destructively', which immediately separates it from any hard-delete behavior. It also positions itself as a write tool alongside mnemostack_graph_add_triple, so an agent knows the operation class without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: calling out that invalidated facts 'drop out of default recall but stay reachable via include_invalidated / as_of' tells the agent this is the soft-retirement path, and the multi-root note tells it when to pass index_root. However, there is no explicit routing against the listed siblings (remember, feedback, resolve), so the agent must infer when invalidation beats those alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemostack_rememberMnemostack RememberA

Store a memory for the caller's tenant (the write counterpart of mnemostack_search).

A write tool: requires the write scope, embeds server-side, and stamps the process key's tenant so the memory lands in — and is recallable from — exactly this tenant's scope. Ids are deterministic from (source, offset, text): retries and repeated content return duplicate without a second embedding call. Long documents pass chunk=true and yield one result per chunk. Use mnemostack_invalidate to retract a stored memory.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNoOptional tags stored in the payload.
textYesThe memory content to store. Embedded server-side. Plain items are capped at 32768 characters; longer documents need chunk=true.
chunkNoSplit a long document server-side into the same fixed character windows `mnemostack index` uses (identical chunk ids). Requires a non-empty source.
offsetNoPosition within `source` for multi-part documents.
sourceNoLogical origin (e.g. 'chat/2026-08-19'). With `offset` it forms the deterministic id: re-sending the same content is a no-cost duplicate, never a second copy.
metadataNoFree payload fields, filterable at recall. Server-reserved keys (underscore-prefixed, structural ones like tenant_id/source, tags/timestamp — use their dedicated parameters — and the lifecycle marker invalidated_at, settable only via mnemostack_invalidate) are rejected.
timestampNoEvent time of the content (ISO-8601); drives temporal recall.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden and does so well: it discloses the required write scope, that embedding happens server-side, that the process key's tenant is stamped, the deterministic id derivation from (source, offset, text), the duplicate-on-retry behavior with no second embedding call, and the one-result-per-chunk outcome. These are non-obvious runtime traits that materially change how an agent calls the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and safety profile before secondary details, and every sentence carries information. Slightly longer than needed since the deterministic-id/duplicate behavior is repeated from the `source` and `offset` schema descriptions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation tool with no annotations, the description supplies the missing auth scope, tenancy semantics, and idempotency model, while the output schema handles return values. Nothing an agent needs to call it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented with constraints and defaults, which sets the baseline at 3. The description reinforces the chunk/source/offset interaction but adds no format or syntax detail beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Store a memory for the caller's tenant') and immediately anchors it against a sibling ('the write counterpart of mnemostack_search'). An agent can distinguish this from search/invalidate/resolve without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the retraction alternative explicitly ('Use mnemostack_invalidate to retract a stored memory') and gives a conditional usage rule for chunk=true on long documents. It stops short of stating when NOT to use this tool (e.g. when to prefer mnemostack_feedback or resolve for updating existing state).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mnemostack_resolveMnemostack ResolveA

Verify a citation: resolve a chunk id back to its source document.

Re-reads the CURRENT source and returns an honest verdict: intact / source_changed / moved (citation still supported), changed / missing (not supported by the current source), or unresolvable (cannot be verified from this process). Includes the snapshot-hash comparison and the fragment when locatable. Read-only; never mutates stored memory; runs outside the recall path. Resolution is confined to the corpus root recorded at ingest — there is deliberately no way for a caller to point it at another directory.

ParametersJSON Schema
NameRequiredDescriptionDefault
chunk_idYesThe [id:...] value from a recall result or answer citation

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: read-only, never mutates stored memory, re-reads the current source, confines resolution to the ingest-time corpus root with no way to redirect it, and defines the possible verdicts (intact/source_changed/moved vs changed/missing/unresolvable). It also discloses inclusions such as the snapshot-hash comparison and fragment when locatable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose in the first clause, followed by the verdict taxonomy and constraints. It is somewhat long, but each sentence carries distinct information (verdict meanings, read-only guarantee, scope confinement), so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description instead covers what the schema cannot: read-only semantics, corpus-root confinement, and the meaning of each verdict. Nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter's provenance ('the [id:...] value from a recall result or answer citation') is already documented in the schema. The description reinforces the id-to-document relationship but adds no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and resource: 'Verify a citation: resolve a chunk id back to its source document.' This is clearly distinct from siblings like mnemostack_search, mnemostack_answer, or mnemostack_remember, and an agent can route to it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is explicit — use it to verify a citation produced by recall or an answer — and it notes it 'runs outside the recall path,' which clarifies its place in the workflow. It stops short of naming a competing sibling or stating when-not to use it, so no explicit alternative routing is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev2.3.1
    • Addedmnemostack_remember
  2. 1 tool updatev2.0.0
    • Addedmnemostack_resolve
  3. 3 tool updatesv0.8.0
    • Changedmnemostack_answer3 fields changed
      • addedInput schema / properties / as_of
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Point-in-time recall (ISO-8601); same contract as mnemostack_search."
        +}
      • addedInput schema / properties / include_invalidated
        Added value: +{
        +  "default": false,
        +  "description": "Include facts marked stale (default false; same as mnemostack_search).",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / token_budget
        Added value: +{
        +  "anyOf": [
        +    {
        +      "minimum": 1,
        +      "type": "integer"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Hard cap on the total (estimated) text tokens of the memories fed to the answer LLM (same contract as mnemostack_search). Unset uses the server-wide default."
        +}
    • Addedmnemostack_invalidate
    • Changedmnemostack_search3 fields changed
      • addedInput schema / properties / as_of
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Point-in-time recall (ISO-8601): return facts valid at this world-time instant, ignoring later invalidation."
        +}
      • addedInput schema / properties / include_invalidated
        Added value: +{
        +  "default": false,
        +  "description": "Include facts marked stale. Default false: memories with an invalidated_at marker are hidden from recall.",
        +  "type": "boolean"
        +}
      • addedInput schema / properties / token_budget
        Added value: +{
        +  "anyOf": [
        +    {
        +      "minimum": 1,
        +      "type": "integer"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Hard cap on the total (estimated) text tokens of the returned results; the final ranking is cut to the prefix that fits. Unset uses the server-wide default, if any."
        +}
  4. 2 tool updatesv0.6.0
    • Changedmnemostack_answer1 field changed
      • addedInput schema / properties / filters
        Added value: +{
        +  "anyOf": [
        +    {
        +      "additionalProperties": true,
        +      "type": "object"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Payload filters applied inside every retriever (exact match or gte/lte ranges); the answer is generated only from memories inside the filtered scope."
        +}
    • Changedmnemostack_search2 fields changed
      • addedInput schema / properties / filters
        Added value: +{
        +  "anyOf": [
        +    {
        +      "additionalProperties": true,
        +      "type": "object"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Payload filters applied inside every retriever: exact match ({\"tenant\": \"a\"}) or gte/lte range ({\"timestamp\": {\"gte\": \"2026-01-01\"}}). Results never include points outside the filtered scope."
        +}
      • addedInput schema / properties / include_trace
        Added value: +{
        +  "default": false,
        +  "description": "Include the per-retriever recall trace (debug; verbose)",
        +  "type": "boolean"
        +}
  5. 3 tool updatesv0.4.3
    • Changedmnemostack_answer2 fields changed
      • addedInput schema / properties / limit / description
        Added value: +"Maximum number of results to return (default 10)"
      • addedInput schema / properties / query / description
        Added value: +"Natural language question or keyword to search memories for"
    • Changedmnemostack_feedback1 field changed
      • addedInput schema / properties / query / description
        Added value: +"Natural language question or keyword associated with the feedback"
    • Changedmnemostack_search2 fields changed
      • addedInput schema / properties / limit / description
        Added value: +"Maximum number of results to return (default 10)"
      • addedInput schema / properties / query / description
        Added value: +"Natural language question or keyword to search memories for"
  6. 4 tool updatesv0.4.1
    • First observedmnemostack_answer
    • First observedmnemostack_feedback
    • First observedmnemostack_health
    • First observedmnemostack_search

TDQS

A4.1/5.0

Scored across 7 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: health checks, raw search, synthesized answer, citation resolution, invalidation, memory write, and feedback. The descriptions explicitly differentiate overlapping-sounding pairs like search vs. answer and invalidate vs. remember. No two tools appear to do the same thing.

Naming Consistency4/5

All tools use a consistent mnemostack_ prefix and snake_case, making them easy to predict. However, the action words mix verbs (search, resolve, invalidate, remember) with nouns (health, feedback), a minor deviation from a pure verb_noun pattern.

Tool Count5/5

Seven tools is well-scoped for a memory backend covering health, read/write, feedback, and citation verification. Each tool earns its place without redundancy or thinness. The count falls comfortably in the ideal 3–15 range.

Completeness4/5

The set covers core lifecycle operations: store, search, answer, invalidate, resolve citations, and record feedback. Minor gaps include no direct update tool (must invalidate and re-remember) and the referenced mnemostack_graph_add_triple is not exposed, but these are workable or intentionally out of scope.

Maintenance

ActivityActive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Persistent memory MCP server for AI coding agents (Claude Code, Codex, Gemini CLI). Hybrid retrieval (vector + BM25), cross-encoder reranking, knowledge graph, session checkpoint/resume, and multi-scope isolation. Local-first with LanceDB.
    30
    85 npm
    15
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Persistent, semantically-searchable memory for AI agents using local PostgreSQL, pgvector, and Ollama embeddings, exposed via MCP with hybrid retrieval, knowledge graph, and auto-recall hook.
    3 npm
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server that enables persistent, hybrid, local memory for LLM agents, with vector + BM25 search, knowledge graph, and policy-driven retention, providing token-budgeted context injection for AI assistants.
    MIT