Skip to main content
Glama
yliuai

Spomory

Spomory

English | 中文

Spomory gives Claude Desktop, Cursor, and Codex CLI a memory that persists across sessions — and is shared between all three. Tell one of them something once (a project detail, a preference, a fact about yourself) and any of them can recall it later, without you repeating yourself.

It runs as an MCP server, either locally on your own machine (free, nothing leaves your computer) or as a hosted cloud service (free signup, memory follows you across devices). The rest of this README covers the local path; for the cloud path see spomory.yliuai.com/get-started/remote.

Curious what it looks like before installing anything? There's a no-signup demo at spomory.yliuai.com — paste in some text, see the entities and relations it extracts.

Get started

Requires Python 3.11+.

1. Install

pip install "spomory[llm,embedding,mcp]"

2. Give it an LLM. Spomory uses an LLM to turn what you tell it into structured facts. Any OpenAI-compatible API works — OpenAI, DeepSeek, Qwen, etc.:

export LLM_API_KEY=sk-...
export LLM_BASE_URL=https://api.deepseek.com   # optional, defaults to OpenAI
export LLM_MODEL=deepseek-chat                 # optional, defaults to gpt-4o-mini

3. Connect it to your client. Claude Desktop, Cursor, and Codex CLI each need a few lines added to a config file, pointing at the spomory-mcp command. The exact steps, a config file example for each client, and a real gotcha we hit (macOS blocking a venv that lives under ~/Documents) are in docs/mcp_quickstart.en.md.

The one thing that trips people up: the config needs the absolute path to spomory-mcp (run which spomory-mcp to find it) — the client doesn't necessarily launch it with your shell's PATH set.

That's it. Data lives locally in ~/.memory-core/ by default (override with MEMORY_CORE_DATA_DIR). A Dockerfile is also included, for MCP directories/hosts that deploy from a container image instead.

Want the LLM calls local too, instead of a cloud API? Point LLM_BASE_URL at any local OpenAI-compatible server — vLLM, Ollama, llama.cpp, or MLX all work — and set EMBEDDING_PROVIDER=openai_compatible to route embedding the same way instead of the local sentence-transformers model, dropping the torch dependency entirely. Details and per-engine examples in docs/mcp_quickstart.en.md.

Related MCP server: Memory Engine MCP

What it can do

Once connected, six tools become available inside the client:

Tool

What it does

add_memory

Remembers something you tell it

search_memory

Recalls whatever's relevant to a question

forget_memory

Deletes the one thing that best matches what you asked to forget

forget_all_memory

Wipes the entire memory graph in one call

get_graph

Shows the memory graph around something, for inspection

export_memory

Exports everything you've stored, as JSON — your data, portable

All six are verified working end-to-end against real Claude Desktop, Cursor, and Codex CLI sessions, both local and remote — see docs/mcp_quickstart.en.md for what "verified" means for each client.

How it works, for the curious

You don't need any of this to use Spomory — it's here for people who want to know what's actually happening underneath.

  • Pluggable LLM / embedding providers: defaults to any OpenAI-compatible API (including Chinese-market LLM providers) + local sentence-transformers (default bge-m3, bilingual Chinese/English).

  • Dual-layer incremental knowledge graph: entities and relations are modeled as independent layers; new data is only extracted and merged in, never a full rebuild. Exact-match filler input ("thanks", "好的", "ok", ...) is skipped before it ever reaches the extraction LLM call, since it can't contain an extractable fact — relevant cost protection on any unauthenticated endpoint. Defaults to a local LocalGraphStore (networkx + SQLite); a PostgresGraphStore cloud implementation also exists, and both share the same behavioral contract test suite.

  • HippoRAG 2-style retrieval: the query is matched directly against triples rather than only against entity nodes; the matched seed nodes are diffused via personalized PageRank for multi-hop association, then assembled into a natural-language context (with source timestamps, so "when did I mention X" is answerable). Ranking on top of that decays a relation's relevance the longer it's gone without being retrieved, and boosts it back up (log-dampened, so it can't dominate PPR rank) the more times the same fact has been restated — a passive signal alongside the active ADD/UPDATE/DELETE/NOOP decisions below.

  • Memory management: an ADD/UPDATE/DELETE/NOOP action space, with a rule-based default policy (RuleBasedPolicy) and a full GRPO training pipeline (memory_manager/train_grpo.py, actually run and verified on a real GPU).

  • Memory passport export + true delete: a JSON-LD style export format, physical deletion, and an audit log.

  • Multimodal image verification: image captioning → reuses the text extraction pipeline → CLIP cross-checks candidate triples. Honestly positioned as "verification," not "native cross-modal extraction."

  • Cloud skeleton: FastAPI user auth/API keys/quotas, a Stripe webhook billing scaffold (skeleton-level only, not production-deployed).

Measured results

Real runs against DeepSeek on 84 QA pairs from LoCoMo-10 (conv-26, first 150 turns) — not cherry-picked, and not competitive with the bigger players' published numbers yet:

Metric

Value

Recall@10 (did the right evidence turn make it into context)

52.4%

Accuracy — strict substring match

19.0%

Accuracy — LLM-judged (looser, wording-tolerant)

44.0%

A prior run (before a fix that folds dates into extracted predicates so "when" questions are answerable) scored lower on accuracy but higher on recall (62.0%) — the fix traded some retrieval recall for a real +14.3-point accuracy gain, and we went and found out exactly why instead of just reporting the accuracy number: the date-folding instruction sometimes misfires on content-free small talk ("Thanks!" → "thanked on 2023-07-03"), and those extra low-value triples crowd out relevant ones out of the fixed top-10 retrieval window. Full numbers, per-category breakdown, and the side-by-side extraction comparison that found this are in docs/benchmark_smoke_test.md.

LongMemEval (xiaowu0162/longmemeval-cleaned oracle variant, first 10 of 500 questions):

Metric

Value

Recall@10

100% (10/10)

Accuracy — strict substring match

30%

Accuracy — LLM-judged

80%

The limitations here matter as much as the numbers:

  1. Only 10 questions, not the full 500 — each question ingests ~27 turns on average (~27 real extraction calls plus one generation and one judge call), and this environment's LLM API calls go through a proxy with real latency; the full dataset would take tens of hours. This is a real run, not a mock, but it's a small sample and shouldn't be read as generalizing to the full dataset.

  2. All 10 happen to be temporal-reasoning type — the dataset also has a multi-session type; load_longmemeval(limit=10) takes the first 10 entries in file order with no stratified sampling, so this sample isn't representative of the dataset as a whole.

  3. Recall@10 = 100% is largely an artifact of the oracle variant's design, not a strong retrieval claim — the oracle variant pre-filters each question's haystack down to only the relevant sessions (no distractor sessions), which is considerably easier than a real deployment's memory store (hundreds/thousands of unrelated turns). This isn't the same task as the full (non-oracle) LongMemEval benchmark and shouldn't be compared directly against numbers other products report on that harder variant.

  4. Strict-match accuracy (30%) is far below LLM-judged accuracy (80%), consistent with the same pattern seen in the LoCoMo results — substring matching systematically undercounts answers that are correct but worded differently.

Raw data: benchmarks/results/longmemeval_oracle_subset.json; the run script is benchmarks/run_longmemeval_subset.py.

For contributors: building from source

Everything below is for people who want to hack on the internals, run the test suite, or reach dependency groups beyond what running the MCP server needs (GPU training, the cloud API skeleton, multimodal verification). If you just want to use Spomory, you don't need any of this — see "Get started" above.

Project layout

src/
├── memory_core/
│   ├── graph/            # entity/relation models, storage adapters (local SQLite / cloud Postgres), incremental writes
│   ├── retrieval/        # query→triple matching, personalized PageRank, context assembly
│   ├── memory_manager/   # action space, reward functions, GRPO training script, policy inference
│   ├── multimodal/       # image captioning + CLIP verification
│   ├── mcp_server/       # MCP Server (the distribution entry point)
│   ├── export/           # memory passport export format + true delete
│   ├── llm/              # pluggable LLM/embedding providers
│   ├── audit.py          # deletion audit log
│   └── usage.py          # retention/usage tracking
└── cloud_api/            # FastAPI cloud service skeleton (auth, quotas, billing)
benchmarks/                # LoCoMo/LongMemEval evaluation harness + multimodal comparison experiments
tests/                     # 94+ tests, from unit tests to real LLM/GPU/Postgres end-to-end verification
docs/                      # per-epic design notes, verification reports, runbooks (see index below)

Set up a dev environment

Prerequisites: Python 3.11+, uv (no uv? python -m venv + pip install -e works as a substitute for the uv commands below).

git clone <this repo's URL> memory-core && cd memory-core
uv venv --python 3.11 .venv
source .venv/bin/activate          # Windows: .venv\Scripts\activate

# Pick dependency groups as needed — they can be combined, no need to install everything:
uv pip install -e ".[dev]"                 # required to run tests/lint
uv pip install -e ".[llm,embedding]"       # required for the "minimal working memory system" (see the demo below)
uv pip install -e ".[mcp]"                 # extra: connecting to Claude Desktop/Cursor
uv pip install -e ".[rl]"                  # extra: GRPO training (requires a GPU + CUDA)
uv pip install -e ".[cloud]"               # extra: cloud API / Postgres backend
uv pip install -e ".[multimodal]"          # extra: image + CLIP verification

The embedding group downloads the default model BAAI/bge-m3 from HuggingFace on first use (~2.2GB) — make sure huggingface.co is reachable (if you're behind the Great Firewall, export HF_ENDPOINT=https://hf-mirror.com routes through a mirror). You can also swap in a smaller model via export EMBEDDING_MODEL=<any sentence-transformers model name>.

Run a minimal example (no MCP, plain Python calls)

With dev + llm + embedding installed and the three env vars from "Get started" set, this script exercises the full "write a memory → retrieve it" pipeline directly (the same logic behind mcp_server/server.py's add_memory/search_memory tools, just calling the library directly instead of going through the MCP protocol layer):

# demo.py
from memory_core.graph.local_store import LocalGraphStore
from memory_core.graph.incremental import IncrementalIngestor
from memory_core.llm.openai_compatible import OpenAICompatibleProvider
from memory_core.llm.local_sentence_transformer import SentenceTransformerProvider
from memory_core.memory_manager.policy import RuleBasedPolicy
from memory_core.retrieval.ppr import personalized_pagerank, rank_entities
from memory_core.retrieval.query_match import match_query_to_triples
from memory_core.retrieval.ranker import build_context

store = LocalGraphStore("demo.sqlite3")          # a local file; delete it to reset
llm = OpenAICompatibleProvider()                 # reads LLM_API_KEY etc. from the environment
embedder = SentenceTransformerProvider()         # downloads bge-m3 on first run

# 1. Write a memory: the LLM extracts triples, incrementally merged into the graph
ingestor = IncrementalIngestor(store, llm, policy=RuleBasedPolicy())
result = ingestor.ingest("I do AI research at CAS, mostly in Python.", source_id="demo")
print(f"added {result.new_entities} entities, {result.new_relations} relations")

# 2. Retrieve: match the query against triples -> PPR diffusion -> assemble a natural-language context
query = "Where do I work?"
entities, relations = store.all_entities(), store.all_relations()
entities_by_id = {e.id: e for e in entities}
matches = match_query_to_triples(query, relations, entities_by_id, embedder, top_k=10)
seed_ids = {r.relation.subject_id for r in matches} | {r.relation.object_id for r in matches}
scores = personalized_pagerank(entities, relations, seed_entity_ids=list(seed_ids))
ranked_ids = [eid for eid, _ in rank_entities(scores)]
print(build_context(relations, entities_by_id, ranked_ids, top_k=10))
python demo.py

Here's real output from a live run against DeepSeek with the exact input shown above (not fabricated, not cleaned up — this is what actually came back):

added 3 entities, 2 relations
I do AI research at CAS (recorded at 2026-09-05 10:40:00).I do AI research mostly in Python (recorded at 2026-09-05 10:40:00).

Exact wording and entity/relation counts depend on the LLM's own extraction and will vary between runs, but as long as the env vars are set correctly, non-empty output means the pipeline works end to end. retrieval/ranker.py detects whether a relation's text is CJK or not and renders it accordingly (no spaces + a Chinese timestamp label for CJK, spaced words + an English timestamp label otherwise), so English input no longer comes out as one run-on word like earlier versions of this demo did.

Testing

pytest                    # everything
pytest -m "not slow"      # skip tests that download models / train — runs in seconds

Most of the "slow" tests aren't mocked — they're real calls (real LLM API, real local embedding model, real CLIP model) and need the corresponding env vars (LLM_API_KEY, etc.) or an already-downloaded model cache.

Documentation index

Doc

Content

mcp_quickstart.en.md (中文)

MCP Server install, configuration, connecting Claude Desktop/Cursor/Codex CLI, real-world gotchas

graph_store_interface.en.md (中文)

Storage adapter interface design

export_format.en.md (中文)

The "memory passport" export format

dataset_format.en.md (中文)

GRPO training data format and how the real dataset was generated

methodology.en.md (中文)

Technical methodology: what's actually verified vs. still open

benchmark_smoke_test.en.md (中文)

Real LoCoMo benchmark results and failure-case analysis

memory_manager_eval.en.md (中文)

Rule-based vs. GRPO-trained policy comparison, including the debugging process

multimodal_verification.en.md (中文)

Image + CLIP verification experiment results

gpu_training_runbook.en.md (中文)

GPU training environment setup log (including real gotchas hit)

postgres_setup.en.md (中文)

Cloud Postgres backend deployment log

leaderboard_submission.en.md (中文)

Third-party leaderboard research

mvp_scope.en.md (中文)

MVP scope definition

privacy_policy_draft.en.md (中文) / product_copy_memory_passport.en.md (中文)

Draft privacy policy / external-facing product copy

eng_note_cjk_rendering_bug.en.md (中文)

Engineering note: a real discover→fix→verify trace for a CJK rendering bug

Every doc above now has both a Chinese and an English version.

Known limitations

  • Multimodal verification for voice input (ASR + audio embedding) isn't implemented yet.

  • The GRPO training dataset (140 real samples) and the number of training steps are still small; memory_manager_eval.md honestly documents how that limits training effectiveness.

  • The cloud API/billing is skeleton-level only and hasn't been connected to a real production environment.

License

See LICENSE.

Available Tools

6 tools
add_memoryBInspect

Extract facts from text and write them into the memory graph.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
source_idNomcp-session

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, indicating a write operation, and destructiveHint=false. The description adds that it extracts facts from text, which is the core behavior. However, it does not disclose potential side effects like deduplication, failure handling, or whether existing memories are overwritten, leaving some behavioral ambiguity beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no filler. It front-loads the verb and resource, making the purpose immediately clear. Every word contributes to the core message.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with an output schema, the description covers the main action but lacks usage guidance and source_id semantics. The output schema exists, so return format is not needed. However, an agent would not know when to use this over alternatives or what source_id does, making the description incomplete for confident invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explicitly mentions 'text' as the source of facts, which adds meaning. However, it completely ignores 'source_id', leaving its purpose and default behavior unexplained. With two parameters and no schema descriptions, the description only partially covers the parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Extract facts from text and write them into the memory graph.' It clearly conveys the add operation, distinct from siblings like search_memory, forget_memory, and get_graph, which are implied by the memory graph context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. It doesn't mention conditions, prerequisites, or compare with export_memory or search_memory. An agent must infer the usage from the sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_memoryB
Read-only
Inspect

Export the full memory graph as a JSON memory passport.

ParametersJSON Schema
NameRequiredDescriptionDefault
subject_idNodefault

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, covering the safety profile. The description adds the scope ('full memory graph') and output format, which is useful context. However, it does not disclose any potential limitations, auth requirements, or behavior beyond the annotation baseline, so it adds moderate value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that front-loads the primary action and output. It contains no redundant phrases and every word adds meaning, making it highly efficient for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While the output schema exists (so return format is covered structurally), the description leaves the optional subject_id parameter unexplained and provides no guidance on when to use this tool over get_graph. For a simple tool with one optional parameter and read-only annotations, the description is adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage for the single parameter subject_id, and the tool description does not mention it at all. The parameter is optional with a default, but the agent has no explanation of what it controls (e.g., which memory subject to export). The description completely fails to compensate for the missing schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Export'), a clear resource ('the full memory graph'), and the output format ('JSON memory passport'). It distinguishes itself from siblings like get_graph, which likely retrieves a graph in a different manner, and from add/search/forget tools that mutate memory.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly state when to use this tool versus alternatives, nor does it mention exclusions or prerequisites. An agent must infer that this is for a complete export, but no guidance is given about when to choose it over get_graph or other memory tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

forget_all_memoryA
Destructive
Inspect

Permanently delete the entire memory graph -- every entity and relation.

forget_memory is deliberately one-fact-at-a-time; this is its bulk counterpart for a user who wants a clean slate (e.g. before re-testing, or a genuine full data wipe) instead of narrowing and retrying a query N times.

destructive_hint=True on this tool's annotations is only a hint -- the MCP spec doesn't require a host to gate on it, so a host that treats it as advisory (or an agent auto-approving destructive tools) could otherwise wipe everything on the first call with no human in the loop at all. confirm is the actual enforcement: the first call (confirm left at its default, False) never deletes anything -- it only reports what a real call would remove -- so triggering a wipe needs an explicit second call with confirm=True, regardless of what the host's UI does or doesn't show.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide destructiveHint=true but the description adds essential behavior beyond that: the first call with confirm=false is a dry run that only reports what would be removed, and only a second call with confirm=true actually deletes. It also explains that destructive_hint is just a hint and may not be enforced by hosts, making the confirm parameter the real enforcement. This is substantial non-annotation context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four sentences and each carries new information: core action, sibling contrast, host enforcement caveat, and confirmation protocol. It is front-loaded with the primary purpose, but it is slightly longer than strictly necessary; the caveat about destructive_hint could arguably be condensed. Overall, well-structured but not maximally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is one parameter, and the description explains it exhaustively, including its confirmation semantics. Return values are covered by the output schema. The tool's destructive nature, safety mechanism, and relation to siblings are all described. No essential calling information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the only parameter is a boolean `confirm` with default false. The description fully explains its meaning: the default leaves it a no-op reporting call, and confirm=true triggers the actual wipe. This completely compensates for the missing schema description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Permanently delete the entire memory graph -- every entity and relation,' which states the exact verb and resource. It then explicitly contrasts with the sibling `forget_memory` ('deliberately one-fact-at-a-time') by calling itself the 'bulk counterpart,' so an agent can disambiguate without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names the alternative tool (`forget_memory`) and gives the specific condition for choosing this one: a user who wants a clean slate (e.g., before re-testing or a genuine full data wipe) instead of narrowing and retrying a query. This directly answers when to use it versus the sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

forget_memoryA
Destructive
Inspect

Find the single fact that best matches query and permanently delete it.

This is the user-facing counterpart to the "true delete" backing export_memory's data-ownership promise (Epic 7.3): without it, that capability existed at the storage layer but a user had no way to actually invoke it from a conversation (e.g. "forget that I work at X"). Deletes at most one relation per call, on purpose -- a query vague enough to match many facts should be narrowed and retried rather than risk deleting the wrong ones silently.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, so the description doesn't need to restate that. It adds valuable behavioral context: the operation is permanent, it deletes at most one relation per call, and it is intentionally conservative to avoid silent mass deletion. It also explains the design rationale (user-facing counterpart to the storage-layer capability). This goes beyond the annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the first sentence states the core action and scope. The second sentence provides essential context and a safety constraint. Every sentence earns its place, and there is no filler or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a single parameter, a clear destructive annotation, and an output schema, so the description doesn't need to explain return values. It covers the key behavioral constraints (permanence, single deletion, narrowing vague queries) and the rationale. The only minor gap is that it doesn't explicitly state what happens when no match is found, but that is a small omission given the overall completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does: it explains that `query` is the fact to match and that the tool finds the single best match. It also implies the query should be specific enough to match exactly one fact, which is critical usage semantics. It doesn't provide format examples, but for a single free-text string parameter, the description gives sufficient meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('find' and 'permanently delete') and a specific resource ('the single fact that best matches `query`'). It clearly distinguishes itself from siblings by emphasizing it is the user-facing counterpart to the 'true delete' backing `export_memory`'s data-ownership promise, and it explicitly contrasts with `forget_all_memory` by deleting at most one relation per call. This makes the tool's purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly explains when to use this tool: when a user wants to forget a specific fact from a conversation (e.g., 'forget that I work at X'). It also provides a clear exclusion: queries vague enough to match many facts should be narrowed and retried rather than risk deleting the wrong ones silently. This gives an agent actionable guidance on when to invoke it and when to avoid it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_graphB
Read-only
Inspect

Return the subgraph around entity_name as JSON.

ParametersJSON Schema
NameRequiredDescriptionDefault
hopsNo
entity_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already declaring readOnlyHint=true and destructiveHint=false, the description does not need to re-state safety. It adds that the tool returns a subgraph in JSON, which is useful behavioral context beyond the annotations. However, it does not explain what the 'hops' parameter affects or whether there are any limits or side effects. This is a modest addition, so a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence of nine words. It is front-loaded with the key action and resource, and contains no filler. Every word earns its place. This is the ideal level of conciseness for a simple read-only query tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return values do not need explanation. However, the description is incomplete for a graph traversal tool with two parameters: `hops` is entirely missing, and the relationship to sibling tools (e.g., search vs. graph retrieval) is not clarified. An agent using this tool would need to understand that `hops` controls depth to call it correctly, and that information is absent. This is a significant contextual gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for both parameters. It only mentions `entity_name` in the main text, clarifying that it is the center of the subgraph, but it says nothing about `hops`, which is an integer with a default of 1 and clearly controls traversal depth. With one of two parameters entirely unexplained and the other only minimally, the description fails to provide adequate semantics for the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Return the subgraph around `entity_name` as JSON' clearly states a specific verb (Return), resource (subgraph), and the output format (JSON). It does not explicitly name sibling tools or differentiate itself, but the function is obvious from the phrasing. Without sibling differentiation, it lands at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like search_memory, export_memory, or add_memory. It neither gives context for the typical use case nor excludes alternatives. The only implicit hint is the phrase 'subgraph around entity_name,' but that is not enough to count as guidance. This is a clear gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_memoryCInspect

Retrieve and assemble a natural-language context relevant to query.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
top_kNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.6/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description says 'Retrieve and assemble a natural-language context,' which describes a read-only operation, yet annotations declare readOnlyHint: false. This direct contradiction makes the behavior unclear. No additional behavioral context such as side effects, permissions, or output assembly process is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no filler and key information front-loaded. It is concise, though it omits useful detail about parameters and usage that would make it more helpful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so return value structure is covered. However, the description does not explain the meaning of `top_k`, does not offer usage guidance, and is contradicted by the annotation. For a simple two-parameter search tool, the gaps are notable but not fatal.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description only references `query` and does not mention `top_k` at all. With schema description coverage at 0%, the schema offers no parameter definitions, so the optional parameter's purpose, behavior, and default are left unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses specific verbs ('Retrieve and assemble') and identifies the resource (natural-language context relevant to the query). It is distinguishable from siblings such as add_memory, forget_memory, and export_memory, though it does not explicitly name memory as the source or compare itself to alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus export_memory, get_graph, or other sibling tools. There are no exclusions, prerequisites, or contextual signals to help an agent decide between alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.2.1
    • Changedforget_all_memory1 field changed
      • addedInput schema / properties / confirm
        Added value: +{
        +  "default": false,
        +  "title": "Confirm",
        +  "type": "boolean"
        +}
  2. 6 tool updatesv0.1.0
    • First observedadd_memory
    • First observedexport_memory
    • First observedforget_all_memory
    • First observedforget_memory
    • First observedget_graph
    • First observedsearch_memory

TDQS

A3.7/5.0

Scored across 6 tools

Disambiguation5/5

Each tool targets a clear, distinct operation: add, graph lookup, semantic search, single-fact delete, full wipe, and export. The two delete tools are explicitly differentiated by scope, and get_graph vs search_memory are distinguished by input type (entity vs natural-language query) and output format.

Naming Consistency5/5

All tool names follow the same lower_snake_case verb_noun pattern: add_, get_, search_, forget_, forget_all_, export_. The shared 'memory' base and the obvious 'forget_memory' / 'forget_all_memory' pair make the naming highly predictable.

Tool Count5/5

Six tools is a well-scoped size for a memory-graph server, covering creation, retrieval, deletion, and export without redundancy or bloat. Each tool earns its place and the count sits comfortably in the ideal 3–15 range.

Completeness4/5

The core memory lifecycle is covered: add, retrieve (graph and semantic), delete single, delete all, and export. Minor gaps exist—there is no explicit update operation and no import counterpart to the export—but agents can work around these via add_memory and re-adding facts.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Gives AI persistent personal memory with hybrid search, temporal decay, and knowledge graph. Works with Claude and any MCP client.
    73 npm
    8
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Provides persistent, graph-based memory for AI agents via MCP, enabling semantic search, wikilink traversal, reminders, and injection protection.
    9
    31
    Apache 2.0
  • A
    license
    Not graded
    quality
    A
    maintenance
    Provides AI agents with persistent, human-like memory infrastructure via MCP, enabling them to store, search, summarize, and forget episodic, semantic, procedural, and working memories across sessions.
    752 npm
    MIT