memory-arbiter-mcp
Summary: Memory Arbiter MCP is a local-only SQLite fact layer for AI agents, exposing four MCP tools for storing, recalling, reviewing, governing, and repairing traceable memories with evidence-based conflict detection.
memory— daily operations:remembera fact (requiresworkspace),find/batch_find(up to 8 merged queries) as a metered index page withpreview/hits/fullcontent modes,readexact source windows viaspan,updatein place (patches or full rewrite),judgea conflict,status,help.memory_review— read-only inspection: overview, doctor health, conflict groups/details, version history, expired memories, audit log, entities.memory_govern— explicitly user-authorized governance: retire, merge near-duplicates, apply/replan/resolve conflict plans, confirm memories, and manage workspaces (rename, migrate, move, confirm pending/registry).memory_repair— maintenance: evidence rebuild, conflict scans (scan_pipeline/scan_queue/scan_candidates/scan_duplicates), history cleanup, entity assignment, pending activation, backup replay, notice lifecycle, semantic runtime control.Recall — hybrid lexical (FTS5) + local GGUF evidence retrieval fused by reciprocal rank, with trust, recency, filter, and workspace adjustments.
Trust & provenance — every memory keeps full original text plus
source_type,source_ref,event_time,ingest_time; protection levels (normal/protected/locked),user_confirmedlocking, and full append-only version history with supersede chains.Conflict handling — one immutable
conflictstable per one-to-many event, lifecycleopen → applying → resolved(ornot_a_conflict), with an optional local mDeBERTa three-class judge producing advisory notices; the judge never picks a winner or edits memory.Workspaces & isolation —
none/weak/strictmodes, canonical normalization, guarded vector admission understrict,defaultpool insulation, and a liverecall_blacklist.jsonlfor unscoped recall.Deployment — stdio by default or opt-in localhost
streamable-httpsharing one server across clients; identity required, config file-only with 21 keys, everything local except an optional PyPI update check.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@memory-arbiter-mcpsave that the user prefers dark mode"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Memory Arbiter MCP
English | 中文
Memory Arbiter is a trustworthy local fact layer for AI agents — not just shared memory, but shared facts that are current, trusted, traceable, and safe to use. It is a local SQLite service exposed over MCP: four product tools, evidence-based recall, advisory conflict notices, and user-authorized governance. Every fact is stored once in local SQLite and every model it can call runs locally.
Current release:
0.17.1(mDeBERTa three-class judge with the claims channel retired; workspace required on remember + workspace-govern actions,workspacesview; automatic row-vector backfill after upgrade; full-project review fix wave).
Why trust it
One complete source of truth. Every memory keeps its full original text. Evidence vectors, full-text search, and rankings are all derived indexes — rebuildable, never the only copy.
Provenance on every write. Each memory carries
source_type,source_ref,event_time, andingest_time. Theuser_confirmedlabel is reserved by convention for facts the user explicitly verified; technically enforced protection is what happens after labeling — auser_confirmedmemory is locked against silent edits.Trust levels.
normal/protected/lockedprotection levels prevent an agent from silently overwriting what is locked;memory_govern(confirm)promotes a memory touser_confirmedonly with per-action user authorization.Full version history. Every edit appends to
memory_historywith a version bump, and supersede chains keep old facts traceable instead of silently replaced.One conflict record per event. A single
conflictstable holds the immutable detection snapshot, value groups, decision, and application results for each one-to-many conflict event. The judge proposes no winner and never edits memory.Authorized governance. Every state-changing
memory_governaction requires per-actionauthorized=trueafter the user confirms that specific action.Local-only. Embeddings run on a local GGUF model; the optional conflict judge is a local mDeBERTa checkpoint (CPU). The single outbound call is an optional PyPI update check, disabled with
update_check.enabled=false.
Related MCP server: memory-mcp
Benchmark: LoCoMo-Refined
Evaluated on LoCoMo-Refined
(all 1,382 questions, official judge Qwen3-14B):
Configuration | Overall | Text-only |
mema + writer-agent (deepseek-v4.1-flash writer, public rules) | 61.00% | 65.62% |
mema verbatim (zero-LLM write path, utterances + | 43.13% | 49.48% |
Context: Mem0's full pipeline scores 48.91% (text-only re-scored by the benchmark authors) — bare mema with no LLM in the write path matches it; with a writer-agent, mema sits second on the leaderboard's text-only basis. The write path also detected 552 semantic conflicts across the 20 replayed conversations (self-trained mDeBERTa judge, running as shipped).
Full methodology, ablations (including answerer-independence and the 82%
information ceiling), reproduction steps and submission files:
evals/locomo/RESULTS.md.
Install & quickstart
Install with your AI Agent
Paste this into Codex, Claude Code, Cursor, or another coding agent with terminal access:
Read the latest README at https://github.com/billy12151/memory-arbiter-mcp.
Install and configure the latest mema release for my operating system and current AI client.
Preserve any existing config and database; do not overwrite or delete existing data.
Ask me before choosing between materially different install modes, changing existing config,
or performing any destructive or privileged action. When finished, run mema doctor and report
the install method, config path, database path, client integration, and verification result.The agent should treat this README as the source of truth, inspect the local environment before choosing uvx, core, vec, or semantic-local, and stop for user input when a safe choice cannot be inferred. A successful install is not complete until mema doctor has run and any warning has been reported.
Install manually
# One command: package + deps + embedding model + config.json (resumable downloads, ModelScope fallback)
curl -fsSL https://memarbiter.cn/install.sh | bashOr step by step:
pip install "memory-arbiter-mcp[vec,semantic-local]" # core + sqlite-vec + local GGUF runtime
mema setup --install # downloads the embedding GGUF (~330MB, resumable); the mDeBERTa judge is a separate manual install
mema doctor # verifymema setup without --install stays guidance-only: it writes ~/.config/memory-arbiter/config.json and self-checks the environment without touching pip or the network; --install is the execution mode (pip installs the extras, downloads the embedding GGUF model with resume/mirror fallback, and writes the finished config itself). Since 0.15.0 configuration is file-only and the whole user surface is 21 keys (see Configuration): paths, identity, workspace/isolation, update_check.enabled, include_size, the embedding model, the optional semantic-conflict mDeBERTa judge (mdeberta_ckpt, see below), and MCP transport/host/port. The reference examples/memory-arbiter.config.example.json shows the same slim surface with per-key notes. Then wire your MCP client from examples/*.mcp.json and start the server with mema.
When a capability is missing (e.g. the models were never downloaded), every tool response carries a persistent degraded-mode banner with the mema setup --install remediation, and each agent's first call includes a capability health card — an incomplete install cannot pass for a complete one silently.
The server requires an explicitly configured identity: set client and agent_id in config.json or the MEMORY_ARBITER_CLIENT/MEMORY_ARBITER_AGENT_ID launch-context environment variables (the stdio examples/*.mcp.json entries do this via env). There are no built-in defaults — the server refuses to start when either is blank. Under stdio this configured identity is the process-level caller identity used for attribution; memory(action="remember") does not accept agent_id/client in data. streamable-http takes caller identity from the per-request headers described below.
stdio remains the default. For one local server shared by several clients, set mcp.transport to streamable-http (or MEMORY_ARBITER_MCP_TRANSPORT=streamable-http, one of the six retained launch-context variables) and connect to http://127.0.0.1:8000/mcp. Each client's MCP server entry must set fixed X-Mema-Client and X-Mema-Agent-Id headers; see examples/streamable-http.mcp.json. The client sends them automatically on every HTTP MCP request—agents should not add identity to individual tool calls. Missing, empty, invalid, duplicated, or conflicting identity is rejected instead of falling back to defaults. Community HTTP mode binds only to localhost, and these headers are advisory provenance, not authentication or multi-tenant isolation.
The daily loop is four calls — remember a reusable fact, find to recall, read for exact lookup, update when a newer source replaces an existing current memory (never create a second active copy of one source of truth). remember requires workspace (0.17.1): first call memory_review(view="workspaces") to list existing buckets, then pass an existing canonical name — 'default' is the global pool; pass a new name only when deliberately creating a new bucket. Point any agent at the packaged rule:
{"action":"help","data":{"topic":"agent_onboarding"}}The four tools
memory:remember,find,batch_find,read,batch_read,update,judge,status,helpmemory_review: read-only health, conflict groups/details, history, expired memory, audit, and entitiesmemory_govern: explicitly authorized retirement, conflict-plan application/resolution, confirmation, and workspace governancememory_repair: evidence rebuild, broad conflict scanning/recording, history cleanup, entity assignment, pending activation, backup replay, semantic runtime control, and notice lifecycle
Every product call returns the envelope {ok, mode, warnings, degraded, data}. Operation-specific action_required, next_action, replan, and records live under data; successful calls may additionally carry a top-level notices array. Each notice has its own action_required and machine-readable call under the notice object. Do not look for a generic top-level action_required.
batch_find runs up to 8 queries in one call and merges the pages (dedup by memory_id; each item carries matched_query_ids/best_query_id; per_query reports per-query stats) — for multi-topic tasks this replaces 5-10 tool round-trips with one. Since 0.15.9 find is also honest about emptiness: a query that recalls nothing returns an empty page with a reword hint (no recent-memory stuffing), and candidates below the calibrated relevance floor (8.1) never enter a query-recall page at all — fewer "looks related, isn't" citations.
find is an index page: by default (content_mode="preview", 0.15.10) each result carries metadata plus content_chars (the full-text length — what a read would cost) and a bounded outline of up to 8 {head, offset} segments whose offsets share read's span coordinate system, so span=[offset, offset+N] slices that exact segment. Keyword query style (0.17.0): a query of space-separated short CJK words (each ≤4 chars, e.g. 向量 唯一键 冲突) is understood as keywords — memories semantically close to the query whose content/subject contains one of the keywords (whole-word substring) rank higher; topic-generic words that flood the candidate pool are ignored, and no forced matching happens when a keyword appears nowhere (reword instead). Content depth is a single-choice enum: content_mode="hits" adds hit_spans — the vector-matched units as {text, start_offset, end_offset} sliced from the source (read span=[start,end] returns exactly that text). Hit spans are never truncated: when merged hits cover ≥50% of the content the item upgrades to full text with hit_spans kept as an annotation; items without vector hits keep the plain preview shape. hit_window=N (default 0) extends each hit with ±N neighbouring complete sentences (neighbours marked matched=false); hit_spans appears only on query-recall pages — browse/filter pages carry none — and stale-version hits are dropped with a stale_hit_spans marker plus a re-query warning (the evidence index may lag right after an edit). content_mode="full" returns whole texts (the old include_content=true, removed in 0.15.10 — the call fails loudly with a migration pointer). Scores compare only within the page, and if the top page misses you should reword the query or add tags_filter rather than deep-page — unfiltered query-recall reports total_estimate=null/has_more=false, while filtered recall keeps the exact count. The size block meters the page as actually returned: returned_chars/returned_count and a tokens_estimate from a deterministic bucket-table estimator (heuristic_v1) calibrated against a Qwen2.5 tokenizer at design time (calibration reference only — no Qwen runtime ships since 0.17.1); it runs ~30% high on pure Chinese prose and ~17% high on pure English — the estimate and the estimated share one yardstick, so savings comparisons stay valid. Since 0.15.6 the same size block rides every recall surface — read (meters the record as returned, span windows included), memory_review expired and history (meter their result lists) — under one global config key include_size (default true); each block's display_hint repeats the token number with a report-this-recall-cost instruction, include_size=false turns all of them off together, and find's old per-call include_size parameter is ignored with a warning. unresolved_conflict_count appears only when page items directly hit an open/applying conflict group, and counts those page items.
How recall works
Lexical and evidence channels recall independently and merge per memory with reciprocal-rank fusion, then trust, recency, filter, and workspace adjustments.
Lexical: FTS5 over content plus subject/tags LIKE; the bounded content-LIKE anchor channel runs only when vectors are unavailable (degradation path since 0.16.10).
Evidence: the write-path job derives row segments from the
subject(leading subject row), sentences, and table rows (0.17.0: the former unit tables retired; row vectors are the one evidence channel). The indexer never extracts facts, infers entities, or calls a model — it only slices the stored source. Evidence hits carry source offsets.memory(action="read", data={"memory_id": 42, "span":{"start":120,"end":640}})returns only that clipped source window plusdata.span.{start,end,total_chars}; omitspanto read the complete source. Span bounds are strict integers with0 <= start < end, andendclips at content length. Table segments longer than 100 rows are exempt as a whole (no row vectors, no pair detection — visible astable_rows_exemptedin receipts).
Conflict groups and notices
Evidence KNN recalls sentence-level neighbours; it does not decide conflict truth. Since 0.17.1, each candidate pair that survives the deterministic funnel (sentence prefilter, memory-level screen, cosine band, rule evidence) is judged by an optional local mDeBERTa encoder (three classes: conflict / no_conflict / possible_conflict, plus a mechanism head). The judge sees the two bare row texts — no metadata, no prompt engineering — and returns calibrated probabilities:
conflictat P ≥semantic_conflict.mdeberta_notice_min_prob(default 0.5) → a normal-severity notice;possible_conflict→ aninfo-severity notice (grey-zone reference, noaction_required);no_conflict→ silent clear (counted in the receipt).
The judge never chooses a winner, suppresses the scheduled scan, or edits memory.
There are deliberately two gates:
Scheduled scan is broad.
memory_repair(task="scan_candidates")retains deterministic KNN/rule candidates; the scan path runs no model — the Agent reviews candidates there.include_duplicates=truestays a single-page spot check (full record_conflict-compatible members for that page); for a full duplicate sweep use the separatememory_repair(task="scan_duplicates")task, which aggregates every page server-side under one global 200-pair cap. The external reviewer records every triaged candidate withrecord_conflict(status="open"|"not_a_conflict")to obtain snapshot dedupe.Write-time notice is judge-gated. A user-visible notice requires the judge to say
conflictabove the configured probability floor (or the deterministic same-key/value-difference direct verdict). The slot identity ridesworkspace + subjectplus a pair-hash difference anchor (the deterministic direct path keeps the real attribute). Anything less fails open into later scan review; internal (same-memory) findings always land pending — the scan-side strong model owns negative verdicts. The synchronous-delivery wait (configurablesemantic_conflict.notice_sync_wait_ms, default3000, clamp0–5000) only gates when a notice is attached; it does not change detection.
Set up the scheduled tasks. mema does not ship an internal timer by design: the pipeline's value loop ends in agent-side judgment, so the external scheduler is what wakes the agent. 0.16.0 spec v2 replaces the v1 page-driven triage loop: an hourly task kicks memory_repair(task="scan_pipeline", data={"action": "kick"}) repeatedly until complete=true (the server walks the library itself, decides full-vs-incremental from per-memory watermarks, and lands every suspected item in an agent-only judgment queue — never in notices or user-facing conflict lists), then drains it with memory_repair(task="scan_queue", data={"action": "page"}) / action="submit" (server-side land-from-reference: dismiss → suppression source, confirm → the agent supplies only slot_key + display values). Keep a daily memory_review(view="doctor"). scan_candidates remains as a manual/diagnostic channel only; conflicts.spec_drift flags v1-era tasks for rebuilding. Each completed full-scan boundary (a page returning next_anchor_memory_id=null with anchors scanned) appends one lightweight audit line to scan_log.jsonl; until then agents receive a scan_never_run/scan_stale guidance notice, and doctor flags a still-owed rebuild (conflicts.scan_required) or a scan idle beyond 14 days (conflicts.scan_stale) — once the tasks run, both fall silent on their own. The full platform-agnostic spec is memory(action="help", data={"topic": "scheduled_tasks"}).
The single conflicts table stores one one-to-many event and its immutable member/value snapshot. Its public lifecycle is open → applying → resolved, with not_a_conflict as a terminal triage result. memory(action="judge") CAS-pins the conflict revision, records the chosen value and plan, and moves it to applying; execute each returned memory_govern(action="apply_conflict_action") sequentially with explicit authorization and the latest revision, then call authorized resolve_conflict only after every planned member action completes. use_as_resolution must land on a member that holds the chosen value group. The default judgment corrects the wrong data in memory — plan update_current_claim or append_superseded_context for members holding superseded claims, and use preserve_historical_record only when the user explicitly asks to keep the historical record. Partial failures remain applying: when data.action_required="replan_conflict", re-read the group/members and call authorized memory_govern(action="replan_conflict") with the current revision and replacement plan — a grounding-failed update_current_claim is recovered either by re-editing so the chosen value appears verbatim in the member content (the stored machine-normalized form; for multi-word phrases that means the normalized string itself), or by passing a replacement chosen_value drawn from the group's value_groups (0.15.15; replacing chosen_value also moves the resolution holder to the new value group unless pinned explicitly). Replanning preserves prior plan history; never retry stale precomputed steps.
Workspaces
Workspace canonical normalization runs in every isolation mode and is separate from access control. none applies no workspace ACL: an omitted workspace spans the library, while an explicitly supplied workspace is canonicalized and scopes that read. weak adds a soft ranking/hint signal (a fixed binary nudge — the continuous vector-distance weighting is no longer a knob). Under strict, the system never silently merges a near-match: a new workspace stays pending until authorized memory_govern(confirm_pending_workspace) activates it. Strict visibility uses guarded vector admission (always on since 0.15.0, a frozen constant): workspace-sensitive recall/read/repair operations, conflict/notice workflows, and console content/count views share one admitted set: the caller canonical plus every canonical at or below a 0.25 cosine cutoff after default-pool, short-name, and generic-substring guards. Process-global maintenance (for example semantic runtime control, backup replay, doctor, and settings) is not a workspace-scoped content view. Missing vectors or sqlite-vec degradation fall back to the exact caller canonical. The reserved default pool is insulated and is not visible from a strict project scope. Automatic vector normalization affects only the memory's workspace_canonical; supported workspace governance uses rename, migrate, move-by-id (move_memories_workspace), pending confirmation, and full-registry confirmation. Internal redirect/negative-decision state prevents old names from re-splitting and suppressed candidates from reappearing, but is not a user-facing workflow.
The first successful write that registers a canonical workspace returns a non-blocking top-level workspace_review notice in none/weak, plus data.write_hints.new_workspace_detected. Review possible duplicates before running authorized confirm_workspaces. strict instead returns the existing blocking action_required=confirm_new_workspace flow and does not emit the duplicate non-blocking notice.
Recall blacklist (0.15.5)
Workspaces listed in recall_blacklist.jsonl (next to the database file) are excluded from unscoped find recall — the default ambient pool. One bare workspace name per line; blank lines and # comments are ignored; edits are live on the next find. No file → the built-in default applies (mema-twin, the mema-twin preference bucket); an empty file → nothing is excluded; a created file replaces the default entirely.
What is not filtered: an explicit workspace argument (even a blacklisted one), strict isolation, filter-driven recall (empty query + tags_filter/time/source_type), the expired-audit path, write-time dedup, conflict scans, and id-based reads. Exclusion matches both the canonical and the raw workspace column, so rows whose canonical drifted are still caught. doctor reports the effective list (recall.blacklist).
Operating mema
mema doctor [--json|--deep]— read-only health checks;--deeploads the GGUF model and probes the live embedding dimension.workspace.reviewwarns (CLI exit 1) for canonicals missing from the reviewed snapshot. Rename/merge duplicates first, then call authorizedmemory_govern(confirm_workspaces)without an explicit list to snapshot the current registry and return this check to pass. The overall CLI exits 0 only when no other warning remains.mema console— read-only local console on 127.0.0.1.memory(action="status")— surfaceslocal_text_evidencecoverage,vec_index_state, the process-local index queue, andsemantic_conflictruntime including queue drops/restarts andcheck_degradation.last_reason.Maintenance tasks on
memory_repair:rebuild_evidence(dry-run then batched execute; after an embedding-model change the index reportsstate=mismatchand rebuild flips it back toreadyautomatically),semantic_control(status/pause/resume/enable/unload/disable),replay_backup(dry-run then authorized execute),cleanup_history,set_entity,activate_pending, andscan_duplicates(a one-call full-library near-duplicate sweep bounded at 200 lightweight pairs;include_quotes=trueadds the triggering evidence quotes).
Evidence/semantic queues are process-local, so a crash or forced shutdown can lose queued work. Do not infer durable coverage from queue depth. After a restart or an evidence-side busy/discard signal, inspect local_text_evidence coverage and run rebuild_evidence until its dry-run is empty and the vector state is ready; semantic-worker queue drops are recovered by the scheduled scan_candidates pass, not by rebuild_evidence. Rebuilding evidence is idempotent derived-index repair; scanning is what recovers conflict candidates/notices that were never processed.
Upgrading from an older database
Upgrade warning for 0.14.8: current runtime startup accepts only schema generation workspace_state_v1. Both conflict_groups_v2 and local_text_evidence_v1, plus older claim/memory-vector/section-vector databases, are refused without modification. Run the public side-by-side mema upgrade. Every schema migration declares vector_effect=preserve|rebuild; the migrations from the two previous evidence generations preserve vector payloads regardless of current model availability. Compatibility is evaluated separately: a different configured embedding space records state=mismatch, disables vector reads, and is repaired later with memory_repair(rebuild_evidence). Both paths compact current workspace redirect/negative-decision state and discard the obsolete workspace decision event ledger.
The side-by-side copy retains memory content/history, backup replay receipts, workspace canonicals and current redirect/negative-decision state, and audit. The obsolete workspace decision event ledger is not copied. Preserve migrations clone FTS/evidence/vector payloads unchanged and transactionally rebuild only the conflict domain; vector health or space mismatch never changes the structural migration result. Rebuild migrations regenerate evidence and vectors. Both paths intentionally start with empty new conflicts/notice state and do not copy old conflicts, append-only conflict_judgments, or semantic_notices history. Current contradictions must be rediscovered by a scheduled full-library scan.
After rebuild, status/doctor reports conflict_scan_required=true with a persistent scan epoch. Only a successful full scan covering the upgrade-time active-memory set with the matching detector version may CAS-clear that flag; partial pages, failed scans, and older-detector scans do not. The target is published only after row/fingerprint checks, a successful PRAGMA wal_checkpoint(TRUNCATE), and removal of target WAL/SHM sidecars — the full-rebuild path additionally requires complete eligible evidence coverage; the source database is never deleted.
# Preview only.
mema upgrade --dry-run
# Stop every mema MCP client/worker. Make a WAL-safe rollback backup:
sqlite3 /absolute/path/to/memory.sqlite3 "PRAGMA wal_checkpoint(TRUNCATE);"
cp /absolute/path/to/memory.sqlite3 /absolute/path/to/memory.pre-0.14.sqlite3
# Migrate and switch the standard JSON config.
mema upgrade
# Restart the MCP client and verify.
mema doctor --json
After the upgrade, the first server start launches a daemon thread that backfills sentence row vectors automatically; watch progress in `mema doctor` (`rows.coverage`) or `memory(action="status")`. With no embedding model configured the backfill stays pending and resumes once the model is available — nothing to run by hand.The full evidence-rebuild path requires sqlite-vec, a configured/readable local GGUF embedding model, llama-cpp-python (install the semantic-local extra because it also runs GGUF embeddings), a writable target directory, and enough free disk. A preserve migration does not load either model and does not require vector completeness; it reports vector compatibility independently and marks incompatible preserved data mismatch. The optional conflict-judge checkpoint itself is never a migration prerequisite. The command reports its selected mode, vector effect/compatibility, memory count, estimated vector work, free disk space, source, and target before asking for confirmation.
The explicit checkpoint above matters because copying only the main .sqlite3 file while live WAL frames exist is not a complete backup; alternatively use SQLite's online .backup command before stopping. Abort if wal_checkpoint(TRUNCATE) reports a non-zero busy count. mema upgrade also checkpoints/verifies the new target before switching, but it does not create the operator's rollback copy of the source.
The old database is never deleted. Standard JSON configuration is backed up and switched only after full verification; environment-variable db_path overrides are reported as a manual action. Use --no-switch to build and verify without editing configuration. --yes skips both the interactive confirmation and its acknowledgement that all writers/workers are stopped and old conflict/judgment/notice history will be permanently omitted; it does not stop processes, checkpoint the source, or create a backup. The lower-level mema migrate-vnext command remains available for diagnostics. Keep the old database until the new one has run successfully in normal use; if the new database has accepted writes, do not switch back without first accounting for those newer records.
Configuration
Configuration is file-only since 0.15.0. Everything tunable lives in ~/.config/memory-arbiter/config.json (or the file the MEMORY_ARBITER_CONFIG launch-context variable points at; mema setup writes the starter template). Engine parameters, timeouts, thresholds, and caps are frozen constants (memory_arbiter/constants.py).
The complete user surface is 21 keys (0.15.14: added semantic_conflict.n_gpu_layers, removed semantic_conflict.max_notice_pairs and policy_path; 0.17.1: removed claims.required, semantic_conflict.model_path and n_gpu_layers, added the four semantic_conflict.mdeberta_* keys):
{
"db_path": "~/.local/share/memory-arbiter/memory.sqlite3",
"backup_jsonl": "~/.local/share/memory-arbiter/memory.backup.jsonl",
"client": "your-client",
"agent_id": "your-agent-id",
"workspace": "default",
"isolation": "none",
"update_check": { "enabled": true },
"include_size": true,
"embedding": {
"model_path": "~/.local/share/memory-arbiter/models/embedding.gguf",
"auto_query": true,
"auto_write": true
},
"semantic_conflict": {
"enabled": true,
"mdeberta_ckpt": "~/.local/share/memory-arbiter/models/mdeberta-v52_ep3_fp16.pt",
"on_write": "async",
"notice_sync_wait_ms": 3000
},
"mcp": {
"transport": "stdio",
"http": { "host": "127.0.0.1", "port": 8000 }
}
}Setting | Purpose |
| Current SQLite database |
| Append-only fallback when SQLite cannot write |
| Required caller identity; no built-in defaults — the server refuses to start when either is blank |
| Default workspace and |
| Optional one-shot background PyPI discovery (default |
| Global switch for the recall size block (v0.15.6): on = |
| Local GGUF embedding model — pointing at it is the sole intent to enable sqlite-vec evidence recall |
| Auto-embed at query/write time (default |
| Local mDeBERTa checkpoint (v52_ep3_fp16, ~531MB, ships separately) for write-time conflict judging. Configured → auto-enabled, loaded at startup, kept resident (CPU fp32). Install the |
| Explicit off-switch; unset + |
| Write-time detection: |
| How long the write response waits for the post-commit check so its result rides along (v0.15.8, default |
| config.json + tokenizer directory; default = |
| P(conflict) floor for a normal-severity notice (default |
| Judge batch size (default |
|
|
| Local HTTP endpoint; host is restricted to loopback, defaults to |
See examples/memory-arbiter.config.example.json.
Semantics worth knowing: the embedding dimension comes from the model itself — the database records the active dimension, and switching to a model with a different dimension automatically drops and rebuilds the vector tables at the new dimension at startup (a one-time full evidence rebuild). Ranking is fixed hybrid (lexical + evidence fusion); there is no ranking-mode knob. HTTP request handling is stateless with a 4 MB request-body cap.
Six environment variables remain as launch context: MEMORY_ARBITER_CONFIG, MEMORY_ARBITER_DB_PATH, MEMORY_ARBITER_BACKUP_JSONL, MEMORY_ARBITER_MCP_TRANSPORT, MEMORY_ARBITER_CLIENT, MEMORY_ARBITER_AGENT_ID. They select process context (which config file, which DB, which transport, which identity), and a config-file value wins over the matching variable. Every other MEMORY_ARBITER_* variable is no longer read — a stale export surfaces a "no longer read" warning in mema doctor, the console settings page, and memory(action="status"). Removed file keys similarly warn "no longer configurable" and are ignored; docs/INTEGRATION.md carries the 0.14 → 0.15 key-migration table.
HTTP mode: sharing one local server
stdio (the default) needs no background process: each MCP client launches mema as its own short-lived child process. Switch to streamable-http only when you want one long-lived local server that several clients connect to.
stdio (default) | streamable-http | |
Who starts mema | each client spawns a child process | you run one persistent process; clients connect to it |
Background process needed | no | yes — otherwise it dies when the terminal closes |
Client config | command + args | url + two fixed request headers |
Good for | one person, one client | several clients on one machine sharing one memory store |
Setting it up:
Config: set
mcp.transportto"streamable-http"in~/.config/memory-arbiter/config.json(orMEMORY_ARBITER_MCP_TRANSPORT=streamable-http).Keep it running: mema has no built-in daemon — use a process manager. On macOS, the launchd template at
examples/com.memory-arbiter.mema.plistruns it at load, restarts on crash, and logs to/tmp/mema.{out,err}.log(replace__MEMA_BIN__with the absolute pathwhich memaprints; put it in~/Library/LaunchAgents/thenlaunchctl load). For a quick try,tmux new -d -s mema 'mema'works.Client: copy
examples/streamable-http.mcp.json, filling inX-Mema-ClientandX-Mema-Agent-Id.
Notes: HTTP request handling is stateless (a frozen constant since 0.15.0) because mema keeps memory and semantic-notice state in SQLite, not in an MCP session. A service restart therefore does not leave clients holding an expired server session. Semantic notices created asynchronously are claimed from SQLite and attached to a later successful tool response as before; only a worker job that has not yet persisted its notice can be interrupted by a process restart.
The client sends the fixed headers automatically on every HTTP MCP request — agents must not add identity to individual tool data, or it is rejected. Missing/empty/duplicate/conflicting identity fails closed (400), never falling back to defaults. The service binds to loopback only; these headers are provenance, not authentication. Because launchd does not inherit your shell PATH or expand ~, put absolute paths in ProgramArguments and for any GGUF model_path in config.json.
Claude Desktop / Claude Code through localhost HTTP
Claude's local MCP configuration launches stdio commands. To reuse one running mema HTTP service instead of spawning another mema process, put this single entry under mcpServers in ~/.claude.json (current Claude Desktop/Cowork and Claude Code installations may share this user-level file):
{
"mcpServers": {
"memory-arbiter": {
"command": "/opt/homebrew/bin/npx",
"args": [
"-y",
"mcp-remote@0.1.43",
"http://127.0.0.1:8000/mcp",
"--allow-http",
"--transport", "http-only",
"--header", "X-Mema-Client:claude",
"--header", "X-Mema-Agent-Id:claude",
"--silent"
],
"env": {
"PATH": "/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin",
"NO_PROXY": "127.0.0.1,localhost"
}
}
}
}Use the absolute npx path from which npx on your machine. Remove any older memory-arbiter entry that directly launches mema/memory-arbiter-mcp, otherwise Claude may start a second server process. Fully quit and reopen Claude Desktop, and restart Claude Code sessions after changing the file. mcp-remote is a third-party bridge; the pinned version above is the configuration tested with mema. If your Claude installation uses a separate Desktop MCP file, place the same single entry there instead, but do not register both copies.
Conflict judge model (mDeBERTa, optional)
The write-time conflict judge needs a one-time manual install (the 1.1GB checkpoint does not ship on PyPI):
pip install memory-arbiter-mcp[mdeberta](torch CPU wheel ~200MB + transformers);download
mdeberta-v52_ep3_fp16.pt(~531MB, fp16 weights) plus themdeberta-basedirectory (config.json + tokenizer, MB-sized) from the mini-clash release artifacts;point
semantic_conflict.mdeberta_ckptat the.ptfile in config.json (auto-enables + preloads at startup;mdeberta_model_dirdefaults to amdeberta-basefolder next to the checkpoint).
mema doctor verifies the dependency, the checkpoint digest, and the judge's label contract. Unconfigured = write-time arbitration disabled (scan/Agent fallback unaffected). ONNX+INT8 (smaller, no torch) is planned for a later release.
Degradation
Without sqlite-vec or an embedding model, lexical recall and memory governance continue; evidence indexing is unavailable.
Without the mDeBERTa judge (extra/checkpoint/config), write-time arbitration is disabled (notices pause); scheduled scan continues returning its deterministic KNN/rule baseline candidates.
If SQLite is unavailable or unwritable, writes use the append-only JSONL envelope only when that write succeeds. JSONL contains memory records and their selected canonical, not internal redirects or negative decisions; preserve or upgrade the SQLite database to retain workspace decision state.
Development
uv run pytest -q
python scripts/sync_version.py --checkDuring development, package/docs may describe an unreleased dev version while server.json (the MCP Registry manifest) stays at the last released version; scripts/sync_version.py advances it as part of release preparation, and --check keeps it from drifting.
中文摘要
Memory Arbiter(迷码)是面向 AI Agent 的本地可信事实层:每条事实只存一份完整原文,向量与检索均为可重建的派生索引。冲突 scan 走宽门召回,write-time notice 走 mDeBERTa 三分类判定(0.17.1 起,漏斗门全保留);单一 conflicts 表保存一对多事件,生命周期为 open → applying → resolved 或 not_a_conflict。裁决后按 judge → apply_conflict_action → resolve_conflict 顺序治理。none/weak/strict 都做 workspace 归一;strict 使用 guarded vector admission,default 池不进入项目 scope。workspace_state_v1 升级会清除旧 conflict/judgment/notice 历史和旧 workspace decision event ledger,并要求完成带 epoch 的全库 scan。完整中文文档见 README.zh-CN.md;另见 INTRO.md 与 docs/INTEGRATION.zh-CN.md。
Available Tools
4 toolsmemoryA
Daily memory operations: remember, find, batch_find, read, update, judge, status, help.
workspace is required on remember — discover buckets once with memory_review(view="workspaces") (e.g., at session start) and reuse that list; re-query only when unsure or before deliberately creating a new bucket ('default' is the global pool).
Call memory(action="help") to discover accepted fields, judge requirements, value enums, update modes, and action_required paths before relying on a result that requests attention.
update edits in place: 1..8 local edits use patches=[{old_text,new_text}] (atomic all-or-nothing, one version bump); full rewrites use new_content; a single small edit may use old_text/new_text.
batch_find runs up to 8 queries in one call and returns one merged index page (dedup by memory_id, matched_query_ids per item).
find is an index page: results carry metadata + content_chars + a bounded outline (offsets usable directly as read span starts), not full content — content_mode (v0.15.10, preview default) is a single-choice enum: "hits" adds hit_spans (vector-matched unit text + span coordinates, never truncated; >=50% coverage upgrades an item to full text; hit_window=N (default 0) extends each hit with +/-N neighbouring complete sentences, neighbours marked matched=false; hit_spans appears only on query-recall pages — browse/filter pages carry none; hits whose evidence row lags the memory's current version are dropped with a stale_hit_spans marker plus a re-query warning), "full" returns whole texts. Space-separated short CJK words (each <=4 chars, e.g. "向量 唯一键 冲突") are understood as keywords: semantically close memories whose content/subject contains a keyword rank higher (0.17.0; whole-word matching, generic words ignored). Offsets are 0-based Unicode code-point offsets into the content as indexed for the item's version; the evidence index may lag right after an edit (see vector_lag). Score compares only within the page; if the top page misses, reword the query or add tags_filter instead of deep paging. The size block meters the returned page (tokens_estimate + display_hint).
| Name | Required | Description | Default |
|---|---|---|---|
| data | No | ||
| action | No | help |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: update is atomic all-or-nothing with one version bump, batch_find caps at 8 queries and dedups by memory_id, find returns an index page rather than full content, stale hit spans are dropped with a marker, and the evidence index may lag after an edit. It omits permissions/auth requirements and error semantics, and never states whether any action deletes data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded with the action list, but the find paragraph is one enormous nested parenthetical run-on that is hard to scan, and version tokens ('v0.15.10', '0.17.0') add noise without helping invocation. Most sentences do carry real information, so it is overloaded rather than padded, but the structure is weak for its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter dispatcher with no annotations and no output schema, the description is thorough about find, update, and batch_find but silent on the behavior of remember, read, judge, and status, which is a real gap given the tool's breadth. Pointing to memory(action="help") for field/judge/enum discovery partially covers the omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% with a free-form 'data' object and an undocumented 'action' string, so the description must compensate, and it largely does: it names the action values, the patches/new_content/old_text shapes, content_mode enum and default, hit_window default, tags_filter, and view="workspaces". It also delegates residual discovery to memory(action="help"), which is a legitimate mitigation for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line gives the resource (memory) and enumerates eight actions (remember, find, batch_find, read, update, judge, status, help), and 'Daily memory operations' loosely distinguishes it from the maintenance-flavored siblings memory_review/govern/repair. However, four of the actions (remember, read, judge, status) are never explained, so an agent cannot tell what half the tool actually does. It is a partial statement of purpose rather than a precise one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It routes the agent to a sibling for a specific need ('discover buckets once with memory_review(view="workspaces") ... re-query only when unsure'), gives alternative parameter strategies ('patches' for 1..8 local edits vs 'new_content' for rewrites), and warns against deep paging ('reword the query or add tags_filter instead'). No explicit when-not-to-use-this-tool guidance exists, but the operational context is unusually concrete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_governA
Authorized governance: retire, merge near-duplicates, apply/replan/resolve conflicts, confirm, and manage workspaces.
Every state-changing action requires explicit user authorization for that action, then authorized=true. Workspace-mutating actions (confirm_pending_workspace, rename_workspace_canonical, migrate_workspace, separate_workspace_alias, move_memories_workspace) also require workspace — call memory_review(view="workspaces") first to list existing buckets. Call memory_govern(action="help") for exact actions, accepted fields, impact notes, and confirmation semantics.
| Name | Required | Description | Default |
|---|---|---|---|
| data | No | ||
| action | No | help |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden; it does disclose the authorization gate and the confirmation-semantics pointer, which is real value. However it says nothing about what retire/merge actually destroy, reversibility, or side effects, deferring impact notes to action="help" rather than stating them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Short and front-loaded with the governing action list, then prerequisites, then the help pointer. The hard-wrapped fragments break the sentences awkwardly but there is no filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity, multi-action tool with no annotations, no output schema and a freeform data object, the description leaves the agent unable to construct a real invocation without an extra action="help" round-trip. The deferral is explicitly documented, which is honest, but the definition is not self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and `data` is an untyped freeform bag, so the description must compensate. It partially does by naming concrete action values (retire, confirm_pending_workspace, rename_workspace_canonical, migrate_workspace, separate_workspace_alias, move_memories_workspace) and by stating which actions need workspace, but accepted data fields are left entirely to the help call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete family of governance verbs (retire, merge near-duplicates, apply/replan/resolve conflicts, confirm, manage workspaces) rather than a tautology, so the agent knows this is the state-changing memory-curation surface. It stops short of a 5 because it never distinguishes itself from memory_repair or memory, which cover adjacent curation work.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives actionable usage rules: every state-changing action needs explicit user authorization plus authorized=true, workspace-mutating actions need workspace, and the agent should list buckets via memory_review(view="workspaces") first. It lacks any when-not-to-use or sibling-choice guidance, keeping it below 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_repairB
Maintenance: evidence rebuild, conflict scans (scheduled-task spec under help topic scheduled_tasks), full-library duplicate sweeps (scan_duplicates), history cleanup, entity assignment, pending activation, backup replay, notices, and semantic runtime control.
Use memory_repair(task="help") for notice handling and semantic_control actions. Semantic notices are advisory; read both memories before dismiss or resolve, and never pass a notice directly to judge or resolve_conflict.
| Name | Required | Description | Default |
|---|---|---|---|
| data | No | ||
| task | No | help |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that semantic notices are advisory and warns to read both memories before dismiss/resolve, which is valuable. Yet most maintenance operations like history cleanup and backup replay are not described in terms of destructiveness, reversibility, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and information-dense, with the maintenance category and action list front-loaded. The second sentence adds crucial handling guidance without unnecessary filler. It is slightly list-heavy but earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a broad maintenance tool with no annotations and no output schema, so the description must supply enough context for correct invocation. It does not explain how to invoke most tasks, what data they require, or what they return. The reference to a help topic is useful but not sufficient for the full scope.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain both parameters. It mentions task='help' and implies various task actions, but the data parameter is entirely unexplained. An agent cannot determine what payload is needed for most listed maintenance tasks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies memory_repair as a maintenance tool and enumerates concrete operations (evidence rebuild, conflict scans, duplicate sweeps, history cleanup, etc.). It distinguishes itself from review/govern siblings by the 'Maintenance' category, though the long list makes the primary purpose somewhat diffuse.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says to use memory_repair(task='help') for notice handling and semantic_control actions, and gives a clear exclusion: never pass a notice directly to judge or resolve_conflict. However, it does not broadly state when to prefer memory_repair over memory_review, memory_govern, or memory.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_reviewB
Read-only inspection: overview, doctor, conflicts, conflict_detail, history, expired, audit, entities, workspaces, help.
Use memory_review(view="help") for accepted fields. Inspect conflict_detail before judging a conflict so its members, value groups, revision, and apply state are visible.
| Name | Required | Description | Default |
|---|---|---|---|
| data | No | ||
| view | No | help |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It does clearly label the operation as 'read-only inspection,' which is a meaningful safety claim. Beyond that, it does not describe output shape, pagination, permissions, or side effects, leaving gaps for a tool with ten modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loads the read-only nature and view list. The second sentence about conflict_detail is slightly orphaned without context, but overall there is little wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no annotations and no output schema, and the description covers only the read-only nature, view names, and one conflict_detail tip. It lacks guidance on what 'data' should contain per view and how this tool differs from its three siblings, so the definition is adequate but incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only names view values rather than describing the 'data' parameter at all. For a two-parameter tool with a free-form 'data' object, the description should explain what data expects for each view; it does not, so it fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Read-only inspection' and enumerates ten view modes, which conveys the resource and scope. However, the verb is not made specific, and the sibling tools (memory, memory_govern, memory_repair) are not referenced, so the agent cannot easily tell why memory_review is distinct from those tools. Purpose is implied but not sharply defined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers one procedural guideline: inspect conflict_detail before judging a conflict, and use view='help' for accepted fields. This is helpful context for one view but does not explain when to choose memory_review over memory, memory_govern, or memory_repair, nor when each of the other nine views is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.16.5- Changed
memory1 field changed- changed
Output schema / (root)Previous value: -{ - "additionalProperties": true, - "title": "memoryDictOutput", - "type": "object" -}New value: +null
- Changed
memory_govern1 field changed- changed
Output schema / (root)Previous value: -{ - "additionalProperties": true, - "title": "memory_governDictOutput", - "type": "object" -}New value: +null
- Changed
memory_repair1 field changed- changed
Output schema / (root)Previous value: -{ - "additionalProperties": true, - "title": "memory_repairDictOutput", - "type": "object" -}New value: +null
- Changed
memory_review1 field changed- changed
Output schema / (root)Previous value: -{ - "additionalProperties": true, - "title": "memory_reviewDictOutput", - "type": "object" -}New value: +null
35 tool updates
v0.11.0- Added
memory - Removed
memory_activate - Removed
memory_arbitrate - Removed
memory_audit_summary - Removed
memory_cleanup_history - Removed
memory_cleanup_inactive_vectors - Removed
memory_compare - Removed
memory_confirm - Removed
memory_correct_conflict_judgment - Removed
memory_doctor_overview - Removed
memory_edit - Removed
memory_get - Added
memory_govern - Removed
memory_history - Removed
memory_list_conflict_judgments - Removed
memory_list_conflicts - Removed
memory_list_entities - Removed
memory_rebuild_claims - Removed
memory_rebuild_embeddings - Removed
memory_recent - Removed
memory_record_conflict - Added
memory_repair - Removed
memory_resolve_conflict - Removed
memory_resync_vec_parent_status - Added
memory_review - Removed
memory_scan_conflict_candidates - Removed
memory_search - Removed
memory_search_expired - Removed
memory_set_entity - Removed
memory_split - Removed
memory_status - Removed
memory_store_embedding - Removed
memory_submit_conflict_judgment - Removed
memory_supersede - Removed
memory_write
1 tool update
v0.9.9- Added
memory_activate
13 tool updates
v0.9.4- Changed
memory_arbitrate4 fields changed- added
Input schema / properties / apply / anyOfAdded value: +[ + { + "type": "boolean" + }, + { + "type": "null" + } +] - changed
Input schema / properties / apply / defaultPrevious value: -falseNew value: +null - removed
Input schema / properties / apply / typeRemoved value: -"boolean" - added
Input schema / properties / authorizedAdded value: +{ + "default": false, + "title": "Authorized", + "type": "boolean" +}
- Added
memory_cleanup_inactive_vectors - Changed
memory_confirm1 field changed- added
Input schema / properties / authorizedAdded value: +{ + "default": false, + "title": "Authorized", + "type": "boolean" +}
- Added
memory_correct_conflict_judgment - Added
memory_list_conflict_judgments - Added
memory_list_entities - Added
memory_rebuild_claims - Changed
memory_resolve_conflict1 field changed- added
Input schema / properties / statusAdded value: +{ + "default": "resolved", + "title": "Status", + "type": "string" +}
- Added
memory_resync_vec_parent_status - Changed
memory_search3 fields changed- added
Input schema / properties / include_superseded / anyOfAdded value: +[ + { + "type": "boolean" + }, + { + "type": "null" + } +] - changed
Input schema / properties / include_superseded / defaultPrevious value: -falseNew value: +null - removed
Input schema / properties / include_superseded / typeRemoved value: -"boolean"
- Added
memory_search_expired - Added
memory_set_entity - Added
memory_submit_conflict_judgment
TDQS
Scored across 4 tools
The four tools are mostly separated by lifecycle stage: daily operations, read-only review, authorized governance, and maintenance. There is some conceptual overlap around conflict inspection, judgment, and resolution across memory, memory_review, and memory_govern, but the descriptions define handoffs well enough for an agent to choose correctly.
All tool names use lowercase snake_case and share a clear memory_ prefix except the base memory tool. The suffixes review, govern, and repair make each tool's role predictable.
Four top-level tools is appropriately compact for a complex memory-management server. Bundling many sub-actions into each tool avoids tool sprawl while keeping high-level categories distinct.
The surface covers memory creation, retrieval, update, review, governance, conflict handling, maintenance, and repair. No obvious lifecycle gaps or dead ends are apparent for the stated memory-arbiter domain.
Maintenance
Related MCP Connectors
One memory, every AI. A shared, user-owned markdown memory your AI clients read and write over MCP.
Shared memory and mail for your AI agents. Verified with Claude Code; other MCP clients in testing.
Persistent AI memory shared across Claude, ChatGPT, coding agents, and compatible MCP clients.
An MCP memory server. One memory your agents share — across models, devices and apps.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA self-hosted MCP server that provides AI assistants with a shared, persistent SQLite-backed memory for storing and retrieving project context, decisions, and discoveries. It enables cross-session continuity and team-wide knowledge sharing to keep AI coding tools aligned and informed.3MIT
- FlicenseNot gradedqualityDmaintenanceA persistent, conflict-aware memory MCP server for AI coding assistants (Cursor, Claude Code).-

threadctx-mcpofficial
AlicenseAqualityCmaintenanceShared memory MCP server for AI coding agents, enabling context sharing across sessions with local SQLite or cloud-based semantic search, compatible with Claude Code and Cursor.264 npm1MIT- AlicenseAqualityCmaintenanceA local-first MCP server that provides a shared Markdown-based memory for AI coding agents, enabling cross-agent context persistence via tools like memory_search and memory_capture.101MIT