Unified Memory MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Unified Memory MCPsave a memory: API key rotation due Friday"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Memory
Local, single-tenant memory and handoffs for MCP clients.
The core uses Python's standard library and SQLite FTS5: no account, hosted service or runtime dependency is needed. Optional embeddings are a separate extra and may download model weights; the default installation uses text search.
Install
Python 3.11 or newer; candidate qualification uses Python 3.12.
python -m venv .venv
# Activate .venv using your platform's standard command.
python -m pip install .
python -m unittest discover -s tests -v
unified-memory-admin --helpThe distribution name is unified-memory-mcp. Version 2.5.1 in this snapshot
is the public-safe candidate, not a claim that a package was uploaded.
Related MCP server: mcp-memory-rs
Connect a stdio client
Use your MCP client's configuration syntax, with explicit executable and database paths. For a client accepting a mcpServers object, the shape is:
{
"mcpServers": {
"memory": {
"command": "/absolute/path/to/.venv/bin/unified-memory-mcp",
"env": {
"UNIFIED_MEMORY_DB": "/absolute/private/path/memory.db",
"MEMORY_AUTOSYNC": "0"
}
}
}
}On Windows the installed executable is under .venv/Scripts. The same user can point multiple local MCP processes at one SQLite store. Do not share a database between mutually untrusted users or tenants.
Tools include memory_save/search/recall, memory_bootstrap, handoff_save/load, memory_policy and explicit-confirmation pruning. Scopes organize records; they are not user identities or RBAC boundaries.
Public-safe defaults
No local assistant memory is imported automatically.
Explicit import needs MEMORY_AUTOSYNC=1 and CLAUDE_MEM_DIR or CLAUDE_MEM_GLOB pointing to a reviewed directory. Importing private text is your decision.
Personal private-store integration and the workstation-specific sync utility are not included.
The optional HTTP gateway requires --workspace-root explicitly. It offers authenticated file operations as well as memory; use stdio when those capabilities are unnecessary.
HTTP is for a trusted loopback/tunnel deployment, not public Internet exposure. Configure MEMORY_GATEWAY_TOKEN or MEMORY_GATEWAY_TOKEN_FILE; do not commit it.
No real database, token, deployment settings, autostart task or SSH material is shipped or activated.
The history-free snapshot contains no workstation paths, personal assistant memory, private records, database files, tokens, raw logs or private settings.
The source manifest records original/exported hashes and AST transformations. The private live server is not modified by building this snapshot.
Backup and restart check
Use only a fresh synthetic directory for a first test:
unified-memory-admin restart-drill --work-root ./restart-demo
unified-memory-admin write-deployment-profile --output ./deployment-profile.jsonBackups and restore-copy require explicit paths; restore-copy never overwrites an existing destination. Protect database and backup permissions yourself. The strict deployment preflight includes Windows-specific storage/ACL checks; do not infer Linux ACL qualification from portable stdio tests.
See SECURITY.md. This candidate is released under MIT; third-party dependencies and optional model weights retain their own terms.
Available Tools
16 toolshandoff_completeC
Mark a handoff complete once the transferred task is finished.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| notes | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It implies a state-changing mutation but says nothing about what completion does (archive, delete, status flip), whether it is idempotent, what permissions are needed, or what the idempotency/error behavior is on a repeated call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the action front-loaded and no filler. It is well-structured but under-specified rather than genuinely concise-with-substance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations, no output schema, and undocumented parameters, the description should explain the resulting state and how notes are used. It leaves the agent without enough to invoke it confidently beyond the happy path.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions either field. It gives no hint that "id" is the handoff identifier or what "notes" records, leaving the agent to infer both from the property names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Mark a handoff complete") plus the triggering condition ("once the transferred task is finished"). An agent can tell it apart from handoff_save/handoff_load, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause "once the transferred task is finished" implies the right moment to call it, but there is no explicit when-not guidance and no mention of alternatives such as handoff_save or handoff_list. Usage is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handoff_listC
Audit recent handoff checkpoints for a project.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| status | No | ||
| project | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It says nothing about ordering, pagination (despite a limit param), what a checkpoint contains, or whether auditing mutates anything. 'Audit' implies read-only but this is not confirmed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with no waste, but it is so terse that conciseness comes at the cost of the information the tool needs. Structure is fine; content is thin.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Three-parameter tool with no annotations, no output schema, and no parameter documentation in either schema or description. The one-sentence description is inadequate for an agent to invoke this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description mentions no parameters. The agent gets no meaning for limit (default? max?), status (what values are valid?), or project format. With three undocumented params the description fails to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (audit) and resource (handoff checkpoints) with a clear modifier (recent, for a project). It is distinguishable from siblings like handoff_save and handoff_load, though the exact relationship to those isn't spelled out.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance. It doesn't say when to prefer this over handoff_load, nor what 'recent' means, nor any prerequisites. The agent must infer everything from the name and the one-line description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handoff_loadC
Load recent unfinished handoffs for a project, normally from another client.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| project | Yes | ||
| consumer | Yes | ||
| include_own | No | ||
| mark_resumed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden, and it omits the most consequential detail: the mark_resumed parameter implies loading can mutate state (claiming/resuming a handoff), yet the description describes only a read-like 'load'. It also doesn't say whether loading consumes the handoff, what happens on concurrent loads, or what is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no waste, but its brevity comes from omission rather than discipline — for a five-parameter, unannotated tool it is under-specified, not efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With five parameters, no annotations, no output schema, and a mutation-capable flag (mark_resumed) left undisclosed, the description is not complete enough for an agent to invoke this confidently. The cross-client framing is the only substantial context provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across five parameters, and the description only loosely gestures at 'project' and 'another client' (consumer). It says nothing about limit, include_own, or mark_resumed, leaving the majority of parameters semantically opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Load'), resource ('handoffs'), and scope ('recent unfinished ... for a project'), which is enough to distinguish it from the write-oriented handoff_save and handoff_complete. It never explicitly contrasts itself with the closely related handoff_list, so the boundary between 'load' and 'list' is left to inference rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Normally from another client' hints at the cross-client scenario but stops short of guidance: no statement of when to use handoff_load rather than handoff_list, no prerequisites, and no conditions under which it should not be used. An agent has to guess the intended workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
handoff_saveC
Create or update a structured work checkpoint so another Claude/Codex session can continue from the same state. Reuses project+source+session_id.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | ||
| task | No | ||
| files | No | ||
| notes | No | ||
| tests | No | ||
| source | Yes | ||
| status | No | ||
| target | No | ||
| project | Yes | ||
| summary | Yes | ||
| blockers | No | ||
| metadata | No | ||
| decisions | No | ||
| next_steps | No | ||
| session_id | No | ||
| completed_work | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses that the operation is an upsert keyed by project+source+session_id, but omits permissions, overwrite/merge semantics, side effects, and return behavior for a mutation tool with 16 parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words, and the operation and its intended outcome are front-loaded. It is appropriately sized for a short definition, even though it omits important details elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 16 parameters, 0% schema description coverage, no annotations, and no output schema, two sentences cannot make the tool safely invocable. Critical parameter meanings, update semantics, and return behavior are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, but it only names project, source, and session_id and gives minimal meaning ('Reuses' them as a key). The other 13 parameters, including required 'summary', are left entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Create or update'), resource ('structured work checkpoint'), and intended outcome ('so another Claude/Codex session can continue from the same state'). It does not differentiate from siblings such as handoff_complete or handoff_load, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a use case (enabling continuation by another session) but gives no explicit when-to-use guidance, prerequisites, or alternatives to sibling tools like handoff_load, handoff_list, or handoff_complete. The reader must infer usage entirely from purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_bootstrapA
Start or resume work across Codex and Claude. Returns the latest active handoff from another client plus relevant durable memory chunks. Call once at the beginning of substantial work.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Current task/topic for memory recall. | |
| scope | No | ||
| project | Yes | Stable repo/project name or path. | |
| consumer | Yes | codex, claude, or claude-remote. | |
| mark_resumed | No | ||
| memory_limit | No | ||
| handoff_limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses what is returned (latest active handoff + memory chunks), which is useful, but says nothing about the mark_resumed side effect, auth/permissions, or whether the operation mutates state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences, front-loaded with the purpose and return payload, with the call-timing instruction last. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with no annotations and no output schema, the description covers purpose adequately but leaves parameter semantics and side-effect behavior (mark_resumed) unexplained. It is minimally sufficient but has clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 43% across 7 parameters, so the description should compensate but does not. It omits any guidance on query, scope, mark_resumed, memory_limit, or handoff_limit, leaving critical recall/limit semantics undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: start/resume work across clients, returning the latest active handoff plus durable memory chunks. This distinguishes it from the individual handoff_load and memory_recall siblings, though it never names them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear timing guidance: 'Call once at the beginning of substantial work.' This tells the agent when to reach for it, but there is no explicit when-not guidance or comparison against the sibling handoff_load/memory_recall tools it supersedes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_deleteC
Delete a memory by id.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and fails to state whether deletion is permanent, reversible, or requires elevated permissions. For a mutation tool, the absence of any consequence disclosure is a meaningful gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler or redundancy, and the action and selector are front-loaded. It is efficient, though arguably under-specified rather than tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a destructive mutation with no annotations, no output schema, and an entirely undocumented parameter, yet the description adds almost nothing beyond the bare action. It leaves the agent without the consequence and identity semantics needed to call it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the sole required parameter has no description in the schema. The phrase 'by id' only restates the parameter's purpose without adding type, format, or valid-range detail, so the description does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (delete) and resource (a memory) scoped by id, so the core action is unambiguous. It does not, however, differentiate itself from siblings like memory_prune, which likely also removes memories and could be confused with this one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus memory_prune or the other deletion-adjacent siblings. The agent must infer from names alone that this is the single-item delete path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_getC
Fetch a single memory by id.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not state that this is a safe read, whether it mutates last-access metadata, or what happens when the id does not exist — a meaningful gap for a lookup tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence, front-loaded with the verb and resource. No waste, though it is arguably too terse to carry the missing context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read with no output schema, brevity is defensible, but the description omits the not-found/error behavior and how an id is obtained, which an agent needs to chain this after a list or search call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single required id parameter, so the description is the only source of semantics. 'By id' confirms id is the lookup key, which adds some meaning, but nothing about format, origin, or whether it is the numeric index from memory_list/memory_search results.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (fetch) and resource (a single memory) with the keying mechanism (by id). This clearly separates it from list/search siblings, though the description never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this over memory_list, memory_search, or memory_recall, all of which could plausibly retrieve memories. The agent must infer the distinction from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_listC
List most recently updated memories, optionally filtered by scope or tag.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | ||
| limit | No | ||
| scope | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It states ordering by most recent update and optional filtering, but it does not disclose pagination behavior, the effect or default of the limit parameter, authentication requirements, or what the returned records contain. For a read-oriented listing tool with no annotations, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It efficiently communicates the core operation and optional filters, which is exactly what a concise tool description should do.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no annotations, no output schema, and 0% schema description coverage, so the description is the only source of behavioral and parameter detail. It omits the limit parameter, pagination behavior, default ordering guarantees beyond 'most recently updated', and return-value expectations, leaving the agent with an incomplete operational picture.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all three undocumented parameters. It mentions 'scope' and 'tag' as filters, giving partial semantic meaning, but it omits 'limit' entirely and provides no value syntax, defaults, or constraints for any parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: 'List most recently updated memories'. It also names the optional filters, so the agent understands the tool's basic function. It does not explicitly distinguish itself from sibling tools like memory_search or memory_recall, so sibling differentiation is missing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says the tool supports optional scope and tag filters, which implies a browsing/list use case, but it provides no explicit guidance on when to use this tool versus memory_search, memory_recall, or memory_get. There are no stated prerequisites, exclusions, or alternative-routing instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_policyB
Report content-free scope authorization, quota, retention, and usage state.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden, and it does make one genuinely useful disclosure: "content-free" tells the agent the response excludes stored memory content, implying a safe, non-destructive read. It does not, however, state read-only status explicitly, auth requirements, or whether the output is cached/per-tenant, so the safety profile is still partly inferred.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the verb front-loaded and no filler. It is efficient, though the dense noun cluster ("content-free scope authorization, quota, retention, and usage state") trades a little clarity for brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with no annotations and no output schema, the description usefully enumerates the categories of state returned, effectively serving as a lightweight return-value summary. It still omits whether the call is side-effect-free and whether it requires an authorized scope, which matters given the absence of annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to clarify in the input; the baseline of 4 applies. No parameter-related claims conflict with the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a verb ("Report") and enumerates the reported state categories (scope authorization, quota, retention, usage), which separates it from read/write siblings like memory_get or memory_list. However, "content-free scope authorization" is opaque jargon that an agent may not parse, and it never says plainly that this returns policy/config metadata. Purpose is inferable but not crisp.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No indication of when to call this versus alternatives such as memory_get, memory_list, or memory_search, and no prerequisites or trigger conditions are given. The nature of the tool is only implied by the noun "policy". A reader must guess its role from the description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_pruneC
Delete expired memories and completed handoffs. Requires confirm=true.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses that this is destructive and that confirm=true is required as a guard, but says nothing about irreversibility, permission requirements, scope (all workspaces vs. current), or what 'expired' means operationally.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the destructive scope before the prerequisite. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an irreversible bulk-deletion tool with no annotations and no output schema, the definition leaves too much unstated: recovery options, affected scope, and the meaning of 'expired' are all missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 0% schema coverage, so the description must compensate. 'Requires confirm=true' does explain that the boolean must be set to true for the call to proceed, which is the key semantic — but it omits the default/behavior when false.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Delete') and a precise scope: expired memories and completed handoffs. This distinguishes it from the sibling memory_delete, which presumably targets individual memories, though the description never names that distinction explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only guidance is the confirm=true prerequisite. There is no statement of when to prune versus calling memory_delete or handoff_complete, nor any warning about the conditions under which pruning is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_recallA
Semantic recall from memory: returns the most RELEVANT chunks by meaning (cosine over embeddings), falling back to keyword search if embeddings are off. Prefer this over memory_search when you want the best context for a topic.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | How many chunks (default 5). | |
| query | Yes | What you want to recall. | |
| scope | No | Optional scope filter. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. It discloses the ranking method (cosine over embeddings) and the server-side fallback to keyword search when embeddings are off, which is important non-obvious behavior. It does not discuss permissions, side effects, or return format, but for a recall tool the disclosed ranking and fallback behavior are the central traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the tool's primary behavior and then the sibling routing condition. Every clause earns its place; there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has three fully documented parameters, no output schema, and no annotations. The description states that it returns chunks, how relevance is computed, and what happens when embeddings are off, which covers the main calling context. It could say more about the returned chunk representation, but is adequate for a recall operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents query, k, and scope. The description adds no parameter-level detail beyond the schema, which makes the baseline of 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: semantic recall from memory that returns relevant chunks by meaning. It explicitly contrasts itself with the sibling memory_search, so an agent can select between them without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit use condition: 'Prefer this over memory_search when you want the best context for a topic.' This names the alternative and the circumstance that selects this tool, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_reindexA
Rebuild chunks + embeddings for all memories. Run once after enabling embeddings (fastembed).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full behavioral burden. It does disclose the operation's scope (ALL memories) and idempotency intent ("run once"), which is useful, but it says nothing about cost/duration, whether existing custom chunks are overwritten, or what happens if invoked repeatedly — meaningful gaps for a bulk rebuild.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, purpose front-loaded and the operational trigger immediately after. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless maintenance tool with no output schema and no annotations, purpose plus trigger is close to sufficient, but an agent still lacks the failure/repeat-safety picture for a global rebuild. Adequate minimum-viable coverage rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters and 100% schema coverage, so there is nothing for the description to document. Baseline of 4 applies; the description adds no claim that would justify going lower.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: rebuild chunks and embeddings for memories. An agent can distinguish this maintenance operation from the CRUD/search siblings in the list (memory_save, memory_search, memory_sync) without opening a schema. It stops short of naming an actual alternative, so it is clear but not sibling-differentiating in an explicit way.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Run once after enabling embeddings (fastembed)" gives a concrete trigger condition and even implies the when-not (don't run repeatedly). It does not compare against memory_sync or memory_prune, which is the only thing keeping it from 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_saveA
Save a memory (fact, preference, decision, context) to the shared local store that both Codex and Claude read. Auto-chunked and embedded for semantic recall. Use for durable info worth recalling.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Optional tags. | |
| scope | No | Bucket, e.g. 'global' or a project. Default 'global'. | |
| source | No | Who is writing, e.g. 'claude' or 'codex'. | |
| content | Yes | The fact to remember. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose real behavior beyond the name – auto-chunking, embedding for semantic recall, and the shared cross-agent store – which is genuinely useful. But it says nothing about deduplication, overwrite semantics, permissions, or what identifier is returned for later update/delete by the sibling tools.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and destination, then the mechanism, then the usage cue. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-annotation, no-output-schema mutation tool it covers the essentials (what it stores, where, how it is indexed). It omits the practical follow-up information an agent needs in a 15-tool memory suite – the returned handle and whether repeated saves overwrite or duplicate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters, including defaults for scope and examples for source. The description adds no syntax, format, or constraint detail beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Save) and resource (a memory) and enumerates the content categories it accepts, plus the destination ('shared local store that both Codex and Claude read'). It does not differentiate itself from siblings like memory_update or handoff_save, which could be confused for similar write operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use for durable info worth recalling' gives a positive usage condition but is fuzzy and names no alternatives or exclusions. An agent still has to infer why it would pick memory_save over handoff_save or memory_update.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_searchC
Full-text (keyword) search of whole memories. Returns ranked matches.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results (default 10). | |
| query | Yes | ||
| scope | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full burden. It notes keyword (not semantic) matching and ranked results, which is some signal, but says nothing about pagination, ranking basis, scope behavior, or whether scope filters by workspace/user.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, no waste, and the matching mode is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A search tool with no annotations, no output schema, and two of three parameters undocumented. Missing scope semantics and the ranking/return shape leaves real ambiguity for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 33% — only 'limit' is described. 'query' and 'scope' are undocumented in both schema and description, so the meaning of 'scope' (workspace? user? memory type?) is entirely left to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Full-text (keyword) search of whole memories') and distinguishes the search mode (keyword rather than semantic, implied by contrast with memory_recall). Doesn't explicitly differentiate from memory_recall or memory_get by name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance. With siblings like memory_recall (likely semantic/associative) and memory_get (likely direct retrieval), an agent must guess which retriever fits a given need.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_syncA
Pull Claude's file memories (*.md) into the store. Idempotent and cheap: only files whose content changed are rewritten and re-embedded. Runs automatically on server start.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it discloses idempotency and cost ("only files whose content changed are rewritten and re-embedded") and the automatic-on-start trigger. It stops short of describing failure modes, permissions, or what the operation returns, but the core behavioral traits are exposed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler, front-loading the action (pull file memories into the store) before qualifying cost and trigger. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter, no-annotation tool this is nearly complete: what it does, its incremental behavior, and when it fires are all covered. The only missing element is any indication of outcome or error behavior, which matters slightly more given there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There are no parameter semantics to explain and the description does not fabricate any.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Pull"), resource ("Claude's file memories (*.md)"), and destination ("the store"). It distinguishes itself from write tools like memory_save, but does not explicitly contrast with the closest siblings memory_reindex or memory_bootstrap, leaving some ambiguity about how they differ.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The note that it "Runs automatically on server start" implies an agent rarely needs to call it manually, which is useful implied guidance. However, it never states when an agent should invoke it versus memory_reindex or memory_bootstrap, nor any exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_updateB
Update a memory's content, tags, or scope by id (re-chunks on content change).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| tags | No | ||
| scope | No | ||
| content | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one meaningful side effect: content changes trigger re-chunking. However, it omits whether unspecified fields are preserved, whether the update is destructive/irreversible, and permission requirements. The single side-effect note is real added value but incomplete for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the verb and operation, parenthetical reserved for the behaviorally important side effect. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-param mutation tool with no annotations, no output schema, and 0% schema description coverage, the description covers the target fields and one side effect but leaves partial-update semantics, reversibility, and error behavior unstated. Adequate minimum, but clear gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description names all four fields (content, tags, scope, id), so it does compensate somewhat by telling the agent which fields are updatable. It still adds no format or semantics for tags (array of strings) or scope (allowed values), and does not explain that id is the only required field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Update) and resource (a memory) plus the exact fields it can change (content, tags, scope) and the lookup key (id). It does not explicitly contrast with siblings like memory_save or memory_get, but the operation is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus memory_save (which may also modify) or memory_get/memory_delete. The 'by id' phrasing implies an existing record is required, but prerequisites and alternatives are left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
16 tool updates
v2.5.1- First observed
handoff_complete - First observed
handoff_list - First observed
handoff_load - First observed
handoff_save - First observed
memory_bootstrap - First observed
memory_delete - First observed
memory_get - First observed
memory_list - First observed
memory_policy - First observed
memory_prune - First observed
memory_recall - First observed
memory_reindex - First observed
memory_save - First observed
memory_search - First observed
memory_sync - First observed
memory_update
TDQS
Scored across 16 tools
Most tools have clearly distinct purposes, but memory_search and memory_recall overlap as search mechanisms, and memory_bootstrap overlaps with handoff_load/handoff_list by also returning handoff data. Descriptions help clarify preferred usage, so misselection is unlikely for careful agents.
All tool names use consistent snake_case with a predictable domain prefix: memory_ or handoff_. Actions are generally verbs and the pattern is stable across the set.
16 tools is slightly above the typical 3-15 range, but each tool serves a distinct role in memory management or handoff lifecycle. The set is rich but not bloated.
The surface covers memory CRUD, search, recall, admin tasks, and handoff lifecycle, so core workflows are complete. Minor gaps include no direct handoff_get by id, though handoff_load/list and memory_prune cover most needs.
Maintenance
Related MCP Connectors
Private-by-default, local-first memory/context/task orchestrator for MCP apps and agents.
- KogniteOAuthdev.kognite
Hosted agent memory: store, search, and recall facts across sessions from any MCP client.
Persistent memory, hybrid search and a goal graph for AI agents, over stdio or remote HTTP.
Person-owned AI memory that learns, not just stores — portable context for any MCP client.
Related MCP Servers
- AlicenseNot gradedqualityFmaintenanceLocal-first, auditable memory for AI agents. Provides durable context for MCP hosts with SQLite storage, CLI, and MCP tools for memory management.2Apache 2.0
- AlicenseNot gradedqualityCmaintenanceLocal-first MCP memory server providing persistent, versioned, queryable memory for AI agents using JSON categories and SQLite FTS5, with optional fleet sync.2Apache 2.0
- AlicenseAqualityBmaintenanceProvides persistent, searchable memory for AI agents across any MCP-compatible client, storing project context, user preferences, and session learnings locally in SQLite with tools to save, retrieve, search, and manage them.128 npmMIT
- AlicenseNot gradedqualityAmaintenanceProvides a drop-in MCP memory server over stdio using a local SQLite store, enabling agents to store, search, and audit memory with provenance, scoped access, and contradiction tracking without requiring API keys, LLMs, or embedding providers.551 npm16Apache 2.0