@kireo/mcp-server
This server provides long-term memory and context management for MCP-compatible AI tools, letting agents save, search, recall, update, and organize memories across namespaces, plus archive/load conversation context and index local code.
Memory management:
memory_savestores facts, decisions, preferences, events, goals, insights, relationships, or other items;memory_getfetches by id;memory_updatepatches fields;memory_deletesoft-deletes with a 30-day recovery window.Search & recall:
memory_searchperforms hybrid semantic + keyword search with filters (namespace, type, tags, entities, date range, score threshold);memory_recallreplays recent or important memories chronologically or by importance with pagination.Organization:
memory_list_namespacesenumerates the user's namespaces with counts; memories can be bucketed into logical namespaces (e.g., per project).Context handoff:
context_saveuploads a conversation summary/handoff,context_loadretrieves it, andcontext_archivemanages archived context.Project awareness:
project_infoexposes project details;kireo indexextracts functions/classes/methods from a local codebase into acode-<repo>namespace searchable viamemory_search.Health & diagnostics:
memory_healthprobes local server version and remote API status.Privacy controls: telemetry can be disabled via
KIREO_TELEMETRY=0; only explicitly supplied content is uploaded.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@@kireo/mcp-serversave a memory: deployment branch is main, not master"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@kireo/mcp-server
Long-term memory for any MCP-compatible AI tool (Claude Code, Cursor, Windsurf, Cline, Zed, Continue …).
What is Kireo memory MCP?
Kireo memory MCP is a Model Context Protocol server that gives Claude Code, Cursor, Cline, Windsurf and any other MCP client long-term memory. Save a decision once; recall it in any later session, on any machine. Hybrid semantic + keyword search over LanceDB, 12 MCP tools, plus local code indexing. Free beta — an API key is all you need.
How do I install Kireo memory MCP?
One line for Claude Code, one JSON block everywhere else. Both need a free key from
https://app.kireo.app/app/api-keys (ki_sk_…).
# Claude Code — add --scope user to get it in every project
claude mcp add kireo --scope user --env KIREO_API_KEY=ki_sk_xxx -- npx -y --package=@kireo/mcp-server kireo-mcpEvery other client takes the same server entry; only the file it goes in differs:
{
"mcpServers": {
"kireo": {
"command": "npx",
"args": ["-y", "--package=@kireo/mcp-server", "kireo-mcp"],
"env": { "KIREO_API_KEY": "ki_sk_xxx" }
}
}
}Client | Where that block goes |
Claude Code |
|
Cursor |
|
Cline | MCP Servers → Configure MCP Servers ( |
Claude Desktop |
|
Windsurf |
|
Zed / Continue / any MCP host | Whatever that host calls its MCP server list — same three fields |
Restart the client afterwards. Node.js ≥ 20 must be on PATH for npx.
Which tools does it expose?
Twelve, over MCP stdio: memory_save, memory_search, memory_recall, memory_get,
memory_update, memory_delete, memory_list_namespaces, memory_health, context_save,
context_load, project_info, and context_archive. Every client sees the same set. Call memory_health first to confirm the key works.
Does it bloat my prompt?
No — memory is pulled, not pushed. Nothing is injected into the system prompt. The agent calls
memory_search only when it decides prior context is worth retrieving, and gets back a bounded
ranked set (default 10 hits, hard cap 50), so tokens are spent per-query rather than per-turn.
How is it different from a CLAUDE.md / .cursorrules file?
A rules file is static text re-read in full every session and shared by nothing. Kireo memory MCP is queried on demand, is written by the agent as work happens, is searchable semantically, and is shared across projects, sessions and machines through namespaces.
Can it index my codebase?
Yes. npx -y -p @kireo/mcp-server kireo index ./ --repo my-app extracts functions/classes/methods
into a code-<repo> namespace that memory_search can reach. Indexing is incremental — re-runs
only send changed files, and the server dedupes identical symbols, so retrying is safe.
Is my code uploaded?
The compact command uploads the conversation summary you ask it to archive. context_save
uploads a handoff; memory_save uploads supplied content; kireo index uploads extracted code
symbols. These actions send their content to your Kireo account. Set KIREO_TELEMETRY=0 to also drop the X-Device-Id header.
What does it cost?
Free beta. Sign up at https://app.kireo.app, create a key, done — no card.
Related MCP server: Synapse Memory
Conversation compression (0.3.0)
The Kireo plugin adds /kireo:compact in Claude Code and $kireo-compact in Codex:
it summarizes the visible conversation, writes <directory-name>.md, and uploads
it to your Kireo account with offline retry and versioned file metadata.
Install the plugin.
The MCP now also provides project_info, context_save, context_load, and
context_archive, alongside the eight general memory tools below.
Quickstart
Get an API key at https://app.kireo.app/app/api-keys (
ki_sk_…).Add this MCP server to your host. Claude Code — run:
claude mcp add kireo --scope user --env KIREO_API_KEY=ki_sk_xxx -- npx -y --package=@kireo/mcp-server kireo-mcpTwo details in that line are load-bearing, both verified against claude 2.1.220 and npm 11 on 2026-08-03:
--package=@kireo/mcp-server kireo-mcp, not@kireo/mcp-server. This package ships two binaries (kireo,kireo-mcp), neither named after the package, sonpx -y @kireo/mcp-servercannot pick one and fails withcould not determine executable to run.--package=, not the short-p. A bare-pafter--gets swallowed by theclaude mcp addoption parser, which then rejects its own flag:claude mcp add kireo --env … -- npx -y -p @kireo/mcp-server kireo-mcperrors withunknown option '--env'. The long form parses cleanly.
Drop --scope user if you only want it in the current project. Alternatively, check a project-scoped .mcp.json into your repo root with the same shape (inside JSON args the short -p is fine — it goes straight to npx and never reaches the claude parser):
// .mcp.json (project root)
{
"mcpServers": {
"kireo": {
"command": "npx",
"args": ["-y", "--package=@kireo/mcp-server", "kireo-mcp"],
"env": { "KIREO_API_KEY": "ki_sk_xxx" }
}
}
}Restart the host. You now have 12 tools available to the AI:
Tool | Purpose |
| Persist a long-term memory |
| Hybrid semantic + keyword search |
| Replay recent/important memories |
| Fetch by id |
| Patch fields |
| Soft/hard delete |
| Enumerate namespaces |
| Probe service |
Configuration
Sources are merged in order: CLI args > env > ~/.kireo/config.json.
ENV / CLI | Default | Description |
| required | Bearer token ( |
|
| Override for self-host. |
|
| Per-request timeout in ms, max |
|
| 5xx/429 retries (alias: |
|
| Exponential backoff base. |
|
| Set to |
|
|
|
| none | HTTP(S) proxy. |
|
| Locale for error hints. |
Logs land in ~/.kireo/logs/ on all platforms (macOS, Linux, Windows).
Indexing local code
Index a repository's symbols (functions / classes / methods) into a
code-<repo> namespace so the AI can recall them via memory_search:
export KIREO_API_KEY=ki_sk_xxx
npx -y -p @kireo/mcp-server kireo index ./ --repo my-appIndexing is incremental — only changed files are re-sent on subsequent runs.
Flag | Default | Description |
| directory basename | Repo name → |
|
| Symbols per upload batch ( |
|
| Per-request timeout (max |
| — | Same as the config table above; CLI flags override env. |
Run kireo --help for the full usage text. --help and --version never touch
the network or the filesystem and don't require an API key. If a batch upload
times out, re-running the same command is safe: the server dedupes identical
symbols, so retries won't create duplicates.
Host setup
Privacy
Set KIREO_TELEMETRY=0 to drop the X-Device-Id header. We never read your code; only the explicit content you pass to memory_save reaches the API.
Troubleshooting
AUTH_INVALID_KEY→ rotate your key at https://app.kireo.app/app/api-keys.QUOTA_EXCEEDED→ upgrade or wait for next billing cycle.Tools missing in your host → run
npx @modelcontextprotocol/inspector node $(npm root -g)/@kireo/mcp-server/bin/kireo-mcp.cjsto verify locally.
License
MIT
Available Tools
12 toolscontext_archiveA
Save and upload a semantically compressed conversation as .md. When to use: after /kireo:compact has summarized the requested conversation rounds, and the user has asked to upload them. The host model does the compression; this tool redacts, writes the Markdown file, splits it losslessly and uploads it to the memory API. Always pass the conversation directory as cwd. Directory archives use their own namespace (different worktrees and same-named directories stay separate) and never change relay project keys. dry_run:true previews without side effects; otherwise invocation saves AND uploads. A failed or incomplete upload stays queued locally. Run compact again to retry. Returns the project, filename, local_path, namespace, snapshot_id, stored, deduped, outbox_pending, memory_ids and dashboard_url. Never call a queued archive uploaded.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | Yes | Absolute directory of THIS conversation; always pass it explicitly. | |
| host | Yes | ||
| dry_run | No | Preview exact redacted upload without disk or network writes. | |
| summary | Yes | Markdown semantically compressed by the host from all requested conversation rounds. Preserve goals, decisions and reasons, constraints, results, corrections and unfinished work. Do not pass raw logs or invent missing history. | |
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does so thoroughly. It discloses side effects ('otherwise invocation saves AND uploads'), the dry_run safety path, local queuing of failed uploads, namespace isolation, and that it 'never change[s] relay project keys.' It even lists returned fields so the agent knows what to expect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-organized: action first, then when-to-use, then side effects, then return fields. Most sentences earn their place. The final sentence 'Never call a queued archive uploaded' is awkward and somewhat cryptic, which prevents a top score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description is unusually complete: it explains the workflow, retry behavior, namespacing, and return fields. The main gaps are the undocumented host and session_id parameters and the exact meaning of the queued-upload warning, but overall the agent can invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 60%, and the description adds useful meaning for cwd ('Always pass the conversation directory as cwd') and dry_run ('previews without side effects'). However, host and session_id receive no additional explanation in either the schema or the description, leaving their semantics partly to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Save and upload a semantically compressed conversation as <directory-name>.md.' It clearly distinguishes itself from siblings by describing its namespace behavior, lossless splitting, and upload-to-memory-API purpose, which sets it apart from context_save and memory_save.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use it: 'after /kireo:compact has summarized the requested conversation rounds, and the user has asked to upload them.' It also gives retry guidance ('Run compact again to retry'), but it does not explicitly mention alternatives or when not to use the tool, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
context_loadA
Load previously saved context for the current project and render it.
When to use: at the start of work on a project — especially on a machine or in a tool where you have not worked on it before.
The first line reports the project identity and the code index freshness. Relay both to the user; never assume the code index is current.
Returns rendered text, grouped constraints → open threads → decisions → gotchas → key files → preferences.
A repo with a .kireo/disabled file, or the KIREO_DISABLED env var,
disables this tool entirely — it returns immediately and does nothing: no
outbox flush, no request, no CONTEXT.md.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | ||
| outbox_dir | No | ||
| token_budget | No | ||
| audit_log_path | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the burden, and it does a solid job: it describes rendered output groups, the first-line identity/index-freshness warning, and the full disabled behavior. It does not fully clarify normal side effects such as whether an outbox flush occurs when enabled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is organized into purpose, when-to-use, output format, and disabled behavior. Every sentence earns its place and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, usage, output structure, and disabled behavior well, but it leaves all four parameters unexplained and does not describe normal-operation side effects or error behavior. For a tool with no annotations and no output schema, those gaps matter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds no meaning for cwd, outbox_dir, token_budget, or audit_log_path. The name hints are not enough, and token_budget in particular needs behavioral context that is absent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Load previously saved context for the current project and render it.' It also names the output grouping, making its role distinct from sibling tools like context_save, context_archive, and the memory_* family.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit guidance: use at the start of work on a project, especially on an unfamiliar machine or tool. It also warns when the tool is disabled, though it does not contrast itself with memory_get or context_save alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
context_saveA
Persist distilled session context for the current project.
When to use: ONCE per /kireo:save, after you have distilled the session. Always call project_info first and show the user the resolved project.
Each entry needs non-empty evidence — a concrete basis from THIS session
(a file path, a command you ran, something the user said). If you cannot
cite one, do not write the entry.
Do NOT store things re-derivable from the repo in 60 seconds — that is the code index’s job. Store decisions, constraints, gotchas, open threads, a small map of key files, and stated preferences.
Privacy: content, evidence, and files are pattern-redacted for known
credential shapes, but that is NOT a privacy guarantee — it cannot catch
business secrets ("customer A’s contract is worth $X") that got paraphrased
into evidence/files from things the session read (.env, docker logs, a
pasted SQL result). The first run is therefore FORCED into a preview by the
tool itself: until a .first-run-acknowledged marker exists next to the
outbox directory, every call returns the exact (post-redaction) text and
writes/uploads nothing, no matter what dry_run says. Show that text to
the user, and only after they confirm there is nothing business-sensitive
call again with an explicit dry_run: false — that call performs the real
save and records the acknowledgement. From then on dry_run is an ordinary
optional preview flag (omitted = real save). A repo with a .kireo/disabled
file, or the KIREO_DISABLED env var, disables this tool entirely — it
returns immediately and does nothing.
Returns: { stored, deduped, failed, outbox_pending, namespace, dashboard_url }
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | ||
| host | Yes | Which host produced this: "claude-code" or "codex". | |
| dry_run | No | Preview only: apply redaction and show the exact text that would leave this device, without writing the outbox or uploading anything. Deliberately has NO schema default, so after parsing three states stay distinguishable: omitted (undefined), explicit false, explicit true. On the very first run (no `.first-run-acknowledged` marker next to the outbox dir) omitting it forces a preview; only an explicit `dry_run: false` performs the real save and records the acknowledgement. Once the marker exists, omitted and false both perform a real save; `true` always previews. | |
| entries | Yes | ||
| outbox_dir | No | ||
| session_id | Yes | ||
| audit_log_path | No | ||
| transcript_path | No | Absolute path to THIS session's transcript (.jsonl). Used for the literal evidence check: file paths and commands cited in `evidence` that never appear in the session mark that entry [uncertain]. Optional — for claude-code it is located automatically from `session_id`; verification is skipped (never blocked) when no transcript can be read. | |
| uncertain_indexes | No | Indexes of entries you are not fully confident in. These are UNIONed with whatever the transcript evidence check flags, and render as [uncertain] on the resume side. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers exceptional disclosure: forced first-run preview keyed on a .first-run-acknowledged marker, three distinguishable dry_run states, pattern-redaction limits with explicit business-secret caveats, disable conditions (.kireo/disabled, KIREO_DISABLED), and the transcript-based evidence verification that marks entries [uncertain].
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Long but perfectly dense — every sentence carries operational weight: purpose, usage trigger, evidence rule, exclusion, privacy caveat, state machine, disable conditions, and return shape. It is front-loaded with the 'When to use' block ahead of the deep dry_run mechanics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 9-parameter mutation tool with no annotations or output schema, the description is complete: it covers call sequencing, content constraints, the first-run state machine, privacy caveats, disable conditions, and even lists the return fields (stored, deduped, failed, outbox_pending, namespace, dashboard_url) in lieu of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 44%, so the description must compensate, and it adds real meaning: it maps the bucket enum to intent ('decisions, constraints, gotchas, open threads, a small map of key files, and stated preferences') and defines evidence as 'a concrete basis from THIS session (a file path, a command you ran, something the user said),' far beyond the schema's minLength:1. Path params like cwd, outbox_dir, and audit_log_path remain unexplained, keeping this a step below full coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource: 'Persist distilled session context for the current project.' The save operation is clearly distinct from siblings like context_load, context_archive, and memory_save by combining the verb (persist) with the scoped resource (distilled session context).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use: 'ONCE per /kireo:save, after you have distilled the session,' plus a hard prerequisite: 'Always call project_info first and show the user the resolved project.' It also gives an exclusion boundary ('Do NOT store things re-derivable from the repo in 60 seconds — that is the code index's job').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_deleteA
Delete a memory. Defaults to a soft delete (30-day recovery window).
When to use:
The user explicitly asks to forget something ("forget that I said …").
A memory is clearly wrong AND not worth correcting.
GDPR / privacy removal request.
When NOT to use:
The memory is just outdated — prefer memory_update.
You're unsure — ask the user first.
Note: deletes are always soft — the API has no immediate hard delete. Soft-deleted memories are purged permanently after the 30-day window.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| hard | No | Accepted for compatibility, but the API only supports soft deletes — the memory is recoverable for 30 days, then purged permanently. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. Discloses soft delete behavior, 30-day recovery window, permanent purge after 30 days, and clarifies that the 'hard' parameter only emulates behavior but still performs soft delete. Could mention auth requirements or side effects, but covers key behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very efficient: first sentence states core action and default behavior, then bullet points for usage guidelines. No wasted words, front-loaded with key information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of output schema, the description covers essential behavioral context (soft delete, recovery window) and usage scenarios. Could mention what the API returns (e.g., success/failure) but not critical for a delete operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% (only 'hard' has a description). The description adds minimal value for 'id' (no additional detail), but clarifies the 'hard' parameter's behavior beyond schema. Partially compensates but does not fully address missing id description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description explicitly states 'Delete a memory' with soft delete behavior. It distinguishes from siblings like memory_update by specifying when not to use (outdated memories) and when to use (explicit forget, wrong memory, GDPR).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear when-to-use and when-NOT-to-use scenarios with concrete examples, and references an alternative tool (memory_update). This offers excellent guidance for agent decision-making.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_getA
Fetch a single memory by id.
When to use:
You already have a memory id from a previous memory_search / memory_recall / memory_save call.
The user references "that memory you saved" and you stored the id.
When NOT to use:
You only have a vague query — use memory_search or memory_recall instead.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The memory id (looks like "mem_01HXVK..."). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Description indicates a read-only operation ('Fetch'), but does not discuss error handling or auth requirements. For a simple fetch with no annotations, this is nearly sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is concise (5 lines), well-structured with bullet points, and every sentence adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema is needed for a simple fetch; the description covers purpose, usage, parameter context, and sibling differentiation completely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers 100% of parameters, but the description adds valuable context: the id comes from previous calls and relates to user references, enhancing understanding beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states 'Fetch a single memory by id' (verb+resource) and differentiates from siblings by specifying when to use vs. not use memory_search/memory_recall.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear 'When to use' and 'When NOT to use' sections, including specific alternatives like memory_search and memory_recall, giving excellent usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_healthA
Probe the Kireo service and report local + remote health.
When to use:
User reports "memory tools aren't working".
You hit unexplained errors and need to confirm the API is reachable.
First-time setup verification.
Returns: { local: { server_version, node_version, platform }, remote: { status } }.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so description carries full burden. It explains the tool probes health and returns local and remote status. No side effects mentioned, appropriate for a read-only health check.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very concise: first sentence defines purpose, then usage guidelines, then return format. No fluff, well-organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a simple health check tool. Explains return structure (local and remote fields). No output schema, but description covers what is returned.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist, so schema coverage is trivially 100%. Baseline for 0 params is 4. Description doesn't add param info because none needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it probes Kireo service health. Distinguishes from sibling memory tools, which all involve data operations (get, save, delete, etc.).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit scenarios: user reports tools not working, unexplained errors, first-time setup. This helps decide when to call memory_health instead of other memory tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_list_namespacesA
List all namespaces the current user has, with per-namespace counts.
When to use:
Before deciding which namespace to save to, when the user has multiple projects.
To answer "what projects do I have memories for?".
For diagnostics.
Returns: array of { name, created_at }.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It implies read-only behavior and specifies the return structure. However, it does not mention authentication or rate limits; for a simple list tool, this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with three clear sections: core function, usage guidance, and return type. Every sentence adds value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description claims 'per-namespace counts' but the return structure only includes name and created_at, not counts. This inconsistency reduces completeness. Also, no mention of pagination or limits, though the input schema has no parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist, so schema coverage is 100% vacuously. The description does not need to add parameter info. Baseline 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it lists all namespaces with per-namespace counts, using a specific verb and resource. It distinguishes from sibling tools that operate on individual memories.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'When to use' section with three specific scenarios (pre-save decision, answering user questions, diagnostics). Provides clear context for usage, though no explicit exclusion or alternative tools mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_recallA
Replay recent or important memories from a namespace, without a query.
When to use:
At the start of a new chat — to load recent context.
The user says "what have we been working on" / "remind me".
You want a chronological feed rather than a search match.
When NOT to use:
The user has a specific question → use memory_search.
You already have the id → use memory_get.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| order | No | recency | |
| since | No | Only return memories with occurred_at >= since. | |
| cursor | No | Opaque cursor from a previous recall response, for pagination. | |
| namespace | No | default |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, description carries full burden. It mentions 'without a query' and chronological feed, but lacks details on memory boundaries, ordering behavior, pagination handling, or effect on memory state. Adequate but not deep.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Highly structured with clear headers, bullet points, and no redundant sentences. Purpose is front-loaded, and every section adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has 5 parameters and no output schema or annotations. Description covers usage context well but omits parameter details and behavioral specifics about ordering or pagination. Adequate given complexity but incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 40%. Description adds no parameter-specific guidance beyond the schema. It does not explain limit, order, since, cursor, or namespace semantics, leaving agent to infer from defaults/enums.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it replays recent/important memories from a namespace without a query. It distinguishes itself from siblings like memory_search (query-based) and memory_get (by id).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'When to use' and 'When NOT to use' sections with specific alternatives (memory_search, memory_get) and context cues like 'new chat' or user phrases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_saveA
Persist a long-term memory for the current user.
When to use:
The user states a preference, decision, fact, plan, or goal that should outlive this chat.
You learn a stable property of the user, their project, or their tooling.
The user explicitly says "remember", "记住", "save this".
When NOT to use:
Ephemeral chat turns ("ok", "thanks").
Sensitive secrets (API keys, passwords) — refuse and warn the user.
Information you can re-derive from the codebase on demand.
Returns: { id, created_at, schema_version, embedding_status } — use memory_get for the full record.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | ||
| type | No | "preference" = user likes/dislikes, "decision" = chosen approach, "fact" = stable truth, "event" = time-bounded happening, "goal" = intent, "insight" = derived learning, "relationship" = link between entities. | fact |
| content | Yes | The memory text to persist. Be concrete and self-contained — future-you should understand it without surrounding chat. | |
| entities | No | Named entities mentioned (people, products, repos). Helps later retrieval. | |
| metadata | No | ||
| namespace | No | Logical bucket (e.g. project name). Use "default" unless the user has multiple isolated contexts. | default |
| importance | No | 0..1 priority hint. Defaults to 0.5 server-side. | |
| occurred_at | No | ISO-8601 timestamp when the memory factually happened. Defaults to server now(). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries full behavioral burden. It discloses persistence nature and return fields, but lacks mention of deduplication, overwrite behavior, or rate limits. The advice to use memory_get for full record is helpful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, front-loaded with purpose, and structured into clear sections. Every sentence adds value without repetition. The bullet lists are efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters and no output schema, the description adequately explains the return structure and references sibling tools. It could elaborate on interactions with memory_update or memory_delete for completeness, but the current coverage is sufficient for most agents.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so the description compensates with extra context like 'Be concrete and self-contained' for content and detailed interpretation of type enum values. This adds value beyond the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Persist a long-term memory for the current user,' which clearly states the action and target resource. It distinguishes itself from sibling tools like memory_get by instructing to use that tool for the full record. The purpose is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'When to use' and 'When NOT to use' bullet points provide concrete scenarios, including when to store preferences, facts, etc., and when to avoid ephemeral chat or secrets. The guidance on refusing sensitive secrets adds valuable safety advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_searchA
Hybrid semantic + keyword search over the user's memories.
When to use:
The user asks "do you remember…" or "what did I say about X".
Before answering project-specific questions, search for stored preferences and decisions.
You need to ground your answer in user-supplied facts.
When NOT to use:
You already have the memory id → use memory_get.
You just want the most recent items in a namespace → use memory_recall.
Returns: ranked list of MemoryRecord with a raw RRF relevance score (typically 0.01–0.04).
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | ||
| type | No | ||
| limit | No | ||
| query | Yes | Natural language query. Hybrid (semantic + keyword) search. | |
| entities | No | ||
| min_score | No | Optional raw RRF score threshold. Typical useful values are 0.01–0.04. | |
| namespace | No | Restrict to a single namespace. Omit to search across all of the user's namespaces. | |
| occurred_to | No | ||
| occurred_from | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses the hybrid search behavior, returns a 'ranked list of MemoryRecord with a raw RRF relevance score (typically 0.01–0.04).' This is informative, though it does not cover potential edge cases like empty results or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear headings and concisely covers purpose, usage, and return format. Every sentence is justified and non-redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (9 parameters, no output schema), the description effectively explains the tool's role, when to use it, and what it returns. It could be more thorough on interpreting the score, but it is sufficient for a search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 33% (low), but the description adds some value beyond the schema, such as explaining the query as 'Natural language query. Hybrid (semantic + keyword) search' and noting typical score ranges for min_score. However, many parameters (tags, type, entities, occurred_from, occurred_to) lack additional explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool performs 'Hybrid semantic + keyword search over the user's memories.' It distinguishes from siblings by specifying when to use (e.g., 'do you remember…') and when not to use (use memory_get or memory_recall instead).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'When to use' and 'When NOT to use' sections provide clear context. It lists alternative tools (memory_get, memory_recall) and specific scenarios, offering strong guidance for an AI agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_updateA
Patch an existing memory. Only the fields you provide are changed.
When to use:
The user corrects a previously stored fact ("actually, I use Tailwind v4 not v3").
You need to add tags / entities to an existing memory.
The user lowers/raises priority of a memory.
When NOT to use:
The memory is wrong AND no longer relevant → use memory_delete.
You don't have the id yet → search first.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| tags | No | ||
| type | No | ||
| content | No | ||
| entities | No | ||
| metadata | No | ||
| namespace | No | ||
| importance | No | ||
| occurred_at | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so description carries full burden. It discloses patch semantics but does not mention return value, side effects, or authentication needs. Adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections, no unnecessary words, and front-loaded with the core action. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 9 parameters, no output schema, and no explanation of entities, importance scale, or return structure, the description lacks depth for an agent to use the tool confidently in complex scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description does not elaborate on any parameter beyond the schema definitions. Parameters like 'metadata' or 'namespace' remain opaque, leaving the agent without additional guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is for patching an existing memory with only provided fields changed. It distinguishes from sibling tools like memory_delete and memory_save by context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'When to use' and 'When NOT to use' sections with concrete scenarios (corrections, tag updates, priority adjustments) and alternatives (memory_delete, search first).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_infoA
Resolve the stable identity of the current project and the namespaces its memories live in.
When to use: before every context_save and context_load, and any time the user asks which project or bucket their memories are going to.
ALWAYS call this before context_save or context_load, and ALWAYS show the
user the returned key and source — a misidentified project is the one
failure the user can catch instantly, and cannot catch at all if hidden.
If warn is non-null the identity is NOT stable across machines; relay it verbatim.
Returns: { key, source, display_name, warn, ctx_namespace, code_namespace, home_namespace }
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Absolute path of the project directory. Defaults to the server process cwd. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With zero annotations, the description carries the full behavioral burden, and it delivers: it defines the `warn` field's semantics ('identity is NOT stable across machines; relay it verbatim'), lists the exact return shape, and warns that a misidentified project is a failure the user can catch instantly. It stops short of explicitly confirming the tool is free of side effects, though 'Resolve the stable identity' strongly implies a pure lookup.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first sentence, and every subsequent sentence earns its place: usage rule, mandatory call ordering, warn behavior, and return shape. At roughly 100 words it is dense rather than padded, doing the work that annotations and an output schema would otherwise need to do.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema and no annotations, the description covers usage timing, mandatory ordering, the return shape, and the warn edge case. The only gap is that fields like `key` vs `source` and the three namespaces are named but never defined — though their meanings are largely inferable from the tool's stated purpose.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a complete description of `cwd` (absolute path, defaults to server cwd), so the baseline is 3. The description adds nothing about the parameter, but none is needed since the schema already carries the full semantic load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names a specific verb and resource: it 'resolve[s] the stable identity of the current project and the namespaces its memories live in.' This clearly differentiates it from the memory_*/context_* siblings, which operate on memory contents rather than project identity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage timing is explicit and imperative — 'ALWAYS call this before context_save or context_load' — and it names a rule-based trigger ('any time the user asks which project or bucket their memories are going to'). It ties the tool directly to the sibling flows it precedes, giving an agent unambiguous routing. A 'when not to use' exclusion is absent, but the positive guidance is strong enough to compensate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.3.1- Added
context_archive - Added
context_load - Added
context_save - Added
project_info
8 tool updates
v0.2.1- First observed
memory_delete - First observed
memory_get - First observed
memory_health - First observed
memory_list_namespaces - First observed
memory_recall - First observed
memory_save - First observed
memory_search - First observed
memory_update
TDQS
Scored across 12 tools
Each tool targets a distinct operation: memory_get/search/recall are clearly separated by id, query, or chronological feed; memory_save/update/delete cover lifecycle edits; context_save/load/archive handle project session context. Even the potentially overlapping memory and context tools are differentiated by their descriptions and usage guidance.
Memory tools consistently use memory_<verb> (memory_save, memory_search, memory_delete), and context tools use context_<verb>. The only deviation is project_info, which breaks the verb-first pattern, and memory_list_namespaces which mixes noun phrasing, but overall the convention is predictable and readable.
12 tools is well-scoped for a memory and context persistence server. Each tool earns its place: CRUD for memories, search/recall/health/namespace utilities, and context save/load/archive. No redundant tools or excessive surface area.
The tool surface covers the full memory lifecycle: create, read, update, delete, search, recall, and namespace management. Context persistence is complete with save, load, and archive. Project identity resolution and health probing fill the operational gaps, leaving no obvious dead ends.
Maintenance
Related MCP Connectors
An MCP memory server. One memory your agents share — across models, devices and apps.
Persistent personal memory for AI assistants — save, search, and recall across every MCP client.
Persistent memory for AI agents to retain, retrieve, and recall conversation context through MCP.
- mcpOAuthai.butlerbrain
Persistent memory for AI assistants. Save once; recall from Claude, ChatGPT, or any MCP client.
Related MCP Servers
- AlicenseAqualityCmaintenanceA production-grade long-term memory MCP server that enables AI agents to persist and recall memories across sessions with importance weighting, confidence calibration, and efficient context window management.955 PyPI1MIT
- AlicenseNot gradedqualityCmaintenanceA lightweight MCP server that provides long-term memory for LLMs by storing and retrieving important facts, decisions, and preferences through smart semantic search and automatic organization.12MIT
- AlicenseAqualityAmaintenanceA long-term memory MCP server for AI agents that stores memories (facts, decisions, etc.) in a single SQLite database with hybrid search and full edit history, ensuring consistency across sessions.381MIT
- AlicenseAqualityDmaintenanceMCP server for persistent, semantic memory across AI sessions; store context, decisions, and learnings and recall them with natural language search.227 npmMIT