Aelios
Aelios exposes a long-term memory server for AI assistants, enabling full CRUD, recall, ingestion, and curation of memories, glossary, and diary entries.
Search and list memories with filters for type, namespace, status, and pagination.
Export all memory records as JSON.
Get, delete, archive, or pin specific memories by ID; pinned memories are precious and exempt from dedup/decay/delete.
Ingest chat messages and optionally auto-extract memories.
Cold-start boot: retrieve yesterday's log, top pinned precious memories, and full glossary.
Explicit recall with complete hit records, glossary, active memories, world facts, and longtail; can include superseded history.
Upsert refined memories by fact_key with tags, confidence, importance, validity, hand-authored signature, and response tendency.
Supersede old memories with new content, linking the supersede chain.
Add or update glossary terms with aliases and examples.
Read diary entries (daily/weekly logs).
Integrates with Cloudflare services including Workers AI, AI Gateway, D1, and Vectorize for memory storage, embeddings, optional full chat forwarding, and overnight consolidation.
Supports OpenAI-compatible chat completions and the OpenAI Responses API, forwarding to OpenAI models and enabling memory for Codex.
Provides a Telegram integration branch (tg-bot), though it is noted as stuck at 2026-.
Aelios
A long memory for your AI. Switch windows, switch clients, switch models — the memory follows you.
English | 简体中文 · QQ group: 1091783659
Aelios is a memory gateway that runs on Cloudflare. Your Chatbox / Claude Code / Codex connects to Aelios first; Aelios forwards the request to the model. Every conversation is kept, and the relevant parts come back automatically next time.
your client → Aelios (identify, remember, recall) → the modelQuick start — three steps
1. Deploy
Click the button and sign in to Cloudflare. The only required field is CHATBOX_API_KEY — make up a password, e.g. sk-my-aelios. Everything else can stay empty.
For the Vectorize index, copy these values: Dimensions 1024, Metric cosine. Build command npm ci, deploy command npm run deploy.
You'll get an address like:
https://companion-memory-proxy.<your-subdomain>.workers.devPrefer to control every step? Fork this repo → connect it in Cloudflare Workers → build npm ci, deploy npm run deploy → add the CHATBOX_API_KEY secret in Worker Settings. Don't run a bare wrangler deploy — it won't create the databases.
2. Add an assistant
Open:
https://<your Worker address>/adminEnter your Worker address and the key from step 1, then go to Settings.
Upstream: your Cloudflare account ID (32 characters). For full chat forwarding, also add a
CLOUDFLARE_API_TOKENWorker secret.Assistant: click Add assistant, for example:
Field
What it means
Example
Name
The segment in the URL, lowercase English
coderPrimary model
Only these models remember and recall
anthropic/claude-opus-4-5Keys
Who may use this assistant
check the primary key
Click Save.
One assistant, one address. Name it coder and the address is:
https://<your Worker address>/coder/v1Only registered primary models get memory. Unregistered models (like the small utility models inside Claude Code) just pass through — nothing is recorded or recalled for them.
3. Point your client at it
Use CHATBOX_API_KEY as the API key everywhere. Write model names as vendor/model, e.g. anthropic/claude-opus-4-5.
Client | Address to use |
Chatbox, Cherry Studio, … |
|
Claude Code |
|
Codex |
|
A bare /v1 without an assistant name routes to the first assistant of that key.
Try it: say "Please remember: my test code phrase is apple-star-0428." Wait a bit, then ask "What is my test code phrase?" If it answers, you're wired up.
Related MCP server: second-brain-cloudflare
Day-to-day: the admin panel
Open /admin. The bottom tabs:
Tab | What it's for |
Today | What you talked about today |
Review queue | Candidate memories consolidated overnight. Each assistant decides its own first; the week's decisions are listed here with an undo |
Important memories | Browse, search, edit, delete |
More | Precious originals, glossary, maintenance tools |
Settings | Upstream, assistants, environment parameters |
Want the AI to remember, forget, or fix something? It's all point-and-click.
How it actually works
Every time you speak, relevant old memories are tucked onto the end of the current conversation.
Raw conversations are stored first; overnight, Aelios consolidates them into long-term memory (the Dream pass). Each assistant decides what it keeps using the model it chats with, and you can undo any of its calls.
Memory lives in your own Cloudflare account (D1 + Vectorize) — never tied to a chat window, never on someone else's server.
One assistant can write to one space and read from several. A fresh conversation can write coder while also reading the old vault coder-old and a shared shared-docs. Leave the recall spaces empty and it only reads its own.
Optional features
Full chat gateway — let Aelios forward to every vendor's models, billed against your own keys:
Cloudflare → AI → AI Gateway: create a gateway and store each vendor's key under it.
Add the
CLOUDFLARE_API_TOKENWorker secret.In
/adminSettings, put your 32-character account ID as the upstream and add assistants.
Skip all of that and memory recall + the overnight Dream still work (via Workers AI).
MCP memory for Claude Code / Codex:
https://<Worker address>/mcp?token=<CHATBOX_API_KEY>For automatic per-message recall and batch write-back in Claude Code, use the Hook in this repo: integrations/claude-code/.
Seeing images (guide-dog): add the GUIDE_DOG_API_KEY secret, then point your client at https://<Worker address>/v1/guide-dog with model companion. The guide-dog only describes images — it writes no memory.
Common pitfalls
Deploy with
npm run deploy, never a barewrangler deploy.Assistant names are part of the URL: use
coder-style English, no spaces.Model names need the vendor prefix:
anthropic/claude-opus-4-5, notclaude-opus-4-5.Only registered primary models get memory. Keep small utility models out.
Never hand-delete the Vectorize index (
memo-kb, 1024-dim cosine).
Digging deeper
How the gateway routes, remembers, and recalls: docs/memory-gateway.md
What happens to requests, how thinking is handled: docs/request-contract.md
Which secrets are required: SECRETS.md
For maintainers: endpoints, MCP tools, and the memory lifecycle — see the sections below.
For maintainers
A memory gateway on Cloudflare Workers. When helping a user deploy, always link the user's own fork; all secrets live in the user's own account.
Resource | Value |
Worker |
|
D1 |
|
Vectorize |
|
Queue |
|
Embedding |
|
Entry points: /<assistant>/v1/chat/completions, /<assistant>/v1/messages, /<assistant>/v1/responses. A bare /v1/... uses the first assistant of that key.
Upstreams: chat goes through compat (all vendors, BYOK); messages / responses hit each vendor's native endpoint. Model names pass through untouched — if the vendor doesn't recognize one, the upstream error says so. Custom OpenAI-compatible base URLs are forwarded as-is.
Configuration has three layers, all under /admin Settings: upstream, assistants, environment parameters. Precedence: panel-saved > GATEWAY_CONFIG > empty config.
Common endpoints
Method | Path | Notes |
GET |
| Health check |
GET |
| Admin panel |
GET |
| Model list |
POST |
| OpenAI-compatible chat |
POST |
| Anthropic messages |
POST |
| OpenAI responses |
GET / POST |
| MCP memory tools |
GET / POST |
| List / create memories |
POST |
| Dynamic recall (preferred by hooks) |
POST |
| Raw search (unmetered) |
POST |
| Ingest raw chats |
GET |
| Cold-start pack |
GET |
| Diary |
GET / POST / DELETE |
| Precious originals / glossary |
GET / POST |
| Review queue; |
Everything except /health and /admin requires Authorization: Bearer <key>. The gateway authorizes by assistant key; memory APIs by memory:read / memory:write scopes.
MCP tools
memory_search memory_list memory_get memory_delete memory_ingest memory_boot memory_recall memory_upsert memory_supersede memory_archive memory_pin glossary_set diary_get memory_export
How memory flows
Writes: assistants write directly via memory_upsert; a nightly cron (10 20 * * *) extracts facts from the day's conversations → each candidate is judged remember-or-let-go by its space's own assistant, using the main model it last chatted with (falling back to JUDGE_MODEL, which leaves unsure ones for you; an assistant can be switched off main-model judging in its settings to save quota), and writes diary / weekly / monthly entries.
Recall: your latest message → vector search + lexical match → batch rerank of original passages + rules → the clean original text is tucked onto the end of the current message. No generative LLM by default: at most one memory on a normal turn, two when answering about the past; low scores are never padded in. Rerank failures fall back to lexical. Scores and trade-offs are visible in /admin → Settings. Transport envelopes, hashes, and message IDs never enter the daily prompt; two adjacent messages within 90 seconds in one session are merged. Active search still returns full records and IDs. Diaries are not auto-injected.
Cleanup: messages live ~7 days; expired memories are flagged at 180 days and hard-deleted 30 days later.
Local verification
npm install
npm run verifyTests need Node.js 22+.
Lint is blocking in CI: npm run lint must be zero-error to merge. There are still 17 noNonNullAssertion warnings; the rule is demoted to warning in biome.json and doesn't affect the exit code — clean them up if you're touching that line, and never bulk-fix with Biome's autofix: it rewrites x!.y into x?.y, turning "crash loudly" into "silently return undefined".
License
AGPL-3.0
Free to use and modify; if you run a modified version as a network service, you must open-source your modifications under the same license.
The final v1 point is tag v1-final (MIT at the time). tg-bot is the Telegram integration branch, stuck at 2026-07-16 as of 2026-09-11 with main 154 commits ahead (135 files, ~18k lines changed) while the branch itself only added 17 files — it can't be merged back as-is. Kept as history; resurrect it by rebasing onto main, never by direct merge.
Community
QQ group: 1091783659
Available Tools
14 toolsdiary_getARead-only
Read daily_log diary entries. These are impressions, not verified facts. Omit date for today+yesterday (recent). Use week (e.g. 2026-W29) for weekly_log. Rolled-up daily dates fall back to weekly_log. Fetch explicitly when needed; do not treat diary as ground truth.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | YYYY-MM-DD; omit for recent (today+yesterday) | |
| week | No | ISO week label, e.g. 2026-W29 | |
| namespace | No | Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds real context beyond that: entries are impressions rather than verified facts, and rolled-up dates silently fall back to weekly_log, which is a non-obvious behavioral trait. It stops short of describing output shape or pagination.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five short clauses, none of which waste words, and the core semantics (recent default, weekly label) are front-loaded. The telegraphic style is dense but readable and earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with no output schema and a fully documented 3-param schema, the description supplies the missing domain framing (what a diary entry is, its reliability) and the date/week fallback rule. Return format is undescribed, but no output schema exists, so the remaining gap is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by explaining how date and week interact (the weekly_log fallback for rolled-up dates) and when to omit each. That is genuine semantics the schema does not carry.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Read daily_log diary entries') and characterises the content ('impressions, not verified facts'), which separates it from the memory_* siblings that return authoritative memories. It does not explicitly name a sibling alternative, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete selection guidance: omit date for today+yesterday, use week for weekly_log, and dates rolled up into weekly_log fall back there. 'Fetch explicitly when needed; do not treat diary as ground truth' sets a usage posture. It lacks explicit when-not/alternative routing to the memory_* tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
glossary_setADestructiveIdempotent
Add or update a glossary term: private vocabulary such as nicknames, in-jokes or project names. Terms are matched literally (not by vector) during memory_recall and are all included in memory_boot. Upserts by term: an existing term gets the new definition, and aliases/examples are replaced by what you pass (omitting them clears them). Returns { data: glossary row }.
| Name | Required | Description | Default |
|---|---|---|---|
| term | Yes | The word or phrase exactly as it is used. | |
| aliases | No | Other spellings or names that should match this term. Replaces any existing aliases. | |
| examples | No | Example sentences using the term. Replaces any existing examples. | |
| namespace | No | Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace. | |
| definition | Yes | What the term means. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations declare destructiveHint=true and idempotentHint=true, and the description explains exactly why: upsert-by-term semantics, plus the non-obvious fact that aliases/examples are replaced by what you pass and omitting them clears them. It also discloses literal (non-vector) matching and returns '{ data: glossary row }', adding real behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with no filler: purpose first, then matching/usage context, then mutation semantics and return shape. Each sentence earns its place and the key constraint (clearing aliases) is stated in the same breath as the upsert behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description gives the return shape, covers matching behavior, mutation/upsert semantics, and namespace handling is covered by the schema. Nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented, including the 'Replaces any existing aliases/examples' semantics. The description's upsert sentence mostly restates what the schema says, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb+resource ('Add or update a glossary term') and immediately scopes the resource ('private vocabulary such as nicknames, in-jokes or project names'), which clearly separates it from the memory_* siblings. An agent can identify the tool's job without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains where glossary terms surface ('matched literally during memory_recall', 'all included in memory_boot'), which gives useful context on why to set one. However, it never explicitly states when to prefer glossary_set over a sibling like memory_upsert, nor any exclusions or prerequisites, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_archiveADestructiveIdempotent
Soft-archive a memory: sets status='archived' and removes it from search and recall, but keeps the record in storage (memory_list with status='archived' still shows it). There is no MCP tool to un-archive. Does not touch the supersede chain. Prefer this over memory_delete when the memory might be needed later. Returns { data: { id, archived: true } }, or an error result 'Memory not found'.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Memory id (e.g. from memory_search, memory_recall or memory_list). | |
| namespace | No | Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare destructiveHint=true/idempotentHint=true; the description goes well beyond by specifying exactly what is and isn't affected (removed from search/recall, retained in storage and visible via memory_list with status='archived', supersede chain untouched), that un-archiving is impossible, and the exact success and error return shapes. This is unusually rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and its effect, then consequence, then alternative-tool guidance, then return shape. Every sentence carries distinct, decision-relevant information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param mutation tool with no output schema, the description covers effect, irreversibility, persistence behavior, alternative selection, and return/error values. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters (id, namespace) are already documented with examples and defaults in the schema, so the description adds no parameter-level meaning. Baseline 3 applies since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Soft-archive a memory') and immediately defines the semantics: status becomes 'archived', it disappears from search/recall, but the record persists. It also explicitly contrasts itself with the sibling memory_delete, so an agent can differentiate without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing guidance: 'Prefer this over memory_delete when the memory might be needed later,' and notes there is no MCP tool to un-archive, which is a decisive constraint for choosing this tool. Both the when-to-use and the irreversible consequence are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_bootA
Cold-start context package for a new session: yesterday's diary log, the latest weekly and monthly summaries, the 20 most recent precious entries (from memory_pin), every glossary term and any spontaneous perception items, as { data: {...} }. Output is stable and deterministically ordered so the client can cache it. Call once at session start, not every turn; for per-question lookups use memory_recall. Never changes memory content; it only records that the returned precious entries were shown.
| Name | Required | Description | Default |
|---|---|---|---|
| namespace | No | Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations mark this readOnlyHint=false and idempotentHint=false, and the description correctly explains why: it 'never changes memory content' but 'records that the returned precious entries were shown.' That reconciles the non-read-only flag with a content-preserving side effect, and adds useful caching context ('stable and deterministically ordered so the client can cache it'). It stops short of describing cost, latency, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the payload contents, then usage rules, then the side-effect caveat — a logical order with no filler. The first sentence is a long enumeration, but each item listed is genuinely informative rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return contract itself (the { data: {...} } shape, its constituent sections, and its deterministic ordering). Combined with the usage rule and side-effect disclosure, an agent has everything needed to invoke this correctly once per session.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single optional parameter (namespace) and schema description coverage is 100%, including the default and the fixed-namespace override case. The description adds nothing about it, which is the expected baseline when the schema fully documents the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific operation (cold-start context package) and enumerates exactly what it bundles: yesterday's diary log, weekly/monthly summaries, 20 most recent precious entries, glossary terms, and perception items. It is clearly distinguishable from siblings like memory_recall and memory_search, which it explicitly routes around.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the when ('call once at session start, not every turn') and names the alternative for the opposite case ('for per-question lookups use memory_recall'). Both the trigger and the explicit exclusion are present, so routing is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_deleteADestructiveIdempotent
Permanently delete one memory by id, removing it from storage and from the search index. This cannot be undone. To hide a memory but keep the record, use memory_archive; to replace an outdated fact while keeping its history, use memory_supersede. Returns { data: { id, deleted: true } }, or an error result if the id is not found or the memory could not be removed from the search index (the record is then kept).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Memory id (e.g. from memory_search, memory_recall or memory_list). | |
| namespace | No | Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, but the description adds real context beyond them: irreversibility ('cannot be undone'), the dual removal from storage AND the search index, and failure semantics where an index-removal failure keeps the record. That is the exact behavioral detail an agent needs before a destructive call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the destructive action and its irreversibility, followed by routing alternatives and return shape. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description supplies the return shape ({ data: { id, deleted: true } }) plus the error case, and covers the destructive semantics, alternatives, and parameters. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters are documented there, including the namespace fallback rule. The description adds essentially no parameter detail beyond 'by id', so the schema carries the full load and baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb+resource+scope: 'Permanently delete one memory by id, removing it from storage and from the search index.' It also names the two adjacent siblings (memory_archive, memory_supersede), letting an agent distinguish it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-not alternatives with the selecting condition for each: use memory_archive to hide while keeping the record, use memory_supersede to replace an outdated fact while keeping history. No inference is required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_exportARead-only
Export the namespace's active memories (content plus metadata) as one JSON payload, for backup or migration. Archived and superseded memories are not included. The result can be large; use memory_list to page or memory_search to look things up. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Only export one memory type (fact, event, preference, relationship, boundary, habit, decision, note). Omit for all. | |
| format | No | Output format. Only json is supported (default). | |
| namespace | No | Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, and the description reinforces this with 'Read-only' while adding real behavioral context: archived/superseded memories are filtered out and the payload can be large. It goes beyond the annotations, though it says nothing about size limits, auth requirements, or truncation risk.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and scope, then the exclusion rule, then the alternative routing. Every sentence carries distinct information with no repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing the return ('one JSON payload' of content plus metadata) and does so, while also covering filtering behavior and where to go for smaller/paged reads. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the type enum, format enum, and namespace fallback behavior all documented in the schema. The description adds no parameter-level detail beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('Export the namespace's active memories ... as one JSON payload') and immediately bounds scope by excluding archived and superseded memories. An agent can distinguish this bulk dump from memory_list/memory_get without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a use context ('for backup or migration') and names alternatives with an explicit condition: 'The result can be large; use memory_list to page or memory_search to look things up.' It lacks an explicit when-not clause (e.g., don't use for single-record retrieval), but the routing intent is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_getARead-only
Fetch one memory by id with its full record and lifecycle fields (fact_key, version status, superseded_by). Returns { data: record }, or an error result 'Memory not found' if the id does not exist in this namespace. Use it after memory_search, memory_recall or memory_list when you need the complete record before editing. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Memory id (e.g. from memory_search, memory_recall or memory_list). | |
| namespace | No | Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered; the 'Read-only.' trailer is largely redundant. The description still adds real behavioral context beyond annotations: the exact return shape { data: record } and the 'Memory not found' error outcome for a missing id in the namespace.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight, front-loaded sentences: return value first, then error case, then usage guidance. 'Read-only.' duplicates the readOnlyHint annotation and is the one sentence that does not fully earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by naming the return envelope and the not-found error. Usage ordering and read-only status are also covered, leaving only minor things (e.g. behavior when the namespace is fixed by the API key) to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both id and namespace are already documented in the schema (including the fixed-namespace behavior). The description adds meaning for 'id' only indirectly ('in this namespace'), so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Fetch one memory by id') plus the payload it returns (full record with lifecycle fields fact_key, version status, superseded_by). The singular 'one memory' contrasts cleanly with the list/search siblings, so an agent can distinguish it without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly specifies when to call it: 'after memory_search, memory_recall or memory_list when you need the complete record before editing.' It names the sibling alternatives and the condition that selects this tool, leaving no inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_ingestA
Store raw chat messages as conversation history. Memories are not extracted immediately: the nightly background pipeline (dream) reads stored messages and distills them into long-term memories later. To save a fact right away, use memory_upsert. Returns { data: { conversation_id, message_ids } }; message_ids can be passed to memory_pin as context.
| Name | Required | Description | Default |
|---|---|---|---|
| source | No | Label for where the messages came from. Defaults to 'mcp'. | |
| messages | Yes | Messages in chronological order. | |
| namespace | No | Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace. | |
| auto_extract | No | Accepted for compatibility only; it is not stored and does not change processing. | |
| conversation_id | No | Conversation to append to. Defaults to the namespace's shared default conversation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations mark this as a non-read-only, non-idempotent write but say nothing about timing. The description discloses the crucial asynchronous behavior (nightly 'dream' pipeline distills messages later) and the return shape, which is well beyond what the annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: purpose first, deferred-processing caveat second, alternative and return value last. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Though there is no output schema, the description supplies the return shape ({ conversation_id, message_ids }) and how to use message_ids downstream. The deferred-processing model is explained, so an agent has everything needed to call and reason about the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so parameters (source, messages, namespace, auto_extract, conversation_id) are fully documented by the schema. The description adds no per-parameter meaning, so the baseline 3 applies; the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Store raw chat messages as conversation history.' It also clarifies the deferred nature of the operation and explicitly names memory_upsert as the immediate-save alternative, letting an agent distinguish it from siblings like memory_upsert or memory_search without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: 'To save a fact right away, use memory_upsert,' which gives both the when-to-use and the alternative. It also points to memory_pin as the downstream consumer of message_ids, so the usage flow is fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_listARead-only
Page through stored memories without a query, pinned first, then by importance, then most recently updated; optionally filtered by type or status. Returns { data, paging: { limit, has_more, next_offset } }; pass next_offset back as offset for the next page. Use memory_search or memory_recall to find memories by meaning. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Only list one memory type (fact, event, preference, relationship, boundary, habit, decision, note). | |
| limit | No | Page size. Defaults to 100. | |
| cursor | No | Legacy paging cursor; only used when the lifecycle store is disabled. Prefer offset. | |
| offset | No | Number of records to skip. Use paging.next_offset from the previous page. | |
| status | No | Only list memories with this status: active (default), archived or superseded. | |
| namespace | No | Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace. | |
| include_ids | No | Legacy mode only: also return the bare id list. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the 'Read-only' clause adds little. The real added value is behavioral: the deterministic ordering (pinned, then importance, then most recently updated) and the paging contract including next_offset round-tripping, neither of which is derivable from annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core behavior and ordering, then paging mechanics, then the alternative-tool pointer. Dense but every clause carries information; the only near-redundancy is the trailing 'Read-only', already implied by annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by disclosing the return shape ({ data, paging: { limit, has_more, next_offset } }) and the offset round-trip, which is what an agent needs to page. Ordering and alternative routing are covered; minor gaps like default page size or namespace binding effects are left to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents type, status, offset, cursor, namespace, limit and include_ids in detail. The description only restates type/status filtering and the offset-for-next-page pattern, adding no format or edge-case detail beyond the schema, so baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Page through stored memories') and adds the exact ordering semantics (pinned, importance, recency) plus optional type/status filtering. It also names sibling tools (memory_search, memory_recall) as the meaning-based alternatives, so an agent can route correctly without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent to use memory_search or memory_recall to find memories by meaning, which cleanly carves out this tool's niche as unfiltered enumeration. It stops short of covering the other read siblings (memory_get, memory_export) or stating when-not to use this tool, so it is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_pinA
Save a moment as a precious entry: pinned, offered through memory_boot (which shows the 20 most recent), and never deduplicated, decayed or deleted by the automatic pipeline. Each call creates a new entry. Write the content so it still makes sense on its own later, and attach the surrounding messages via context_message_ids. For ordinary facts use memory_upsert. Returns { data: precious record }.
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | The moment to keep, self-contained and readable on its own. | |
| namespace | No | Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace. | |
| context_message_ids | No | Ids of stored messages (e.g. message_ids from memory_ingest) to keep as surrounding context. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: pinned, surfaced through memory_boot (20 most recent), never deduplicated/decayed/deleted by the automatic pipeline, and each call creates a new entry. This discloses durability and lifecycle behavior that annotations (which only say not read-only, not idempotent) do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the core behavior, then usage guidance, then the alternative tool, then return shape. Every clause carries information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter write tool with no output schema, it covers lifecycle (durable, pinned), discoverability (memory_boot), non-idempotency (new entry each call), and return shape. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds meaning beyond the schema by explaining the purpose of context_message_ids (attach surrounding messages) and the self-contained writing requirement for content, though it doesn't expand namespace semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Save/pin) and resource (a moment/precious entry) and distinguishes itself from siblings by naming memory_upsert as the tool for ordinary facts. An agent can tell this apart from memory_upsert or memory_ingest without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it vs the alternative: 'For ordinary facts use memory_upsert.' It also gives a usage instruction for the content ('write so it still makes sense on its own later') and the routing via context_message_ids.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_recallA
Answer-oriented recall: matches glossary terms literally, runs hybrid search over active memories, reranks the hits and drops anything below min_score. Returns { data: { hits, glossary_hits, week_blocks, meta } }, each hit with id, score and source so you can cite or edit it. Precious entries are not searched here (they come from memory_boot). Prefer this over memory_search when you want only the relevant memories for the current question. Never changes memory content, but marks the returned memories as recently injected, which briefly lowers their rank in later automatic recall.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | Maximum number of memory hits. Defaults to 20. | |
| query | Yes | The question or topic to recall memories for. | |
| types | No | Only recall these memory types (fact, event, preference, relationship, boundary, habit, decision, note). Omit for all types. | |
| min_score | No | Relevance floor applied after reranking. Defaults to the server setting (0.15). | |
| namespace | No | Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace. | |
| include_history | No | When true, include superseded memory versions (status/version_status=superseded). Default false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and non-idempotent, which is surprising for a 'recall' tool; the description resolves this by explaining it 'marks the returned memories as recently injected, which briefly lowers their rank in later automatic recall' — a genuine side effect beyond structured data. It also states the return shape and that precious entries aren't searched. It does not mention auth namespace binding edge cases, but the side-effect disclosure is the key non-obvious behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences front-load the core mechanism and return shape, then handle routing and side effects. No wasted words, though the return-shape enumeration and side-effect clause make it slightly long; every sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers mechanism, return structure, sibling routing, exclusion of precious entries, and the rank-injection side effect — comprehensive for a 6-param tool with no output schema. Minor gap: it doesn't explain namespace versus API-key binding interplay, but that is adequately handled by the schema description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so every parameter including k, types, min_score, namespace, and include_history is already documented in the schema with defaults and constraints. The description adds the pipeline context (min_score applied after reranking) but no syntax or format details the schema lacks. Baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb+resource ('recall memories') and details the exact pipeline: literal glossary matching, hybrid search over active memories, reranking, and min_score filtering. This is far more specific than siblings like memory_search, and the explicit 'Prefer this over memory_search when...' clause makes the distinction operational.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it ('when you want only the relevant memories for the current question') and names the alternative (memory_search). It also carves out the precious-entry exclusion and notes memory_boot as their source, so the agent knows not to expect them here.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_searchA
Hybrid search (vector + keyword) over the user's active long-term memories. Returns up to top_k full records (id, content, type, status, scores) after a relevance filter but without reranking, as { data: [...] }. Use it to find memories you want to inspect or edit by id. To answer a question with the most relevant, reranked memories plus glossary hits, use memory_recall instead. Never changes memory content; it only updates recall counters.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | What to look for, in natural language or keywords. | |
| top_k | No | Maximum number of records to return. Defaults to the server setting (50). | |
| types | No | Only return these memory types (fact, event, preference, relationship, boundary, habit, decision, note). Omit for all types. | |
| namespace | No | Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, which could mislead an agent into thinking this is a pure read; the description reconciles this by stating 'Never changes memory content; it only updates recall counters.' It also discloses that results are filtered but not reranked and are returned as { data: [...] }. It does not cover pagination, error behavior, or auth requirements, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, zero filler. The core operation is front-loaded, the return shape follows, and the sibling routing plus the mutation caveat come last as qualifications.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still specifies the return shape and fields (id, content, type, status, scores) and the { data: [...] } envelope. Between that and the 100%-covered input schema, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (query, top_k, types, namespace) is already documented in the schema. The description restates top_k and the returned fields but adds no syntax, default, or format detail beyond the schema, which is the expected baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific mechanism and resource: 'Hybrid search (vector + keyword) over the user's active long-term memories.' It explicitly scopes the corpus to active memories and distinguishes itself from memory_recall and memory_list, so an agent can route without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the use case ('find memories you want to inspect or edit by id') and names the concrete alternative with its selecting condition: 'To answer a question with the most relevant, reranked memories plus glossary hits, use memory_recall instead.' Explicit when/when-not/alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_supersedeADestructive
Replace an outdated memory while keeping history: marks old_id as superseded (kept, and visible via memory_recall with include_history) and inserts a new active memory linked to it. The new entry inherits the old fact_key unless new_fact_key is given. If old_id does not exist, the new memory is still created and oldStatus is 'missing'. Use it when a fact changed over time; to fix a mistake in place use memory_upsert, and to remove a memory use memory_archive or memory_delete. Returns { data: { oldStatus, newId } }.
| Name | Required | Description | Default |
|---|---|---|---|
| old_id | Yes | Id of the memory being replaced. | |
| reason | No | Short note on why the old memory is outdated; stored with the chain. | |
| new_type | No | One of fact, event, preference, relationship, boundary, habit, decision, note. Other values become fact (default). | |
| namespace | No | Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace. | |
| authored_by | No | E-axis signature for the new entry (hand sources only) | |
| new_content | Yes | The up-to-date memory text. | |
| valid_as_of | No | When the new fact became true (ISO date or datetime). | |
| new_fact_key | No | fact_key for the new entry. Defaults to the old entry's fact_key. | |
| response_tendency | No | E-axis: how to respond when the new memory fires |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the destructiveHint annotation by disclosing the exact mutation mechanics: the old entry is kept and remains visible via memory_recall with include_history, a new active memory is linked to it, and the fact_key is inherited unless overridden. It also documents the edge case that a missing old_id still creates the new memory with oldStatus 'missing'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core behavior and return shape in a compact sequence of sentences with no filler. Slightly dense at five sentences, but each one carries distinct information (mechanism, edge case, alternatives, return).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter destructive mutation tool, the description covers the operation, its side effects on history, the edge case, sibling routing, and the return shape. Nothing an agent needs to invoke it correctly is missing, even without an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds genuine meaning: new_fact_key defaults to the old entry's fact_key, and old_id's non-existence yields a 'missing' status rather than failure. These behavioral notes exceed what the schema documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Replace an outdated memory while keeping history') and immediately contrasts it with the sibling tools memory_upsert, memory_archive, and memory_delete. An agent can distinguish it from all siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it ('when a fact changed over time') and names the alternatives for adjacent scenarios ('to fix a mistake in place use memory_upsert, and to remove a memory use memory_archive or memory_delete'). No inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_upsertADestructive
Write a refined memory immediately (no waiting for the nightly dream), keyed by fact_key. If an active memory with the same fact_key exists, its content and fields are overwritten in place and the old text is not kept; otherwise a new memory is created. To keep the old version as history, use memory_supersede. Pass authored_by (with the default source) to mark the memory as hand-authored: it ranks above distilled memories and the automatic pipeline cannot overwrite it. Returns { data: { id, created } }.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Free-form labels. | |
| type | No | One of fact, event, preference, relationship, boundary, habit, decision, note. Other values become fact (default). | |
| source | No | Who is writing. Leave unset (defaults to 'mcp'). authored_by is only stored when source is mcp, manual, api or remember_now. | |
| content | Yes | The memory text, one self-contained statement. | |
| fact_key | Yes | Stable key for this fact, e.g. 'user:favorite_coffee'. Reusing a key updates that memory. Facts about the outside world rather than the user use the prefix 'world_fact:'. | |
| namespace | No | Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace. | |
| confidence | No | 0 to 1, how sure you are it is true. Defaults to 0.8. | |
| importance | No | 0 to 1, how much this matters. Defaults to 0.6. | |
| authored_by | No | E-axis signature; only honored on hand sources (mcp/manual/api/remember_now) | |
| valid_as_of | No | When the fact became true (ISO date or datetime). Used as the event date in recall. | |
| response_tendency | No | E-axis: how to respond when this memory fires |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the destructiveHint annotation by stating exactly what is lost ('the old text is not kept'), that matching fact_key overwrites fields in place, and that hand-authored memories outrank distilled ones and resist pipeline overwrites. It even discloses the return shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and the key consequence, and each subsequent sentence carries distinct information (supersede alternative, authored_by behavior, return value). The '(no waiting for the nightly dream)' aside is slightly informal but still informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter mutation tool with no output schema, the description covers the destructive semantics, the key alternative, the special-parameter behavior, and the response shape. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is already 100%, so baseline is 3, but the description adds real meaning for authored_by (ranking behavior, pipeline protection) and clarifies the fact_key overwrite contract that the schema only hints at.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Write a refined memory'), its keying mechanism (fact_key), and its timing semantics versus the nightly distillation pipeline. It is clearly distinguishable from siblings like memory_supersede and memory_ingest.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: 'To keep the old version as history, use memory_supersede,' and explains the authored_by path for hand-authored memories. It gives a clear context and one named alternative, though it doesn't contrast against memory_ingest or spell out when not to write at all.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
14 tool updates
v0.1.2- Changed
diary_get1 field changed- added
Input schema / properties / namespace / descriptionAdded value: +"Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace."
- Changed
glossary_set5 fields changed- added
Input schema / properties / aliases / descriptionAdded value: +"Other spellings or names that should match this term. Replaces any existing aliases." - added
Input schema / properties / definition / descriptionAdded value: +"What the term means." - added
Input schema / properties / examples / descriptionAdded value: +"Example sentences using the term. Replaces any existing examples." - added
Input schema / properties / namespace / descriptionAdded value: +"Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace." - added
Input schema / properties / term / descriptionAdded value: +"The word or phrase exactly as it is used."
- Changed
memory_archive2 fields changed- added
Input schema / properties / id / descriptionAdded value: +"Memory id (e.g. from memory_search, memory_recall or memory_list)." - added
Input schema / properties / namespace / descriptionAdded value: +"Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace."
- Changed
memory_boot1 field changed- added
Input schema / properties / namespace / descriptionAdded value: +"Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace."
- Changed
memory_delete2 fields changed- added
Input schema / properties / id / descriptionAdded value: +"Memory id (e.g. from memory_search, memory_recall or memory_list)." - added
Input schema / properties / namespaceAdded value: +{ + "description": "Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace.", + "type": "string" +}
- Changed
memory_export3 fields changed- added
Input schema / properties / format / descriptionAdded value: +"Output format. Only json is supported (default)." - added
Input schema / properties / namespace / descriptionAdded value: +"Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace." - added
Input schema / properties / type / descriptionAdded value: +"Only export one memory type (fact, event, preference, relationship, boundary, habit, decision, note). Omit for all."
- Changed
memory_get2 fields changed- added
Input schema / properties / id / descriptionAdded value: +"Memory id (e.g. from memory_search, memory_recall or memory_list)." - added
Input schema / properties / namespaceAdded value: +{ + "description": "Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace.", + "type": "string" +}
- Changed
memory_ingest7 fields changed- added
Input schema / properties / auto_extract / descriptionAdded value: +"Accepted for compatibility only; it is not stored and does not change processing." - added
Input schema / properties / conversation_id / descriptionAdded value: +"Conversation to append to. Defaults to the namespace's shared default conversation." - added
Input schema / properties / messages / descriptionAdded value: +"Messages in chronological order." - added
Input schema / properties / messages / items / properties / content / descriptionAdded value: +"Message text, or an OpenAI-style content parts array." - added
Input schema / properties / messages / items / properties / role / descriptionAdded value: +"One of user, assistant, system, tool. Other roles are dropped." - added
Input schema / properties / namespace / descriptionAdded value: +"Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace." - added
Input schema / properties / source / descriptionAdded value: +"Label for where the messages came from. Defaults to 'mcp'."
- Changed
memory_list7 fields changed- added
Input schema / properties / cursor / descriptionAdded value: +"Legacy paging cursor; only used when the lifecycle store is disabled. Prefer offset." - added
Input schema / properties / include_ids / descriptionAdded value: +"Legacy mode only: also return the bare id list." - added
Input schema / properties / limit / descriptionAdded value: +"Page size. Defaults to 100." - added
Input schema / properties / namespace / descriptionAdded value: +"Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace." - added
Input schema / properties / offset / descriptionAdded value: +"Number of records to skip. Use paging.next_offset from the previous page." - added
Input schema / properties / status / descriptionAdded value: +"Only list memories with this status: active (default), archived or superseded." - added
Input schema / properties / type / descriptionAdded value: +"Only list one memory type (fact, event, preference, relationship, boundary, habit, decision, note)."
- Changed
memory_pin3 fields changed- added
Input schema / properties / content / descriptionAdded value: +"The moment to keep, self-contained and readable on its own." - added
Input schema / properties / context_message_ids / descriptionAdded value: +"Ids of stored messages (e.g. message_ids from memory_ingest) to keep as surrounding context." - added
Input schema / properties / namespace / descriptionAdded value: +"Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace."
- Changed
memory_recall5 fields changed- added
Input schema / properties / k / descriptionAdded value: +"Maximum number of memory hits. Defaults to 20." - added
Input schema / properties / min_score / descriptionAdded value: +"Relevance floor applied after reranking. Defaults to the server setting (0.15)." - added
Input schema / properties / namespace / descriptionAdded value: +"Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace." - added
Input schema / properties / query / descriptionAdded value: +"The question or topic to recall memories for." - added
Input schema / properties / types / descriptionAdded value: +"Only recall these memory types (fact, event, preference, relationship, boundary, habit, decision, note). Omit for all types."
- Changed
memory_search4 fields changed- added
Input schema / properties / namespace / descriptionAdded value: +"Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace." - added
Input schema / properties / query / descriptionAdded value: +"What to look for, in natural language or keywords." - added
Input schema / properties / top_k / descriptionAdded value: +"Maximum number of records to return. Defaults to the server setting (50)." - added
Input schema / properties / types / descriptionAdded value: +"Only return these memory types (fact, event, preference, relationship, boundary, habit, decision, note). Omit for all types."
- Changed
memory_supersede8 fields changed- added
Input schema / properties / namespace / descriptionAdded value: +"Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace." - added
Input schema / properties / new_content / descriptionAdded value: +"The up-to-date memory text." - added
Input schema / properties / new_fact_key / descriptionAdded value: +"fact_key for the new entry. Defaults to the old entry's fact_key." - added
Input schema / properties / new_type / descriptionAdded value: +"One of fact, event, preference, relationship, boundary, habit, decision, note. Other values become fact (default)." - added
Input schema / properties / old_id / descriptionAdded value: +"Id of the memory being replaced." - added
Input schema / properties / reason / descriptionAdded value: +"Short note on why the old memory is outdated; stored with the chain." - added
Input schema / properties / response_tendency / descriptionAdded value: +"E-axis: how to respond when the new memory fires" - added
Input schema / properties / valid_as_of / descriptionAdded value: +"When the new fact became true (ISO date or datetime)."
- Changed
memory_upsert10 fields changed- changed
Input schema / properties / authored_by / descriptionPrevious value: -"E-axis signature; only honored on hand sources (mcp/manual/api)"New value: +"E-axis signature; only honored on hand sources (mcp/manual/api/remember_now)" - added
Input schema / properties / confidence / descriptionAdded value: +"0 to 1, how sure you are it is true. Defaults to 0.8." - added
Input schema / properties / content / descriptionAdded value: +"The memory text, one self-contained statement." - added
Input schema / properties / fact_key / descriptionAdded value: +"Stable key for this fact, e.g. 'user:favorite_coffee'. Reusing a key updates that memory. Facts about the outside world rather than the user use the prefix 'world_fact:'." - added
Input schema / properties / importance / descriptionAdded value: +"0 to 1, how much this matters. Defaults to 0.6." - added
Input schema / properties / namespace / descriptionAdded value: +"Memory space to use. Defaults to 'default'. Ignored when the API key is bound to a fixed namespace." - added
Input schema / properties / source / descriptionAdded value: +"Who is writing. Leave unset (defaults to 'mcp'). authored_by is only stored when source is mcp, manual, api or remember_now." - added
Input schema / properties / tags / descriptionAdded value: +"Free-form labels." - added
Input schema / properties / type / descriptionAdded value: +"One of fact, event, preference, relationship, boundary, habit, decision, note. Other values become fact (default)." - added
Input schema / properties / valid_as_of / descriptionAdded value: +"When the fact became true (ISO date or datetime). Used as the event date in recall."
14 tool updates
v0.1.0- First observed
diary_get - First observed
glossary_set - First observed
memory_archive - First observed
memory_boot - First observed
memory_delete - First observed
memory_export - First observed
memory_get - First observed
memory_ingest - First observed
memory_list - First observed
memory_pin - First observed
memory_recall - First observed
memory_search - First observed
memory_supersede - First observed
memory_upsert
TDQS
Scored across 14 tools
The retrieval trio (memory_search, memory_recall, memory_list) overlaps at a glance, but the descriptions explicitly cross-reference each other and clarify when to use which, and the write tools (upsert, supersede, archive, delete) are carefully distinguished. memory_ingest vs memory_upsert is also well delineated. Only the subtle search/recall split keeps it from a perfect score.
Every tool follows a consistent noun_verb pattern (memory_search, memory_upsert, glossary_set, diary_get). The memory_/glossary_/diary_ prefixes correspond cleanly to distinct resources rather than mixing arbitrary conventions.
14 tools sit comfortably in the well-scoped range for a memory system, and each tool has a distinct lifecycle role (search, recall, ingest, boot, pin, upsert, supersede, archive, delete, export, glossary, diary) without obvious redundancy.
The memory lifecycle is thoroughly covered from ingest through search, edit, supersede, archive and delete, plus glossary and diary reads. Minor gaps remain: no un-archive operation (explicitly noted) and no glossary deletion or diary write tool.
Maintenance
Related MCP Connectors
Persistent memory for AI agents across Claude, ChatGPT and any MCP client.
- mcpOAuthai.butlerbrain
Persistent memory for AI assistants. Save once; recall from Claude, ChatGPT, or any MCP client.
One memory, every AI. A shared, user-owned markdown memory your AI clients read and write over MCP.
Persistent, portable memory for AI assistants — your private memory graph, from any MCP client.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceProvides cross-device access to a persistent knowledge graph via Cloudflare Workers, enabling memory storage and retrieval through both MCP protocol and REST API with full-text search capabilities.-
- AlicenseNot gradedqualityAmaintenanceSelf-hosted semantic memory layer for Claude and MCP-compatible AI clients. Store notes, search by meaning not keywords, and recall relevant context automatically across sessions. Runs free on Cloudflare Workers, D1, Vectorize, and Workers AI4 npm800MIT
- AlicenseNot gradedqualityDmaintenanceSelf-hosted semantic memory for AI agents. Save worklogs, decisions, and notes via MCP, then recall them across sessions by meaning rather than keyword. Backed by Postgres + pgvector with local embeddings (multilingual-e5-base).1MIT
- FlicenseNot gradedqualityBmaintenanceA self-hosted AI memory service that shares durable project context between MCP-compatible clients such as Claude Code and Devin, via a REST API and a Streamable HTTP MCP server with tools to remember, recall, list, and forget memories.-