Remembra
Remembra is an external memory server for AI assistants that stores, searches, and manages persistent memories across sessions via MCP, HTTP, and SDKs.
Store memories of 11 types (fact, decision, role, history, observation, etc.) with trust, scope, importance, retention, and provenance metadata.
Batch operations: store, update, delete, or export multiple memories with per-item validation.
Search memories by keyword, type, scope, filters (archived, expired, quarantined), with ranked results and optional explanations.
List memories with pagination and filtering by type/scope/archive status.
Retrieve a single memory by ID, including related links and backlinks.
Create or remove typed relationships between memories (supports, contradicts, supersedes, refines, duplicates, related).
Update existing memories with optimistic concurrency and history snapshots.
Archive and revive memories to manage lifecycle without permanent deletion.
Permanently delete memories.
Digest a conversation transcript to automatically extract and store facts, decisions, roles, and history (requires LLM provider).
Run maintenance: decay old memories, auto-delete expired archived ones, and backfill embeddings.
Generate bounded, token-limited context for a session via memory_context / /api/v1/context.
Manage tenant-scoped data and enforce host-minted authorization in strict mode.
Export/import portable snapshots with validation and idempotent restore.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@RemembraRemember that we chose PostgreSQL 16 for this project."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Remembra 🧠
External memory for AI assistants that remains useful after the context window ends.
Remembra stores facts, decisions, preferences, roles, constraints, relationships, and project history outside a model context window. It then retrieves only the authorized, relevant subset for the current session. The same memory service is available through MCP, HTTP, a TypeScript SDK, and a web dashboard.
Current release: @hilbras/remembra@5.0.2 · V5.0.2 release notes
Why Remembra?
Long-running assistants fail in predictable ways: they repeat questions, lose decisions, confuse project context, and treat untrusted transcript text as trusted instructions. Remembra separates memory storage from the conversation and gives the host a small, inspectable control plane.
conversation ──► memory_store / memory_digest ──► durable memory
│
new session ──► memory_context / memory_search ◄───────┘
│
└─► bounded, relevant context onlyV5 adds the production boundaries needed for shared deployments:
Deterministic context assembly with explicit token and candidate budgets.
Host-resolved tenant identity; public callers cannot select an organization.
Strict and legacy modes with fail-closed migration/readiness checks.
Bounded hybrid retrieval across keyword, vector, trust, scope, time, and relations.
Verified recovery with signed snapshots, migration manifests, checkpoints, and rollback.
Compatibility first: V4.9 clients, Markdown, the eleven memory types, and the original thirteen MCP tools remain available.
Related MCP server: mcp-memory
Choose an integration
Integration | Best for | Identity/auth model |
MCP over stdio | OpenCode, Claude Code, Cline, Kimi Code | Local process; no network API key required |
HTTP API | Web apps, ChatGPT actions, service-to-service calls |
|
TypeScript SDK | Node and edge-compatible fetch clients | API key; SDK rejects caller-supplied identity fields |
Host TypeScript APIs | Multi-tenant services and operators | Opaque, host-minted |
Web dashboard | Browsing, editing, auditing, graph, operations | Served by the HTTP process; data calls remain authenticated |
Quick start
Requirements
Node.js
18.14.1or newer.No API key is needed for keyword-only local MCP or HTTP-on-loopback use.
SQLite is used by the CLI when
better-sqlite3is available; the Markdown backend remains supported for compatibility and library integrations.
Install
npm install -g @hilbras/remembraRun the MCP server
MCP is the default mode. Configure your client to launch the remembra binary over stdio:
remembraFor Claude Code:
claude mcp add remembra -- remembraFor OpenCode, add the following to ~/.config/opencode/opencode.json:
{
"mcp": {
"remembra": {
"type": "local",
"command": ["remembra"]
}
}
}See client setup for Cline, Kimi Code, and other clients.
Run HTTP mode and the dashboard
export REMEMBRA_API_KEY="$(openssl rand -hex 32)"
export REMEMBRA_HOME="$HOME/.remembra"
remembra --http --port 8787The server binds to loopback when no key is configured. A non-loopback host without a key is refused. Put a TLS reverse proxy in front of any public deployment.
curl http://127.0.0.1:8787/healthOpen http://127.0.0.1:8787/ for the dashboard. The shell is static and public on a keyed server; all data requests made by the page still require the API key.
Store and search over HTTP
curl -X POST http://127.0.0.1:8787/api/v1/memories \
-H "content-type: application/json" \
-H "x-api-key: $REMEMBRA_API_KEY" \
-d '{"type":"decision","content":"The project uses the /api/v1 namespace","scope":"project/demo","importance":4}'
curl "http://127.0.0.1:8787/api/v1/memories/search?query=api&scope=project/demo&limit=5" \
-H "x-api-key: $REMEMBRA_API_KEY"New integrations should use /api/v1/*. Existing unversioned routes remain supported for V4.9 compatibility.
Use the TypeScript SDK
The SDK is a side-effect-free fetch client; importing it does not start the CLI.
import { Remembra } from "@hilbras/remembra/sdk";
const memory = new Remembra({
endpoint: "http://127.0.0.1:8787",
apiKey: process.env.REMEMBRA_API_KEY,
});
await memory.store({
type: "fact",
content: "The project uses /api/v1",
scope: "project/demo",
});
const results = await memory.search({ query: "api", scope: "project/demo", limit: 5 });
const context = await memory.context({
query: "What architecture decisions should I remember?",
scope: "project/demo",
maxTokens: 4000,
limit: 50,
});
console.log(context.context);
console.log(context.tokenCount, context.retrievalMetadata.selectedCount);The SDK supports pagination, cancellation, structured errors, lifecycle operations, history, relations, batch operations, and the trusted tenant entity methods documented in docs/sdk.md.
A dependency-free Python client with synchronous and asynchronous surfaces is
available as hilbras-remembra; see docs/python.md.
The memory model
Every memory has a type, scope, provenance, trust level, importance, retention policy, and optional typed relations. The model deliberately separates what was observed from whether it should be trusted.
Eleven semantic types
Type | Use it for | Example |
| Stable knowledge | “The project uses PostgreSQL 16.” |
| User or team preferences | “Prefer concise answers.” |
| A choice already made | “Chose JWT over server sessions.” |
| A hard limit or prohibition | “Never commit secrets.” |
| Standing behavior | “Run tests before opening a PR.” |
| Persona or operating rule set | “Act as the staff engineer.” |
| A named person, service, or repository | “billing-service belongs to payments.” |
| A typed connection between entities | “billing-service depends on ledger-db.” |
| A dated occurrence | “The events table was migrated.” |
| Condensed chronology of work | “Authentication was redesigned in March.” |
| Raw signal awaiting validation | “p95 spiked after deployment.” |
Trust is a gate, not a score
system and verified content is treated as stronger evidence than trusted content; unverified content remains searchable but does not receive the standing-instruction boost. Conversation-derived roles and instructions land unverified until explicitly approved.
Treat retrieved memories as data with provenance, not automatically as commands. In particular, a global role or instruction is intentionally powerful and should be audited before it is trusted in a shared deployment. See security.
Scope and retrieval
A scope narrows relevance; it never replaces tenant authorization. In V5, global means global within the authenticated organization, not global across all tenants. Retrieval combines:
authorized tenant/project filters;
standing role and instruction gates;
keyword and optional embedding candidates;
trust, provenance, pinned retention, importance, and recency;
temporal filters and bounded one-hop relation expansion;
deterministic tie-breaking and diversity selection.
Oversized context candidates are skipped and counted. Remembra never silently truncates a memory or returns an over-budget context.
V5 context retrieval
memory_context and POST /api/v1/context use the same ranked, authorized retrieval path as search, then apply a deterministic token budget while walking the ranked results.
curl -X POST http://127.0.0.1:8787/api/v1/context \
-H "content-type: application/json" \
-H "x-api-key: $REMEMBRA_API_KEY" \
-d '{"query":"release and migration decisions","scope":"project/demo","maxTokens":4000,"limit":50,"explain":true}'The response contains:
{
"memories": [
{
"id": "memory-id",
"type": "decision",
"content": "The project uses the /api/v1 namespace",
"scope": "project/demo",
"trust": "trusted"
}
],
"context": "[memory-id] DECISION (scope: project/demo, trust: trusted)\nThe project uses the /api/v1 namespace",
"tokenCount": 812,
"retrievalMetadata": {
"query": "release and migration decisions",
"scope": "project/demo",
"maxTokens": 4000,
"tokenCounter": "conservative-estimate-v1",
"candidateCount": 50,
"selectedCount": 3,
"omittedCount": 2
}
}The default budget is 4000 tokens, the hard maximum is 100000, and the candidate cap is 100. Internal embedding vectors are never included in context responses. The context API is read-only and does not refresh recency.
Full contract: V5 context specification.
HTTP API
The HTTP process exposes the dashboard, health/readiness, metrics, memory operations, snapshots, administration, and V5 context/tenant routes.
Surface | Examples | Purpose |
Health and metrics |
| Readiness, liveness, Prometheus metrics |
Memory CRUD |
| Store, inspect, patch, and forget memories |
Retrieval |
| Ranked search and bounded context |
Lifecycle |
| Archive, restore, decay, and vector backfill |
Graph/history |
| Diffs, typed relations, and backlinks |
Batch/digest |
| Bounded writes/read-only search and LLM extraction |
Snapshots |
| Portable backup and idempotent restore |
Administration |
| Audit and operational visibility |
Tenant entities |
| Trusted organization, user, project, agent, and membership administration |
When REMEMBRA_API_KEY is set, data routes accept x-api-key or Authorization: Bearer. /health and the static dashboard shell are intentionally public; the shell contains no memory data. /api/v1 responses include X-Remembra-API-Version: v1.
The SDK preserves legacy response shapes and exposes server-managed identity rejection. For the complete route/error compatibility contract, see public API and stability.
MCP tools
The V4.9 manifest contains thirteen tools. V5 manifest version 2 adds memory_context without renaming or removing an existing tool.
Tool | Purpose |
| Persist one of eleven memory types |
| Bounded store, update, delete, selected export, or read-only search |
| Patch fields with optional optimistic concurrency |
| Park a memory without deleting it |
| Return an archived memory to active storage |
| Extract and store memories from a transcript |
| Retrieve relevant memories |
| Build deterministic, token-bounded V5 context |
| Browse stored memories with filters/pagination |
| Fetch one memory, relations, and backlinks |
| Add, remove, or retype graph relationships |
| View version history and line diffs |
| Run decay, deletion, and embedding backfill |
| Permanently delete one memory |
Full argument schemas and session-flow guidance are in docs/tools.md.
Tenant-safe deployments (V5)
V5 models an organization as the security boundary:
organization
├── users
├── projects
├── agents
└── memoriesThe host authenticates the caller, resolves current membership, and mints an opaque TenantContext. Remembra does not trust a tenant ID, organization ID, user ID, project ID, or agent ID supplied in an ordinary HTTP body, query parameter, MCP argument, or SDK payload.
Legacy and strict modes
Mode | Behavior |
| Explicit V4.9 compatibility. Reads/writes only the legacy namespace and refuses mixed tenant data. |
| Requires a current host-minted context for every data-plane operation; missing, stale, or mixed data fails closed. |
V4.9 memories do not silently acquire a tenant. A migration must explicitly assign them to an organization and produce a signed, checksummed manifest before strict rollout.
Local operator binding
The CLI can bind a local process to one trusted tenant through environment variables:
export REMEMBRA_TENANT_MODE=strict
export REMEMBRA_TENANT_ID=org-demo
export REMEMBRA_TENANT_MEMBERSHIP_VERSION=membership-42
export REMEMBRA_TENANT_PROJECT_ID=project-demo
export REMEMBRA_SNAPSHOT_KEY="$(openssl rand -hex 32)"
remembra --httpREMEMBRA_TENANT_ID is the opaque organization selector. The membership version must match the host's current directory state. Raw HTTP headers, CLI arguments, and SDK fields cannot replace these bindings.
Host integration
import { createTenantContext } from "@hilbras/remembra/tenant";
const tenant = createTenantContext({
organizationId: "org-demo",
membershipVersion: "membership-42",
projectId: "project-demo",
scopes: ["global", "project/project-demo"],
capabilities: ["tenant:read", "tenant:write"],
});
// Pass this opaque object only from trusted host code.
// A public transport resolver must return it after authentication.For organization administration, inject a TenantDirectory and TenantEntityService into the host application. Organization provisioning is default-deny unless an explicit authorizeBootstrap hook is supplied. Membership changes are versioned, bounded, and audited.
The complete isolation and migration contract is in V5 tenant specification. The security evidence requirements are in V5 threat model.
Storage, lifecycle, and recovery
Storage backends
The CLI prefers SQLite for local durability and search:
$REMEMBRA_HOME/
├── data.sqlite # default SQLite database
├── data.sqlite-wal/-shm # only while SQLite is open
├── .history/ # file-backend history, when applicable
└── .remembra.lock # cross-process mutation lockMemoryStore provides the readable Markdown backend and remains compatible with legacy data and explicit Markdown export/import. SQLite adds FTS5 when available, bounded candidate SQL, WAL mode, vector blobs, audit tables, and online backup support.
Memory IDs are globally unique. Relations, superseded references, snapshot references, and compressed references are validated against the same store before publication.
Lifecycle
Store: durable write with atomic file/database semantics.
Update: optimistic concurrency through
expectedVersion; a stale writer receivesCONFLICTand writes nothing.Archive/revive: reversible lifecycle transitions.
Decay: unused memories may be archived; expired archived memories may be deleted.
History: content changes retain bounded pre-images and line diffs.
Maintenance:
remembra maintainperforms decay and embedding backfill.
Snapshots and migration
# Portable snapshot; legacy mode accepts unsigned V4 snapshots.
remembra export backup.json
# Validate the entire snapshot without writing.
remembra import backup.json --dry-run
# Idempotent restore; existing IDs/duplicates are skipped.
remembra import backup.jsonIn strict mode, exports and imports use a canonical HMAC envelope and require REMEMBRA_SNAPSHOT_KEY. The complete snapshot/reference preflight happens before any write; per-record restore is idempotent but an operational failure after preflight can leave a partial application, so keep a verified backup. Tenant migration adds a signed manifest, checksum preflight, durable checkpoints, verified resume, failure records, and an explicit publication marker. For the V5.0.1 analyze/plan/apply commands, see the security and migration guide.
SQLite operators can use the verified recovery helpers from @hilbras/remembra/sqlite-recovery for online backup, integrity/schema checks, atomic restore, and retained-previous rollback. Close the live service before restoring and reject active SQLite sidecars.
See storage, migration, and public API.
Providers and optional intelligence
Storage and keyword retrieval work without any provider key. Configure providers only when you need digest extraction or semantic search.
# Hosted providers
export OPENAI_API_KEY=...
export REMEMBRA_LLM=openai
export REMEMBRA_EMBEDDINGS=openai
# Fully local Ollama
export REMEMBRA_LLM=ollama
export REMEMBRA_LLM_MODEL=llama3.2
export REMEMBRA_EMBEDDINGS=ollama
export REMEMBRA_EMBEDDING_MODEL=nomic-embed-textProvider calls have per-attempt timeouts, bounded retries/backoff, a wall-clock budget, cancellation, and normalized errors. Embedding failure degrades to keyword search; LLM digest failure does not partially store a transcript. The provider adapter contract is public through @hilbras/remembra/providers.
See provider configuration for all variables and adapter examples.
Security and operations
Remembra is designed to fail closed at the boundaries that matter:
API keys are compared with timing-safe equality; a non-loopback server without a key refuses to start.
Request bodies, batches, queues, provider calls, retrieval candidates, context budgets, and page sizes are bounded.
Tenant predicates are applied inside backend queries before limits, counts, ranking, and cursor creation.
Caches and provider work use tenant-safe partitions.
Public identity fields and tenant headers are rejected rather than trusted.
Snapshots, migrations, directories, and SQLite restores reject symlinks, oversized inputs, tampering, and invalid references.
Audit, structured logs, and Prometheus metrics are available for operational review.
Optional PII redaction (
REMEMBRA_REDACT=1) and AES-256-GCM file-backend encryption (REMEMBRA_ENCRYPT_KEY) are available for higher-risk deployments; SQLite, snapshots, and transport need separate volume/backup/TLS controls.
Deployment checklist
Use a long random
REMEMBRA_API_KEYfor every non-loopback HTTP deployment.Terminate TLS at a reverse proxy; do not expose plain HTTP directly to the internet.
Keep
REMEMBRA_HOME, snapshots, migration manifests, and encryption keys access-controlled.Use strict mode with a host-resolved identity system for shared or multi-tenant deployments.
Audit global roles and instructions regularly.
Test encrypted/signed backups and restore procedures before relying on them.
Monitor
/health,/metrics, structured logs, and queue/provider failures.Use filesystem encryption and OS process isolation in addition to application-level controls.
Read security, self-hosting, V5.0.2 authorization, the final V5 threat model, and observability before exposing a deployment.
Configuration reference
Variable | Default | Purpose |
|
| Local data root |
| unset | HTTP API/UI data authentication |
| loopback-safe | HTTP bind address |
|
| HTTP port; |
| enabled | Set |
| disabled | Set |
| disabled | 32-byte hex key for AES-256-GCM file-backend memory/history encryption (not SQLite/snapshots/transport) |
|
| Digest provider |
|
| Embedding provider; |
|
|
|
| — | Required organization selector in strict local mode |
| — | Required current membership version in strict local mode |
| — | 64-character hex HMAC key for strict snapshots |
Additional limits and provider controls are documented in self-hosting, security, and providers.
Architecture
┌──────────────────────────────────────────────────────────────┐
│ Transport adapters │
│ MCP stdio · HTTP /api/v1 · TypeScript SDK · dashboard │
└──────────────────────────┬───────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────┐
│ MemoryService │
│ authorization · ranking · context · lifecycle · snapshots │
│ digest · relations · jobs · provider policy │
└──────────────────────────┬───────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────┐
│ MemoryBackend │
│ SqliteBackend (default CLI) · MemoryStore (file backend) │
└──────────────────────────┬───────────────────────────────────┘
│
┌───────────────────┴───────────────────┐
▼ ▼
┌───────────────┐ ┌────────────────────┐
│ Provider │ │ Tenant directory │
│ adapters │ │ in-memory / file │
└───────────────┘ └────────────────────┘Key design properties:
transports are thin; policy and authorization live in the service/backend boundary;
V5.0.2 makes project/user/agent selectors conjunctive and requires explicit export authority;
SQLite and file backends implement the same tenant-aware contract;
atomic writes, advisory locking, and crash recovery protect local durability;
relation/history/audit/job paths apply the same tenant filter as primary reads;
provider failures are bounded and do not weaken storage correctness.
See architecture, memory model, and storage format for the detailed contracts.
Development
git clone https://github.com/Hilbras/Remembra.git
cd Remembra
npm install
npm run build
npm test
npm run docs:checkUseful commands:
Command | Purpose |
| Compile TypeScript and copy the dashboard |
| Run the full test suite |
| Validate relative documentation links |
| Run the tenant/security adversarial matrix |
| Run migration, snapshot, and SQLite recovery tests |
| Run deterministic 10K/50K scale benchmarks |
| Run isolated strict-tenant 10K/100K benchmarks |
| Run the complete fail-closed release gate |
| Run TypeScript in watch mode |
The release gate includes build, tests, security/recovery matrices, documentation, audit, package contents, and benchmarks. See V5 release gates before publishing.
Documentation map
Area | Documentation |
First install | |
Clients and MCP setup | |
HTTP and SDK | |
V5 context | |
Tenants and migration | Tenant contract · V5.0.2 authorization · V4.9 migration · V5.0.1 migration guide |
Security | |
Storage and recovery | |
Providers | |
Operations | |
Project process | |
Future architecture |
Compatibility and roadmap
V4.9 remains supported: legacy HTTP routes, Markdown, existing clients, and all thirteen original MCP tools remain available.
V5 is additive:
memory_context, tenant entities, and versioned APIs do not rename or remove the V4.9 surface.Current focus: hardening the production memory platform, operational recovery, and measurable retrieval quality.
V6 direction: see the V6 architecture specification for the security-first policy model, provider independence, offline-first core, migration lifecycle, and release roadmap.
Schema boundary: tenantless V4 records use schema
3; tenant records use schema4.
License
MIT © Hilbras
Available Tools
13 toolsmemory_archiveArchive a memoryA
Move a memory to the archived tree — out of search results but kept (and listed with includeArchived). Prefer this over forgetting when something may be needed again.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Memory id (from memory_store or memory_list) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It states the core non-destructive consequence—removed from search results but still accessible via includeArchived—and implies reversibility by contrasting with forgetting. It does not detail return values, idempotency, or behavior for already-archived memories, but it covers the key behavioral trait for this simple operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler: the first states the action and immediate consequence, the second gives selection guidance. The most important behavioral detail is front-loaded before the alternative reference.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, no-output-schema tool, the description is complete: it explains what happens, how archived memories can still be found, and when to use this over the main alternative. Nothing an agent needs to select or invoke the tool correctly is obviously missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single 'id' parameter is already described as coming from memory_store or memory_list. The tool description adds no extra parameter-level meaning, but since the schema fully documents it, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Move a memory to the archived tree') and clearly states the effect: out of search results but kept and listed with includeArchived. It also differentiates from the sibling memory_forget, so an agent can distinguish the two at a glance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names the alternative (memory_forget) and gives the condition for choosing this tool: 'Prefer this over forgetting when something may be needed again.' This gives the agent a clear decision rule for when to archive instead of delete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_batchRun a memory batchA
Run a bounded store, update, delete, or selected export batch. Items are validated before writes; operational failures are returned per item and are not a transaction.
| Name | Required | Description | Default |
|---|---|---|---|
| ids | No | Delete/export ids | |
| items | No | Store or update items | |
| operation | Yes | Batch operation |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden and does add meaningful execution details: 'Items are validated before writes; operational failures are returned per item and are not a transaction.' This clearly communicates the non-atomic, per-item failure behavior. It stops short of describing return shapes, permissions, or size limits, but the core execution semantics are transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no fluff. The action and scope are front-loaded, and the second sentence adds non-obvious behavioral constraints. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for selecting and using the tool at a high level: it names operations, validation, and non-transactional behavior. However, no output schema exists, and the description does not clarify the item structure for store/update, the meaning of 'bounded,' or the success response shape, so an agent is left with some uncertainty after invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so each parameter already has a description; the description adds little that is not already expressed by 'Delete/export ids' and 'Store or update items.' It does reinforce the operation-to-parameter mapping, but this is modest value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Run') and an explicit resource ('a bounded store, update, delete, or selected export batch'), with the four operation modes enumerated. It clearly differentiates itself from the single-operation sibling tools by being the batch variant.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'batch' and the operation list imply usage for grouped memory operations, but the description never explicitly says when to choose this over memory_store, memory_update, memory_delete, etc. There are no exclusions, prerequisites, or alternative-routing statements; usage is mostly inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_digestDigest a sessionA
Extract facts, decisions, roles and history from a conversation transcript and store them automatically (exact and near-identical duplicates are skipped; changed quantities go to the LLM merge). Call at the end of a session with the transcript or a detailed summary of it. Requires REMEMBRA_LLM + an API key.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No | Scope for extracted memories (default: global) | |
| source | No | Originating session/client | |
| transcript | Yes | Conversation transcript or a detailed summary of the session |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It does disclose deduplication behavior and the REMEMBRA_LLM/API key prerequisite, but the phrase 'changed quantities go to the LLM merge' is vague and unexplained, and side effects or failure behavior are not covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the main action, and includes dedup behavior, timing, and a prerequisite without filler. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The core operational context is present: what to do, when to call, and what input to provide. However, with no output schema, the description does not explain what the tool returns or how the LLM merge behaves, leaving some ambiguity for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds little beyond the schema, only reinforcing that transcript can be a detailed summary. No extra semantics are provided for scope or source.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool extracts facts, decisions, roles, and history from a transcript and stores them, which is a specific verb+resource. It does not explicitly name sibling tools to differentiate, but the session-digest framing distinguishes it from memory_store and memory_batch.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear timing guidance: 'Call at the end of a session' with a transcript or detailed summary. It does not mention when not to use the tool or alternatives, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_forgetDelete a memoryB
Permanently delete a memory by its id.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Memory id (from memory_store or memory_list) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does disclose the most important trait for a destructive operation: the deletion is "Permanently" irreversible. However, it says nothing about deletion side effects (e.g., impact on related memories given memory_relate and memory_digest siblings), behavior on a non-existent id, or required permissions, which leaves gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single six-word sentence with zero filler. The destructive scope "Permanently delete" is front-loaded, and the deletion key "by its id" completes the sentence efficiently. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter delete-by-id tool this is minimally adequate: the agent knows the action, the key, and the permanence. But with no annotations and no output schema, an agent is left guessing about error behavior (e.g., deleting a non-existent id), cascading effects on related memories, and confirmation semantics. A short note on side effects or missing-id handling would make it complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents id as "Memory id (from memory_store or memory_list)". The description only echoes the parameter relationship with "by its id" and adds no new semantics such as id format, validity constraints, or idempotency behavior. Baseline 3 is appropriate since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: "Permanently delete a memory by its id." The word "Permanently" usefully differentiates this from the sibling memory_archive (likely a soft-delete) without naming it. It slightly misses the top tier because it does not explicitly name the sibling it is not, but the intent is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives. The sibling list includes memory_archive, memory_revive, memory_update, and memory_get, any of which could be relevant alternatives for a user considering deletion, yet the description offers no routing or exclusion. The usage context is only weakly implied by the purpose statement itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_getGet a memoryA
Fetch one memory by id with its related links and backlinks (memories that point at it). Use after memory_search when you need the full statement, not the snippet.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Memory id — returns the memory with related links and backlinks |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses that the tool returns the memory plus related links and backlinks, and implies a read-only operation via 'Fetch'. However, it does not explicitly state the absence of side effects, authentication requirements, or error behavior (e.g., if the id does not exist). It adds some behavioral context but leaves gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core operation and return content, followed by a targeted usage note. No extraneous information; each sentence contributes meaningfully.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with no annotations and no output schema, the description covers the essential purpose, return contents, and a usage trigger. It is somewhat incomplete regarding error cases or side-effect confirmation, but these are minimal for a 'get' operation. Overall it provides enough for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes the 'id' parameter with 100% coverage ('Memory id — returns the memory with related links and backlinks'). The description does not add new parameter-level meaning; it repeats the same return behavior. Baseline 3 applies under high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Fetch', names the resource ('one memory by id'), and specifies the return content ('related links and backlinks'). It also distinguishes from memory_search by noting it returns the full statement rather than a snippet, making the tool's role clear relative to siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use after memory_search when you need the full statement, not the snippet,' which names the relevant alternative and the condition for using this tool. This is clear when-to-use guidance, and it also implies when not to use it (when a snippet suffices).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_historyShow memory historyA
Version history of one memory with unified line diffs — every content-changing update (e.g. a contradiction merge) snapshots the previous version. Newest first.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Memory id to show version history for | |
| limit | No | Max past versions to return (newest first) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and it does add value by disclosing snapshots on content-changing updates, unified line diffs, and newest-first ordering. However, it does not state that the operation is read-only, whether non-content changes are excluded, or any side effects or access requirements, leaving notable gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the first sentence states the core purpose, and the em-dash clause adds the snapshot rule without excess. The final 'Newest first' specifies ordering efficiently. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must compensate for return expectations. It gives high-level behavior (diffs, newest first) but does not specify the shape of each version entry (e.g., timestamps, fields, diff format). The tool is simple, but an agent lacks precise output structure knowledge.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the schema already defines each parameter clearly. The description adds no parameter-specific meaning beyond what the schema provides, so a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly names the action ('show version history'), the resource ('one memory'), and the key qualifier 'one memory' that distinguishes it from collection-oriented siblings like memory_list and memory_batch. It also implies it is not the current-state tool (memory_get) by emphasizing 'history'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case—retrieving a memory's historical versions—but does not explicitly state when to choose this over alternatives like memory_get, nor does it provide any exclusions or alternative routing. The usage context is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_listList memoriesB
List stored memories, optionally filtered by scope or type; paginate with offset/limit.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | ||
| limit | No | Pagination: max memories to return | |
| scope | No | ||
| offset | No | Pagination: skip this many matching memories | |
| includeFuture | No | ||
| includeExpired | No | ||
| includeArchived | No | Include archived memories (flagged) | |
| includeQuarantined | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions filtering and pagination but does not state default behavior for the boolean include flags (e.g., whether expired/archived are included by default), ordering, or whether the operation is read-only (though implied). This is a significant gap for a list tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the primary action and immediately covers filtering and pagination. No wasted words or redundant details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters, no annotations, and no output schema, the description is insufficient. It does not explain default filter behavior, response format, or how it differs from sibling list-like tools (memory_search, memory_history). An agent would need to infer too much to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 38% (3 of 8 params have descriptions). The description adds meaning for scope/type and offset/limit (matching the text), but completely omits explanation for includeFuture, includeExpired, includeQuarantined, and only partially covers includeArchived (which the schema already describes). Since coverage is low, the description should compensate more thoroughly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), a clear resource ('stored memories'), and mentions filtering and pagination. It is unambiguous about what the tool does, though it does not explicitly name sibling alternatives like memory_search or memory_get to differentiate itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (listing memories with optional filters) but provides no explicit guidance on when to prefer this over memory_search, memory_history, or memory_get. There are no exclusion criteria or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_maintainRun maintenanceA
Run maintenance: archive memories unused past REMEMBRA_ARCHIVE_AFTER_DAYS (default 90), auto-delete archived memories past REMEMBRA_ARCHIVE_TTL_DAYS (default 365), and backfill missing embedding vectors. Roles never decay. Safe to call anytime.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it delivers. It discloses destructive auto-deletion, archiving conditions, TTL defaults, embedding backfill, and explicitly states that roles never decay. It also reassures that calling it is safe anytime.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one dense but well-organized sentence that front-loads the action and then lists each behavior clearly. There is no filler, and every phrase contributes operational meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless maintenance tool with no annotations and no output schema, the description covers what happens, when it happens, configuration defaults, and safety. Nothing needed to invoke the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema carries no burden. The description adds useful configuration context by naming REMEMBRA_ARCHIVE_AFTER_DAYS and REMEMBRA_ARCHIVE_TTL_DAYS with defaults, which is more than baseline for an empty-schema tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Run maintenance') and enumerates concrete operations: archive unused memories, auto-delete old archived memories, and backfill missing embeddings. This clearly differentiates it from the sibling list/forget/store/digest/search tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not name alternatives, but it gives explicit timing guidance: 'Safe to call anytime.' Since this is a standalone maintenance operation rather than a choice among competing query/update tools, this contextual guidance is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_relateLink memoriesA
Create or remove directed links between memories (the relationship graph): e.g. tie a decision to the facts it depends on, or a history entry to the decision it records. Targets must exist; backlinks are visible via memory_get.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Source memory id | |
| kind | No | Edge kind (plan §4.7): supports | contradicts | supersedes | refines | duplicates | related | related |
| action | No | add (default) creates links, remove deletes them | add |
| related | Yes | Target memory ids to link to / unlink from |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden. It does add useful behavioral context: links are directed, targets must pre-exist, and backlinks are observable via memory_get. But it omits other behavioral traits an agent may need — whether add is idempotent when a link already exists, whether a partially failing batch is transactional, or error behavior on missing targets. Adequate but with clear gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The core action and domain are front-loaded in the first sentence, while the second adds examples plus two constraints (target existence, backlink visibility). Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a graph-modifying tool with a rich schema (enums fully enumerated, 100% param coverage) and no output schema, the description covers the essential operational facts: what it does, a prerequisite, and how to observe results. Minor omissions like idempotency and partial-failure behavior exist, but the essentials an agent needs to invoke it correctly are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters and the two enums (kind with six values, action with two). Per the baseline for high schema coverage, the description need not repeat parameter details. It does add genuine value by explaining the semantics of the link kinds ('tie a decision to the facts it depends on'), but that is contextual enrichment rather than compensation for a documentation gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (create/remove), a resource (directed links between memories), and names the domain ('relationship graph'). The examples — tying a decision to its facts, or a history entry to the decision it records — make the resource concrete and clearly distinguish this from sibling tools like memory_store (storing a memory) and memory_get (reading one). An agent can tell what this does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The examples imply when to use it (relating decisions, facts, history entries) and it states a concrete prerequisite ('Targets must exist'). It also routes the agent to memory_get for viewing backlinks. However, it never explicitly names alternatives or states when NOT to use this tool versus memory_update or memory_store. The usage context is present but implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_reviveRevive an archived memoryB
Bring an archived memory back to active search.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Memory id (from memory_store or memory_list) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It indicates a state change (archived to active) but does not specify whether the operation is idempotent, what happens if the memory is already active, whether it is reversible, or any potential side effects. For a mutation tool, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no wasted words. It is front-loaded with the action and result, making it immediately clear. For a tool with only one parameter and no complexity, this is appropriately sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, no output schema, no annotations), the description is minimally sufficient but lacks important operational context. It does not mention error conditions, idempotency, or prerequisites (e.g., that the memory must already be archived). An agent could call it correctly based on the schema, but would benefit from a note about failure cases or state constraints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides 100% coverage of the single 'id' parameter with a description indicating it comes from memory_store or memory_list. The description adds no additional meaning beyond the schema, so the baseline of 3 applies. No extra context about the parameter is needed, as the schema is fully descriptive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a clear verb ('bring back') and specific resource ('archived memory') plus the result ('active search'). It clearly indicates the tool unarchives a memory, but it does not explicitly distinguish from siblings like memory_archive (which archives) or memory_update (which could modify status). However, the purpose is unambiguous enough for an agent to infer its role among the sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. For instance, it does not mention that this should be used when a memory needs to become searchable again, or that memory_archive is the inverse. The description simply states the action without context on conditions or exclusions, leaving an agent to infer usage from the name and siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_searchSearch memoriesB
Retrieve relevant memories from external storage. Call this at the start of a session (or whenever prior context might exist) to recover facts, decisions, roles and history.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | ||
| limit | No | ||
| query | No | Keywords to match (omit to get a scope/recency-ranked list) | |
| scope | No | Current project path or workspace id to filter by | |
| explain | No | Include per-memory score breakdown (V4.2.0+) | |
| includeFuture | No | ||
| includeExpired | No | ||
| includeArchived | No | ||
| includeQuarantined | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden, but it only says memories are 'relevant' and recoverable. It does not explain matching semantics, ordering, return format, or how filters like includeExpired/includeArchived/includeQuarantined affect results. This is a significant transparency gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core retrieves role, and the usage note is valuable. However, the phrase 'facts, decisions, roles and history' largely duplicates enum values already present in the schema, so not every word adds unique value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 parameters, no annotations, and no output schema, this description is insufficient for an agent to fully understand how to invoke the tool correctly. It tells when to call it but not how filtering flags, limits, or explain scoring behave, leaving key invocation details unresolved.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, so the description must compensate for the many undocumented parameters. It does not; it only lists memory categories like facts, decisions, roles, and history, which partially overlap the 'type' enum. The six undocumented include flags and limit behavior remain unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Retrieve') and resource ('relevant memories from external storage'), and its focus on relevance and session-start recovery distinguishes it reasonably from memory_list or memory_get. However, it does not explicitly differentiate from siblings, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Call this at the start of a session (or whenever prior context might exist)' is explicit, actionable guidance on when to use the tool. It lacks any mention of when NOT to use it or which sibling tools are better suited for specific cases, so it does not fully earn a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_storeStore a memoryB
Persist a fact, decision, role or history entry so it survives context window resets. Use type 'fact' for stable knowledge, 'decision' for choices already made, 'role' for standing instructions/roles, 'history' for condensed chronology of past work.
| Name | Required | Description | Default |
|---|---|---|---|
| meta | No | V4.4/V4.5: security, lifecycle, and compression metadata | |
| tags | No | Keywords that boost retrieval | |
| type | Yes | fact | preference | decision | constraint | instruction | role | entity | relationship | event | history | observation | |
| owner | No | V4.7: user | agent | project | organization | global | |
| scope | No | 'global' for always-relevant memories, or a project path/id for project-scoped ones | global |
| trust | No | Trust classification; default derived from provenance (direct store → trusted, conversation digest → unverified) | |
| access | No | V4.7: private | shared | global | |
| source | No | Originating session or client | |
| content | Yes | The memory itself, written as a standalone statement | |
| retention | No | Decay protection: decaying (default) | pinned (never decays + rank boost) | persistent (archivable, never auto-deleted) | neverExpire (fully exempt) | ephemeral (accelerated clock lands with 4.5.0; today behaves as decaying) | |
| validFrom | No | ISO timestamp when this claim becomes valid | |
| confidence | No | Certainty of this claim 0..1, independent of importance (default 1.0 direct, 0.7 digests) | |
| importance | No | 1=minor, 5=critical (default 3) | |
| observedAt | No | ISO timestamp of the original observation (for backfills) | |
| provenance | No | Provenance (plan §4.3); defaults to { sourceType: manual } | |
| validUntil | No | ISO timestamp when this claim ceases to be valid | |
| supersededBy | No | ID of the memory that supersedes this one |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It states persistence across resets, which implies a write operation, but does not mention side effects, overwrite/deduplication behavior, permission requirements, or what happens on success. For a mutation tool with zero annotation coverage, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise: a single sentence stating purpose followed by one sentence of type guidance. It front-loads the core action and avoids redundancy. However, it could be slightly more structured by grouping related parameters, but for its length it is efficient with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 17 parameters including nested objects, the description is minimal. It does not explain the other type enums (preference, constraint, instruction, entity, relationship, event, observation), nor does it clarify owner/scope/trust/access semantics or the provenance object. There is no output schema, so return value is unaddressed. The description is inadequate for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all 17 parameters. The description adds semantic guidance for the 'type' parameter by mapping specific types to use cases, which adds value beyond the schema's enum listing. However, it does not add meaning to other parameters like owner, scope, trust, or retention, which remain schema-only. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool persists a memory entry to survive context resets, and gives specific examples of memory types (fact, decision, role, history). This distinguishes it from read-oriented tools like memory_search or memory_list, but it does not explicitly compare to other write tools like memory_batch or memory_update. The verb 'persist' and resource are clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on selecting the 'type' field (fact, decision, role, history) with conditions for each. However, it does not say when to use this tool vs. memory_batch (for storing multiple) or memory_update (for modifying existing entries). No exclusion or alternative tool names are mentioned. Usage guidance is partial.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_updateUpdate a memoryA
Patch an existing memory by id — any subset of type/content/scope/tags/importance/confidence/source. Changing scope moves it between trees; a content change refreshes its embedding and snapshots the old version into memory_history.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Memory id to update | |
| tags | No | Keywords that boost retrieval | |
| type | No | fact | preference | decision | constraint | instruction | role | entity | relationship | event | history | observation | |
| scope | No | 'global' for always-relevant memories, or a project path/id for project-scoped ones | global |
| trust | No | Trust classification; default derived from provenance (direct store → trusted, conversation digest → unverified) | |
| reason | No | Why this version supersedes the last — recorded in the history entry (plan §4.6) | |
| source | No | Originating session or client | |
| content | No | The memory itself, written as a standalone statement | |
| retention | No | Decay protection: decaying (default) | pinned (never decays + rank boost) | persistent (archivable, never auto-deleted) | neverExpire (fully exempt) | ephemeral (accelerated clock lands with 4.5.0; today behaves as decaying) | |
| confidence | No | Certainty of this claim 0..1, independent of importance (default 1.0 direct, 0.7 digests) | |
| importance | No | 1=minor, 5=critical (default 3) | |
| expectedVersion | No | Optimistic concurrency: update fails with CONFLICT (HTTP 409) unless the stored version matches |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It successfully discloses two important side effects: changing scope moves the memory between trees, and a content change refreshes the embedding and snapshots the old version into memory_history. This adds meaningful behavioral context beyond the schema. However, it does not mention concurrency control (expectedVersion) or potential conflict behavior, which could be relevant for agents.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero fluff. It front-loads the primary action and allowed fields, then adds the two critical behavioral caveats. Every clause earns its place, and the language is precise and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 12 parameters, a rich schema, and no output schema, the description covers the essential operational aspects: what fields can be updated and the key side effects. It does not explain return values, error scenarios (e.g., 409 conflict with expectedVersion), or permission requirements, but these are either implied by the schema or not critical for correct invocation. The absence of an output schema means the agent cannot expect structured response details, which is fine. Overall, the description is sufficient for calling the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds semantic value by explaining the effect of specific parameters: scope changes cause tree movement, and content changes trigger embedding refresh and history snapshots. This goes beyond the schema's parameter descriptions, which only state defaults and basic meaning. Other parameters like trust, retention, and reason are not elaborated, but the schema descriptions are adequate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the operation (Patch), the target (an existing memory by id), and the supported fields. It distinguishes itself from sibling tools like memory_store (create) and memory_forget (delete) by explicitly limiting to existing memories and listing updateable attributes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no explicit guidance on when to use this tool vs. its siblings. While 'Patch an existing memory' implies it is for updates rather than creation, it does not mention alternatives like memory_store for new memories or memory_get for retrieval, nor does it specify prerequisites (e.g., that the memory must already exist). This leaves the agent to infer usage from the operation nature.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
12 tool updates
v4.8.0- Added
memory_archive - Added
memory_batch - Changed
memory_digest1 field changed- added
Input schema / properties / transcript / minLengthAdded value: +1
- Changed
memory_forget1 field changed- added
Input schema / properties / id / descriptionAdded value: +"Memory id (from memory_store or memory_list)"
- Added
memory_get - Added
memory_history - Changed
memory_list7 fields changed- added
Input schema / properties / includeArchivedAdded value: +{ + "description": "Include archived memories (flagged)", + "type": "boolean" +} - added
Input schema / properties / includeExpiredAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / includeFutureAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / includeQuarantinedAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / limitAdded value: +{ + "description": "Pagination: max memories to return", + "maximum": 500, + "minimum": 1, + "type": "integer" +} - added
Input schema / properties / offsetAdded value: +{ + "description": "Pagination: skip this many matching memories", + "minimum": 0, + "type": "integer" +} - changed
Input schema / properties / type / enumPrevious value: -[ - "fact", - "decision", - "role", - "history" -]New value: +[ + "fact", + "preference", + "decision", + "constraint", + "instruction", + "role", + "entity", + "relationship", + "event", + "history", + "observation" +]
- Added
memory_relate - Added
memory_revive - Changed
memory_search6 fields changed- added
Input schema / properties / explainAdded value: +{ + "description": "Include per-memory score breakdown (V4.2.0+)", + "type": "boolean" +} - added
Input schema / properties / includeArchivedAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / includeExpiredAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / includeFutureAdded value: +{ + "type": "boolean" +} - added
Input schema / properties / includeQuarantinedAdded value: +{ + "type": "boolean" +} - changed
Input schema / properties / type / enumPrevious value: -[ - "fact", - "decision", - "role", - "history" -]New value: +[ + "fact", + "preference", + "decision", + "constraint", + "instruction", + "role", + "entity", + "relationship", + "event", + "history", + "observation" +]
- Changed
memory_store18 fields changed- added
Input schema / properties / accessAdded value: +{ + "description": "V4.7: private | shared | global", + "enum": [ + "private", + "shared", + "global" + ], + "type": "string" +} - added
Input schema / properties / confidenceAdded value: +{ + "description": "Certainty of this claim 0..1, independent of importance (default 1.0 direct, 0.7 digests)", + "maximum": 1, + "minimum": 0, + "type": "number" +} - added
Input schema / properties / content / minLengthAdded value: +1 - added
Input schema / properties / importance / defaultAdded value: +3 - added
Input schema / properties / metaAdded value: +{ + "additionalProperties": false, + "description": "V4.4/V4.5: security, lifecycle, and compression metadata", + "properties": { + "compressedFrom": { + "items": { + "type": "string" + }, + "type": "array" + }, + "compressionAt": { + "type": "string" + }, + "contradicted": { + "type": "boolean" + }, + "injected": { + "type": "boolean" + }, + "quarantined": { + "type": "boolean" + } + }, + "type": "object" +} - added
Input schema / properties / observedAtAdded value: +{ + "description": "ISO timestamp of the original observation (for backfills)", + "type": "string" +} - added
Input schema / properties / ownerAdded value: +{ + "description": "V4.7: user | agent | project | organization | global", + "enum": [ + "user", + "agent", + "project", + "organization", + "global" + ], + "type": "string" +} - added
Input schema / properties / provenanceAdded value: +{ + "additionalProperties": false, + "description": "Provenance (plan §4.3); defaults to { sourceType: manual }", + "properties": { + "agentId": { + "description": "Agent that produced the memory", + "type": "string" + }, + "agentType": { + "description": "Agent role, e.g. researcher or coder", + "type": "string" + }, + "agentVersion": { + "description": "Version of the agent that produced the memory", + "type": "string" + }, + "conversationId": { + "description": "Conversation this memory belongs to", + "type": "string" + }, + "messageId": { + "description": "Message within that session", + "type": "string" + }, + "provider": { + "description": "Provider/model (e.g. the digest LLM)", + "type": "string" + }, + "runId": { + "description": "Execution run this memory belongs to", + "type": "string" + }, + "sessionId": { + "description": "Session that produced the memory", + "type": "string" + }, + "sourceType": { + "description": "manual (default) | conversation | agent | import | system", + "enum": [ + "manual", + "conversation", + "agent", + "import", + "system" + ], + "type": "string" + }, + "taskId": { + "description": "Task this memory belongs to", + "type": "string" + } + }, + "type": "object" +} - added
Input schema / properties / retentionAdded value: +{ + "description": "Decay protection: decaying (default) | pinned (never decays + rank boost) | persistent (archivable, never auto-deleted) | neverExpire (fully exempt) | ephemeral (accelerated clock lands with 4.5.0; today behaves as decaying)", + "enum": [ + "pinned", + "persistent", + "ephemeral", + "decaying", + "neverExpire" + ], + "type": "string" +} - added
Input schema / properties / scope / defaultAdded value: +"global" - added
Input schema / properties / supersededByAdded value: +{ + "description": "ID of the memory that supersedes this one", + "type": "string" +} - added
Input schema / properties / tags / defaultAdded value: +[] - added
Input schema / properties / tags / descriptionAdded value: +"Keywords that boost retrieval" - added
Input schema / properties / trustAdded value: +{ + "description": "Trust classification; default derived from provenance (direct store → trusted, conversation digest → unverified)", + "enum": [ + "unverified", + "trusted", + "verified", + "system" + ], + "type": "string" +} - added
Input schema / properties / type / descriptionAdded value: +"fact | preference | decision | constraint | instruction | role | entity | relationship | event | history | observation" - changed
Input schema / properties / type / enumPrevious value: -[ - "fact", - "decision", - "role", - "history" -]New value: +[ + "fact", + "preference", + "decision", + "constraint", + "instruction", + "role", + "entity", + "relationship", + "event", + "history", + "observation" +] - added
Input schema / properties / validFromAdded value: +{ + "description": "ISO timestamp when this claim becomes valid", + "type": "string" +} - added
Input schema / properties / validUntilAdded value: +{ + "description": "ISO timestamp when this claim ceases to be valid", + "type": "string" +}
- Added
memory_update
6 tool updates
v0.4.0- First observed
memory_digest - First observed
memory_forget - First observed
memory_list - First observed
memory_maintain - First observed
memory_search - First observed
memory_store
TDQS
Scored across 13 tools
Most tools are clearly distinct: store, get, update, forget, archive, revive, search, list, relate, history, and maintain each map to a separate operation. The main ambiguity is memory_batch, which wraps store/update/delete/export and could be confused with the individual write/update/delete tools, though its batch purpose is stated.
The memory_ prefix is consistent and most suffixes are action verbs: store, search, list, forget, get, relate, update, archive, revive, maintain. memory_batch and memory_history break the verb pattern, but the overall convention remains predictable and readable.
Thirteen tools is well within the ideal range and each tool covers a distinct aspect of memory management: CRUD, search, versioning, relationships, archival, batch operations, ingestion, and maintenance. No tool feels redundant or unnecessary.
The surface fully covers memory lifecycle: store, retrieve, list, search, update, delete, archive, revive, version history, relation management, batch processing, and automated digesting from transcripts. There are no obvious dead ends or missing core operations for the stated purpose.
Maintenance
Related MCP Connectors
Cross-tool persistent memory and context for AI assistants over MCP.
- mcpOAuthai.butlerbrain
Persistent memory for AI assistants. Save once; recall from Claude, ChatGPT, or any MCP client.
Persistent personal memory for AI assistants — save, search, and recall across every MCP client.
Persistent memory for AI agents across Claude, ChatGPT and any MCP client.
Related MCP Servers
- AlicenseAqualityDmaintenanceProvides persistent memory for AI assistants via MCP, enabling them to store and recall facts, preferences, and tasks across conversations using either local file storage or a cloud backend with semantic search.55 npmMIT
- AlicenseAqualityDmaintenanceProvides persistent memory with semantic search for MCP-based AI agents, enabling them to store and recall information across sessions using vector embeddings.41MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to have persistent long-term memory by automatically storing and retrieving important information via MCP tools.MIT
- AlicenseNot gradedqualityBmaintenanceProvides a persistent, cross-tool memory layer for AI coding agents via MCP, enabling storage and retrieval of decisions, preferences, and context across different tools and models.1 npm1MIT